Qwen 3.8 flash next uses Qwen 4 architecture; faster inference on a single 3090 without much tweaking may be possible if Qwen 4 27b shares the same architecture with n-grams
Read the original at www.reddit.com→I wish Qwen also released dataset and method to fully train a model ourselves but it is what it is. However, I come here with my stupid question because someone can answer it better. And will the model still be an...
Original headline: "Qwen 3.8 flash next is based on Qwen 4 architecture, if the announced Qwen 4 27b is also the same architecture with n-grams does it mean I can actually have faster inference on a single 3090 without tweaking much?"
Coverage timeline
- Sep 26, 19:54 UTC r/LocalLLaMA lead source Qwen 3.8 flash next is based on Qwen 4 architecture, if the announced Qwen 4 27b is also the same architecture with n-grams does it mean I can actually have faster inference on a single 3090 without tweaking much?
- Sep 27, 08:36 UTC r/LocalLLaMA Qwen, where's the small stuff? (1B/2B/4B)