Alibaba’s Qwen team open-sourced Qwen3.8-Flash-Next on August 26 — a 125B multimodal MoE model that activates just 6B parameters per token. It’s the first public look at the Qwen4 architecture — not a paper, actual weights on Hugging Face (qwen-community-1.0 license).
What the new architecture changes
The story isn’t size, it’s cost. GDN + QSA hybrid attention (three of every four layers run linear attention), Gated Residual, a 51B N-gram embedding table, and the Muon optimizer cut training cost to roughly 1/9 of Qwen3.7-Plus — while beating it across the board. Coding gains the most: 62.5 on SWE-bench Pro, 91.9 on LiveCodeBench v6. Context is 262K native, 1M with YaRN, with up to 7.6× faster prefill at the long end.
Run it now, or take the API
The weights already run on vLLM, SGLang, and llama.cpp — Simon Willison had quantized builds going on day one. The production version hits QwenCloud API at $0.16/1M input and $0.47/1M output, pricing built for agents that burn tokens all day: coding agents, long-document office work, video understanding. It already powers Qwen Code and QwenWork.
If 6B active parameters can match a flagship, everyone else’s inference bill is now a strategy question.
You Might Also Like
- Alibaba Qwen Smart Glasses g1 s1 275 ai Glasses With Swappable Batteries and a Qwen api Backend
- Qwen Image 3 0 Alibaba Renders 10px Text and Swallows 4 5k Token Prompts
- Alibaba Qwen 3 5 Just Dropped and it Brought 10 Million Milk Teas With it
- Ggml Llama cpp Joins Hugging Face and Honestly it was Only a Matter of Time
- 397 Billion Parameters on a 48gb Macbook Flash moe Turns Apples 2023 Research Into Reality

Leave a comment