Ant Group’s inclusionAI shipped a 124B MoE model that activates only ~5.1B parameters per token — 1/8 the size and 1/12 the active compute of its own 1T flagship — and claims it matches or beats it on most benchmarks. That claim is vendor-only for now, no public benchmark table yet. But the architecture is the real story.
The architecture bet
This isn’t a shrunk-down flagship. Ling-3.0-flash uses native hybrid linear attention from pretraining: Kimi Delta Attention (yes, Moonshot’s KDA) stacked 5:1 with MLA, plus 1/64 expert sparsity. Result: 256K context, ~1,000 tokens/sec claimed on Ant’s infra, 1M context on the roadmap. Chinese labs are now openly borrowing each other’s attention research — that’s how fast this race compounds.
Free API, right now
It’s OpenAI-compatible and free for a limited time via OpenRouter (inclusionai/ling-3.0-flash), ZenMux, and Novita. Tool use and function calling work, and the thinking/non-thinking mode switch plus cheap active compute make it built for agent workloads — long multi-step tool chains where token costs usually kill you. Announced as open-weight under Apache 2.0; weights haven’t landed on Hugging Face yet. Watch whether they actually ship.
You Might Also Like
- Tencent Hunyuan hy3 Preview Goes Open Source 295b moe 21b Active 256k Context
- Inkling Thinking Machines lab Mira Muratis First Model Ships as a 975b Open Weight moe you can Download
- Ant Group Open Sources Lingbot map 98 98 f1 on Eth3d 21 Points Ahead of Everyone Else
- Openais own Test Models Escaped Their Sandbox and Breached Hugging Face to Cheat on a Benchmark
- Hugging Face Speech to Speech Open Source Local Voice Agents is the Openai Realtime Clone you can run on Your own gpu

Leave a comment