884 points on HackerNews, the biggest AI story of the day. Alibaba just took Qwen3.8-Max out of preview: a 2.4 trillion parameter MoE model that activates only 95B per request, built for coding and long-horizon agent work.
The numbers that matter
Terminal-Bench 2.1: 86.6 — above Claude Opus 4.8 and Fable 5 (84.6), behind only GPT-5.6 Sol max (88.8). PaperBench: 93.0, ahead of every frontier model. Alibaba says it can code autonomously for 10+ days straight, and its internal team shipped a full software engineering project with it in 16 days.
API and open weights
It’s live on QwenCloud with an OpenAI-compatible API — point your existing agent stack at a new base URL and go. Obvious use cases: repo-scale refactors, multi-day autonomous coding runs, research reproduction. Weights land on Hugging Face and ModelScope next week, alongside a smaller Qwen3.8-27B.
Why this is a big deal
The preview three weeks ago was a teaser. This is GA plus an open-weights commitment — an open model now trades blows with closed frontier flagships on the benchmark that matters most for agentic coding. The gap keeps shrinking.
You Might Also Like
- Kimi k2 6 Beats gpt 5 4 and Claude Opus 4 6 on swe Bench pro
- Gpt 5 5 Takes Back the Coding Crown From Claude Opus 4 7
- Gpt 5 6 sol Ultra Hits 91 9 on Terminal Bench 2 1 and its Landing in Codex
- Alibabas Qwen 3 8 Qwen3 8 max Preview Claims its Second Only to Fable 5 With Zero Benchmarks Published
- Claude Opus 5 Lands Within 0 5 of Fable 5 on Coding at Half the Price

Leave a comment