Alibaba’s Qwen team shipped the third generation of its image model today. HN front page within hours — 74 points, 40 comments before lunch.
What it actually is
One foundation model that generates and edits, no separate editing checkpoint. It takes instructions up to 4.5k tokens (2.0 capped around 1k), renders text down to 10px, does native rendering in 12 languages, and fills a 3×3 grid layout in a single pass. Newspapers, storyboards, exam papers with LaTeX, fake web UIs, livestream overlays — the text-dense stuff every other model turns into soup. The rest of the work went into faces: hair strands, skin pores.
Where the API fits
Qwen models ship on DashScope behind an OpenAI-compatible endpoint, so dropping this into an existing pipeline is a base-URL change. Use cases are the ones nobody wants to hire a designer for: poster variants, slide decks, localized ad creative in 12 languages, infographics straight from numbers.
Qwen-Image-2.0 landed in February — 7B, native 2K, first place on AI Arena for both generation and editing. Every version in this line eventually hit Hugging Face under Apache 2.0. 3.0’s weights aren’t up yet. When they are, Nano Banana and Seedream are facing a free rival that runs on your own GPU.
You Might Also Like
- Google Nano Banana 2 Lite Gemini 3 1 Flash Lite Image a Picture in 4 Seconds for 0 034 per 1000
- Qwen Image 2 0 Just Dropped and i Honestly Wasnt Expecting This
- Alibaba Qwen 3 5 Just Dropped and it Brought 10 Million Milk Teas With it
- Google Nano Banana 2 Just Dropped and its Absurdly Fast
- Alibaba Qwen Smart Glasses g1 s1 275 ai Glasses With Swappable Batteries and a Qwen api Backend

Leave a comment