DeepSeek pushed the V4-Flash API into public beta and quietly swapped in a retrained 0731 build. Same architecture, same size — but the numbers jumped in a way that’s hard to ignore. Terminal Bench 2.1 went from 61.8 to 82.7, past GLM-5.2’s 81.0 and within arm’s reach of Claude Opus 4.8’s 85.0. DeepSWE climbed from 7.3 to 54.4. Cybergym went 38.7 to 76.7. That’s not a tune-up, that’s a different model wearing the same clothes.
What it actually is
A coding-focused agent model you hit over an API — not a chatbot, not a wrapper. The point is autonomous coding: run commands in a terminal, fix repos, close loops. Opus 4.8 lives in that same lane, and this thing is now trading punches with it on the benchmark that matters for agents.
The API and the price
Still $0.14 in / $0.28 out per million tokens, with cached input at $0.0028 — basically free on repeat context. New Responses API support and a Codex adapter mean you can drop it into existing agent stacks. OpenRouter and fal already list it. Opus-tier coding at loose-change prices is the whole story.
You Might Also Like
- Hiveterm Bets on the Multi Agent Workspace Claude Codex and Gemini in one Terminal
- Kimi k2 6 Beats gpt 5 4 and Claude Opus 4 6 on swe Bench pro
- Deepseek tui Tops Github Trending a Claude Code Clone Wired to Deepseeks api
- Claude Sonnet 5 Scores 63 2 on swe Bench pro at a Third of Opus 4 8s Price
- Zcode Zhipu z ai glm 5 2 Coding Agent hit hns 1 Spot Chinas Open Answer to Claude Code

Leave a comment