Google shipped three Flash models on July 21, all aimed at the same problem — agents that run thousands of steps and burn tokens doing it.
3.6 Flash uses 17% fewer output tokens than 3.5 Flash and still scores better where agents live: DeepSWE code editing 49% vs 37%, MLE Bench 63.9% vs 49.7%, OSWorld-Verified computer use 83.0% vs 78.4%. Pricing fell to $1.50/$7.50 per million tokens, down from $9 on output.
3.5 Flash-Lite is the loud one. 350 output tokens/sec, Terminal-Bench 2.1 from 31% to 54%, SWE-Bench Pro 54.2% — at $0.3/$2.5, roughly a fifth of Flash’s price. Nearly doubling an agentic benchmark while cutting cost 80% is the entire pitch.
Where to plug it in
Both are live in the Gemini API through Google AI Studio, Gemini Enterprise, Android Studio, the Gemini app, and GitHub Copilot. The obvious pattern: 3.6 Flash as the orchestration and reasoning layer, Flash-Lite as the cheap worker grinding through repetitive tool calls underneath.
The model Google won’t sell you
3.5 Flash Cyber finds vulnerabilities, paired with the CodeMender agent, at frontier level on CyberGym. It ships only to governments and trusted partners in a limited pilot. First time a major lab has carved offensive security into its own SKU and then declined to list it — a capability admission dressed up as a product.
You Might Also Like
- Mistral Medium 3 5 Scores 77 6 on swe Bench one Point shy of Gemini 3 1 pro
- Gemini 3 1 Flash Lite Hits ga 0 25 m Input Tokens 2 5x Faster Ttft
- Google Nano Banana 2 Lite Gemini 3 1 Flash Lite Image a Picture in 4 Seconds for 0 034 per 1000
- Google Gemini 3 5 pro ga Ships With a 2m Token Context Window the Biggest in any Production Model
- Grok 4 5 Spacexai Solves swe Bench pro Tasks With 4 2x Fewer Tokens Than Opus 4 8

Leave a comment