ARC Prize published the official numbers on September 3: GPT-6 Astra, OpenAI’s new flagship model, hit 99.9% on ARC-AGI-3 Semi-Private — a suite of interactive games built so models can’t lean on training data and must work out the rules themselves. Six months ago the state of the art was 7.8%.
The harness is worth 37 points
Through the standard harness, Astra scores 62.7% for $26,098. Through the Provider Adapter harness, which keeps reasoning state alive between turns, it scores 99.9% — for $18,817. Better score, lower bill. The scaffolding around the model now matters as much as the model.
Beating humans at their own efficiency game
The wilder number: Astra used fewer actions than the human baseline on 96% of levels, averaging 51.7% fewer per level. ARC researchers had bet action efficiency would stay a human advantage. Along the way, Astra invented its own algebraic notation to track game state and wrote per-game tools like maze_solver.py on the fly.
ARC Prize’s own caveat: the games are deterministic and closed-ended, so saturation is not AGI. Still — a benchmark designed to embarrass frontier models lasted under a year.
You Might Also Like
- Two api Settings Tripled gpt 5 6 Sols arc agi 3 Score Openai Says the Harness not the Model was the Problem
- Prime Intellects Prime Agent Scores 95 5 on arc agi 3 First Harness to Beat Human Experts
- Arc agi 3 Turns ai Testing Into a Video Game and Every Frontier Model is Losing
- Gemini 3 Deep Think Scores 84 6 on arc agi 2 Googles Reasoning bet is Paying off
- Claude Fable 5 1 Mythos 5 1 Anthropic 90 on arc agi 2 and Cache Reads Just got 75 Cheaper

Leave a comment