How far is AI from doing AI research on its own? Prime Intellect just put a number on it: 82%.
NanoGPT Speedrun Frontier isn’t a product — it’s the largest open experiment on autonomous AI research yet. 153 fully autonomous runs, 18 frontier models (Fable 5, Opus 5, GPT-5.6 Sol, Kimi K3, Grok 4.6, DeepSeek V4 Pro, Qwen 3.8…), each agent locked in a sandbox with 8xH200 GPUs for up to 8 days. No internet. One task: iterate the optimizer recipe for training a 124M-parameter nanoGPT, and beat the clock.
The 82% number
The human record took dozens of researchers months of collective effort. Fable 5, running alone, closed 81.7% of the gap from baseline to that record — 2,726 steps against the human 2,600. Agents ran up to 963 experiments per run, burning as much as 2.9B tokens. Hypothesis, experiment, error analysis: the loop works. What’s still missing is the genuinely new algorithmic idea.
Why it matters
Everyone claims their model “does research.” This is the first large-scale attempt to measure it — and everything is open: code, results, plus 41 full agent trajectories with tool calls and scratchpads. It hit 62 points on the HackerNews front page for a reason.
You Might Also Like
- Kimi k2 6 Beats gpt 5 4 and Claude Opus 4 6 on swe Bench pro
- Deepseek v4 pro Hits gpt 5 Parity on 5 of 7 Benchmarks at a Fraction of the Cost
- Gpt 5 6 sol on Cerebras 750 Tokens s Openais Frontier Model Gets 15x Faster in July
- Grok 4 5 Spacexai Solves swe Bench pro Tasks With 4 2x Fewer Tokens Than Opus 4 8
- Ramp x Prime Intellect a 500 rl Fine Tune of a 9b Open Model Beat Every Frontier Config

Leave a comment