Top AI Product

Every day, hundreds of new AI tools launch across Product Hunt, Hacker News, and GitHub. We dig through the noise so you don't have to — surfacing only the ones worth your attention with honest, no-fluff reviews. Explore our latest picks, deep dives, and curated collections to find your next favorite AI tool.


GPT-6 Astra scores 99.9% on ARC-AGI-3 — the anti-memorization benchmark just got saturated

ARC Prize published the official numbers on September 3: GPT-6 Astra, OpenAI’s new flagship model, hit 99.9% on ARC-AGI-3 Semi-Private — a suite of interactive games built so models can’t lean on training data and must work out the rules themselves. Six months ago the state of the art was 7.8%.

The harness is worth 37 points

Through the standard harness, Astra scores 62.7% for $26,098. Through the Provider Adapter harness, which keeps reasoning state alive between turns, it scores 99.9% — for $18,817. Better score, lower bill. The scaffolding around the model now matters as much as the model.

Beating humans at their own efficiency game

The wilder number: Astra used fewer actions than the human baseline on 96% of levels, averaging 51.7% fewer per level. ARC researchers had bet action efficiency would stay a human advantage. Along the way, Astra invented its own algebraic notation to track game state and wrote per-game tools like maze_solver.py on the fly.

ARC Prize’s own caveat: the games are deterministic and closed-ended, so saturation is not AGI. Still — a benchmark designed to embarrass frontier models lasted under a year.


You Might Also Like


Discover more from Top AI Product

Subscribe to get the latest posts sent to your email.



Leave a comment