GPT-6 Astra shipped three days ago. Robocurve, a YC-backed independent robotics eval lab, already has it running a pair of bimanual YAM arms and moving real objects. It’s the first independent embodied test of the new model, and the gap is ugly: 181 points and 136 comments on HackerNews say people noticed.
The block-in-bowl blowout
Put a block in a bowl: Astra finished 19 of 20 runs. Claude Fable 5.1 managed 8 of 20. Fable 5, 1 of 20. Astra also did each run in 2.5 minutes at $0.94 — Fable 5.1 took 6.8 minutes and $2.12. More than twice as reliable, at half the time and cost.
Where both models faceplant
The puzzle task — fitting a piece into its slot — flattened everyone: 2 of 20 for Astra and Fable 5.1 alike. Fine-grained manipulation is a shared ceiling, whoever’s logo is on the model.
This isn’t a robot product. It’s a general frontier model steering off-the-shelf arms through Robocurve’s open-source Inspect Robots harness, same agent policy for every model. Independent embodied benchmarks landing 72 hours after a launch — that’s SWE-bench treatment coming for robotics.
You Might Also Like
- Kimi k2 6 Beats gpt 5 4 and Claude Opus 4 6 on swe Bench pro
- Gpt 5 5 Takes Back the Coding Crown From Claude Opus 4 7
- Bytedance ui Tars Desktop Scores 61 6 on Screenspot pro Leaving gpt 4o and Claude Behind
- Claude Fable 5 Global Redeployment Export Controls Lifted From us ban to Worldwide Access in 18 Days
- Satya Nadella Calls Claude Fable 5 Editorially Controlled in Front of his own Copilot Engineers

Leave a comment