Coding agents got SWE-bench. Hardware just got its version. On September 4 the atopile team released EEBench, a benchmark where AI agents design actual PCBs: 13 analog and digital tasks, written as declarative atopile code instead of clicks in a CAD tool, then graded by SPICE simulation plus real-world constraints — component tolerances, cost, whether you can actually buy the parts. It hit the HackerNews front page with 392 points and 209 comments.
The scoreboard
Claude Opus 5 leads at 61.6%. Grok 4.6 scores 57.1%, Claude Fable 5.1 takes 56.4% — Anthropic holds two of the top three. GPT-5.6 Sol trails badly at 39.4%. xAI already cites its EEBench score in Grok’s official model card, so labs are treating this as a number worth competing on.
Where models actually fail
Not on textbook electronics — that knowledge is shockingly deep. They fail on physics meeting reality. One task required 545µF of effective capacitance; ceramic capacitors lose most of their rating under DC bias, and the submitted design delivered 11.4µF at working voltage. Dead circuit. Knowing the formula isn’t the same as shipping a board, and now there’s a scoreboard that proves it.
You Might Also Like
- Kimi k2 6 Beats gpt 5 4 and Claude Opus 4 6 on swe Bench pro
- Gpt 5 5 Takes Back the Coding Crown From Claude Opus 4 7
- Anthropic Claude Design Wiped 7 off Figmas Stock and it Isnt Even a Figma Clone
- Claude Opus 5 Lands Within 0 5 of Fable 5 on Coding at Half the Price
- Gpt 5 6 sol is the Best Vision Model Openai Ever Shipped Roboflow Benchmark

Leave a comment