Top AI Product

Every day, hundreds of new AI tools launch across Product Hunt, Hacker News, and GitHub. We dig through the noise so you don't have to — surfacing only the ones worth your attention with honest, no-fluff reviews. Explore our latest picks, deep dives, and curated collections to find your next favorite AI tool.


OpenAI: Separating signal from noise in coding evaluations — the lab just called its own benchmarks unreliable

Everyone ships a coding model with a shiny SWE-bench number attached. OpenAI just published a technical post arguing most of those numbers are noise. It hit 216 points on Hacker News, and the reason is obvious: it’s OpenAI throwing cold water on the exact scoreboard the whole industry uses to declare victory.

What it actually says

This isn’t a product. It’s a blog post — “Separating signal from noise in coding evaluations” — that dissects why coding benchmarks lie. The uncomfortable finding: run the same model on the same benchmark multiple times and scores can swing by ten-plus points. That means a lot of “new SOTA” announcements sit inside the measurement error. Nobody moved the needle; they got a lucky roll.

OpenAI traces the noise to a few sources. Contamination, because benchmark tasks come from public GitHub repos the models already trained on. Randomness in sampling. And a design flaw: unit tests inside a pull request were written to validate one specific human fix, not to define an implementation-agnostic “solved.” A correct-but-different solution fails. OpenAI even retracted its own earlier nod to SWE-Bench Pro after finding the same rot.

Why it matters

The whole “signal” here is trust. When bad evals feed capability and safety decisions, you’re steering with a broken gauge. Coming from the company that benefits most from big benchmark headlines, that’s a pointed admission — and a challenge to read the next “beats every open model” press release with a skeptic’s eye.


You Might Also Like


Discover more from Top AI Product

Subscribe to get the latest posts sent to your email.



Leave a comment