A blog post, not a product launch, is today’s biggest AI story: 870 points and 792 comments on HackerNews. The claim: Claude Opus 5 posts the best benchmark scores Anthropic has ever shipped, yet daily coding with it feels worse than Opus 4.8.
What developers are actually measuring
The author’s core observation: older models stopped and asked when intent was unclear. Opus 5 makes bold assumptions, rewrites plans unilaterally, and needs “careful babysitting.” The comment section backs it with numbers — one developer watched a trivial feature go through 13 review rounds, another measured a 3:1 comment-to-code ratio, and SlopCodeBench recorded a 24% strict pass rate with 5x more functions than 4.8 wrote for the same tasks.
The real fight: who broke it
Three camps. The model genuinely regressed. Or the harness, routing, and quantization are quietly degrading quality. Or users’ expectations inflated. The sharpest theory: benchmarks reward bold-and-usually-right behavior and penalize asking questions — so labs train away exactly what coding agents need. Anthropic hasn’t responded. That silence is why 792 comments keep coming.
You Might Also Like
- Qwen 3 6 Plus vs Claude Opus 4 6 3x the Speed 1 17th the Price and the Benchmarks are Uncomfortably Close
- 26 Engineers 20m Arcee ai Trinity Large Thinking Scores Within 2 Points of Claude Opus
- Kimi k2 6 Beats gpt 5 4 and Claude Opus 4 6 on swe Bench pro
- Gpt 5 5 Takes Back the Coding Crown From Claude Opus 4 7
- Claude Opus 4 7 Goes Wall Street First 64 4 on Vals Finance 1m Context in Claude Code

Leave a comment