Top AI Product

Every day, hundreds of new AI tools launch across Product Hunt, Hacker News, and GitHub. We dig through the noise so you don't have to — surfacing only the ones worth your attention with honest, no-fluff reviews. Explore our latest picks, deep dives, and curated collections to find your next favorite AI tool.


HANDBOOK.md Benchmark (Surge AI): the top agent still fails 64% of company-policy tasks

Everyone keeps saying you can govern an AI agent by handing it a long policy doc — an AGENTS.md, a fat system prompt, a 100-page handbook. Surge AI’s new HANDBOOK.md Benchmark (Surge AI) says that’s mostly wishful thinking.

What it actually is

It’s an eval suite for agents, not a product you install. 65 tasks, each one a self-contained fake company: real files (PDFs, Excel, Word), internal tools, and live external MCP services — Gmail, Google Calendar, Slack, Jira, Shopify. At the center sits a handbook averaging 43 pages, topping out at 124. The agent gets one instruction: follow the rules. Then it has to let that handbook constrain every email, Slack reply, and Jira move across finance, medical billing, insurance, logistics, and HR.

Why it’s trending

The numbers are brutal. Best config — Claude Fable 5 on max reasoning — passes just 36.2%. GPT-5.6 Sol lands at 23.5%. Most frontier models sit under 25%. Cranking up effort barely helps and sometimes hurts.

The takeaway hit HackerNews front page at 300+ points: a long document does not reliably govern an agent. If your agent-governance plan is “write more markdown,” this is the cold water.


You Might Also Like


Discover more from Top AI Product

Subscribe to get the latest posts sent to your email.



Leave a comment