Everyone keeps saying you can govern an AI agent by handing it a long policy doc — an AGENTS.md, a fat system prompt, a 100-page handbook. Surge AI’s new HANDBOOK.md Benchmark (Surge AI) says that’s mostly wishful thinking.
What it actually is
It’s an eval suite for agents, not a product you install. 65 tasks, each one a self-contained fake company: real files (PDFs, Excel, Word), internal tools, and live external MCP services — Gmail, Google Calendar, Slack, Jira, Shopify. At the center sits a handbook averaging 43 pages, topping out at 124. The agent gets one instruction: follow the rules. Then it has to let that handbook constrain every email, Slack reply, and Jira move across finance, medical billing, insurance, logistics, and HR.
Why it’s trending
The numbers are brutal. Best config — Claude Fable 5 on max reasoning — passes just 36.2%. GPT-5.6 Sol lands at 23.5%. Most frontier models sit under 25%. Cranking up effort barely helps and sometimes hurts.
The takeaway hit HackerNews front page at 300+ points: a long document does not reliably govern an agent. If your agent-governance plan is “write more markdown,” this is the cold water.
You Might Also Like
- Google A2ui Agent to User Interface Finally a Standard way for ai Agents to Show you Things
- Agent Action Protocol aap the Missing Layer Above mcp That Actually Makes Agents Production Ready
- Addy Osmani Open Sources Agent Skills 19 Workflows That Make ai Agents Code Like Google Engineers
- Google Deep Research Agents With mcp Score 93 3 on Deepsearchqa Plug Into Private Data
- Google Search Information Agents Turn 1 Billion ai Mode Users Into Agent Operators

Leave a comment