Harvey, the $5B legal AI unicorn, open-sourced LAB (Legal Agent Benchmark): 1,671 real legal tasks across 24 practice areas plus transactional work. Not multiple choice — actual law-firm assignments like M&A data room review, each with agent instructions, client documents, and expert-written rubrics (75,000+ grading criteria total). The repo just hit GitHub Trending: 744 stars, +87 in a single day.
Why the scores are brutal
All-pass scoring: miss one rubric criterion, fail the whole task. At launch, Claude Opus 4.7 led with 7.1%. GPT-5.5 managed 2.1%. Three months later, the top model on the leaderboard passes barely 20%. Lawyers can relax — for now. And that’s the real story: the company selling legal AI just published the public standard for “can AI do a lawyer’s job,” and today’s answer is mostly no.
Run it yourself
LAB ships with an execution harness plus multi-model adapters — plug in any model, run the tasks, get graded scores. Legal teams use it to fact-check vendor claims; labs use it to test long-horizon agents on messy real documents. It’s the first vertical agent benchmark released by the category leader itself. Whoever writes the exam grades the market.
You Might Also Like
- Microsoft Word Legal Agent Ships at 30 Seat Harvey Charges 1200
- Cloudflare Computer Cloudflare Computer Hits 1 on Github Trending Every ai Agent Gets its own Machine
- Heretic Just hit Github Trending and the ai World has Opinions
- Pentagi Just hit 1 on Github Trending and Yeah its Worth the Hype
- Pageindex Just hit Github Trending and it Might Make you Rethink rag Entirely

Leave a comment