Ramp and Prime Intellect just published a real production result, not a benchmark stunt: they took a 9B open-weight model, ran GRPO reinforcement fine-tuning on it, and beat every frontier configuration they tested on catalog review — deciding whether merchant transactions match the right product category. Total training cost: about $500. It hit HN’s front page at 182 points.
What they actually built
This isn’t a chatbot or an app. It’s a narrow, in-house classifier trained with RL on Ramp’s own private labeled data. The tuned 9B reached 87.7% of the maximum achievable score versus 76.9% for the best frontier setup — a small open model outscoring the big general-purpose ones on the one task that mattered.
Why the numbers sting
The cost gap is the headline. Inference runs $0.50 per 1,000 items — 40× cheaper than the cheapest frontier option and roughly 340× cheaper than the most expensive. So on any narrow task where you own the data, a $500 fine-tune of a small open model can out-accuracy and out-price GPT-class frontier models by two orders of magnitude. The moat was never the model. It was the labeled data.
You Might Also Like
- 0 004 per Task how Atlas Squeezes Frontier Level Coding From a Single 500 gpu
- Prime Intellect lab Hits ga per Token rl Training Across 14 Models
- Entire the 60m bet on Fixing ais Code Review Headache
- Claude Code Security Just Dropped and it Already Found 500 Zero Days Nobody Knew About
- Mercury 2 Just hit 1000 Tokens per Second and its not Even Using Transformers

Leave a comment