Ornith-1.5 is an open-weight model family (9B dense, 35B-A3B MoE, 397B MoE) from DeepReinforce, built for coding and agentic work. The pitch: no human-curated training tasks. The model proposes its own.
The loop: propose, scaffold, solve
Ornith-1.0 learned its own agent scaffolds but still trained on tasks humans picked. 1.5 closes the loop — the model generates progressively harder tasks, builds a task-specific scaffold for each, produces rollouts, and RL reward flows through all three stages. Task rewards multiply validity × difficulty (targeting ~20% success rate) × novelty, so broken or trivial tasks earn nothing.
The 397B version scores 86.1 on Terminal-Bench 2.1 and 56.0 on DeepSWE. Claude Opus 4.8: 85.0 and 59.0. An open model matching a frontier closed one on a self-generated curriculum — that’s the headline.
Open weights, 9B runs locally
All three sizes ship as open weights. The 9B pulls straight from Ollama and still hits 47.0 on Terminal-Bench — a real local coding agent, not a toy. The 35B activates just 3B parameters per token, cheap to serve.
HackerNews (152 points) is split on one question: does a model grading its own homework eventually degrade? Nobody knows yet. That’s exactly why this release matters.
You Might Also Like
- Ornith 1 0 Deepreinforce Self Scaffolding Coding Models Open Weights That Write Their own rl Scaffold 397b Hits 82 4 on swe Bench
- Kimi k2 6 Beats gpt 5 4 and Claude Opus 4 6 on swe Bench pro
- Claude Sonnet 5 Scores 63 2 on swe Bench pro at a Third of Opus 4 8s Price
- Deepseek v4 Flash 0731 Hits 82 7 on Terminal Bench Chasing Opus 4 8 at 0 14 m Tokens
- Qwen3 8 max Goes ga Alibabas 2 4t Model Beats Claude Fable 5 on Terminal Bench Weights Open Next Week

Leave a comment