Fireworks AI shipped Nexus on July 28, a drop-in routing and cost-control layer that sits under your existing coding agents and decides, per request, whether a task actually needs a frontier model. Most of them don’t. Routine edits go to an open-weight model; the hard stuff still passes through to Claude or GPT on your own key. Fireworks says that split delivers a 3–5× cost cut, 33% less per merged PR, and 50–75% off total coding AI spend.
How the routing works
A custom-trained model scores each request’s difficulty. Easy → cheap open-weight model served by Fireworks. Hard → your existing provider, on your key, which Fireworks says it never stores. On top of that is an enterprise control plane: per-seat spend caps, usage tracking, and evals showing cost, latency, and quality per model.
The API angle
The open-source piece is FireConnect (Apache 2.0), a one-line install that hooks into Claude Code, Codex, and OpenCode without touching your workflow. It runs on Fireworks’ Anthropic- and OpenAI-compatible APIs, so most tools connect with just a base URL and a model ID.
Why it matters: open-weight routing turns “which model” into a spend control point — the argument everyone in AI infra is having this week.
You Might Also Like
- Openai Codex Claude md Auto Import Makes Switching From Claude Code a two Click Move
- Kimi Webbridge Plugs Claude Code Cursor and Codex Into Your Browser no Cloud Relay
- Openai Codex Plugin cc Hits 22700 Stars Openai Built an Official Plugin for Rival Claude Code
- Google Labs Releases Stitch Skills Official Stitch Plugins for Claude Code Cursor and Codex
- Claude Code Burns 33k Tokens Before Reading Your Prompt Opencode Needs 7k Systima ai Tested

Leave a comment