Cloudflare just showed its homework: a technical deep-dive on how it serves Moonshot’s Kimi K2.6 and Z.ai’s GLM 5.2 — massive long-context MoE models — on its global edge network. 96 points on HackerNews the day it dropped.
Smaller, faster, safer — with numbers
Three moves. FP8 KV-cache quantization doubles Kimi K2.6’s supported context from 686K to 1.37M tokens and lifts peak throughput 41%. INT4 compression shrinks GLM 5.2’s checkpoint from 705GB to 421GB, with decode up to 55% faster. A cache integrity check guards hundreds of concurrent requests sharing memory pages, at under 1% overhead. Accuracy drift: 0.8 points max on GSM8K and MMLU. Net result: roughly 30% cheaper per token.
Why this matters
Chinese open-weight models already take 30-46% of US enterprise token usage. Now America’s biggest edge company is building custom serving pipelines for them — first-class citizens, not exotic imports. That’s an infrastructure-level endorsement no benchmark can buy.
The API angle
Both models sit behind Workers AI’s inference API: call Kimi or GLM from any Worker, no GPU cluster to rent. Obvious fit — long-context agents and document-heavy pipelines running at edge latency.
You Might Also Like
- Lm Studio Bionic Runs Claude Code Style Agents on glm 5 2 and Kimi k2 7
- Echo Tracerml Routes Across glm 5 2 and Kimi k2 7 to hit Fable Level Quality at a Third the Cost
- Cloudflare Vinext one Engineer one Week and a Whole lot of ai
- Glm 5 Just Dropped and its the Open Source Model Nobody saw Coming
- Mcp2cli the Tool That Cuts mcp Token Costs by 99 Just hit Hacker News

Leave a comment