AI Models & APIs
-
OpenAI GPT-5.6: same smarts, half the price — Sol hits 750 tokens/s on Cerebras
The pitch for OpenAI GPT-5.6 isn’t “smarter.” It’s “same intelligence, cheaper and faster.” That’s a tell about where the frontier fight moved. GPT-5.6 is a family of three text models, all served through the OpenAI API. Sol is the flagship ($5 in / $30 out per 1M tokens), Terra is the value tier ($2.50 /… Continue reading
-
Gemini Robotics 2 gives humanoids whole-body control, from feet to fingertips
Google DeepMind shipped Gemini Robotics 2 on July 30, and the headline is simple: last year’s model only drove a robot’s upper body. This one runs the whole machine — walking, crouching, reaching, gripping — while it reasons through a multi-step task in real time. What it actually does It’s a vision-language-action model, not a… Continue reading
-
OpenAI bets $250M on ChatGPT for Academic Researchers — free GPT-5.6 for 100,000 scientists
OpenAI is handing out its best models for free. On July 29 it launched ChatGPT for Academic Researchers, part of a $250M+ commitment through 2027. This isn’t a new app or API — it’s a gated access program that drops frontier models onto researchers’ desks at no cost. What researchers actually get The full GPT-5.6… Continue reading
-
TurboFieldfare — Gemma 4 26B running in ~2GB RAM on any M-series Mac
A 26B model in 2GB of RAM sounds impossible. That’s exactly why it hit 556 points on Hacker News and 865 GitHub stars in a day. What it actually is TurboFieldfare is a local inference engine written in Swift and Metal that runs Google’s Gemma 4 26B-A4B — 14.3GB of weights — on an 8GB… Continue reading
-
Claude Mythos cracked HAWK and 7-round AES — flaws humans missed for years
An AI model just found real math holes in two cryptosystems that experts had reviewed for years. Not a jailbreak, not a demo — actual novel attacks. Anthropic’s Frontier Red Team published it on July 28. What Claude Mythos actually did Claude Mythos is Anthropic’s unreleased frontier model, and here it worked as a semi-autonomous… Continue reading
-
Kimi Linear (Moonshot AI) decodes 6x faster at 1M tokens by cutting KV cache 75%
Moonshot AI dropped Kimi Linear and it hit the HackerNews front page the same day (145 points). It’s not a chatbot or an agent — it’s an open-weight language model, but the news is the plumbing underneath it: a new attention architecture built to make long context cheap. What it actually is Kimi Linear is… Continue reading
-
Ramp × Prime Intellect: a $500 RL fine-tune of a 9B open model beat every frontier config
Ramp and Prime Intellect just published a real production result, not a benchmark stunt: they took a 9B open-weight model, ran GRPO reinforcement fine-tuning on it, and beat every frontier configuration they tested on catalog review — deciding whether merchant transactions match the right product category. Total training cost: about $500. It hit HN’s front… Continue reading
-
Neutrino-1 8B (Fermion Research) squeezes an 8B model into a single 3.88GB file
Fermion Research took Qwen3-8B, ran ternary quantization-aware training, and distilled out an 8.19B decoder that ships as one 3.88GB file. All 252 linear layers stored in a three-value weight format 8x smaller than fp16, decoded inside the matrix kernels. It hit HackerNews the day it launched, 2026-07-27. What it actually is An open-weights language model,… Continue reading
-
Anthropic: Our position on open-weights models — Dario Amodei’s case for going closed, right as China floods in
This isn’t a product. It’s a policy shot. Dario Amodei signed “Anthropic: Our position on open-weights models,” a blog post that hit Hacker News on July 28 and cleared 580 points in a few hours. It’s Anthropic explaining why the US should think hard about open-weight models — which, conveniently, is also a defense of… Continue reading
-
Kimi K3’s 2.8T open weights are live, but running them takes 64 GPUs
Moonshot just put Kimi K3’s weights online, free to download. At 2.8 trillion parameters it’s the largest open-weight model ever shipped. The catch: even after MXFP4 four-bit quantization, the download is 1.4TB. What you actually get K3 is a mixture-of-experts model — each token fires only 16 of its 896 experts, about 50 billion active… Continue reading
