-
agent-skills (Addy Osmani) hit 8,600 GitHub stars by forcing coding agents to act like senior engineers
Google Chrome’s Addy Osmani open-sourced agent-skills, and GitHub trending noticed fast — over 1,100 stars in a single day, 8,600+ total. It’s not a model or an app. It’s a pack of production-grade engineering skills you drop into a coding agent so its output stops looking like a demo. What it actually is The problem… Continue reading
-
Xiaomi MiMo-V2-Pro is now the 1 model by weekly tokens on OpenRouter — ahead of every US frontier lab
A phone maker just outranked OpenAI and Anthropic on the one metric that reflects real usage: tokens actually served. Xiaomi’s MiMo-V2-Pro, a 1T-parameter mixture-of-experts LLM, sits at the top of OpenRouter’s weekly token chart. What it is and why it’s winning MiMo-V2-Pro is a text model built for agents and code. It ships a 1M-token… Continue reading
-
Anthropic Drug Discovery Program (neglected diseases): the AI lab is now making its own drugs
Anthropic stopped just selling models. On June 30 it launched an internal drug discovery program aimed at “neglected diseases” — conditions where the biology is well understood but no big pharma bothers, because there’s no money in it. What it actually is Two things. First, real preclinical drug programs run inside Anthropic. Second, Claude Science,… Continue reading
-
GPT-5.6 Sol Ultra hits 91.9% on Terminal-Bench 2.1 — and it’s landing in Codex
OpenAI’s new flagship, GPT-5.6 Sol, is a reasoning model built for agentic coding, and it just set a new SOTA on Terminal-Bench 2.1 — the benchmark that measures planning, iteration, and tool-juggling inside a real command line. Base Sol scores 88.8%. The “Sol Ultra” config pushes that to 91.9%, ahead of Claude Mythos 5 (84.3%)… Continue reading
-
CircleChat runs your AI agents like a company — with an LLM judge signing off every task
Most agent frameworks let your bots talk. CircleChat makes them ship. It’s team chat where AI agents are first-class members: give the workspace a goal, and the agents break it into tasks on a kanban board, claim the work, and report progress in channels you can actually read — instead of one endless chat transcript.… Continue reading
-
Meetily (Zackriya-Solutions/meetily): 17K stars, +1,400 in a day, and not a single byte hits the cloud
Otter.ai and Granola record your meetings by shipping the audio to their servers. Meetily’s whole pitch is the opposite: nothing leaves your laptop. That nerve is clearly raw — the repo pulled +1,400 GitHub stars in a single day, past 17K total, on 180K+ downloads. What it actually is A desktop AI meeting note-taker for… Continue reading
-
Meta Watermelon claims a GPT-5.5 tie — but Wang won’t say on which benchmarks
Meta’s superintelligence chief Alexandr Wang told staff in early July that Watermelon, Meta’s next flagship LLM, has already caught up to OpenAI’s GPT-5.5 on “key benchmarks.” Watermelon is still in training. There’s no release date, no API, and — conveniently — no word on which benchmarks. What Watermelon actually is A frontier text model, successor… Continue reading
-
OpenWiki (LangChain): the docs aren’t for you — they’re for your coding agent
Every coding agent hits the same wall: it doesn’t know your codebase, and stuffing everything into one bloated CLAUDE.md doesn’t scale. LangChain just open-sourced OpenWiki, a CLI that fixes this by writing docs for the machine instead of the human. What it actually does Point OpenWiki at a repo and it scans the code, then… Continue reading
-
OpenAI acquires Ona (formerly Gitpod) to keep Codex agents coding while your laptop sleeps
OpenAI announced on June 11, 2026 it’s buying Ona — the German company most developers still know as Gitpod, the cloud dev environment with 750,000+ users. This isn’t a talent grab. It’s infrastructure. What OpenAI actually bought Today’s coding agents die when the session ends. Close the tab, lose the context, start from zero. Ona’s… Continue reading
-
DeepSeek-V4-Pro & V4-Flash ship with 1M context and $0.28/M output
DeepSeek is back, and it’s aiming straight at the agent crowd. The new V4 series is two open-weight MoE models — V4-Pro (1.6T total, 49B active) and V4-Flash (284B total, 13B active) — both shipping with a 1-million-token context window by default. Not a premium tier. The floor. What it actually is These are API-callable… Continue reading
