AI Research & Analytics
-
Pangram 4 + Pangram Image: one false positive per 24,000 documents, and now it reads pictures too
Most AI detectors are a coin flip that ruins a real student’s day. Pangram Labs just raised $9M led by Menlo Ventures and shipped two models the same day to argue it doesn’t have to be that way. What it actually is Pangram 4 is an AI-detection model (not a chatbot, not a writing tool)… Continue reading
-
HANDBOOK.md Benchmark (Surge AI): the top agent still fails 64% of company-policy tasks
Everyone keeps saying you can govern an AI agent by handing it a long policy doc — an AGENTS.md, a fat system prompt, a 100-page handbook. Surge AI’s new HANDBOOK.md Benchmark (Surge AI) says that’s mostly wishful thinking. What it actually is It’s an eval suite for agents, not a product you install. 65 tasks,… Continue reading
-
Claude Mythos cracked HAWK and 7-round AES — flaws humans missed for years
An AI model just found real math holes in two cryptosystems that experts had reviewed for years. Not a jailbreak, not a demo — actual novel attacks. Anthropic’s Frontier Red Team published it on July 28. What Claude Mythos actually did Claude Mythos is Anthropic’s unreleased frontier model, and here it worked as a semi-autonomous… Continue reading
-
Kimi Linear (Moonshot AI) decodes 6x faster at 1M tokens by cutting KV cache 75%
Moonshot AI dropped Kimi Linear and it hit the HackerNews front page the same day (145 points). It’s not a chatbot or an agent — it’s an open-weight language model, but the news is the plumbing underneath it: a new attention architecture built to make long context cheap. What it actually is Kimi Linear is… Continue reading
-
Ramp × Prime Intellect: a $500 RL fine-tune of a 9B open model beat every frontier config
Ramp and Prime Intellect just published a real production result, not a benchmark stunt: they took a 9B open-weight model, ran GRPO reinforcement fine-tuning on it, and beat every frontier configuration they tested on catalog review — deciding whether merchant transactions match the right product category. Total training cost: about $500. It hit HN’s front… Continue reading
-
Neutrino-1 8B (Fermion Research) squeezes an 8B model into a single 3.88GB file
Fermion Research took Qwen3-8B, ran ternary quantization-aware training, and distilled out an 8.19B decoder that ships as one 3.88GB file. All 252 linear layers stored in a three-value weight format 8x smaller than fp16, decoded inside the matrix kernels. It hit HackerNews the day it launched, 2026-07-27. What it actually is An open-weights language model,… Continue reading
-
last30days-skill hit 54K stars: mvanhorn built a research agent that reads Reddit, X and YouTube so you don’t have to
Ask any AI “what happened with X lately” and it hallucinates or hands you a stale training-data answer. Matt Van Horn’s fix is a skill, not another chatbot. What it actually is last30days-skill is a plug-in skill for coding agents — mount it in Claude Code or Codex and it works. Give it a topic… Continue reading
-
Kronos — open-source foundation model for financial markets hits 32.9k stars, gaining ~400 a day
Someone finally built GPT for candlesticks. Kronos, from developer shiyu-coder, treats K-line charts (OHLCV bars) as a language: a dedicated tokenizer quantizes continuous price and volume data into discrete tokens, then a decoder-only Transformer predicts what comes next. Same recipe as an LLM, pointed at markets instead of text. Pre-trained on 45+ global exchanges, MIT-licensed,… Continue reading
-
WorldMonitor (koala73) hit 70K GitHub stars — +4,000 in one day
Fastest-climbing AI repo on GitHub Trending right now. WorldMonitor is an open-source, self-hosted dashboard that turns the entire planet into one screen: geopolitics, 29 stock exchanges, commodities, crypto, energy, aviation, military, and cyber — 65+ data sources and 500+ feeds pulled into a single situational-awareness view. What it actually is Not a SaaS, not an… Continue reading
-
—
title: “$25 of GPT-5.6 Sol Ultra found a pre-auth WordPress RCE — Searchlight Cyber’s wp2shell” Exploit brokers pay around $500,000 for a WordPress pre-auth RCE. Adam Kues at Searchlight Cyber got one for $25 in API spend — half a day of GPT-5.6 Sol Ultra, no human bug hunting. What it actually is Not a… Continue reading
