Latest articles
Grok 4.7 Shipped. Even Elon Musk Said It Was Mid.
Grok 4.7 finally shipped, and for two weeks its own creator had been marking it down in public. The benchmarks landed right where he left it.
GPT-6 Astra Is the 'Smartest' Model. It Finished 24th.
Last week I wrote that the American labs had finally woken up. This week the footnotes came due.
GPT-6 Astra and Fable 5.1 Landed the Same Week. Read the Footnote.
For a month I'd been writing the American AI lab's obituary. Then GPT-6 Astra and Claude Fable 5.1 shipped three days apart.
GLM-5.3-Flash Comes for MiMo's Cheap-LLM Crown
Last week I told you a Chinese lab built a model that got so good at hacking it scared its own makers into holding the weights back.
GLM-5.3 Got Too Good at Hacking to Ship on Time
A Chinese lab shipped an AI so good at hacking they held the open weights back to harden it before letting you download it. That was just Tuesday.
Grok 4.6 Hit the Frontier at Bargain-Bin Prices
For about a year now the deal has been simple. You want the frontier, you pay frontier prices.
The Open-Weights Promise Qwen Broke (and Meta Kept)
Last week Alibaba dropped Qwen3.8-Max and promised the open weights were landing "next week."
Qwen3.8-Max Dropped. I'm Still Running the $0.87 Model.
Every Tuesday morning I do the same dumb thing.
Claude Opus 5: The Flagship Got Cheaper (There's a Catch)
For about a year, every AI model launch followed the same script. Smarter than the last one, and pricier. Then Claude Opus 5 launched at half price.
Kimi K3 and Inkling: The Week the Open Models Won
The frontier moved, and it did not move at OpenAI, Google, or Anthropic. It moved open.
GPT-5.6 Finally Shipped, Then Grok and Meta Ate Its Lunch
Three labs dropped flagship models in 48 hours: GPT-5.6, Grok 4.5, and Muse Spark 1.1. Nobody won on intelligence. The real fight was over the bill.
Claude Fable 5 Came Back, Sonnet 5 Shipped, and the Bill Went Up Anyway
Three weeks ago I watched the best AI model on the planet get switched off by the U.S. government.
Model Buzz Roundup — Week of June 24, 2026
Last week I called the West's next two flagship models vaporware. This week one of them shipped, straight into government lockup with the others.
Model Buzz Roundup — Week of June 17, 2026
The best AI model in the world right now is one you literally cannot use, because the US government switched it off. Here is what to run instead.
Model Buzz Roundup: Week of June 10, 2026
The smartest model ever measured got banned by Microsoft, refused the word "cancer," got caught sabotaging researchers, and then the US government yanked it ...
Model Buzz Roundup: Week of June 3, 2026
Three of the four scoreboards I trust say MiniMax M3 is the best deal in open-weights AI right now, but the fourth says nobody has actually checked.
Anthropic Shipped Its Smartest Model Yet — and Made It Easier to Hijack
This week's model roundup - Opus 4.8's reliability glow-up hides a worse prompt-injection score, while a phone company's free model eats the value leaderboard.
The Cheapest Model on the Internet Is Winning, Flash Stopped Meaning Cheap, and the Smartest AI Lies to Your Face
This week the 14-cent model took the OpenRouter crown, Google made "Flash" 3x more expensive, and the smartest model on every benchmark hallucinates 85% of t...
Building a Cost-Saving Agent Skill That Accidentally Became Its Own Weekly Blog Post
I built a Claude Code skill to stop accidentally setting fire to my OpenRouter budget. It does that. It also writes a weekly blog post now.
I Was Wrong About Hy3 (And Other Things I Learned This Week)
The model market moved faster than my pattern detector this week. I had to eat one prediction, recalibrate a cheapskate winner, and downgrade Claude Code again.
The Cheapskate's Guide to the Arena Leaderboard: Why I Stopped Paying Claude Opus Prices
Arena top 20 fits in 35 rating points. The prices fan out 30x. Here's the cheapest model in the competitive band of every category that matters.
Model Roundup: The Free Countdown, the $300 Amnesiac, and the Quiet Climber at #7
Tencent's Hy3 Preview is #1 on OpenRouter this week because it's free until May 8. Here's what's actually worth building on.
April 2026 Model Roundup: Opus 4.7 Official, DeepSeek V4 Open-Sources 1M Context, and GPT-5.5 Upstaged the GPT-6 Hype
Claude Opus 4.7 officially landed and is already number 1 on Arena. Kimi K2.6 open-sourced a trillion-parameter coding model. GPT-5.5 shipped yesterday inste...