GPT-6 Astra and Fable 5.1 Landed the Same Week. Read the Footnote.
For about a month now I’ve been writing the same obituary. American AI lab ships nothing. Google ships a fourth Flash instead of the model it actually promised. Cheap Chinese open-weight models quietly eat the frontier from the bottom while everyone waits for a flagship that keeps slipping. I had the funeral playlist queued up.
Then over three days the corpse sat up and shipped two flagships. Anthropic dropped Claude Fable 5.1 on September 1. OpenAI dropped GPT-6 Astra on September 3. They landed tied for the top of the intelligence charts, both priced at a flat ten dollars in and fifty out, and just like that the premium ceiling the cheap models spent all summer erasing was back.
And here’s the part I didn’t see coming. In the exact same week, down at the other end of the market where nobody films the keynote, the cheapest-good-model crown finally changed hands. Both ends of the board moved at once. Let me walk you through it, and then give you the table you actually came here for.
- The West remembered it makes models
- Meanwhile, down in the bargain bin, a crown changed hands
- The cheapskate picks
- Google shipped another Flash while the real thing pretrains
- What is coming
- The honest version
The West remembered it makes models
Start with Fable 5.1, because Anthropic set the tone.
This is the follow-up to Fable 5, and Anthropic did the thing where they beat their own more expensive model with it. Fable 5.1 finishes ahead of Opus 5 on every category they published. Terminal-Bench-Science jumped to 52.6 percent against Opus 5’s 29.0. Their knowledge-work benchmark, GDPval-AA v2, went to 1853 against Opus 5’s 1824. Same sticker as before, ten and fifty per million, but cache reads got cut 75 percent to a quarter per million tokens, which matters more than it sounds if you run anything with a big fixed context.
On the Artificial Analysis Intelligence Index it lands at number one. And here’s where you have to read the label. Every Fable 5.1 score on that leaderboard is tagged “with fallback.” That means the headline number quietly blends in a weaker model’s answers whenever a safety classifier refuses the real one, at a frequency nobody discloses. Anthropic pulled the same move on the Opus 5 chart back in July. The number is real. The footnote is load-bearing. The number-one model on the board isn’t purely the number-one model.
Two days later OpenAI answered with GPT-6. Yes, six. GPT-6 Astra, ten and fifty per million, a 1.05 million token context, staged rollout to a handful of orgs first and then the ChatGPT tiers and the API “over the coming days.” It ties Fable 5.1 at the top of the intelligence index. On computer use it does 72.6 percent on OSWorld 2.0 at roughly 47 percent less time per task than GPT-5.6 Sol, which is the stat that actually matters for agent workloads.
Then it gets weird. Astra “saturates” FrontierMath Tier 4 at 97.6 percent, ARC-AGI-3 at 99.9 percent, and ExploitBench at a clean 100 percent. That last one is an offensive-security benchmark. A frontier model that fully solves the “can you write a working exploit” test. And the 99.9 on ARC is measured “under OpenAI’s provider adapter harness,” which is the kind of phrase you learn to slow down and read twice, because a benchmark run inside the vendor’s own harness is graded homework. To round it out, OpenAI is shipping a program called Daybreak that loosens safeguards for vetted organizations. So the model that aced the exploit test also gets an official channel to relax its guardrails. If that gives you the same feeling the GLM-5.3 “too good at hacking to ship on time” story gave you last month, you’re paying attention.
The thing to take away isn’t “which one wins.” They’re basically tied and it’ll take Arena weeks to sort them out, because Arena always lags a launch by a couple weeks while votes pile up. The thing to take away is that both American labs shipped their best model in one window at the same premium price, and the summer story about the West being asleep is, for now, dead. With an asterisk on each.
Meanwhile, down in the bargain bin, a crown changed hands
Here’s the part I’ve been tracking for six weeks and finally get to close out.
For six straight roundups the answer to “what is the cheapest model that is actually good” never moved. It was MiMo v2.5 Pro from Xiaomi. Eighty-seven cents per million output, open weights, and it kept turning up as the cheapest model inside the competitive band of four different Arena categories, week after week. It was the anchor. I could set my watch by it.
Then last week GLM-5.3-Flash from Z.ai showed up cheaper and, on the hard benchmarks, smarter. The only reason I didn’t crown it on the spot was votes. Arena marks a rating “preliminary” until enough people have voted on it, and GLM-5.3-Flash was sitting on a couple thousand shaky votes while MiMo had fifty thousand solid ones. So I called it the emerging pick, kept MiMo as the printed anchor, and wrote down the tripwire: if GLM-5.3-Flash holds cheapest-in-band once the votes firm up, the spine has genuinely shifted.
This week the votes firmed up. GLM-5.3-Flash is now the cheapest model in the competitive band for Overall, Coding, Instruction Following, and Hard Prompts, on real vote counts, and where it overlaps MiMo it out-rates it. In Overall it sits at 1474 with 4,672 votes for fifty cents per million output. MiMo is at 1468 for eighty-seven cents. Cheaper and higher. The six-week reign is over. MiMo’s the runner-up now, and it’s still a perfectly good runner-up with ten to twenty times the vote count, which is exactly why I keep it in the table.
The catch, because there’s always a catch. GLM-5.3-Flash scores mid on the hard-reasoning index, a 42 where the frontier models are in the fifties. It’s preference-strong and cheap, not a deep-reasoning machine. It’s also slow, about 60 output tokens a second against a median north of 70, which compounds in agent loops where every step waits on the last. And the cheapest price you’ll see quoted for it, seven and a half cents in and a quarter out, is a launch promo that expires September 9. List is fifteen and fifty cents. Build your budget on the list price, not the promo, unless you enjoy surprises on the tenth.
The cheapskate picks
Same method as always. For each Arena category I take the leader’s rating, draw a band 50 points below it, and find the cheapest model that still sits inside that band. The whole premise is that Arena ratings cluster tight at the top, so the category leader is usually only a rounding error better than something 20 to 100 times cheaper. I compute the band from the full table in code, not by eyeballing the first screen, because eyeballing it is how you delete the entire cheap tail and crown the cheapest expensive model by accident. Ask me how I know.
Bands this week ran 56 deep in Overall, about 63 in Coding, and thin in Math where the whole board is still preliminary. Arena data dated around September 2.
| Category | Leader | $ leader out | Cheapskate pick | $ pick out | Δ rating | Cheaper by |
|---|---|---|---|---|---|---|
| Overall | claude-fable-5 (1507) | $50 | GLM-5.3-Flash (1474, #29) | $0.50 | −33 | ~100× |
| Coding | claude-opus-4-7-high (1552) | $25 | GLM-5.3-Flash (1534, #9) | $0.50 | −18 | 50× |
| Creative Writing | claude-fable-5 (1504) | $50 | Gemini 3-Flash (1459, #27) | $3 | −45 | ~16.7× |
| Instruction Following | claude-opus-4-6-high (1514) | $25 | GLM-5.3-Flash (1465, #34) | $0.50 | −49 | 50× |
| Hard Prompts | claude-opus-4-6-high (1533) | $25 | GLM-5.3-Flash (1496, #31) | $0.50 | −37 | 50× |
| Math | claude-fable-5 (1529, prelim) | $50 | Gemini 3.7-Flash-high (1524, #4) | $3.75 | −5 | ~13× |
A few notes on reading that. The Coding line is the loudest one on the board: GLM-5.3-Flash is ranked ninth in coding, top-ten, for fifty cents per million. The catch there is 1,274 votes, still on the thin side, so if you want certainty over savings, MiMo at rank 26 with 15,813 votes for eighty-seven cents is the steadier bet. In Hard Prompts, GLM-5.3-Flash and MiMo are literally tied at 1496; GLM is cheaper, MiMo has twelve times the votes. Take your pick based on whether you trust the number or the sample size.
Creative Writing stays a Gemini story because GLM-5.3-Flash never cracked that band, and Math is a low-confidence mess this week, thin preliminary votes top to bottom and no pick under $3.75, so treat that row as a suggestion and not a promise.
The one thing I won’t do is pretend both value picks are fast. They’re not. GLM-5.3-Flash and MiMo are both slow. If you’re wiring one into an autonomous agent that chains dozens of calls, the fifty-cent price tag can balloon into a fifty-cent-per-call wall-clock tax. Cheap and slow is a real trade, not a free lunch. Know which one your workload cares about.
Google shipped another Flash while the real thing pretrains
Quick check-in on Google, who continue to run the strangest release cadence in the business.
On September 2 they shipped Gemini 3.8 Flash. That’s the third Flash release in six weeks. It costs exactly what 3.7 Flash cost, 75 cents in and $3.75 out, and beats it on every benchmark they published, plus there’s a locked-down 3.8 Flash Cyber sibling for security work. It’s a genuinely good, cheap coding-and-agent workhorse, and it already shows up at number eight overall on Arena.
But it’s built on the 3.7 base, not a new one, and it’s shipping instead of the Gemini 3.5 Pro they promised back in the spring, which has been quietly shelved after missing so many dates I lost count. The real model, Gemini 4, just “cleared pretraining” with strong preliminary results and the largest training run in Google’s history. No benchmarks, no price, no date, late 2026 if you believe the tea leaves. So Google’s strategy remains: ship a steady drip of excellent Flash models to stay in the headlines while the actual next-generation model bakes in the background. It’s working, in the sense that they’re still in the conversation. It’s also the fourth time this summer I’ve written that exact paragraph.
What is coming
Two things worth watching.
Grok 4.7 is the loud one. Musk said on September 2 it would be out “in ten days,” which points at roughly September 12. It’s a 2.1 trillion parameter model, up 40 percent from Grok 4.6, and part of its training data comes from SpaceX internal engineering records, which is either the most interesting or the most concerning detail depending on your mood. There’s no model card, no price, no benchmark table, and no API id yet, so treat the date as a tweet and not a commitment.
Gemini 4, covered above, is the other. Pretraining done, everything else unknown.
The honest version
The clean narrative would be that the West is back and the story is over. It’s not that clean.
Yes, Fable 5.1 and GPT-6 Astra are real, and yes they’re good, and yes the premium ceiling exists again. But both of them shipped with a footnote you have to read before you trust the headline. One blends a weaker model into its own benchmark and doesn’t tell you how often. The other aced an exploit-writing test and comes with a program to loosen its own safety rails. That’s not a reason to dismiss them. It’s a reason to read the harness section before you quote the number.
And the more durable story is the one nobody put on a stage. The cheapest genuinely good model on the board is a fifteen-cent open-weight model from Z.ai that just took a crown a Xiaomi model held for a month and a half. The frontier got a loud, expensive reload this week. The floor got quietly cheaper. If you’re actually shipping something and paying the bill yourself, guess which one changes your life more.
I’ll be back next week to see whether Grok 4.7 shows up on the twelfth or whether “ten days” means what it usually means.
