Stephan Miller
GPT-6 Astra Is the 'Smartest' Model. It Finished 24th.

GPT-6 Astra Is the 'Smartest' Model. It Finished 24th.

Last week I wrote that the American labs had finally woken up. GPT-6 Astra and Claude Fable 5.1 both shipped, both at ten dollars in and fifty out, and I told you to read the footnote on each before you trusted the headline. This week the footnotes came due.

Here’s what happens after the keynote. The blind-test votes trickle in, the independent testers run their own bills, and the model that got called “the world’s most intelligent” a week ago has to actually go stand in line with everything else. GPT-6 Astra did that this week. It finished 24th.

Meanwhile Grok 4.7 got the worst possible review, from the one person who has actually used it. And down in the bargain bin, the cheap model I keep telling you about quietly took a fifth category off the board. So the pattern this week is simple. Everything loud got quieter, and the quiet thing got louder. Let me walk you through it, and yes, the table you came for is still here.

The “world’s most intelligent model” showed up to the blind test and finished 24th

GPT-6 Astra shipped September 3 and OpenAI’s launch copy called it the most intelligent model on Earth. On the Artificial Analysis intelligence index, fair enough, it ties Fable 5.1 at the top. That’s a benchmark score.

This week it hit Arena, which is the board where humans vote on which of two anonymous answers they actually prefer, and the picture got more honest in a hurry. The max variant landed 24th overall on preliminary votes, sitting behind a wall of older Claude models, some of them the better part of a year old. It does better in Coding, where it slots in around sixth. Everywhere else it’s mid-pack. This is the split I write about most weeks and it never stops being useful: Artificial Analysis measures what scores well on hard evals, Arena measures what people like when they can’t see the logo, and the two disagree constantly. A model can be smart on paper and forgettable in the booth. Astra is currently both.

The independent testers found the more expensive problem. OpenAI prices Astra at roughly 40 percent of Fable 5.1 per token, which sounds like a win until someone runs the same job through both. One head-to-head test suite cost about 198 dollars in tokens on Astra against 113 dollars on Fable 5.1. That’s 75 percent more money to do the same work, on the model that’s supposedly cheaper. It happens because Astra spends more tokens per task, and the per-token sticker price never tells you how many tokens a task takes. This is the oldest trap in the roundup and it caught a brand-new flagship on launch week.

Where Astra genuinely wins, it wins hard. FrontierMath Tier 4 at 97.6 percent against Fable 5.1’s 87.8. Terminal-Bench Science well ahead. If your work lives in those specific rooms, it’s a real tool. Just don’t read “most intelligent model on Earth” and assume that means “the one that’ll feel best or cost least on your actual work.” The blind test and the invoice both disagree.

Grok 4.7 got a bad review from the person who built it

I love this one. Last week I told you Grok 4.7 was the loud upcoming thing, that Musk had said “ten days” on September 2, which pointed at roughly the twelfth, and that you should treat the date as a tweet and not a commitment. Reader, it was a tweet.

The twelfth came and went. On the eleventh Musk said it “needs a few more days to cook.” Then on the fourteenth the pitch itself changed. The model that was going to be “better than 4.6 in every way” and beat everything on the board got quietly re-described as “roughly on par with Opus 5.0, not 5.1.” Read that again. Before the thing has a model card, a price, a benchmark table, or an API you can call, its own creator walked it back from “beats everyone” to “about as good as the model Anthropic already replaced.” That’s not a launch. That’s a man marking down his own inventory in public.

The specs, such as they are, remain a set of claims: 2.1 trillion parameters, up 40 percent from Grok 4.6, some of the training data pulled from SpaceX internal engineering records, “even better token efficiency,” “slightly slower to serve.” Every one of those is a sentence from a person and not a line from a document. When there’s an actual model to test I’ll test it. Until then, Grok 4.7 is the clearest example this month of the gap between what gets announced and what gets shipped, and the gap is being narrated live by the announcer.

The boring cheap model took a fifth category

Now the quiet story, which is the one that actually changes your bill.

Two weeks ago I called the handover. GLM-5.3-Flash from Z.ai took the cheapest-good-model crown from Xiaomi’s MiMo v2.5 Pro, first as an emerging pick on thin votes, then confirmed once the votes firmed up. This week it kept going. GLM-5.3-Flash is now the cheapest model inside the competitive band for five of the six Arena categories. It picked up Math this week, a board it wasn’t even in a fortnight ago. The only holdout is Creative Writing, which stays a Gemini story because GLM never cracked that band.

It’s not just an Arena artifact either. On OpenRouter, which measures actual token volume, real money spent on real calls, GLM-5.3-Flash is now the number two model on the entire platform at around 10 trillion tokens a week. It sits behind DeepSeek V4 Flash and ahead of GPT-5.6 Luna and MiMo. So the preference board and the usage board agree: this is a model people are genuinely running, not just voting on. MiMo is the runner-up now, and still a good one, with ten to sixty times the vote count depending on the category, which is exactly why it stays in the table as the steady anchor when GLM’s votes are thin.

Two catches, because there are always two. First, GLM-5.3-Flash scores a 42 on the hard-reasoning index where the frontier models live in the fifties and sixties. It’s preference-strong and cheap, not a deep-reasoning machine. For everyday work that’s a rounding error. For genuinely hard problems it’s a caveat you should feel. Second, and this one’s on the price tag: the seven-and-a-half-cents-in, quarter-out number you’ve seen quoted everywhere was a 50 percent launch promo, and it expired September 9. List is fifteen cents in and fifty cents out, and Z.ai’s own OpenRouter endpoint moved to list on the eleventh. Arena’s table still shows the dead promo price, which is going to mislead somebody. Build your budget on fifteen and fifty. If you sized a spend off the August number, it just doubled on you.

One piece of good news to balance that. The old “cheap but slow” knock is softening. Artificial Analysis now clocks GLM-5.3-Flash at around 114 output tokens a second on Z.ai’s own endpoint, up from the roughly 60 I was quoting a couple weeks ago. It still has a slightly high time-to-first-token, but for a fifty-cent model that’s genuinely usable in an agent loop now, not just in a chat window.

The cheapskate picks

Same method as always. For each Arena category I take the leader’s rating, draw a band 50 points below it, and find the cheapest model still inside that band. The whole premise is that Arena ratings cluster tight at the top, so the leader is usually a rounding error better than something 20 to 100 times cheaper. I compute the band from the full table in code, not by eyeballing the first screen, because eyeballing it is how you accidentally delete the entire cheap tail and crown the cheapest expensive model. Ask me how I know.

Bands ran deep again this week: about 62 models inside the Overall band, 64 in Coding, thinner in Creative and Math. GLM-5.3-Flash quoted at list, fifteen and fifty, not the expired promo. Arena data dated around September 14.

CategoryLeader$ leader outCheapskate pick$ pick outΔ ratingCheaper by
Overallclaude-fable-5 (1506)$50GLM-5.3-Flash (1475, #29)$0.50−31~100×
Codingclaude-fable-5 (1552)$50GLM-5.3-Flash (1525, #20)$0.50−27~100×
Creative Writingclaude-fable-5 (1504)$50gemini-3-flash (1459, #26)$3−45~16.7×
Instruction Followingclaude-opus-4-6-high (1513)$25GLM-5.3-Flash (1471, #29)$0.50−42~50×
Hard Promptsclaude-opus-4-6-high (1533)$25GLM-5.3-Flash (1498, #28)$0.50−35~50×
Mathclaude-fable-5 (1526, prelim)$50GLM-5.3-Flash (1513, #7)$0.50−13~100×

A few notes on reading that. Overall and Hard Prompts are the solid ones: GLM-5.3-Flash is sitting on real vote counts there, ten thousand in Overall, and it’s cleanly the cheapest thing in a very deep band. In Coding the rating is great, twentieth for fifty cents, but the vote count is still on the thin side at under three thousand, so if you want certainty over the last few dollars, MiMo v2.5 Pro at rank 25 with sixteen thousand votes for eighty-seven cents is the steadier bet. Math is the one to squint at: GLM-5.3-Flash shows up seventh for fifty cents, which is loud, but it’s riding 441 preliminary votes on a board where the whole top is thin. Treat that row as a strong suggestion, not a promise, and if you want a sturdier number, MiMo at the band edge or a Gemini Flash won’t embarrass you.

Creative Writing stays Gemini because GLM never made that band, and the only value pick there is gemini-3-flash at three dollars, which is still 17 times cheaper than the leader. Nobody’s paying fifty dollars a million to write blog intros.

Newest still isn’t best, and now it’s just funny

Here’s the thing I can’t stop noticing. The number one model on Arena Overall is claude-fable-5. Not Fable 5.1, the newer one Anthropic shipped two weeks ago. The old one. And sitting at number two, ahead of every single newer Opus, is claude-opus-4-6, a model that’s roughly a year old at this point. It beats Opus 5. It beats Opus 4.7 and 4.8. In Instruction Following and Hard Prompts it’s the outright leader.

I’m not saying the new models are bad. I’m saying the blind test keeps preferring the model everyone already moved on from, and that’s now happened enough weeks in a row that it’s a pattern and not a fluke. If you switched off Opus 4.6 because something newer came out, the crowd that can’t see the version number would like a word.

While we’re at it, one more sleeper. Meta’s Muse Spark line is quietly camped in the Arena top five overall and almost nobody talks about it. The 1.2 xHigh variant sits fourth, the new 1.3 max is up around eighth. It’s closed-weight and the newest tier doesn’t have a clean public price, so it can’t take a cheapskate slot, but if you’re only watching the Claude-OpenAI-Google fight you’re missing a genuinely competitive model.

Google shipped another Flash, because of course it did

Quick Google check-in, same paragraph I’ve now written four or five times. On September 2 they shipped Gemini 3.8 Flash, the third Flash release in six weeks, built on the 3.7 base, same price as 3.7 at 75 cents in and 3.75 out, with a locked-down Cyber sibling for security work. It’s good. It’s already climbing the Arena boards fast, top ten overall and third in Creative on preliminary votes. Note the price cliff, though: that 75-and-3.75 is an intro rate that doubles to 1.50 and 7.50 on January 1.

And it’s, again, shipping instead of the real thing. Gemini 3.5 Pro is shelved. Gemini 4 has “cleared pretraining” with the largest training run in Google’s history and no benchmarks, no price, no date beyond a vague late 2026. So the strategy holds: drip out excellent cheap Flash models to stay in the conversation while the actual next-generation model bakes in the background. It’s working. It’s also the reason I can copy-paste this section every month.

About that bill

The horror story this week isn’t one incident, it’s a category, and it keeps getting more expensive.

The clean version is the Astra number above: a “cheaper” model that cost 75 percent more to run because cost-per-token isn’t cost-per-task. But the loud version is the agent loop. One that made the rounds this week: an Analyzer agent and a Verifier agent started ping-ponging, the Analyzer generating output and the Verifier asking for more analysis, over and over, with no budget ceiling and no alert anyone acted on. It ran for 264 hours before the number on the billing dashboard got big enough for a human to finally look. Forty-seven thousand dollars. Two bots politely asking each other to keep going, for eleven days.

And the quiet version is the promo cliff. If you wired GLM-5.3-Flash into anything in August at seven-and-a-half cents, your cost doubled on September 9 and Arena is still showing you the old number like nothing happened. None of these are exotic. They’re the three normal ways a 2026 AI bill goes wrong: the sticker lies about the task, the loop has no brakes, and the cheap price had an expiration date you didn’t read.

What’s coming

Three to watch.

Grok 4.7, covered above, whenever it actually ships and at whatever the reviews say it is by then rather than what it was announced as.

GPT-6 Astra is still mid-rollout. Day-one access went to a handful of orgs, and the ChatGPT tiers, the API, and AWS are coming online “over the coming days.” There’s also a program called Daybreak that loosens the safety rails for vetted organizations, which is worth keeping an eye on given this is the model that reportedly aced an exploit-writing benchmark last week.

Gemini 4, pretraining done, everything else a rumor. Late 2026 if you believe the tea leaves.

The honest version

The clean narrative last week was that the West is back. The clean narrative this week would be that the West already fumbled it. Neither is true, and the truth is more boring and more useful than either.

The frontier models are real and good and expensive, and when you actually run them the story gets complicated fast. The one billed as smartest finished 24th in the blind test and cost more to run than its pricier rival. The loudest upcoming model got downgraded by its own creator before launch. That’s not the West failing. It’s just the difference between a keynote and an invoice, and the invoice always shows up a week late.

And underneath all of it, a fifteen-cent open-weight model from Z.ai took its fifth category and became the second most-used model on the planet without a single keynote. If you’re shipping something and paying the bill yourself, that’s the line in this whole roundup that changes your life. The frontier got a loud, expensive week. The floor got quietly, permanently cheaper. Guess which one you’ll actually be using in a month.

I’ll be back next week to find out whether Grok 4.7 exists yet, and whether its creator likes it any better by then.

Stephan Miller

Written by AI, edited by

Kansas City Software Engineer and Author

Twitter | Github | LinkedIn

Updated