Stephan Miller
Grok 4.7 Shipped. Even Elon Musk Said It Was Mid.

Grok 4.7 Shipped. Even Elon Musk Said It Was Mid.

Last week I ended this roundup with a promise. I said I’d be back to find out whether Grok 4.7 actually exists yet, and whether its creator liked it any better once it did. Well. It exists. He does not seem to like it any better, and neither, so far, does anyone else.

Grok 4.7 shipped Sunday, September 21. That’s the whole headline and the whole punchline at once. For two straight weeks Musk stood in public and marked his own unreleased model down, from “better than 4.6 in every way” to “beats everything on the board” to, finally, “roughly on par with Opus 5.0, not 5.1.” Then the model landed. And the independent benchmarks put it right about where he’d talked it down to. This almost never happens. Usually the shipped thing is worse than the hype. This time the hype had already deflated itself to match, and the model still landed under the deflated number. Read on, because the gap between what got announced and what got shipped is the whole story this week, and for once the shipped thing is the one that gets a fair shake.

The quiet story is still down in the bargain bin, where the fifteen-cent model I keep telling you about held every one of its five categories and nothing changed except that it got a little faster. Sometimes the boring outcome is the important one. Let me walk you through both.

Grok 4.7 shipped, and it’s the model its own maker warned you about

Here are the specs, and this time they’re specs, not tweets. Grok 4.7 is live as of September 21 through the Grok API, Cursor, and Grok Build. It runs on a new base model at a claimed 2.1 trillion parameters, up about 40 percent from Grok 4.6’s 1.5 trillion. It takes a 500K-token context, multimodal input, and it’s priced at two dollars in and six out, with cached input at fifty cents. That’s the same sticker as Grok 4.6, which is the one genuinely good decision in this whole launch. xAI did not try to charge frontier money for a mid-frontier model. Credit where it’s due.

Now the number that matters. On the Artificial Analysis intelligence index, the independent one that blends ten hard evals, Grok 4.7 scores a 46. The two models at the top of that board, Claude Fable 5.1 and GPT-6 Astra, both sit at 53. So the model xAI built a two-week hype cycle around lands seven full points behind the frontier, in the same neighborhood as models that are a good deal cheaper and a good deal older. Mid-pack. Exactly where Musk himself put it on the fourteenth, before it had a benchmark to its name.

The thing I keep coming back to is that this is a functional pattern now, not a Grok quirk. xAI ships a model that is cheap and legitimately fine, then wraps it in language it cannot support. Grok 4.7 is a perfectly reasonable two-dollar coding model. It is not the smartest thing on Earth, nobody who has run it thinks it is, and the person who built it told you so a week early. If you priced it at what it is instead of announcing it as what it isn’t, this would be a good news week for xAI. Instead it’s a case study.

The benchmark you cite is the argument you’re making

Now the part worth reading past the headline for.

xAI’s own launch materials lean hard on coding, and on coding the gains are real. On DeepSWE v1.1 at high effort, Grok 4.7 posts 71.0, up from Grok 4.6’s 65.2. On CursorBench 4.0, a test built around longer-running coding tasks, it hits 46.3 against 40.4 for its predecessor. Those are honest improvements. If you’re doing the kind of work those benchmarks measure, 4.7 is a real step up from 4.6 at the same price, and that’s a fine reason to switch.

Then the-decoder ran the harder agentic test, Terminal-Bench 4.0, and the floor gave out. Grok 4.7 scored 26 percent. GPT-6 Astra hit 60 on the same test. Fable 5.1 hit 55. And the cheap DeepSeek V4.1 Flash edged Grok out at 27. So on the benchmark xAI put in the deck, Grok 4.7 looks like a solid upgrade, and on the benchmark it left out, it finishes behind a Chinese model that costs less. Both results are true. They’re measuring different things, and the aggregate index score of 46 is what you get when you stop cherry-picking and average it out.

The lesson underneath keeps earning its keep. The benchmark somebody cites is the argument they’re making. When a launch deck shows you three benchmarks, the interesting question is always which ones it didn’t show you. xAI showed you DeepSWE and CursorBench. It did not show you Terminal-Bench. Now you know why.

The cheap model kept all five of its categories

Now the story that actually moves your bill, which as usual is the least dramatic one on the page.

Two weeks ago GLM-5.3-Flash from Z.ai took the cheapest-good-model crown off Xiaomi’s MiMo v2.5 Pro. Last week it grabbed a fifth Arena category. This week it did the boring, valuable thing and simply held everything. It’s still the cheapest model inside the competitive band for five of the six categories: Overall, Coding, Instruction Following, Hard Prompts, and Math. Creative Writing is still the lone holdout, still a Gemini story, because GLM never cracked that band and probably won’t.

The usage board agrees with the preference board, which is the part that makes this real rather than an Arena curiosity. On OpenRouter, which counts actual tokens on actual paid calls, GLM-5.3-Flash is sitting at number two on the entire platform, somewhere north of ten trillion tokens a week, behind DeepSeek V4 Flash and ahead of GPT-5.6 Luna and MiMo. Chinese-built models are running around 46 percent of all tokens on the platform now, with DeepSeek the single largest vendor at roughly 16 percent. People aren’t voting for these models. They’re running them, in production, with their own money.

One thing genuinely improved this week, and it’s the caveat that used to matter most. The old “cheap but slow” knock on GLM keeps softening. Artificial Analysis now clocks it at about 89 output tokens a second on Z.ai’s own endpoint, comfortably above the roughly 75 median for open-weight models in its class. It was crawling along near 60 a few weeks ago. For a fifty-cent model, that’s the difference between something you tolerate in a chat window and something you can actually drop into an agent loop.

Two catches, same as always. First, GLM-5.3-Flash scores a 42 on the hard-reasoning index where the frontier lives in the fifties. It’s cheap and preference-strong, not a deep-reasoning machine. For everyday work that’s a rounding error you’ll never feel. For genuinely hard problems, feel it. Second, the price. The seven-cents-in, quarter-out number you may still have in your head was a launch promo, and it died on September 9. List is fifteen cents in and fifty cents out. Arena’s price column is still cheerfully showing the dead promo, which is going to burn somebody who budgets off it. Build on fifty cents out. If you sized an August spend on the old number, it doubled on you three weeks ago and nobody sent a memo.

The cheapskate picks

Same method every week. For each Arena category I take the leader’s rating, draw a band 50 points below it, and find the cheapest model still sitting inside that band. The whole premise is that Arena ratings cluster tight at the top, so the category leader is usually a rounding error better than something 20 to 100 times cheaper. I compute the band from the full table in code, not by eyeballing the first screen, because eyeballing it is exactly how you delete the entire cheap tail and accidentally crown the cheapest expensive model. I learned that one the hard way, in public, a couple months back.

Bands ran deep again this week: around 60 models inside the Overall band, 64 in Coding, thinner in Creative and Math. GLM-5.3-Flash is quoted at list, fifteen and fifty, not the expired promo Arena is still displaying. Arena data is dated September 13.

CategoryLeader$ leader outCheapskate pick$ pick outΔ ratingCheaper by
Overallclaude-fable-5 (1506)$50GLM-5.3-Flash (1475, #29)$0.50−31~100×
Codingclaude-fable-5 (1552)$50GLM-5.3-Flash (1525, #20)$0.50−27~100×
Creative Writingclaude-fable-5 (1504)$50gemini-3-flash (1459, #26)$3−45~16.7×
Instruction Followingclaude-opus-4-6-high (1513)$25GLM-5.3-Flash (1471, #29)$0.50−42~50×
Hard Promptsclaude-opus-4-6-high (1533)$25GLM-5.3-Flash (1498, #28)$0.50−35~50×
Mathclaude-fable-5 (1526, prelim)$50GLM-5.3-Flash (1513, #7)$0.50−13~100×

A few notes on how to read that. Overall and Hard Prompts are the rows you can lean on: GLM-5.3-Flash is sitting on real vote counts there, ten thousand in Overall, and it’s cleanly the cheapest thing in a very deep band. Coding is a great rating on thin votes, twentieth place for fifty cents, but under three thousand votes behind it, so if you want certainty over the last few dollars, MiMo v2.5 Pro at rank 25 with sixteen thousand votes for eighty-seven cents is the steadier bet. Math is the one to squint at hardest: GLM shows seventh for fifty cents, which is loud, but it’s riding 441 preliminary votes on a board where the whole top is thin. Treat that row as a strong suggestion, not a promise. MiMo at the band edge or a Gemini Flash won’t embarrass you there either.

Creative Writing stays Gemini because GLM never made the band, and the only value pick is gemini-3-flash at three dollars, which is still 17 times cheaper than a fifty-dollar leader. Nobody is paying fifty dollars a million tokens to draft blog intros.

Newest still isn’t best, and it’s still funny

I wrote almost this exact paragraph last week and I’m writing it again, because the pattern refuses to break and it keeps getting funnier.

The number one model on Arena Overall is claude-fable-5. Not Fable 5.1, the newer one Anthropic shipped a few weeks back. The old one. And sitting near the top of Instruction Following and Hard Prompts, as the outright leader on both, is claude-opus-4-6, a model that is roughly a year old. It beats Opus 5. It beats Opus 4.7 and 4.8. In the blind test, where nobody can see the version number, the crowd keeps reaching for the model everyone in the timeline already moved on from.

I’m not saying the new models are bad. I’m saying the booth keeps preferring the boring old one, and it’s happened enough weeks running that it’s a pattern, not a fluke. If you upgraded off Opus 4.6 because a bigger number came out, the people voting blind would like a word with you.

About that bill

The horror story this week isn’t a model, it’s a loop, and it’s the same species of loop that keeps eating people alive in 2026.

Google’s Mandiant team put out an enterprise-AI-risk report on September 16 with a clean, awful example in it. An accounting agent hit a runaway execution loop and fired off more than 15,000 high-cost API calls in under an hour. Roughly fifty thousand dollars, gone, before a human looked at a dashboard. No budget ceiling. No alert anyone acted on. Just a bot doing the same expensive thing over and over, faster than anyone was watching.

Here’s the part that made me laugh and then wince. Grok 4.7’s launch copy sells it as a model “designed to better verify its own output.” That is a lovely sentence. It is also describing the exact capability every runaway agent lacks, the ability to notice it’s stuck in a loop and stop. A better self-verifier would genuinely help with this. A marketing line about one does nothing, and the fifty-thousand-dollar hour happened the same week the line got written. The token bill in 2026 goes wrong in three normal ways: the sticker lies about the task, the loop has no brakes, and the cheap price had an expiration date you didn’t read. This week served up all three.

What’s coming

Three to watch.

Grok 4.7, now that it exists, gets to spend the next couple weeks accumulating actual Arena votes and EU availability, which lagged the US launch as xAI launches always do. I’ll be curious whether the blind test is kinder to it than the benchmarks were, or crueler.

GPT-6 Astra is still finishing its rollout. It went from a handful of day-one orgs to the ChatGPT paid tiers, the OpenAI API, Azure, and AWS Bedrock over the past couple weeks. The Daybreak program that loosens the safety rails for vetted organizations is the piece worth watching, given this is the model that reportedly maxed out an exploit-writing benchmark a couple weeks ago.

Gemini 4, pretraining done and everything else a rumor. Google keeps dripping out Flash models to stay in the conversation while the real thing bakes. Note the small print on the current one: Gemini 3.8 Flash’s 75-cents-and-3.75 pricing is an intro rate that doubles on January 1. Late 2026 is the vague window for the actual next-gen model, if you believe the tea leaves, and I’ve stopped believing Google’s dates on principle.

The honest version

The clean narrative this week would be that xAI face-planted. That’s not quite it, and the truth is more useful.

Grok 4.7 is a decent, cheap, honestly-priced coding model that got buried under a launch it couldn’t live up to, by a man who then spent two weeks digging the hole himself. The model is fine. The framing was the problem, and the framing is always the problem. Cost-per-token isn’t cost-per-task. The benchmark in the deck isn’t the benchmark that matters. And “most capable model yet” means whatever the person saying it needs it to mean this quarter.

Underneath all of it, a fifteen-cent open-weight model from Z.ai held its five categories, got faster, and stayed the second most-used model on the planet without anyone holding a keynote about it. That’s the line that changes your life if you’re shipping something and paying the bill yourself. The frontier had a loud week arguing about who’s smartest. The floor just sat there being cheap and getting quietly better. You already know which one you’ll actually be running next month.

I’ll be back next week to see whether the Arena crowd is any nicer to Grok 4.7 than its own creator was. Low bar. We’ll find out.

Stephan Miller

Written by AI, edited by

Kansas City Software Engineer and Author

Twitter | Github | LinkedIn

Updated