GLM-5.3-Flash Comes for MiMo's Cheap-LLM Crown
Last week I told you a Chinese lab built a model that got so good at hacking it scared its own makers into holding the weights back. That was GLM-5.3. The 744B flagship. The one Z.ai benched for a couple weeks of safety hardening.
This week its little brother showed up and took four of the value leaderboards.
I have been tracking the cheapest-good-model question every week for a while now, and for six straight cycles the answer never moved. It was MiMo v2.5 Pro. Then GLM-5.3-Flash landed on August 26, priced lower than MiMo and scoring higher on the one benchmark that actually measures whether a model can think, and the streak was over.
- The six-cycle king finally got shoved off the throne
- The asterisk that keeps me from crowning it
- The Cheapskate Picks
- The bill always comes due
- What’s coming, and what’s stalling
- The takeaway
The six-cycle king finally got shoved off the throne
Here is the pattern that held from late July all the way through last week. Every time I pulled the Arena boards, computed the competitive band for each category, and sorted by price, MiMo v2.5 Pro sat at the bottom of the price column in Overall, Coding, Instruction Following, and Hard Prompts. Same $0.43 in, $0.87 out. Same “yeah it’s slow but who cares at this price” caveat. Six weeks running.
MiMo won on a simple trick. Arena ratings cluster so tight at the top that dozens of models fit within 50 points of the category leader, and MiMo lived deep in that band at a fraction of the price of everything above it. It wasn’t the smartest model. It was the cheapest model that was still good enough to sit next to the expensive ones.
GLM-5.3-Flash does the exact same trick, except it does it better on both axes. It’s cheaper. It’s also smarter.
Here’s the money shot. On Artificial Analysis, which grades models on hard objective benchmarks instead of crowd votes, MiMo scores a 43 on the Intelligence Index. GLM-5.3-Flash scores a 57. A 57 puts it at number 4 out of 111 open-weight models, in the same neighborhood as Kimi K3 and GLM’s own 744B flagship. And it does that at $0.15 in, $0.50 out.
When the cheapest-in-band model on Arena is also near the top of the capability chart on a completely different methodology, that’s two independent judges pointing at the same model.
The asterisk that keeps me from crowning it
GLM-5.3-Flash’s Arena ratings are all preliminary. Every single one. In Coding it’s sitting at rank 11 with a 1531 rating, which sounds great, until you notice that rating is built on 663 votes with a plus-or-minus of 24 points. In Overall it has 2,424 votes. In Instruction Following, 860. MiMo, by contrast, has 56,020 votes behind its Overall rating and tens of thousands in every other category.
GLM-5.3-Flash is the emerging value pick. MiMo is still the established one. If you want the cheapest number on the board and you’re fine being an early tester whose vote helps firm up a rating that might wobble, take GLM-5.3-Flash. If you want the answer you can defend when the bill comes, MiMo is still sitting right there at $0.87 with fifty thousand votes vouching for it.
Two more asterisks while I’m here. First, both of these value picks are slow. GLM-5.3-Flash pushes about 41.7 tokens per second and MiMo about 37, and the median for models this size is closer to 65. That matters a lot if you’re running an agent loop where latency stacks up over hundreds of calls.
Second, the cheapest price you’ll see quoted for GLM-5.3-Flash is a lie with an expiration date. OpenRouter shows it at $0.075 in, $0.25 out right now. That’s a 50 percent launch promo that dies on September 9. The list price is $0.15 and $0.50.
The Cheapskate Picks
The method is the same as always. Take each category leader’s Arena rating, draw a 50-point band below it, throw out anything more expensive, and whatever’s cheapest and still in the band wins. Bands this week ran from 20 models deep in Math up to 61 in Coding, so no, “the cheapest of the top 20” does not cut it. You have to pull the full band or you’ll crown the wrong model, which this newsletter has done before and isn’t about to repeat.
Prices are output dollars per million tokens, since output is what dominates most real bills.
| Category | Leader | $ out | Cheapskate pick | $ out | Δ rating | Cheaper by |
|---|---|---|---|---|---|---|
| Overall | claude-fable-5 | $50 | GLM-5.3-Flash (prelim) | $0.50 | -38 | ~100x |
| Coding | claude-opus-4-7-high | $25 | GLM-5.3-Flash (prelim) | $0.50 | -21 | 50x |
| Creative Writing | claude-fable-5 | $50 | gemini-3-flash | $3 | -47 | ~16.7x |
| Instruction Following | claude-opus-4-6-high | $25 | GLM-5.3-Flash (prelim) | $0.50 | -49 | 50x |
| Hard Prompts | claude-opus-4-6-high | $25 | GLM-5.3-Flash (prelim) | $0.50 | -38 | 50x |
| Math | claude-opus-5-max | $25 | gemini-3.7-flash-high (prelim) | $3.57 | -12 | ~7x |
A few things worth saying out loud about that table.
The Coding row is the wild one. GLM-5.3-Flash is only 21 rating points behind the best coding model on the planet, and it costs one fiftieth as much. If those 663 votes hold up, that’s not a value pick, that’s a heist. If you’d rather trust votes, MiMo is at $0.87 in the same band with 15,413 of them, and hy3 from Tencent is at $0.53 with 1,813. Take your pick of cheap.
Creative Writing stays out of the GLM story entirely. GLM-5.3-Flash didn’t even make the band there. The Gemini Flash line still owns creative value, with plain old gemini-3-flash at $3 and a healthy 4,622 votes. Google’s Flash prices bounce around week to week like a rental car company, so always re-check them, but for now that’s your pick.
Math is the category I trust least this week. The whole board is thin. The leader, Opus 5 max, is preliminary on 681 votes, and the cheapest in-band pick, gemini-3.7-flash-high at $3.57, is running on 314 votes with a plus-or-minus of 33. If you want something less wobbly, gemini-3.5-flash-high at $4.50 has 1,564 votes behind it. Neither is going to save you real money the way the GLM picks do. So is there an actual bargain hiding in Math this week? No. Sometimes a category just doesn’t have one, and Math is it right now.
The bill always comes due
Fortune ran a piece on August 22 about a CEO who was out to dinner when he caught one of his AI agents quietly burning through about a thousand dollars in tokens while nobody was watching. His take was that the scary part isn’t the cost, it’s the “insecurity,” his word for an agent that can’t tell it’s stuck in a loop and will happily thrash forever because it has no idea it’s failing. That rhymes with everything else I’ve read this year. Uber reportedly torched its entire annual AI coding budget in four months and capped its engineers at $1,500 a month. Amazon supposedly spent half a billion dollars in a single month after it handed out access with no caps at all.
This is the context that makes a $0.50 model interesting instead of just cheap. When your agent might loop for a weekend unsupervised, the difference between $0.50 and $50 per million tokens is the difference between an annoying bill and a company-wide incident.
One more gotcha, and this one is pure naming nonsense. “Flash” used to mean something specific in GLM land. The old hierarchy went Full, then Air, then Flash, with Flash being the trimmed-down budget runt of the family. GLM-5.3-Flash breaks that rule completely. It is not a shrunk version of the 744B flagship. It’s a separate model built on a new base with a different architecture, and the r/LocalLLaMA crowd spent a good chunk of launch day untangling that confusion. So if you assumed “Flash” meant “the dumb cheap one,” you were wrong.
What’s coming, and what’s stalling
The frontier, meanwhile, is weirdly quiet, and the quiet is all coming from the American labs.
Grok 4.7 is the big pending item. It’s a 2.1 trillion parameter model that xAI has been training with supplemental SpaceX data, and Musk said on August 12 it needed “three to four weeks,” which would have put it in early September. It’s slipping. Training is reportedly done, the release keeps sliding, and I’ll believe it when it’s on a leaderboard.
Gemini 3.5 Pro is, for all practical purposes, dead. Google missed date after date on it, and the word now is that they’ve shelved it and pivoted straight to pretraining Gemini 4, which they’ve publicly confirmed is a full new foundation model. Prediction markets give Gemini 4 something like an 85 percent chance of shipping by the end of November. In the meantime the 3.6 and 3.7 Flash models are Google’s actual working flagships, which is a strange place for the company to be.
And the GLM-5.3 story from last week got its ending. After the two-week safety delay, Z.ai did release the full 744B flagship weights on Hugging Face around August 28. So both halves of the family are now out in the open. The flagship you’ll need roughly eight GPUs to run, and the Flash you can rent for a nickel. Guess which one more people will actually use.
The takeaway
Six weeks of “the answer is still MiMo” and then the answer changed, which is the most interesting thing that’s happened in the value tier since I started tracking it. GLM-5.3-Flash is cheaper and smarter than the model that held the crown for a month and a half, and it’s the open-weight sibling of a flagship that got benched for being too dangerous. That’s a hell of a family.
But I’m not going to sit here and tell you the streak is definitively over on 663 votes. What I’ll tell you is this. If you’re already paying MiMo prices for a workload where a wrong answer costs you an annoyed shrug instead of a lawsuit, go throw some traffic at GLM-5.3-Flash and help firm up those ratings. If you need the safe answer today, MiMo hasn’t gotten worse, it just got company. Either way, quote the list price, not the promo, and remember that both of them are slow enough to matter if you’re running them in a loop.
Check back next week. If GLM-5.3-Flash is still cheapest-in-band with a few thousand more votes under it, we’ll call the reign officially over. If it wobbles back out of the band, well, MiMo will be right where it always is, at the bottom of the price column, being boring and cheap and good enough. I could do worse than a king like that.
