Stephan Miller
GLM-5.3 Got Too Good at Hacking to Ship on Time

GLM-5.3 Got Too Good at Hacking to Ship on Time

A Chinese AI lab shipped a model that got so good at hacking, mid-training, that they decided to hold the open weights back and spend two more weeks hardening it before they let you download it. The model started writing full exploitation chains on its own and they got a little spooked by their own homework.

That’s Z.ai’s GLM-5.3. And it is the cleanest sign yet that the interesting stuff in AI is no longer happening where it used to. The frontier is leaking into open weights.

Meanwhile the cheap model I keep telling you to use won its sixth straight week as the best value on the board, Google reportedly gave up on the flagship it promised at I/O and started over, and Elon says Grok 4.7 is “3 to 4 weeks” out, which in Musk time means sometime before the heat death of the universe.

The model that outgrew its own training

GLM-5.3 came out of Z.ai on August 14. Same base model as GLM-5.2, they say. Every gain came from scaled-up post-training, which is already a slightly unnerving sentence when you read the rest of it.

Here’s the part that made me put my coffee down. On CyberGym, a cybersecurity benchmark, GLM-5.3 scored 84.5%. That is the leading published number on that benchmark. It beats Claude Fable 5 (83.8%) and GPT-5.6 Sol (83.6%), which are the two most safety-gated, most lawyered, most “we take this very seriously” frontier models the West has. An open-weights model from a company most American developers can’t name just posted the best offensive-security score anyone has published.

And the way Z.ai describes how it got there is the actual story. They fed the model vulnerability-discovery data expecting it to get better at finding isolated bugs. Instead, in their own words, “capability continued compounding as training scaled, and the model began reasoning across multiple stages of exploitation, forming coherent plans for complete exploitation chains rather than isolated bug-finding.”

Read that again. They wanted a better bug-finder. They got something that plans whole attacks. ExploitBench doubled, from 24.4% to 54.4%. On a practical exploitation task, the old model solved 29 problems in two hours and GLM-5.3 solved 105. That is not a nudge. That is a different animal.

So Z.ai is now delaying the open-weights release, which they’ve promised for “about two weeks,” specifically to harden a model they publicly admit developed offensive capability faster than they expected. Sit with the shape of that. For a year the story was: the West locks its models down, China ships open. Remember Claude Fable 5 getting yanked offline by a US export-control order back in June because a jailbreak might unlock cyber capabilities? Now a Chinese lab is voluntarily holding its own weights over the exact same fear.

But every one of those cyber numbers is on Z.ai’s own evaluations. Nobody independent has replicated CyberGym at 84.5% yet. Treat it like every vendor launch table, which is to say: interested, but skeptical, until someone with no skin in it runs the eval. The exact decimal is marketing until proven otherwise.

The cheapskate table, and why it’s boring on purpose

Every week I pull the Arena leaderboards, take each category leader’s rating, draw a line 50 points below it, and find the cheapest model that still sits inside that band. Not the cheapest of the top 20. The cheapest inside the whole 50-point window, because that window runs 50-plus models deep and the good cheap stuff lives way down in the ranks where nobody scrolls.

I do it in code now, because I got burned eyeballing the first screen once and it’s a mistake you only make in public one time.

Here is where the value actually is this week:

CategoryLeader (output $/1M)Cheapskate pick (output $/1M)Cheaper byRating gap
Overallclaude-fable-5 ($50)mimo-v2.5-pro ($0.87)~57x-40
Codingclaude-opus-4-7-high ($25)mimo-v2.5-pro ($0.87)~29x-33
Creative Writingclaude-fable-5 ($50)gemini-3.6-flash-high ($1.88)~27x-38
Instruction Followingclaude-opus-4-6-high ($25)mimo-v2.5-pro ($0.87)~29x-43
Hard Promptsclaude-opus-4-6-high ($25)mimo-v2.5-pro ($0.87)~29x-37
Mathclaude-opus-5-max ($25)gemini-3.6-flash-high ($1.88)~13x-35

MiMo v2.5 Pro from Xiaomi takes four of the six categories again. Overall, Coding, Instruction Following, Hard Prompts. Eighty-seven cents per million output tokens versus fifty dollars for the Fable 5 leader in the Overall board. That’s not a discount. That’s a different pricing universe, for a rating that lands 40 points back on a scale where the entire competitive top end fits inside 50.

This is the sixth straight week MiMo has held this spine. Six weeks. If you’re waiting for me to announce some exciting new value pick, I don’t have one, and that’s the point. The honest recommendation hasn’t changed since mid-July. When you want cheap and genuinely competitive for everyday work, this is still the answer. Artificial Analysis backs it up from the other direction too: MiMo sits on their Intelligence-vs-Cost Pareto frontier, which means two totally different methodologies point at the same model. That’s about as strong a signal as this job produces.

The caveats are the same ones I gave you last month, because they haven’t changed either. MiMo is slow, about 52 tokens per second, below the median. And its raw intelligence score on Artificial Analysis is 43 against a frontier around 60. So for genuinely hard reasoning, this isn’t your model. For the 90% of work that’s “summarize this, rewrite that, follow these instructions,” you will not feel the missing 40 rating points and you will very much feel the 57x price cut.

Creative Writing and Math both go to Gemini 3.6 Flash-high at $1.88, same as last week. Two notes on the table. In Coding there’s a cheaper gamble sitting right at the band edge: hy3, Tencent’s Hunyuan 3, at 53 cents, two points above the cutoff line on only 1,655 votes. It undercuts MiMo but it’s riding the very edge of the band on thin data, so I keep MiMo as the anchor and file hy3 under “for people who like living dangerously.” And the Math pick is shaky: the leader there is preliminary on 592 votes and the whole band is only 12 models deep. Take it as provisional.

Newest is not best, and the labs know it

Here’s a thing that should embarrass more people than it does. On the Arena Overall board, Anthropic’s Claude Opus 4.6-high sits at #2. Its own newer sibling, Opus 5-high, sits at #7. The older model beats the newer one from the same company. Opus 5 only clearly wins Math, and that’s on preliminary votes.

We have fully arrived at the era where “just use the latest one” is bad advice, and it’s bad advice inside a single lab’s own lineup. The version number went up. The blind-comparison preference went down. That happens more than the launch blog posts would ever tell you.

The flip side is Grok 4.6, which is the opposite trap. On Artificial Analysis it’s up in the top cluster at 60-61, right there with Opus 5 and GPT-5.6 Sol. On Arena it’s languishing around #46 Overall on 3,400-ish fresh votes. That’s not a contradiction, it’s just Arena being slow to accumulate votes on a new model. When the hard-benchmark score and the crowd-vote score disagree this hard on a fresh release, believe the benchmark for now. Grok 4.6 is smarter than it looks in the blind booth. At $2/$6 flat, it’s also cheap for what it is.

The token bill comes due

If you run agents, you already know the horror story, because you’ve either lived it or you’re one unattended weekend away from it.

The numbers coming out this cycle are genuinely stupid. One company reportedly ran up something like a $500 million Claude bill after forgetting to set usage limits for employees. Uber apparently torched its entire 2026 AI-coding budget by April. Teams routinely report hitting 3x their annual token budget by spring. And the recurring nightmare, the one that’ll get you personally: a handful of agents stuck in a recursive loop, running unattended over a weekend, ringing up tens of thousands of dollars because every step re-sends the full context and every step gets billed.

The mechanism never changes. An agent doesn’t make one call. It makes hundreds, in loops, across tools, while you’re asleep. The cost model of a chatbot and the cost model of an autonomous agent are not the same species, and a lot of budgets are still priced like it’s 2024.

Then there’s slopsquatting, which is my new favorite piece of dystopia. LLMs hallucinate package names. Attackers noticed. So they pre-register the fake names the models invent, as malware. Your coding agent confidently runs pip install on a package that didn’t exist until a bad actor created it to catch exactly this mistake. It’s a cost bug and a supply-chain attack wearing the same coat. Around one in five AI-suggested packages don’t exist, and now some of the ones that “do” are traps.

And loop it back to GLM-5.3 for a second, because the dual-use thing isn’t abstract. A model that plans full exploitation chains, shipped as open weights you can run locally with no logging and no off switch, is the defensive-security teams’ entire worry expressed in a single download. The vendor said the quiet part out loud this time. That’s progress of a sort.

What’s coming, and what quietly died

Grok 4.7 is “training done, SpaceX data going in, 3 to 4 weeks out,” per Musk, a window that already slipped from late July. Pencil in early September, and use a pencil with a good eraser. GLM-5.3’s open weights are the two-weeks-out release I mentioned above, assuming the safety hardening goes to plan. The delay is the story there, not the date.

Then there’s Gemini 3.5 Pro, which is, according to SemiAnalysis, shelved. Dead. Google reportedly gave up after missing late June, then July 17, then early August, and pointed its engineers at pretraining Gemini 4 instead, the largest single training run Pichai says the company has ever attempted. So Gemini 3.6 Flash is the stopgap flagship now. That’s a strange sentence to type about a Flash model. When your fast, cheap tier is holding the fort because the real flagship got scrapped, you can guess how the pretraining bet is going. One more for the pile: GLM-5.2 Turbo shipped August 17, a speed-tuned variant, not on Arena yet. Filed for next week.

The honest read

I’ve been writing this roundup long enough to watch a genuine trend line form, and here it is: the capability that used to justify locking models in a vault keeps showing up in open weights you can download, mostly from China, at a fraction of the price, a few weeks later. First it was cost. Now it’s frontier-grade cyber capability. The export controls, the safety gates, the “limited release via vetted partners” stuff, all of it assumes the good models live in a few buildings you can regulate. The board says otherwise. Chinese models are about 45% of all OpenRouter traffic now, up from a world where US models were 70% a year ago.

None of this is a reason to run GLM-5.3 as your daily driver, and it’s definitely not a reason to skip the boring correct answer. If you take one practical thing from this week, take the table: MiMo v2.5 Pro for the everyday grind, Gemini 3.6 Flash-high when you need creative or math on a budget, and a very short leash on any agent you let run unattended. The exciting model this week is the one that scared its own creators. The one you should actually be using is still the unglamorous 87-cent workhorse that’s won six weeks running and will probably win a seventh.

See you next week, assuming an agent hasn’t bankrupted either of us by then.

Stephan Miller

Written by AI, edited by

Kansas City Software Engineer and Author

Twitter | Github | LinkedIn

Updated