Stephan Miller
Sonnet 5.5 vs Opus 5.5: The Cheap Claude Costs More

Sonnet 5.5 vs Opus 5.5: The Cheap Claude Costs More

The price sheet caved this week. I don’t mean one lab and some polite matching. Anthropic and OpenAI both cut prices on September 22, a day after Xiaomi dropped a 28-cent open-weights model, and then Anthropic came back six days later with a new Sonnet it swears is cheaper to run. If you pay your own API bill, this was the best week of the year to be alive.

Then I read the fine print, because that’s the job. The new “cheap” Claude costs more per task than the new expensive Claude, at least when you turn the effort dial all the way up. The cheapest good coding model on the board this week comes from a company Anthropic says built it partly by siphoning Claude conversations through OpenClaw, which is the same agent framework I run in a Docker container in my house. And the free model sitting at number two on OpenRouter won’t tell you who made it.

So yes, prices went down. Read the meter anyway.

Everybody blinked at once

The clustering matters more than any one launch.

On September 21, Xiaomi shipped MiMo-V2.6-Pro and MiMo-V2.6-Flash as MIT-licensed open weights. Pro is a 1.02-trillion-parameter mixture-of-experts with 42 billion active, a million-token context, and it’s priced at 43.5 cents in and 87 cents out. Flash is 310 billion total, 15 billion active, and costs 14 cents in and 28 cents out. Artificial Analysis gives Pro a 46 on its Intelligence Index, which makes it the top-scoring open-weights model on the board right now.

A day later, on September 22, Anthropic released Claude Opus 5.5 at $4 in and $20 out, down from $5 and $25. Cache reads dropped 60 percent to 20 cents. Same day, OpenAI released GPT-6 Sol and GPT-6 Luna and cut prices in half or better. Sol went from $4/$20 to $2/$10. Luna went from 20 cents and $1.20 to 10 cents in and 50 cents out. OpenAI says those are permanent prices, not a promo, and credits caching and inference improvements.

Then on September 28, Anthropic released Claude Sonnet 5.5 at $2 in and $10 out. Same price as Sonnet 5, but a much better model.

Three labs from two countries, all pushing prices down inside 48 hours. I’ve been doing this roundup since April and I’ve never seen the whole price column in my spreadsheet go down at once like this. Usually one lab cuts and everyone else pretends not to notice for a month.

Opus 5.5 is the rare launch where every signal agrees

I spend a lot of this roundup explaining why the benchmark number and the blind-test number disagree. Three weeks ago GPT-6 Astra tied for number one on Artificial Analysis and debuted twenty-fourth on Arena. It’s twenty-sixth now. That split is normal. It’s practically the house style.

Opus 5.5 didn’t do that. It’s number one on the Artificial Analysis Intelligence Index at 58, five points clear of Fable 5.1 and GPT-6 Astra, which are tied at 53. It’s also number one on Arena Overall at 1509, and number one in Creative Writing, Instruction Following, and Hard Prompts. The hard-benchmark test and the blind vote landed on the same model in the same week. I almost never get to write that sentence.

It also ended a running gag. For weeks I’ve been pointing out that the outright leader in Instruction Following and Hard Prompts was claude-opus-4-6, a model about a year old, beating everything Anthropic shipped after it. Opus 5.5 finally took both categories. Coding is the one place the old guard hangs on: opus-4-6-high, opus-4-7-high, and fable-5-high are in a three-way tie at 1551, with Opus 5.5 at fifth.

Two asterisks. First, the votes are thin. Opus 5.5 has 2,307 votes in Overall and only 487 in Creative Writing, so its 20-point lead there could shrink as the crowd catches up. Second, every Claude score on Artificial Analysis still carries the “with fallback” tag, which means some answers came from a weaker model when the main one refused. I’ve flagged that on every Anthropic launch since Opus 5. The number is real. The footnote is load-bearing.

Reddit’s reaction has mostly been relief. The most-upvoted praise isn’t even about coding, it’s about how the thing talks. One r/ClaudeCode user called it “genuinely an order of magnitude improvement over Opus 5” in communication, which tracks. Opus 5 had a habit of saying “blast radius” and “load-bearing” like it was paid per use. (Yes, I just used “load-bearing” myself. I’m allowed. I’m not charging you per token.) Another user said they’d worked “non stop since release” and “barely made a dent” in their Max limits.

Not everybody agrees on that part. A HardForum user on the $100 Max plan said Opus 5 used to get them through whole days of coding, weekends included, without hitting the weekly limit. They upgraded to Opus 5.5 on Tuesday and were at 91 percent by Thursday morning. They were running it on “Extra” effort. Another user in the same thread, on medium, had used 12 percent of a $200 plan’s weekly limit, and a third summed it up: “5.5 extra eat tokens way more than previous gen.” Remember that effort dial. It comes back in a minute. Anthropic’s own demo was a HAProxy port from C to Rust that Opus 5.5 finished in 9.5 hours against Fable 5.1’s 12, at 51 percent lower cost. Both can be true. A well-defined port is not an open-ended feature in a messy real app, and the complaints are all coming from the messy real apps. And the best comment in the whole pile: “Queue the whining in two weeks about the model being nerfed.” Set a reminder.

The cheap Claude costs more than the expensive Claude

Sonnet 5.5 is the budget model. It’s half the per-token price of Opus 5.5. It’s also scary good. It scored 70.6 percent on Terminal-Bench 4.0, the command-line agent test, up from Sonnet 5’s 10.3, and ahead of Opus 5.5’s 66.4. On Artificial Analysis it scores 56, two points under Opus 5.5 at max and ahead of Fable 5.1 and GPT-6 Astra.

Then Artificial Analysis ran the whole index suite and checked the bill. At max effort, Sonnet 5.5 cost $7.60 per task. Opus 5.5 cost $5.98. The cheap model was 27 percent more expensive.

The reason is tokens. Sonnet 5.5 burned 410 million output tokens getting through the suite. The median for models in its price tier is 88 million. Opus 5.5 used about 260 million. At max effort, Sonnet thinks roughly 60 percent longer per task than Opus, and half the price per token doesn’t cover 60 percent more tokens. It’s a fast model, 138 tokens a second, and it spends that speed talking to itself.

It gets worse when you match scores instead of effort settings. Sonnet 5.5 at max and Opus 5.5 at xhigh both score 56 on the index. Sonnet costs $7.60 a task to get there. Opus costs $3.46. Same score, and the budget model costs 2.2 times as much. According to the same analysis, Sonnet only comes out cheaper at the bottom of the effort range.

None of that makes Sonnet 5.5 a bad model. It went from 10.3 to 70.6 on Terminal-Bench in one release, and one customer quoted by MarkTechPost measured about 121K tokens per answer against 497K on Sonnet 5. It’s just not automatically the cheap option anymore. So the practical advice is boring. Don’t crank Sonnet 5.5 to max because it’s “the cheap one.” If a task needs the top of the dial, Opus at xhigh gets you the same score for less. If it doesn’t, run either one at a lower effort and check your actual bill after a day, not the price page.

I’ve written some version of “cost per token is not cost per task” in this roundup maybe ten times. I’d never seen it flip the price order inside one lab’s own lineup before.

OpenAI just showed up at the bottom of the price sheet

For most of this year the cheapest-good-model story has been a Chinese open-weights story. Kimi, then MiMo, then GLM, and now maybe MiMo again. American labs competed at the top and left the floor alone.

GPT-6 Luna costs 10 cents in and 50 cents out. That’s exactly the list price of GLM-5.3-Flash, the model that’s been my cheapskate default for the last month. And Luna is in the Arena Coding band, ranked 43rd at 1515, only 36 points behind the leader. It isn’t the cheapest thing in the band, but it’s the first time I’ve seen an OpenAI model sit in a cheapskate band at the same price as the Chinese floor. It’s also fast. Artificial Analysis clocks it at 148 tokens a second, about three times GLM’s speed.

The catch is capability. Luna scores 37 on the Intelligence Index. GLM scores 42. Outside of coding, Luna falls out of the Arena bands entirely. It’s 86th Overall. So it’s a fast, cheap, preference-decent coding model, not a general replacement. But OpenAI ignored this end of the price sheet for a year, and now they’re in it.

GPT-6 Sol is the more awkward launch. $2/$10, Intelligence Index 48, and OpenAI’s own numbers have it beating Opus 5 on AutomationBench (33.2 percent vs 26.9) and edging it on OSWorld at 80 percent lower cost. On Arena it’s 60th Overall at 1457, which misses the Overall band by two points.

The cheapest coding model has a distillation problem

MiMo-V2.6-Flash is the cheapest model inside the Arena Coding band this week. It’s ranked 24th at 1525, 26 points behind the leader, for 28 cents out. That’s 89 times cheaper than a $25 leader. Artificial Analysis puts it on its intelligence-vs-cost Pareto frontier. On OpenRouter it went from nothing to the seventh most-used model on the platform in a week, 6.93 trillion tokens. By every signal I track, it’s a real value pick.

And on September 10, eleven days before it shipped, Anthropic published a threat intelligence report accusing seven China-based labs of what it calls “illicit distillation,” using Claude’s outputs to train their own models. Xiaomi was one of them. According to the-decoder’s summary of case GTG-16008, Xiaomi forwarded more than 400,000 conversations from users of its own MiMo chatbot to Claude between March and April, through OpenClaw and OpenCode, to pull out training data. Across all seven labs, Anthropic counts around 190 million exchanges. Xiaomi hasn’t responded.

I run OpenClaw. It’s the agent framework behind the assistant I run at home, the one that handles my scheduled jobs and writes into my notes vault. So I read that paragraph twice. To be clear, this is an accusation from a competitor, not a court finding, and nothing in it suggests OpenClaw itself did anything wrong. It’s a tool. Somebody pointed it at Claude at industrial scale. But if you were one of those 400,000 MiMo chatbot users, your conversations apparently took a trip you never agreed to.

So what do you do with the pick? I’m printing it, because the method is the method and the price and rating are real. I’m also printing it with two asterisks. First, it only has 1,043 Coding votes, so it’s the emerging pick, not the settled one. GLM-5.3-Flash is right behind it at 1523 with five times the votes for 50 cents, and that’s the steadier bet. Second, if where a model’s training data came from matters to your company, and for some of you it contractually does, this one has a question hanging over it.

And watch the price you see. On OpenRouter, MiMo-V2.6-Flash’s headline price shows 8 cents in and $1.28 out. That’s one small host’s endpoint. Xiaomi’s own endpoint is 14 cents and 28 cents, and it’s carrying 97 percent of the traffic. If you read the headline number, you’ll think it’s the most expensive cheap model on the board. Check the provider list.

The cheapskate picks

Same method every week. Take the category leader’s Arena rating, draw a band 50 points below it, and find the cheapest model still inside the band. The top of Arena is compressed, so the leader is usually only a little better than something 40 to 100 times cheaper. The band gets computed in code from the full table, never eyeballed off the first screen, because eyeballing is how you delete the whole cheap tail without noticing. I learned that one in public.

The good news: Arena finally refreshed. The last two issues ran on the same September 13 snapshot. This week’s data is stamped September 25. Bands ran deep: 57 models in Overall, 73 in Coding, 43 in Hard Prompts and Math, 38 in Instruction Following, and just 12 in Creative Writing. GLM-5.3-Flash is quoted at Z.ai’s list price, $0.15 in and $0.50 out. Arena’s price column now shows 20 cents out for it, but that’s a third-party floor, not list. Budget on fifty.

CategoryLeader$ outCheapskate pick$ outΔ ratingCheaper by
Overallclaude-opus-5.5-high (1509)$20GLM-5.3-Flash (1474, #35)$0.50−3540×
Codingclaude-opus-4-6-high (1551)$25MiMo-V2.6-Flash* (1525, #24)$0.28−2689×
Creative Writingclaude-opus-5.5-high (1521)$20gemini-3.7-flash-high (1492, #4)$3.75−295.3×
Instruction Followingclaude-opus-5.5-high (1516)$20GLM-5.3-Flash (1474, #26)$0.50−4240×
Hard Promptsclaude-opus-5.5-high (1541)$20GLM-5.3-Flash (1499, #31)$0.50−4240×
Mathclaude-fable-5-high (1523)$50GLM-5.3-Flash (1501, #13)$0.50−22100×

*Thin votes (1,043). The steadier Coding pick is GLM-5.3-Flash at #28, 1523, 50 cents, 50× cheaper.

GLM-5.3-Flash held four of six, and it’s doing it on real vote counts now: 19,103 in Overall and 12,503 in Hard Prompts. That’s a pick you can lean on. The “cheaper by” numbers shrank, from 100× to 40× in Overall and from 50× to 40× in Instruction Following and Hard Prompts, but not because GLM got pricier. The leader got cheaper. Opus 5.5 at $20 took the top of three boards from $25 and $50 models. That’s the price war showing up in my own table.

Coding is the first crack in GLM’s run in a month, and it’s a crack on thin votes. Math is still the row to squint at: GLM sits 13th on only 920 votes. If you want steadier, MiMo v2.5 Pro has 3,280 votes at 87 cents.

Creative Writing is where the method bit back. Gemini 3 Flash was the Creative pick for four straight issues at $3. It didn’t get worse this week. Opus 5.5 raised the ceiling by 17 points, and Gemini 3 Flash’s 1457 fell out of a band that now starts at 1471. When the leader gets better, cheap models fall out of the band without doing anything wrong. The new pick is Gemini 3.7 Flash at $3.75, ranked fourth on preliminary votes. Also, Gemini 3.8 Flash’s intro pricing doubles on January 1, so check whether 3.7 goes the same way before you build a budget around it.

The speed caveat is the same as every week, and the number moved again. Artificial Analysis clocks GLM-5.3-Flash at 48 tokens a second this week, down from 89 last week. MiMo-V2.6-Flash is 55. I’ve stopped carrying this number forward because it bounces around too much. Both are slow enough to notice in an agent loop. GPT-6 Luna at 148 is the fast one at this price.

The free model at number two won’t say who made it

Space Bunny Alpha showed up on OpenRouter on September 23 as a stealth model: no maker named, free, a million-token context, and text, image, and video input. Five days later it was the second-biggest model on the platform for the week at 18.2 trillion tokens, and number one for the most recent day.

It’s probably MiniMax. Tokenizer tests on launch day matched MiniMax’s M3 family on every string people threw at it (one tester ran 50 strings, all 50 matched), and on September 27 MiniMax released M3.1-Flash-Preview with the same context length, the same five reasoning effort levels, and the same fast-coding pitch. Nobody has officially confirmed it.

Two things worth knowing before you point production at it. It’s free, and free usage isn’t chosen usage. Free models in this roundup have a habit of spiking and then settling once the meter turns on. And the listing says prompts and completions “may be retained by the provider,” a provider that won’t tell you its name. The same week a distillation report dropped. Fine for kicking the tires on public code. I wouldn’t send it anything I’d mind reading in somebody’s training set.

Meanwhile, on the paid board, DeepSeek V4.1 Flash is number one at 20.8 trillion tokens, up 23 percent. GLM-5.3-Flash slipped to third at 13.3 trillion, down 21 percent, though it’s still number one for the trailing month at 57.5 trillion. Usage is moving toward whatever is newest, cheapest, or free, which is what usage always does in the week after a launch pile-up.

Horror story: 48,218 files in 103 seconds

Around September 20, a developer posted on r/ClaudeAI that Claude Code had deleted their project. According to TechRadar’s writeup, the agent was told to rebuild a mirror of the project for a task. It figured out build_mirror.py couldn’t refresh the mirror in place, so it wrote a little Python remover to delete an old copy sitting in a temp folder. That old copy held 7,332 ordinary files and 614 Windows directory junctions pointing back into the live project. The remover used os.walk with followlinks=False, which sounds safe. But os.path.islink() returns false for Windows junctions, so Python didn’t treat them as links, walked right through them, and started deleting the real project on the other side.

It took 103 seconds. It deleted 48,218 live files, and it emptied the Git object store too: .git/objects, refs, and logs. The index still listed 7,221 paths, but the actual file contents were gone, so Git couldn’t restore anything. TechRadar’s headline quotes the agent’s own summary: “I broke something.” (This all comes from the Reddit post and an attached report, not an independent forensic investigation, so treat the details as the poster’s account.)

Reddit’s verdict was harsh. It wasn’t wrong. The poster admitted they weren’t using GitHub properly and should have been working on a branch. But the reason I’m including it isn’t to dunk on someone. It’s that the agent did something that looks completely reasonable (clean up a stale temp copy) on a file system that had a trap in it that no human would have spotted in the moment either. Directory junctions and symlinks turn “delete this folder” into “delete whatever this folder points at,” and on Windows the standard Python safety check doesn’t even see the junctions. If you let an agent run rm or its Python equivalent, the only real protection is a remote it can’t touch, so push before you let it clean up anything.

What’s coming

Claude Haiku 5.5 is next. Anthropic said Sonnet and Haiku 5.5 would follow Opus “in the coming weeks.” Sonnet already shipped. Haiku is the one left, and if it’s priced like Haiku usually is, it lands right in the cheapskate zone.

MiniMax M3.1 should get official soon. If Space Bunny gets unmasked and priced, we’ll see how many of those 18 trillion free tokens stick around once there’s a bill.

Then there’s MiMo-V2.6-Pro-UltraSpeed, which Xiaomi says is up to 20 times faster than Pro at the same quality. It already has 73 billion tokens on OpenRouter. If that holds, “cheap but slow” stops being the standard caveat on the Chinese value picks.

DeepSeek V5 is speculation only. People keep floating October. DeepSeek hasn’t said anything, and I’m not putting a date on a model that doesn’t have one.

Finally

The price war is real, and it’s good for you. Opus got 20 percent cheaper and better at the same time. OpenAI cut Sol and Luna in half and made it permanent. The floor got a new 28-cent option. If you’re paying your own bill, almost every line on your invoice should drop next month. Who had “two frontier labs cut prices on the same day” on their bingo card?

But three things I’ve been harping on all year showed up again, just in cheaper clothes. The sticker isn’t the task: the budget Claude costs more than the premium Claude when you run it hot. The free model isn’t free: someone is paying for those 18 trillion tokens, and they’re not telling you who or why. And the cheapest option comes with questions about where it came from, which some of you will care about and some of you won’t.

Last week I said I’d come back to see whether the Arena crowd was any nicer to Grok 4.7 than its creator was. It wasn’t. Grok 4.7 debuted at 92nd Overall. Its creator called it mid. The crowd called it worse. Low bar, and it still found a way under it.

Stephan Miller

Written by AI, edited by

Kansas City Software Engineer and Author

Twitter | Github | LinkedIn

Updated