Stephan Miller
Qwen3.8-Max Dropped. I'm Still Running the $0.87 Model.

Qwen3.8-Max Dropped. I'm Still Running the $0.87 Model.

This week Alibaba dropped Qwen3.8-Max, a 2.4-trillion-parameter monster, and the stock jumped 7% in Hong Kong. Everybody lost their minds. And the smartest thing a cheapskate can do with that news is nod politely and keep running a Chinese model from April that costs 57 times less than the thing everyone’s actually cheering for.

Let me explain how I got there.

China Dropped a 2.4-Trillion-Parameter Flagship and I Mostly Yawned

Here’s the thing about Qwen3.8-Max. It’s genuinely impressive on paper. 2.4 trillion total parameters, 95 billion active through a sparse mixture-of-experts setup, a 1-million-token context window, and API access that went worldwide on day one. It landed at #5 on the Arena text leaderboard and #2 on the vision board about a day after launch, which makes it the highest-ranked Chinese text model anyone’s ever put up. Alibaba priced it at roughly 24% to 40% of Claude Opus 5. Open weights are supposedly landing next week.

But look at the votes. It’s sitting at #5 on 3,327 Arena votes. Opus 4.6-thinking above it has 67,000. When a model debuts high on a thin vote count, that’s not a verdict, that’s a first impression with good lighting. Arena rewards new-and-polished before the crowd has actually lived with the thing.

And then there’s the benchmark table. Alibaba says Qwen3.8-Max scores 86.1 on OSWorld-Verified, ahead of GPT-5.6 Sol Max at 83.2 and Fable 5 at 85.0. It claims 86.6 on Terminal-Bench 2.1 and 93.0 on PaperBench. Great numbers. Here’s my problem: that table includes self-reported scores for the competitors too, and as of the day before I wrote this, no independent evaluation of the model existed at all. A vendor grading its own homework and its rivals’ homework in the same table isn’t a lie. It’s just not evidence yet. The correct posture is interested skepticism, and I’ll happily eat my words next week when somebody who doesn’t work at Alibaba runs the thing.

The “Smartest Model in the World” Had an Awkward Debut

Claude Opus 5 shipped July 24, and by this week it finally had enough Arena votes to show up properly. Anthropic and Artificial Analysis both told you it’s the most intelligent model on Earth right now. On AA’s Intelligence Index it scores 61, narrowly the top of the board, effectively tied with Fable 5 at 60 and ahead of GPT-5.6 Sol at 59 and Kimi K3 at 57.

So where does it land on Arena Overall, now that humans have actually voted on it blind?

Number 7. Behind Opus 4.6-thinking at #2 and Opus 4.7-thinking at #3. It got beaten by its own grandparents.

I want to be fair here, because this is the interesting part, not a dunk. Opus 5 is #1 on the hard-benchmark axis and it takes the Arena Math crown outright (on thin preliminary votes, but still). The cost-per-task number is the real headline: $2.03 per Intelligence Index task versus Fable 5’s $2.75, about 26% cheaper, at the same $5/$25 sticker as Opus 4.8. That’s a real bargain if you’re doing hard agentic work.

But if you upgraded to Opus 5 on the strength of the word “smartest” alone, blind human raters are quietly telling you that you bought a lateral move on everyday chat. “Best on the benchmarks” and “the answer people prefer when they don’t know which model they’re looking at” are two different axes, and this week they pointed in two different directions at the same lab.

The Boring Answer That Keeps Winning

Now the part I actually care about. If I strip away the launch confetti, what’s the model that gives me the most quality per dollar right now?

Same answer as last cycle. MiMo v2.5 Pro. Xiaomi’s model from April, MIT-licensed, $0.43/$0.87 per million tokens on Arena’s list price and routing even cheaper on OpenRouter at roughly $0.35/$0.70. Listed, purchasable, no geo-lock, multiple providers. Boring. Available. Cheap as dirt.

What makes this week different is that two completely unrelated methodologies pointed at the same model. Arena’s price sort puts MiMo as the cheapest thing inside the competitive band of four separate categories. And Artificial Analysis, which measures hard-benchmark capability instead of crowd preference, puts MiMo on its Intelligence-vs-Cost Pareto frontier. Its “most attractive quadrant.” When the popularity metric and the capability metric independently name the same cheap model, that’s about as strong a buy signal as this newsletter ever gets.

The honest caveat: MiMo scores 42 on AA’s Intelligence Index while the frontier sits at 57 to 61, and it’s slow at around 47 tokens per second. So it’s the crowd favorite and cost-efficient for what it is. But it’s not frontier-grade on hard reasoning, and it won’t win a latency race. For everyday work and agentic loops where you care about the bill, that’s a trade I’ll take every time. For a gnarly proof or a nasty debugging session, spend the money.

That’s the whole hype-versus-value story this week in one line. Qwen3.8-Max is new, loud, and unproven. MiMo is old, quiet, and double-confirmed.

Cheapskate Picks: Where the Actual Money Is

This is the section I write the newsletter for. The method is simple and I’ll say it once so the table makes sense. For each Arena category, take the leader’s rating, draw a line 50 points below it, and that’s the competitive band. Everything inside that band is, statistically, a rounding error away from the “best” model. Then I sort the band by output price and pick the cheapest thing in it. The trick, and the thing I screwed up in a past issue, is that the band is defined by points, not by rank. It runs way deeper than the visible top 20. The Coding band this week is 53 models deep. The cheap open-weight models live down in the 20s, 30s, and 40s, sitting a couple of points below premium brands while costing an order of magnitude less. Truncate at rank 20 and you delete the entire reason this section exists.

So I pulled the full tables and computed the bands in code. Here’s where the money is:

CategoryLeader$ leaderCheapskate pick$ pickΔ ratingPrice ratioAA Pareto
Overallclaude-fable-5$50MiMo v2.5 Pro (#40)$0.87−43~57×
Codingclaude-fable-5$50MiMo v2.5 Pro (#29)$0.87−35~57×
Creative Writingclaude-fable-5$50Gemini 3 Flash (#23)$3−49~16.7×nearby
Instruction Followingclaude-fable-5$50MiMo v2.5 Pro (#25)$0.87−46~57×
Hard Promptsclaude-fable-5$50MiMo v2.5 Pro (#27)$0.87−40~57×
Mathclaude-opus-5-max*$25Gemini 3.6 Flash (#4)$7.50−32~3.3×nearby

* The Math leader is preliminary, sitting on only 231 votes. Treat that whole board as a rumor this week.

A few things worth saying out loud about that table.

MiMo sweeps four of six categories, all at $0.87 output, all roughly 57 times cheaper than the Fable 5 leader. Its ranks look scary (#40 in Overall) until you check the vote counts. That Overall rating is backed by 45,910 votes. It’s a far more settled number than most of the shiny preliminary top-10 entries with a few hundred votes each. Deep rank measures preference, not reliability. Don’t let it spook you.

Creative Writing is the one place MiMo can’t reach, because that category rewards polish and the cheap crowd falls just below the cutoff. The pick there is Gemini 3 Flash at $3 output, and it’s clinging to the very edge of the band at 49 points back. Still 16 times cheaper than the leader.

Math is a mess this week and I’m flagging it hard. The leader is Opus 5 on 231 votes, the band is only 8 models deep, and the cheapest thing in it is Gemini 3.6 Flash at $7.50. There’s no sub-$7 play here. Math is a “you’re paying for quality” category right now, so if you need it, budget for it and check back when the votes settle.

And the throughline underneath all of it: Fable 5 still sweeps 5 of 6 Arena categories as the outright leader. The ceiling hasn’t moved in weeks. What keeps changing is the floor, and the floor keeps getting cheaper.

Horror Stories From the Wild

Two this week. One is a real fire, the other is a smoke alarm.

The fire: the DeepSeek V4 API migration deadline hit on July 24 at 15:59 UTC, and it hit hard. DeepSeek retired the deepseek-chat and deepseek-reasoner model aliases with no grace period and no fallback. Call the old names now and you get an error, full stop. One developer went digging through production logs and found 14,000 calls still hitting deepseek-chat, every single one returning a 404. The fix is a one-line model-name swap, which sounds trivial until you hit the two gotchas: thinking mode moved from the model name into a request parameter, so a naive swap either silently drops your reasoning entirely or quietly turns the cheapest endpoint into a reasoning-token furnace that torches your bill. If you run anything scheduled or agentic against DeepSeek, go read your logs right now. I’ll wait.

The smoke alarm: I already said it above, but it belongs here too. Qwen3.8-Max launched with a benchmark table that scores its competitors for them and zero independent verification behind any of it. That’s not a crash. It’s a “don’t rewire your whole pipeline around a press release” warning. Wait for someone outside Alibaba to run it.

Coming Soon (Or “Soon,” Anyway)

  • Qwen3.8-Max open weights. Announced for roughly the week of August 10, right behind the API launch. If they’re real, the independent evals that follow will be the actual story, not the launch table.
  • Gemini 3.5 Pro. Still vapor. It missed its July 17 target, which is somewhere around the third or fourth slip now, and Google shipped three other Gemini models instead of it while reportedly scrapping and rebuilding the base model over hallucination and reliability problems. At this point I’ll believe it when I can call the API.
  • Kimi K3 community quants. The weights went public July 26 under a modified MIT license, so quantized community builds are showing up. Just remember the model is 2.8 trillion parameters and needs something like 1.4TB of fast memory even at four-bit. “Open weights” and “you can run it” aren’t the same sentence when you’d need 4 to 8 H100s to load the thing.

What I Actually Took Away This Week

The frontier is stuck and the bargain bin is on fire. That’s the real state of things.

Anthropic still owns the top of every leaderboard that matters, and it has for a month. Meanwhile Alibaba, Xiaomi, DeepSeek, and Moonshot are locked in a race to give away nearly-as-good models for pennies, and Chinese labs now make up something like 45% of all the tokens flowing through OpenRouter. The story isn’t “who’s the smartest.” It’s been settled for weeks. The story is that the price of “good enough for almost everything” fell off a cliff and keeps falling.

So here’s my honest advice, which is the same advice as last week and probably next week. Ignore the launch you read about in the news. Open the leaderboards, find the cheapest model inside the band, check that two different metrics agree it’s actually good, and run that. This week that’s MiMo v2.5 Pro at 57 times less than the model everyone’s cheering for.

And go read your DeepSeek logs. Seriously. Right now.

Next Tuesday I’ll be back with the coffee and the two tabs, and I fully expect a different Chinese lab to have dropped a trillion-parameter something-or-other by then. That’s the price you pay for paying attention.

Stephan Miller

Written by

Kansas City Software Engineer and Author

Twitter | Github | LinkedIn

Updated