Stephan Miller
The Open-Weights Promise Qwen Broke (and Meta Kept)

The Open-Weights Promise Qwen Broke (and Meta Kept)

Last week Alibaba dropped Qwen3.8-Max and promised the open weights were landing “next week.” Well. It’s next week. I went looking for them this morning with my coffee, and there’s nothing on Hugging Face, no license, and no new date. The 2.4-trillion-parameter monster everyone lost their minds over is still API-only, still a black box, still grading its own homework.

Meanwhile, the other side of the world did the exact opposite. Meta shipped a coding model and a terminal agent, then turned around and said it’s going to open the weights. So this week the script flipped. China promised open and didn’t deliver. The US shipped and pledged to open up. And me? I’m still running the same $0.87 model from April that I was running last week, and the week before that, and the week before that.

Let me walk you through it.

The Open-Weights Promise That Evaporated

Here’s where we left Qwen3.8-Max last Tuesday. Genuinely impressive spec sheet. 2.4 trillion parameters, 95 billion active, a million-token context window, priced at roughly a quarter of Claude Opus 5. It rocketed onto the Arena boards on day one. And Alibaba said the weights, plus a smaller Qwen3.8-27B, would go public the week of August 10 on Hugging Face and ModelScope. First open-weight release at Max scale for the Qwen line. Big deal.

That week is now. The weights haven’t appeared, no license has been named, and Alibaba hasn’t given a new date. The 27B is vapor too. No architecture details, no context length, no license, nothing.

The good news for the model, if you’re keeping score, is that it finally got an independent grade. Last week the only benchmarks were Alibaba’s own, in a table that helpfully scored its competitors for them. This week Artificial Analysis actually ran it and landed it at 56 on the Intelligence Index. Which is fine. It’s a good model. But 56 sits below the top cluster of Opus 5 at 61, Fable 5 at 60, GPT-5.6 Sol at 59, and Kimi K3 at 57. So the one hard number we got came in under the launch-day hype, and the open weights that were supposed to let anyone verify the rest didn’t show. If you rewired your stack around that press release, this is your reminder not to.

Meanwhile, Meta Did the Opposite

While Alibaba was quietly not shipping, Meta was loudly shipping. On August 5 it put out Muse Spark 1.2 and a terminal coding agent called Muse Code. The model is a coding specialist with a million-token context, built for the whole plan-execute-validate-fix loop across a big codebase instead of spitting out isolated snippets. It debuted at #4 on Arena Overall and #8 on the Coding board, though both of those ratings are riding preliminary vote counts in the low thousands, so treat them as a first impression with good lighting.

Then, on August 10, Meta announced it’s going to open Muse Spark 1.2’s weights. It already opened a smaller sibling, Muse Glimmer, which is now showing up on the trackers. If the Spark 1.2 release actually lands, it’d be the strongest US open-weight model going right now.

Sit with the reversal for a second, because it’s the real story this week. For a solid year the pattern was: American labs ship closed, Chinese labs give the good stuff away for pennies. This week Alibaba dangled open weights and pulled them back, and Meta is the one queuing up a genuine open release. One missed week isn’t a cancellation, and a pledge isn’t a release, so don’t over-read it. But the roles inverted, and that’s worth noticing.

The Catch in Meta’s Cheap Tier

Now for the part where the cheapskate in me perks up and then immediately gets suspicious. Muse Code has a “contributor” tier. Standard pricing on Muse Spark 1.2 is $1.25 in and $4.25 out per million tokens. The contributor tier is $0.10 in and $0.20 out. That’s roughly 12 times cheaper on input and 21 times cheaper on output. My kind of number.

Except you don’t pay in dollars. You pay in data. Everything you send and everything the model sends back becomes training material for future Meta models. And here’s the part that actually bugs me: Meta hasn’t said whether that data use stops at training, or whether it also covers evaluation, red-teaming, and product analytics. The scope is undefined, which in practice means you assume the worst.

So the honest guidance is the boring guidance. Fine for public code, synthetic code, throwaway experiments, anything you’d have posted to a gist anyway. Absolutely not for client repos, secrets, NDA material, or unreleased product logic. This is the oldest cheapskate trap there is. The sticker price isn’t the real price. Sometimes the discount is the product and you’re the inventory.

The Boring Answer, Week Four

Okay. Strip away the launch confetti and the broken promises. What’s the model that gives me the most quality per dollar right now?

Same answer as last week. And the week before. MiMo v2.5 Pro. Xiaomi’s model from April, $0.43/$0.87 on Arena’s list price, routing across seven providers on OpenRouter with no geo-lock. Listed, purchasable, cheap as dirt. This is now the fourth cycle running where it’s the cheapest thing inside the competitive band of four separate Arena categories, and it’s still sitting on Artificial Analysis’s Intelligence-vs-Cost Pareto frontier. The popularity metric and the capability metric keep independently pointing at the same cheap model. That’s the strongest buy signal this newsletter ever produces, and it just refuses to change.

While I’m here, one thing worth calling out about the top of the boards. Anthropic owns the leader spot in five of six Arena categories this week. But look at which Anthropic model. On Instruction Following and Hard Prompts the leader is Opus 4.6-thinking. Not Opus 5. The newest, “smartest” flagship, the one that tops the hard-benchmark index, only clearly wins the Math board, and it does that on preliminary votes. So even inside one lab, “newest” and “what blind human raters actually prefer” are pointing in different directions. The word “smartest” is doing a lot of unearned work in the marketing.

Cheapskate Picks: Where the Money Actually Is

This is the section I write the whole thing for. The method, one more time, because it’s the only way the table makes sense. For each Arena category, take the leader’s rating, draw a line 50 points below it, and everything above that line is the competitive band. Statistically it’s a rounding error away from the “best” model. Sort that band by output price, take the cheapest thing in it. The trick I’ve screwed up before is that the band is defined by points, not by rank. It runs way deeper than the visible top 20. The Overall band this week is about 53 models deep. The cheap open-weight models live down in the 30s and 40s, a couple of points behind premium brands and an order of magnitude cheaper. Truncate at rank 20 and you delete the entire reason this section exists.

So I pulled the full tables and computed the bands in code. Here’s where the money is.

CategoryLeader$ leaderCheapskate pick$ pickΔ ratingPrice ratioAA Pareto
Overallclaude-fable-5$50MiMo v2.5 Pro (#37)$0.87-38~57×
Codingclaude-fable-5$50MiMo v2.5 Pro (#27)$0.87-35~57×
Creative Writingclaude-fable-5$50Gemini 3 Flash (#21)$3-48~17×nearby
Instruction Followingclaude-opus-4-6-thinking$12.50MiMo v2.5 Pro (#24)$0.87-44~14×
Hard Promptsclaude-opus-4-6-thinking$12.50MiMo v2.5 Pro (#27)$0.87-38~14×
Mathclaude-opus-5-max$25Gemini 3.6 Flash (#5)$3.75-39~6.7×nearby

A few things worth saying out loud.

MiMo sweeps four of six again, all at $0.87 output. Its ranks look scary until you check the vote counts. That #37 Overall rating is backed by 50,000 votes. That’s a far more settled number than most of the preliminary top-10 entries sitting on a few hundred votes each. Deep rank measures preference, not reliability. Don’t let it spook you.

Creative Writing is the one place MiMo can’t reach, same as always, because that board rewards polish and the cheap crowd falls just below the cutoff. The pick there is Gemini 3 Flash at $3 output, clinging to the edge of the band 48 points back, still 17 times cheaper than the leader. Math is the compressed board this week, only about 6 models deep, and the cheapest thing in it is Gemini 3.6 Flash at $3.75. Real discount, just a smaller one, and the leader up top is preliminary anyway.

And there’s a wildcard I have to mention because the method demands it. On the Overall board, a model called hy3, which is Tencent’s Hunyuan 3, sits at rank #52 with an output price of $0.53. That’s cheaper than MiMo. But it’s parked at exactly the leader-minus-50 line, with a rating margin wide enough to dip below the cutoff on any given day. So it’s the cheaper gamble, not the anchor. If you want to ride the very edge of the band to save another thirty cents, it’s there. I’m keeping MiMo as the pick because I like my recommendations boring and my votes plentiful.

Horror Stories From the Wild

Three this week, sliding from “your problem” to “everyone’s problem.”

The first one I already told you: Meta’s contributor tier, where the cheapest coding option on the board is cheap precisely because it’s eating your prompts. If you saw the $0.20 output price and got excited before reading the fine print, that’s the fire. Go check what you piped through it.

The second is Qwen3.8-Max’s disappearing act. Not a crash, just a broken promise with a benchmark table full of self-graded numbers behind it. The lesson is the same one this newsletter keeps preaching. Wait for someone who doesn’t work at the lab to run the model before you build anything on top of it.

The third is the one that should actually scare you, because it’s not about any single model. It’s slopsquatting. Roughly one in five package dependencies that AI coding assistants suggest simply don’t exist. Attackers know this. They watch for the common hallucinated package names and register real malware under those exact names. So your agent confidently invents an import, you run the install without blinking, and now you’ve pulled a payload onto your machine. The fix is unglamorous and non-negotiable: read your dependencies before you install them. Every time. Yes, even the ones that look obviously real.

Coming Soon (Or “Soon,” Anyway)

  • Gemini 3.5 Pro. Still vapor. It was widely expected on August 7, then got pushed again on “deployment issues”. That’s roughly 67 days past Google’s “within a month” promise from I/O, after the team reportedly scrapped and rebuilt the base model over coding and reliability failures. Google keeps shipping Flash tiers instead. I’ll believe it when I can call the API.
  • Qwen3.8-Max open weights. Overdue as of this writing, no license, no date. If they land, the independent evals will be the story, not the launch table.
  • Meta Muse Spark 1.2 open weights. Announced August 10. A pledge, not a release, but if it ships it’s the best US open-weight model out there.
  • Ling 3.0 Flash. This one actually happened. Ant Group’s inclusionAI open-sourced it August 5 under MIT, 124 billion parameters with 5.1 billion active, a 262K context. In a week defined by open weights that didn’t ship, a small Chinese lab quietly shipped some. Worth a look if you’re running your own hardware.

What I Actually Took Away This Week

The frontier is boring and the promises are noise.

Anthropic still owns the top of nearly every board and has for over a month. The interesting action is all down in the bargain bin, where a Chinese lab dangled open weights and yanked them, an American lab pledged to open its own, and a $0.87 model from April kept quietly winning four categories while nobody talked about it. The launches change. The headlines change. The correct move does not.

Same advice as last week, and probably next week. Ignore the launch you read about in the news. Open the leaderboards, find the cheapest model inside the band, confirm two different metrics agree it’s actually good, and run that. This week that’s still MiMo v2.5 Pro at 57 times less than the thing everyone’s cheering for.

And read your dependencies before you install them. I mean it. The slop is coming from inside the house now.

Next Tuesday, same coffee, same two tabs. I fully expect another lab to have promised me something by then. I’ll believe that one when I can download it too.

Stephan Miller

Written by

Kansas City Software Engineer and Author

Twitter | Github | LinkedIn

Updated