Kimi K3 and Inkling: The Week the Open Models Won
Last week I signed off by saying I’d see you when Gemini 3.5 Pro and DeepSeek V4 had presumably set something on fire. Well. Gemini 3.5 Pro missed its launch date again, for the third time. DeepSeek V4 is quietly shipping with a deadline attached that’s going to ruin somebody’s Friday. And while the three biggest AI labs on the planet were busy being late, two models you can just download showed up out of nowhere. One of them is now the fourth-smartest model on Earth, ahead of Claude Opus 4.8, and it came from Moonshot, a lab most people couldn’t have picked out of a lineup a year ago.
That’s the whole week. The frontier moved, and it didn’t move at OpenAI, Google, or Anthropic. It moved open.
- The Open Frontier Crashed the Party
- Kimi K3 Is Genuinely Smart and It Never Stops Thinking
- Inkling: Mira Murati Gives It Away
- Meanwhile, the West Was Late Again
- The Cheapskate Picks
- Horror Stories From the Wild
- What’s Coming
- The Takeaway
The Open Frontier Crashed the Party
Here’s the setup you need. For most of this year, the “open weights are catching up” story has been a Chinese story. DeepSeek, Qwen, GLM, MiMo. Good models you could download, consistently a few points behind the closed Western flagships, consistently way cheaper. The pattern was reliable enough to be boring.
This week two things happened in a 48-hour window that broke the pattern.
On July 15th, Thinking Machines Lab shipped Inkling, a 975-billion-parameter open-weights model, under an Apache 2.0 license, on Hugging Face, right now, today. On July 16th, Moonshot AI shipped Kimi K3, a 2.8-trillion-parameter monster that lands at number four on the Artificial Analysis Intelligence Index. Two frontier-class open (or opening) models from two labs that are not the big three, dropped back to back.
On the Intelligence Index, Kimi K3 lands at 57.1, behind only Claude Fable 5 at 59.9 and GPT-5.6 Sol at 58.9, and ahead of Claude Opus 4.8 at 55.7. Read that again. An open-weight model from a Chinese lab is beating Anthropic’s own Opus 4.8 on raw intelligence. That is not “catching up.” That’s arrived.
Kimi K3 Is Genuinely Smart and It Never Stops Thinking
Let me tell you what makes Kimi K3 the headline and then tell you the part that’ll cost you money.
The good part first. K3 is a 2.8-trillion-parameter mixture-of-experts model, which is the largest open-weight model anybody has announced. It only activates 16 of its 896 experts per token, so it’s big but not insane to run. It’s got a 1-million-token context window and it’s multimodal. The full weights are coming to Hugging Face by July 27th under a modified MIT license that lets you use it commercially.
But the number that matters is Arena. Normally when a model launches, the Arena leaderboard takes a week or two to catch up, because Arena is vote-based and nobody’s voted on the new thing yet. New models are supposed to be invisible on Arena for a bit. Kimi K3 did not get that memo. Three days after launch it’s already in the top 10 of every category I track: Overall #8, Coding #9, Creative Writing #9, Instruction Following #10, Hard Prompts #10. On Arena data dated July 19th. Three days. That doesn’t happen unless a lot of people tried it and a lot of people liked what they saw.
On cost-per-task, Artificial Analysis clocked it at about $0.94 to run one Intelligence Index task. That’s roughly half of Opus 4.8’s $1.80 for a model that scores higher on intelligence. On paper, that’s a steal.
Now the part that’ll cost you money.
K3 always thinks. There is no non-thinking variant. The reasoning_effort is locked to maximum, and every one of those reasoning tokens bills as output at $15 per million. Simon Willison ran his usual “generate an SVG of a pelican on a bicycle” test on launch day and the model burned 13,241 reasoning tokens before it produced a 3,417-token answer. His verdict on the pricing was blunt: “This is expensive. The pelican cost 25 cents.” One pelican. A quarter.
So the cost-per-task number is real, but it’s an average, and the tail on that average is fat. The visible answer you get back is a fraction of what you actually pay for, because you’re renting the model’s internal monologue at output rates whether you want it or not. If you were burned by the whole cost-per-task lesson last week (Grok 4.5 winning on tokens-per-job while sitting fourth on intelligence), K3 is the same lesson wearing different clothes. The sticker says $3/$15. Your invoice will say something with more zeroes.
Inkling: Mira Murati Gives It Away
The other open drop is the one that made me sit up, and not because of the benchmarks.
Inkling comes from Thinking Machines Lab, which is Mira Murati’s outfit. If the name doesn’t ring a bell, she was OpenAI’s CTO. So this is a former top executive of the most closed, most commercial AI lab in the world shipping a 975-billion-parameter model on Hugging Face under Apache 2.0, framed explicitly around low cost and, in their words, “resistance to censorship.” You can legally fine-tune it, ship it in a product, and never pay Thinking Machines a cent.
The specs are interesting. 975B total parameters, only 41B active per token, multimodal across text, images, and audio. It’s got a thinking-effort dial you can turn from 0.2 to 0.99, which is exactly the knob Kimi K3 refuses to give you. Want cheap and fast? Turn it down. Want it to grind? Turn it up. That’s the right design, and Moonshot should take notes.
On the composite Intelligence Index, Inkling only scores 41, which is well down the board, so don’t expect it to beat Fable 5 in a general chat. But dig into the specific benchmarks and it’s a different story: 77.6% on SWE-bench Verified, 97.1% on AIME 2026 math, 87.2% on GPQA Diamond, and 74.1% on MCP Atlas for agentic workflows. Sebastian Raschka called it the best open multimodal generalist out there right now, and the category-level scores back that up. It’s not a great chatbot. It might be a great tool.
Here’s why these two drops matter together. The “own your weights” argument used to be a China thing, and if you had opinions about running Chinese models in production, you could opt out of the whole conversation. You can’t anymore. When Mira Murati is handing out Apache-2.0 weights and Moonshot is shipping the biggest open model ever built, the choice in front of you isn’t “American closed model or Chinese open model.” It’s “rent from three labs, or own from everybody else.” That’s a genuinely different question than it was a month ago.
Meanwhile, the West Was Late Again
While all that was happening, the closed Western frontier did its now-familiar thing: it announced a date and then missed it.
Gemini 3.5 Pro was targeting July 17th. That was already a slip from June, which was itself a slip from Google I/O in May. Google reportedly scrapped the base model entirely and rebuilt it after engineers found structural failures in recursive tool-calling and SVG generation. Ambitious. Also, the 17th came and went and it’s still limited to a handful of enterprise preview customers. No public API, no confirmed pricing, no confirmed 2-million-token context window, none of the rumored Deep Think reasoning layer you can actually touch. Third promised date, third miss. At this point I’m not writing another word about Gemini 3.5 Pro until there’s an endpoint I can hit with a real API key. It’s vapor until it isn’t.
DeepSeek V4, on the other hand, is very real, and it comes with homework. The V4 family is graduating from preview to stable, and if you use DeepSeek’s hosted API, you have a hard deadline: July 24th at 15:59 UTC. After that, the old model names deepseek-chat and deepseek-reasoner stop working. They start returning HTTP 404 and 400 errors. More on that in the horror section, because it’s nastier than it looks.
The Cheapskate Picks
Okay. The part you can actually use.
Same method as always. For each Arena category, I take the leader’s rating and find the cheapest model within 50 rating points of it. Arena’s top is compressed, so “cheapest in the band” is a real choice between models that are genuinely close, not settling for junk.
And this week I have to own something, because I got it wrong the first time I ran these numbers. “Within 50 points” is a band defined by points, not by rank, and I’d been reading down the first screen of the leaderboard and calling that the band. It is not the band. In Coding, the leader is Opus 4.7-thinking at 1553, which puts the cutoff at 1503 — and the leaderboard doesn’t drop below 1503 until rank 46. Forty-six models are inside that window. I was sorting the first twenty and declaring a winner, which is exactly the mistake this section exists to stop you from making. The bands run 44 deep in Overall, 39 in Hard Prompts, 27 in Instruction Following. Only Creative (14) and Math (6) actually fit on one screen.
Fix the band, and the answer changes in four of six categories. Arena data below is dated July 19th. Fable 5 leaders run about $50 per million output tokens; Opus-thinking leaders about $25.
| Category | Leader | $ out | Cheapskate pick | $ out | Δ rating | Cheaper by |
|---|---|---|---|---|---|---|
| Overall | Fable 5 (1507) | $50 | MiMo v2.5 Pro (1466, #33) | $0.87 | −41 | ~57x |
| Coding | Opus 4.7-thinking (1553) | ~$25 | MiMo v2.5 Pro (1519, #23) | $0.87 | −34 | ~29x |
| Creative | Fable 5 (1513) | $50 | Gemini 3.5 Flash (1467, #11) | $9 | −46 | ~5.6x |
| Instruction Following | Fable 5 (1513) | $50 | MiMo v2.5 Pro (1470, #19) | $0.87 | −43 | ~57x |
| Hard Prompts | Fable 5 (1533) | $50 | MiMo v2.5 Pro (1494, #22) | $0.87 | −39 | ~57x |
| Math | Fable 5 (1550) | $50 | Grok 4.5 (1504, #5) | $6 | −46 | ~8.3x |
So the value story of the week is Xiaomi’s MiMo v2.5 Pro at $0.43 in / $0.87 out, sweeping four of six categories. Not narrowly, either. On Coding it sits 34 points off the best coding model in the world at a twenty-ninth of the output price, and it’s been sitting there with 11,355 votes behind that rating — this is not a preliminary number on a model that dropped Tuesday. It shipped in April. It’s MIT-licensed open weights. In a week whose headline is “the open models won,” the cheapest competitive model on four separate leaderboards turning out to be a phone company’s MIT-licensed side project isn’t a coincidence. It’s the same story from a different angle.
Two honest caveats, because a 57x price gap deserves scrutiny.
First, Artificial Analysis does not rank MiMo v2.5 Pro near the top — it scores 42 on the Intelligence Index, against Kimi K3’s 57.1. That’s the classic Arena-loves-it/AA-doesn’t split, and it means what it usually means: people prefer MiMo’s answers head-to-head, but it isn’t doing frontier-grade reasoning on the hard stuff. For everyday coding, instruction-following, and general work, preference is the metric that matches how you’ll actually use it. For genuinely hard problems, pay up.
Second, it’s slow. 58.2 output tokens per second, below median for its class. If you’re running an agent loop where latency compounds across hundreds of calls, that 57x price advantage buys you a wall-clock penalty you should measure before you commit.
Two runners-up worth knowing. If your workload is input-heavy or long-context, Qwen3.7 Plus is cheaper on the input side at $0.32 in / $1.28 out and sits inside the Coding band at #36. And last week’s champion, Meta’s Muse Spark 1.1 at $1.25 / $4.25, is still competitive on rating — but it’s now about 5x more expensive than MiMo and it’s geo-locked to US developers on OpenRouter, so it falls out of the recommendation twice over. More on that below.
Kimi K3 is right there in the Coding band too at #9, but at $15 output with that always-on thinking tax, it is not the cheapskate answer. It’s the “I want the best open model and I’ll pay for it” answer.
Math is the fun one this week, and it’s the one category where the band really is tiny — six models, top to bottom. The pure cheapskate pick is Grok 4.5 at $6 output, sitting at #5, cheapest thing in the band. But look one row up: Gemini 3.5 Flash is at #3 on Math, 1518, only 32 points behind Fable 5, for $9. For the third or fourth week running, Gemini 3.5 Flash is quietly the smart-money answer for math and creative work. If you’re paying Fable 5’s $50 to do math and Flash is sitting 32 points back at a fifth of the price, you’re not buying quality, you’re buying a rating difference smaller than the noise in the measurement.
Also worth flagging: several of these Fable 5 category-leader ratings are sitting on thin vote counts (Math #1 has under 500 votes). Don’t over-trust a fresh #1 with that little data behind it.
Horror Stories From the Wild
Every roundup needs the part where I tell you what’s going to ruin your day. This week there are three, and one of them has a countdown clock.
The DeepSeek migration cliff. July 24th, 15:59 UTC. If you’ve got deepseek-chat or deepseek-reasoner hard-coded anywhere in your codebase, those names get retired and start returning 404 and 400 errors. Here’s the sneaky part: both of those names already route to DeepSeek-V4-Flash under the hood, and have since April. So nothing looks broken today. Your code works fine right up until next Friday afternoon, when it doesn’t. And the fix isn’t a clean find-and-replace, because thinking mode moved from being its own model name to being a request parameter. If you naively swap deepseek-reasoner for deepseek-v4-flash and call it done, you’ll silently drop reasoning on all that traffic and wonder why your outputs got dumber. Go patch this now, not on the 24th.
Kimi K3’s metered brain. I already covered this up top, but it belongs in the horror section on principle. There is no way to turn K3’s thinking off. Every request pays for a full reasoning trace at $15 per million output tokens, whether the question needed it or not. Simon Willison’s pelican cost 25 cents. Scale that across an agent loop making thousands of calls and the “half the cost of Opus” headline evaporates fast. Contrast it with Inkling shipping a thinking-effort dial the same week, and K3’s design choice looks less like a feature and more like a billing strategy.
The value pick you’re locked out of. Meta shipped the cheapest competitive brand-name model of the week and then restricted it to US developers on OpenRouter. Muse Spark 1.1 was last week’s cheapskate champion in three categories, and for most of the planet it’s just a leaderboard entry you get to look at. Cool. Very open. Very connected world we live in. It got dethroned on price this week anyway, which is its own kind of answer: the model you can buy from anywhere is 5x cheaper and ships its weights.
What’s Coming
- Kimi K3 open weights land on Hugging Face by July 27th, modified MIT license. This is the one to actually wait for if you want to own it instead of renting the API.
- Gemini 3.5 Pro, allegedly, eventually. Missed July 17th. I’ll believe it when I can curl it.
- DeepSeek V4 graduates preview to stable, with that mandatory API migration deadline on July 24th. Set a reminder.
- Inkling derivatives. Apache 2.0 weights are already out, so expect fine-tunes and specialized variants to start showing up fast. That’s the whole point of shipping it open.
And the background hum, unchanged and getting louder: Chinese-origin models are now about 46% of all OpenRouter tokens, and as of July 20th they hit a record 58% share among US firms specifically. US labs have fallen from roughly 70% of token volume to 30% year over year. Anthropic’s Claude slid from 29% to 13%. DeepSeek alone is 17.6%, the single largest vendor on the platform, with Qwen right behind. Two more open flagships this week did not slow that trend. They poured gas on it.
The Takeaway
The intelligence race at the very top is still basically a tie: Fable 5, GPT-5.6 Sol, Kimi K3, Opus 4.8, all inside four points. But the interesting thing this week wasn’t who’s smartest. It was who’s ownable.
For the first time, the answer to “I want a near-frontier model I can actually hold in my hands” isn’t a compromise. Kimi K3 is fourth on the whole intelligence board and its weights ship in a week. Inkling is out right now under a license that lets you do whatever you want with it. The former CTO of OpenAI is handing out model weights. Meanwhile the three labs whose whole pitch is “trust us, it’s on our servers” spent the week being late, geo-locking their cheap model, and metering your reasoning tokens.
So here’s what I’m actually doing. I’m patching my DeepSeek model names before Friday, because I’m not getting caught by a 404 on a Friday. I’m keeping Gemini 3.5 Flash on math and creative, because paying 5x more for a rating rounding error is still stupid this week. I’m moving my bulk coding and instruction-following traffic to MiMo v2.5 Pro and eating the latency, because 29x is not a margin you argue with. I’m waiting the seven days for Kimi K3’s weights before I decide whether it’s worth the thinking tax, because owning it changes the math. And I’m downloading Inkling this weekend just to see what a 975-billion-parameter model does on my own hardware, because I can, and a month ago I couldn’t.
Rent everything, own nothing was the mood in June. In July the challengers handed out the keys. Read your bills, and maybe download a frontier model while it’s free. See you next week, when Gemini 3.5 Pro has presumably missed a fourth date.
