Stephan Miller
Grok 4.6 Hit the Frontier at Bargain-Bin Prices

Grok 4.6 Hit the Frontier at Bargain-Bin Prices

For about a year now the deal has been simple. You want the frontier, you pay frontier prices. Twenty-five, fifty bucks a million tokens. Or you go cheap and you accept that you’re running something a tier down, good enough for most jobs but not the thing topping the leaderboards. Premium or bargain. Pick a lane.

This week that deal broke from three directions inside about 48 hours. xAI shipped a model that sits at the actual intelligence frontier and charges bargain-bin prices for it. Google shipped a cheap coding workhorse while its actual flagship stayed exactly as imaginary as it was last month. And the open-weights promise I spent all of last week’s post complaining about? It shipped. The weights are real. You can download them this morning.

And through all of it, the single cheapest model on my boards kept quietly beating things that cost 57 times more. Same model I’ve been recommending since April. Let me walk you through it.

Grok 4.6 Walked Up to the Frontier and Sat Down

Grok 4.6 landed on August 12, about five weeks after 4.5, and it was live in Cursor, Grok Build, and the API the same afternoon. No slow rollout, no waitlist. Here it is, go use it.

The number that matters: it jumped roughly 5 points on the Artificial Analysis Intelligence Index to land at 61. That ties GPT-5.6 Sol and sits two points behind Claude Opus 5. That is the frontier. Not “the frontier for the price,” not “punching above its weight.” The actual top cluster, the models everyone pays a premium to touch.

Now the price. Grok 4.6 is $2 in, $6 out per million tokens. Flat, unchanged from 4.5. GPT-5.6 Sol, the model it just tied on intelligence, is $5 in and $30 out. Do that division. For the cost of running one GPT-5.6 Sol job you can run the frontier-tier Grok five times over. Artificial Analysis put it plainly in their own headline: Grok 4.6 returns xAI to the intelligence frontier and leads on cost efficiency. The context window is still 500k, unchanged, which is the one spec that didn’t move.

Here’s the asterisk, because there’s always an asterisk. On AA-Omniscience, the benchmark that measures whether a model knows what it doesn’t know, Grok 4.6 scores 48.2% accuracy and a 65.7% non-hallucination rate. Translate that: when Grok hits something it genuinely doesn’t know, it declines to answer about two times out of three, and it confidently makes something up the other third. A brand-new model at the top of the intelligence index that fabricates on one in three unknowns is not a dealbreaker. But if you’re about to wire it into an agent loop where it can act on its own confident guesses, you want to know that number before, not after.

Arena hasn’t caught up yet, which is normal. Grok 4.6 is only #44 on the Overall board on a few thousand fresh votes, because Arena is a popularity contest and popularity takes weeks to accumulate. So you get the classic split: the hard benchmarks love it, the blind-taste crowd hasn’t voted. It’s smarter than it looks in the booth. Give the votes a couple weeks.

Google Shipped Another Flash While the Flagship Stays a Ghost

On August 13, one day after Grok, Google put out Gemini 3.7 Flash. That’s 23 days after Gemini 3.6 Flash. Three weeks. Google is iterating the cheap tier faster than most labs update their changelog.

And it’s a real update, not a version-number bump. On Google’s own DeepSWE v1.1 coding benchmark it scores 65.3% against 3.6 Flash’s 49.0%. That’s a 16-point jump on real software-engineering tasks. They didn’t retrain from scratch either, just algorithmic improvements and user feedback layered onto the previous version, then shipped as a full replacement. On the Artificial Analysis index it lands at 56, and it’s one of the fastest models AA tracks, around 344 output tokens a second. Introductory pricing is $0.75 in, $3.75 out through the end of the year. It’s live in AI Studio, Android Studio, Antigravity, and the enterprise platform. On Arena it’s already #3 in Creative Writing and #4 in Math, though both of those ride preliminary vote counts in the hundreds, so treat them as a strong first impression under flattering light.

Here’s the part that makes me laugh, though. Axios noticed it too: Gemini 3.7 Flash arrived before Gemini 3.5 Pro. The Pro. The flagship. The one Sundar Pichai promised at I/O back on May 19 would land “within a month.”

That was roughly ninety days ago. Forbes ran a piece on August 13 with the deeply original title “Gemini 3.5 Pro Delay Continues.” People have started calling it the longest-awaited model of 2026. The reporting says the team hit a structural problem, scrapped the base model, rebuilt it, and that Google has already started pretraining an entirely new flagship it’s calling Gemini 4. So the flagship is so broken they’re skipping ahead to its replacement while shipping Flash after Flash after Flash to keep the lights on.

I’m not complaining, exactly. The Flash tier is where the value is and Google’s clearly good at it. But when a company ships three iterations of its budget model in the time its premium model goes from “next month” to “we started building the next one instead,” that tells you where the actual engineering is landing. It’s landing on cheap and fast.

The Broken Promise From Last Week Actually Shipped

If you read last week’s post, you watched me go looking for Qwen3.8-Max’s open weights with my coffee and come back with nothing. Alibaba had promised them “the week of August 10” and the deadline came and went with no license, no files, no new date. I filed it under broken promises and moved on.

Well. On August 12 they shipped. Qwen3.8-2.4T-A95B went up on Hugging Face and ModelScope, 2.4 trillion parameters with 95 billion active, alongside the smaller Qwen3.8-27B that actually fits on a workstation. First time Alibaba has ever put a Max-class model in the public’s hands. About two days late, which in this business is basically on time. So the “China promised and stalled” beat from last week only half held. They stalled, then they delivered.

The independent grade came in at 58 on the intelligence index, which puts it just under the 60-to-63 frontier cluster. Good model, not a record-breaker, and now that the weights are public the real test starts: what do people build with it and how does it hold up under evals that Alibaba didn’t run itself.

And it wasn’t the only open drop. On August 14, Z.ai quietly released GLM-5.3. It’s not on Arena yet so I can’t rank it, but Z.ai’s GLM line has anchored my cheapskate picks for months, so it goes straight on the watch list. When the boards catch up, we’ll see where it lands.

Newest Still Isn’t Best

Quick detour to the top of the boards, because it keeps being true and it keeps mattering. Anthropic owns the leader slot in five of six Arena categories again. But look at which Anthropic model is winning.

On the Overall board the leader is Claude Fable 5. On Instruction Following and Hard Prompts, the leader is Opus 4.6, the high-effort variant. Not Opus 5. The newest, most expensive, top-of-the-intelligence-index flagship sits at #7 and #10 on Overall, behind its own older siblings. Opus 5 only clearly wins the Math board, and it does that on 437 votes, which is thin enough to wobble.

So even inside a single lab, “newest” and “what blind human raters actually prefer” are pointing at different models. The version number went up. The preference didn’t follow. Every time a lab tells you the new one is smarter, remember that smarter on a benchmark and better in your actual work are two different measurements, and the second one is the one you’re paying for.

Cheapskate Picks: Where the Money Actually Is

This is the section I write the whole thing for. Method first, because the table is meaningless without it. For each Arena category, take the leader’s rating, draw a line 50 points below it, and everything above that line is the competitive band. Statistically it’s a coin-flip away from the “best” model. Then sort that band by output price and grab the cheapest thing in it. The mistake I’ve made before is reading the band off the visible top 20. The band is defined by points, not rank, and it runs way deeper than the first screen. The Overall band this week is 57 models deep. The cheap open-weight stuff lives down in the 30s and 50s, a rounding error behind the premium brands and an order of magnitude cheaper. So I pulled the full tables and computed the bands in code.

CategoryLeader$ leaderCheapskate pick$ pickΔ ratingPrice ratioAA Pareto
Overallclaude-fable-5$50MiMo v2.5 Pro (#38)$0.87-38~57×
Codingclaude-fable-5$50MiMo v2.5 Pro (#25)$0.87-34~57×
Creative Writingclaude-fable-5$50Gemini 3.6 Flash (#12)$1.88-33~27×nearby
Instruction Followingclaude-opus-4-6$25MiMo v2.5 Pro (#25)$0.87-44~29×
Hard Promptsclaude-opus-4-6$25MiMo v2.5 Pro (#28)$0.87-37~29×
Mathclaude-opus-5-max$25Gemini 3.6 Flash (#7)$1.88-41~13×nearby

A few things worth saying out loud.

MiMo v2.5 Pro sweeps three of six again, all at $0.87 output. That’s Xiaomi’s model from April, the one I’ve recommended every week for over a month now. This is the fifth cycle running it’s been the cheapest thing inside the competitive band of multiple categories, and it’s still sitting on Artificial Analysis’s Intelligence-vs-Cost Pareto frontier, index score 43, in the most efficient quadrant. The popularity metric and the capability metric keep independently landing on the same cheap model. That’s the strongest buy signal this newsletter produces, and it simply will not change. On OpenRouter it routes across seven providers with a 30% discount live right now, no geo-lock, purchasable everywhere. The trade is speed: it’s slow, around 56 tokens a second, so for a tight agent loop where latency compounds you might pay up. For everything else, it’s the boring correct answer.

Its ranks look scary until you check the vote counts. That #38 Overall rating is backed by more than 50,000 votes. A #28 Hard Prompts slot sits on 33,000. Those are far more settled numbers than the preliminary top-10 entries riding a few hundred votes. Deep rank measures preference, not reliability. Don’t let it spook you.

Creative Writing and Math flipped this week, and the reason is a price cut. Gemini 3.6 Flash, which I quoted at $3.75 last issue, now shows up on Arena at $0.38 in and $1.88 out. That’s a straight halving, almost certainly Google trimming the old tier the moment 3.7 Flash launched on top of it. At $1.88 it’s the cheapest thing in both the Creative band and the Math band, and it beats every Opus except Fable 5 in Creative. Math is the compressed board as always, only 10 models deep, and the whole thing rides thin preliminary votes, so treat that pick as directional rather than gospel.

And there’s a wildcard I have to name because the method demands it. On the Overall board, hy3, which is Tencent’s Hunyuan 3, sits at #54 with a $0.53 output price. That’s cheaper than MiMo. But it’s parked right at the leader-minus-50 line with a rating margin wide enough to slip below the cutoff on any given day, on only 4,600 votes. So it’s the cheaper gamble, not the anchor. If you want to ride the very edge of the band to save another thirty cents a million, it’s sitting right there. I’m keeping MiMo as the pick because I like my recommendations boring and my vote counts high.

Horror Stories From the Wild

Three this week, running from “new model problem” to “your problem” to “everyone’s problem.”

First, the one I already flagged: Grok 4.6’s fabrication rate. A frontier model that makes something up on one in three of the things it doesn’t know is fine for a chat window where you can eyeball the answer. It is not fine bolted into an autonomous loop where its confident guess becomes the next tool call. New model, top of the index, still lies to you a third of the time it’s cornered. Know that going in.

Second, a story that’s becoming the defining nightmare of agentic coding. A frontend team deployed a new agent that hallucinated a missing dependency, then entered a 400-step resolution loop trying to fix a problem that never existed. Every one of those 400 steps re-sent the entire accumulated context. Every step billed. The model invented a problem and then spent your money in a circle trying to solve it.

Third, the cost story, because it’s the one that gets everyone eventually. One developer kicked off an autonomous refactoring run over a long weekend, on a workload the team hadn’t even validated, and came back to a $4,200 API bill. The broader pattern the piece describes: within about 90 days of switching on coding agents, the AI bill becomes the second-largest line item on the engineering ledger, right after salaries. This is exactly why the cheapskate math isn’t a hobby. When a single unattended session can drain four figures, the difference between $0.87 and $50 a million stops being an abstraction and starts being your quarter.

Coming Soon (Or “Soon,” Anyway)

  • Gemini 3.5 Pro. Still vapor, now roughly ninety days past Google’s “within a month.” Reportedly scrapped and rebuilt over reliability and coding failures, with Google already pretraining Gemini 4 in the background. I’ll believe it when the API returns a token.
  • Qwen3.8 open weights. These actually shipped on August 12. The story now isn’t the launch, it’s the independent evals. Watch what the community builds and benchmarks over the next couple weeks.
  • GLM-5.3. Z.ai dropped it August 14. Not on Arena yet. Given how often GLM models have anchored these picks, it’s the one I’m most curious to see ranked.
  • Grok 4.6 in the EU. As usual, xAI’s European rollout is lagging the US launch. If you’re on that side of the Atlantic, expect it later in the month.

What I Actually Took Away This Week

The wall between “frontier” and “cheap” is coming down, and it’s coming down fast.

A year ago the frontier was a walled garden you paid $30 a million to enter. This week a frontier-tier model shipped at $6 output, a cheap coding model got 16 points better in three weeks, and a Max-class model’s weights went public for anyone to download. Meanwhile the flagship everyone waited all summer for is so broken its maker started over, and the $0.87 model from April kept winning categories nobody talks about.

The launches change. The headlines change. The correct move does not. Open the leaderboards, find the cheapest model inside the competitive band, confirm two different metrics agree it’s actually good, and run that. This week that’s still MiMo v2.5 Pro at 57 times less than the thing at the top of the board. And if you’re feeling brave, Grok 4.6 is now a genuine frontier option at a price that doesn’t require a finance meeting.

Next Tuesday, same coffee, same two tabs. Some lab will have promised me something by then. I’ll believe that one when I can download it too.

Stephan Miller

Written by

Kansas City Software Engineer and Author

Twitter | Github | LinkedIn

Updated