<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>Stephan Miller</title>
    <description>Kansas City Software Engineer and Writer</description>
    <link>https://www.stephanmiller.com/</link>
    <atom:link href="https://www.stephanmiller.com/feed.xml" rel="self" type="application/rss+xml"/>
    <pubDate>Sat, 10 Oct 2026 22:17:08 -0500</pubDate>
    <lastBuildDate>Sat, 10 Oct 2026 22:17:08 -0500</lastBuildDate>
    <generator>Jekyll v4.2.2</generator>
    
      <item>
        <title>Gemini 4 Argon Won Every Arena Category. You Can&apos;t Use It.</title>
        <description>&lt;p&gt;The best model in the world this week is one you can’t use.&lt;/p&gt;

&lt;p&gt;Gemini 4 Argon showed up on Arena and took first place in all six text categories I track: Overall, Coding, Creative Writing, Instruction Following, Hard Prompts, and Math. It’s listed at $2 in and $10 out, which is cheaper than every Claude model it knocked off the top. So I did what any reasonable person with an OpenRouter account would do. I went to try it. 404.&lt;/p&gt;

&lt;p&gt;That’s because Google isn’t selling it to you. Argon is locked to a few hundred cybersecurity organizations (650-plus, if you’re counting) while Google figures out whether it’s safe to let the rest of us have it. Same week, OpenAI went one step further and refused to ship its own next flagship, GPT-6.1 Astra, because in testing it lied about what it had done. So the two biggest frontier stories this week are both about models you can’t have. And the one you can’t have still managed to rearrange my cheapskate table just by existing.&lt;/p&gt;

&lt;ul id=&quot;markdown-toc&quot;&gt;
  &lt;li&gt;&lt;a href=&quot;#gemini-4-argon-won-everything-on-preliminary-votes&quot; id=&quot;markdown-toc-gemini-4-argon-won-everything-on-preliminary-votes&quot;&gt;Gemini 4 Argon won everything on preliminary votes&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;#the-astra-that-lied-and-the-sol-you-can-buy-instead&quot; id=&quot;markdown-toc-the-astra-that-lied-and-the-sol-you-can-buy-instead&quot;&gt;The Astra that lied, and the Sol you can buy instead&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;#a-model-you-cant-call-just-rearranged-my-cheapskate-table&quot; id=&quot;markdown-toc-a-model-you-cant-call-just-rearranged-my-cheapskate-table&quot;&gt;A model you can’t call just rearranged my cheapskate table&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;#space-bunny-is-still-free-still-anonymous-and-now-number-one&quot; id=&quot;markdown-toc-space-bunny-is-still-free-still-anonymous-and-now-number-one&quot;&gt;Space Bunny is still free, still anonymous, and now number one&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;#horror-stories&quot; id=&quot;markdown-toc-horror-stories&quot;&gt;Horror stories&lt;/a&gt;    &lt;ul&gt;
      &lt;li&gt;&lt;a href=&quot;#your-agents-status-report-might-be-fiction&quot; id=&quot;markdown-toc-your-agents-status-report-might-be-fiction&quot;&gt;Your agent’s status report might be fiction&lt;/a&gt;&lt;/li&gt;
      &lt;li&gt;&lt;a href=&quot;#four-trillion-prompt-tokens-to-a-stranger&quot; id=&quot;markdown-toc-four-trillion-prompt-tokens-to-a-stranger&quot;&gt;Four trillion prompt tokens to a stranger&lt;/a&gt;&lt;/li&gt;
      &lt;li&gt;&lt;a href=&quot;#budgeting-off-a-leaderboard-number&quot; id=&quot;markdown-toc-budgeting-off-a-leaderboard-number&quot;&gt;Budgeting off a leaderboard number&lt;/a&gt;&lt;/li&gt;
    &lt;/ul&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;#whats-coming&quot; id=&quot;markdown-toc-whats-coming&quot;&gt;What’s coming&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;#finally&quot; id=&quot;markdown-toc-finally&quot;&gt;Finally&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;h2 id=&quot;gemini-4-argon-won-everything-on-preliminary-votes&quot;&gt;Gemini 4 Argon won everything on preliminary votes&lt;/h2&gt;

&lt;p&gt;Arena’s board, stamped October 2, has Argon at 1525 Overall, 20 points clear of claude-opus-4-6-high at 1505. It leads Coding at 1560, Creative Writing at 1519, Instruction Following at 1528, Hard Prompts at 1551, and Math at 1530.&lt;/p&gt;

&lt;p&gt;Every one of those ratings carries the “Preliminary” tag. It has 4,932 votes in Overall, about 1,100 to 1,250 each in Coding and Creative Writing, and just 255 in Math. For comparison, the Opus 4.6 it displaced has 77,636 Overall votes. A model can win a category on 255 votes. Opus 4.6 has about 300 times that many votes in Overall alone.&lt;/p&gt;

&lt;p&gt;Artificial Analysis is less excited. On its &lt;a href=&quot;https://artificialanalysis.ai/models/gemini-4-argon&quot;&gt;Intelligence Index&lt;/a&gt;, Argon scores 53. That puts it eighth, tied with Fable 5.1 and GPT-6 Astra, behind Claude Opus 5.5 at 58 and Claude Sonnet 5.5 at 56. That’s frontier, no question. It’s just not the top of the frontier. The blind-vote crowd loves it and the hard benchmarks say it’s good. I’ve seen this split before, usually in the opposite direction. GPT-6 Astra tied for first on AA and debuted 24th on Arena. Argon is the mirror image.&lt;/p&gt;

&lt;p&gt;Why can’t you buy it? Google announced Argon last week and &lt;a href=&quot;https://techwireasia.com/2026/10/google-gemini-4-argon-cybersecurity-access/&quot;&gt;rolled it out first through its Fairwind Program&lt;/a&gt;: more than 650 partners, including government agencies, critical infrastructure operators, and security vendors. Those organizations can only hand it to employees in cybersecurity, incident response, or pentesting roles, and they have to use multi-factor authentication. Google’s stated reason is that it’s still testing safeguards against cyberattacks, CBRN misuse, and indirect prompt injection. The self-reported numbers explain the nerves: 68 percent on CWE-bench v1 for vulnerability remediation, 77.9 percent on DeepSWE v1.1, and a story about finding a patient-data bug in hospital software that earlier frontier models missed. Google says it trained Argon specifically to find, validate, and patch software vulnerabilities. Nobody independent has checked any of that yet.&lt;/p&gt;

&lt;p&gt;Next in line are paid API customers and Google AI Ultra subscribers. No date. And the price on Arena’s board is &lt;a href=&quot;https://smartscope.blog/en/blog/gemini-4-argon-access-introductory-pricing-2026/&quot;&gt;introductory&lt;/a&gt;. After the intro period it goes to $4 in and $20 out, and Google hasn’t said when the intro period ends. So the leaderboard’s number one is off limits, priced at a number that won’t last, ranked on votes that are still coming in.&lt;/p&gt;

&lt;h2 id=&quot;the-astra-that-lied-and-the-sol-you-can-buy-instead&quot;&gt;The Astra that lied, and the Sol you can buy instead&lt;/h2&gt;

&lt;p&gt;OpenAI had its own flagship lined up for October. It’s not coming.&lt;/p&gt;

&lt;p&gt;According to &lt;a href=&quot;https://thehackernews.com/2026/09/openai-shelves-gpt-61-astra-after-tests.html&quot;&gt;reporting picked up by The Hacker News&lt;/a&gt;, GPT-6.1 Astra failed OpenAI’s internal safety and alignment audits. The details read like the postmortem nobody wants to write about an agent. It didn’t disclose actions it had taken, and it went ahead without asking permission. Testers also caught it reaching for outside tools in situations flagged as unsafe.&lt;/p&gt;

&lt;p&gt;Saachi Jain, OpenAI’s head of safety systems, put it this way: the model “didn’t quite meet the bar in terms of staying within scope and authorization, and how it communicates back to the user about the type of work it’s done.”&lt;/p&gt;

&lt;p&gt;In plain English, it did stuff you didn’t ask for and then lied to you about it. If you run coding agents for a living, that’s the failure mode that matters most. A dumb model wastes your afternoon. A model that lies about what it did is how you lose a database.&lt;/p&gt;

&lt;p&gt;The model OpenAI does sell got a report card too. The same day, the UK’s AI Security Institute published a report on GPT-6 Astra, the one OpenAI has been selling since early September. In simulated testing, Astra “conducted a range of unsanctioned attack activities” at a higher rate than GPT-5.6 Sol and GPT-5.5, sometimes even after the scope was spelled out for it. The report’s examples: “creating fake identities which it used to deceive developers, posting comments from fake accounts arguing against the results of accurate security reviews, and delivering malicious payloads to open-source codebases.” That’s a simulation, not an incident. But plenty of people have that exact model wired into a coding agent right now.&lt;/p&gt;

&lt;p&gt;So at DevDay on September 29, OpenAI shipped &lt;a href=&quot;https://techcrunch.com/2026/09/29/openai-launches-gpt-6-1-sol-says-it-nearly-matches-gpt-6-astra-and-costs-less/&quot;&gt;GPT-6.1 Sol&lt;/a&gt; instead. OpenAI says it gets close to GPT-6 Astra on agentic coding, computer use, and professional work at a fifth of Astra’s price. On &lt;a href=&quot;https://openrouter.ai/openai/gpt-6.1-sol&quot;&gt;OpenRouter&lt;/a&gt; that works out to $2 in, $10 out, and 10 cents per million for cached input. It’s the number one new model on OpenRouter’s trending list this week at 927 billion tokens. Artificial Analysis gives it 52, one point behind Argon. On Arena it’s 21st Overall at 1483 and eighth in Coding on only 609 votes.&lt;/p&gt;

&lt;p&gt;GPT-6.1 Sol is $2/$10, and so is Gemini 4 Argon’s intro price. Two labs landed on the same sticker for near-frontier capability the same week, and only one of them will actually sell it to you. OpenRouter’s provider table also shows Sol is slow right now, 32 to 50 tokens a second depending on the host. There’s a “Fast” endpoint at double the price if that matters to you.&lt;/p&gt;

&lt;h2 id=&quot;a-model-you-cant-call-just-rearranged-my-cheapskate-table&quot;&gt;A model you can’t call just rearranged my cheapskate table&lt;/h2&gt;

&lt;p&gt;Same method every week. Take each Arena category’s leader, draw a band 50 points below its rating, and find the cheapest model still inside the band. I compute the band in code from the full table, 129 rows per category, never off the first screen. I learned that one the hard way.&lt;/p&gt;

&lt;p&gt;This week the leader in every category is Argon. That changes two things at once. Argon raised every ceiling, so every band moved up. And Argon is cheap for a leader, so every “cheaper by” ratio collapsed. Bands got shallower too: 37 models deep in Overall (57 last week), 60 in Coding (73), 31 in Hard Prompts (43), 23 in Instruction Following (38), 35 in Math (43). Creative Writing went the other way, from 12 to 18, because Argon only beat Opus 5.5 there by three points.&lt;/p&gt;

&lt;table&gt;
  &lt;thead&gt;
    &lt;tr&gt;
      &lt;th&gt;Category&lt;/th&gt;
      &lt;th&gt;Leader&lt;/th&gt;
      &lt;th&gt;$ out&lt;/th&gt;
      &lt;th&gt;Cheapskate pick&lt;/th&gt;
      &lt;th&gt;$ out&lt;/th&gt;
      &lt;th&gt;Δ rating&lt;/th&gt;
      &lt;th&gt;Cheaper by&lt;/th&gt;
    &lt;/tr&gt;
  &lt;/thead&gt;
  &lt;tbody&gt;
    &lt;tr&gt;
      &lt;td&gt;Overall&lt;/td&gt;
      &lt;td&gt;gemini-4-argon-high (1525)&lt;/td&gt;
      &lt;td&gt;$10&lt;/td&gt;
      &lt;td&gt;MiMo-V2.6-Pro (1480, #26)&lt;/td&gt;
      &lt;td&gt;$0.87&lt;/td&gt;
      &lt;td&gt;−45&lt;/td&gt;
      &lt;td&gt;11.5×&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Coding&lt;/td&gt;
      &lt;td&gt;gemini-4-argon-high (1560)&lt;/td&gt;
      &lt;td&gt;$10&lt;/td&gt;
      &lt;td&gt;MiMo-V2.6-Flash* (1521, #33)&lt;/td&gt;
      &lt;td&gt;$0.28&lt;/td&gt;
      &lt;td&gt;−39&lt;/td&gt;
      &lt;td&gt;36×&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Creative Writing&lt;/td&gt;
      &lt;td&gt;gemini-4-argon-high (1519)&lt;/td&gt;
      &lt;td&gt;$10&lt;/td&gt;
      &lt;td&gt;gemini-3.7-flash-high (1493, #5)&lt;/td&gt;
      &lt;td&gt;$3.75&lt;/td&gt;
      &lt;td&gt;−26&lt;/td&gt;
      &lt;td&gt;2.7×&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Instruction Following&lt;/td&gt;
      &lt;td&gt;gemini-4-argon-high (1528)&lt;/td&gt;
      &lt;td&gt;$10&lt;/td&gt;
      &lt;td&gt;deepseek-v4.1-flash-max (1479, #21)&lt;/td&gt;
      &lt;td&gt;$1.20&lt;/td&gt;
      &lt;td&gt;−49&lt;/td&gt;
      &lt;td&gt;8.3×&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Hard Prompts&lt;/td&gt;
      &lt;td&gt;gemini-4-argon-high (1551)&lt;/td&gt;
      &lt;td&gt;$10&lt;/td&gt;
      &lt;td&gt;MiMo-V2.6-Pro (1505, #22)&lt;/td&gt;
      &lt;td&gt;$0.87&lt;/td&gt;
      &lt;td&gt;−46&lt;/td&gt;
      &lt;td&gt;11.5×&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Math&lt;/td&gt;
      &lt;td&gt;gemini-4-argon-high (1530)&lt;/td&gt;
      &lt;td&gt;$10&lt;/td&gt;
      &lt;td&gt;GLM-5.3-Flash (1503, #13)&lt;/td&gt;
      &lt;td&gt;$0.50&lt;/td&gt;
      &lt;td&gt;−27&lt;/td&gt;
      &lt;td&gt;20×&lt;/td&gt;
    &lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;

&lt;p&gt;*1,508 votes. GLM-5.3-Flash sits right next to it at 1523 with 6,157 votes for 50 cents, and that’s the steadier bet.&lt;/p&gt;

&lt;p&gt;Argon is quoted at its intro price. When it goes to $20 out, every ratio in that table doubles. GLM-5.3-Flash is quoted at Z.ai’s list price of 50 cents out. Arena’s price column shows 20 cents, but that’s a third-party floor.&lt;/p&gt;

&lt;p&gt;Last week GLM-5.3-Flash was the pick in four of six categories, and this week it’s the pick in one. GLM didn’t get worse. Its Overall rating went from 1474 to 1473. The band moved. The Overall cutoff is now 1475, so GLM missed it by two points. Hard Prompts’ cutoff is 1501 and GLM sits at 1500, one point short. In Instruction Following it’s six points short. A model nobody outside Fairwind can call, rated on preliminary votes, pushed the most-used cheap model of the last month out of half my table.&lt;/p&gt;

&lt;p&gt;So which table do you actually shop from? Here’s the same one, recomputed without Argon.&lt;/p&gt;

&lt;table&gt;
  &lt;thead&gt;
    &lt;tr&gt;
      &lt;th&gt;Category&lt;/th&gt;
      &lt;th&gt;Purchasable leader&lt;/th&gt;
      &lt;th&gt;$ out&lt;/th&gt;
      &lt;th&gt;Cheapskate pick&lt;/th&gt;
      &lt;th&gt;$ out&lt;/th&gt;
      &lt;th&gt;Cheaper by&lt;/th&gt;
    &lt;/tr&gt;
  &lt;/thead&gt;
  &lt;tbody&gt;
    &lt;tr&gt;
      &lt;td&gt;Overall&lt;/td&gt;
      &lt;td&gt;claude-opus-4-6-high (1505)&lt;/td&gt;
      &lt;td&gt;$25&lt;/td&gt;
      &lt;td&gt;GLM-5.3-Flash (1473)&lt;/td&gt;
      &lt;td&gt;$0.50&lt;/td&gt;
      &lt;td&gt;50×&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Coding&lt;/td&gt;
      &lt;td&gt;claude-fable-5-high (1552)&lt;/td&gt;
      &lt;td&gt;$50&lt;/td&gt;
      &lt;td&gt;MiMo-V2.6-Flash (1521)&lt;/td&gt;
      &lt;td&gt;$0.28&lt;/td&gt;
      &lt;td&gt;179×&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Creative Writing&lt;/td&gt;
      &lt;td&gt;claude-opus-5.5-high (1516)&lt;/td&gt;
      &lt;td&gt;$20&lt;/td&gt;
      &lt;td&gt;gemini-3.7-flash-high (1493)&lt;/td&gt;
      &lt;td&gt;$3.75&lt;/td&gt;
      &lt;td&gt;5.3×&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Instruction Following&lt;/td&gt;
      &lt;td&gt;claude-opus-5.5-high (1517)&lt;/td&gt;
      &lt;td&gt;$20&lt;/td&gt;
      &lt;td&gt;GLM-5.3-Flash (1472)&lt;/td&gt;
      &lt;td&gt;$0.50&lt;/td&gt;
      &lt;td&gt;40×&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Hard Prompts&lt;/td&gt;
      &lt;td&gt;claude-opus-5.5-high (1534)&lt;/td&gt;
      &lt;td&gt;$20&lt;/td&gt;
      &lt;td&gt;GLM-5.3-Flash (1500)&lt;/td&gt;
      &lt;td&gt;$0.50&lt;/td&gt;
      &lt;td&gt;40×&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Math&lt;/td&gt;
      &lt;td&gt;claude-fable-5-high (1522)&lt;/td&gt;
      &lt;td&gt;$50&lt;/td&gt;
      &lt;td&gt;GLM-5.3-Flash (1503)&lt;/td&gt;
      &lt;td&gt;$0.50&lt;/td&gt;
      &lt;td&gt;100×&lt;/td&gt;
    &lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;

&lt;p&gt;(In Hard Prompts and Math, MiMo-V2.6-Flash technically sneaks in at exactly the band edge for 28 cents, but in Math that’s on 330 votes. I’m not crowning anything on 330 votes.)&lt;/p&gt;

&lt;p&gt;That’s basically last week’s table. GLM is back in four categories, and the year-old claude-opus-4-6-high is still the top Overall model you can actually buy.&lt;/p&gt;

&lt;p&gt;MiMo-V2.6-Pro is now the Overall and Hard Prompts pick at 87 cents. Artificial Analysis rates it 46, the best score among open-weights models, at 13 cents per index task. It’s also slow: 46.6 tokens a second with a 5.2-second wait for the first token, well under the 75 tok/s median. And the distillation allegation from last issue, Anthropic’s report claiming Xiaomi routed 400,000 user conversations through Claude, still hangs over the whole MiMo family. That’s an accusation from a competitor, not a ruling. Whether it matters to you depends on your lawyers.&lt;/p&gt;

&lt;p&gt;MiMo-V2.6-Flash held the cheapest-in-band spot in Coding for a second week, and its vote count went from 1,043 to 1,508. I said last week that if it held that spot with a few thousand votes, the Coding crown had moved. It’s not there yet. AA gives it 38, about 6 cents per task, 62 tokens a second, and it’s verbose: 240 million output tokens to get through AA’s suite against a 140 million median.&lt;/p&gt;

&lt;p&gt;DeepSeek V4.1 Flash is the new Instruction Following pick at $1.20, and it’s the only pick this week that isn’t slow. AA clocks it at 222 tokens a second, third fastest among the open-weights models AA tracks, and it starts answering in about a second. It’s right at the band edge, one point above the cutoff, so it could fall out next week. On OpenRouter some third-party hosts list it at 18 to 40 cents out, way under list. Check uptime before you chase those.&lt;/p&gt;

&lt;p&gt;Creative Writing is still the category where you pay for quality. The best value is Gemini 3.7 Flash at $3.75, just 2.7 times cheaper than Argon. Gemini 3.8 Flash costs the same and rates lower, so 3.7 wins the tie.&lt;/p&gt;

&lt;h2 id=&quot;space-bunny-is-still-free-still-anonymous-and-now-number-one&quot;&gt;Space Bunny is still free, still anonymous, and now number one&lt;/h2&gt;

&lt;p&gt;Last week Space Bunny Alpha was the second-biggest model on OpenRouter. This week it’s &lt;a href=&quot;https://openrouter.ai/rankings&quot;&gt;first&lt;/a&gt;, at 38.5 trillion tokens, up 112 percent. Stealth models as a group now account for 10.5 percent of all text requests on OpenRouter, and stealth requests went up 296 percent in a week.&lt;/p&gt;

&lt;p&gt;It’s still free. It still hasn’t told anyone who made it. The tokenizer evidence still points at &lt;a href=&quot;https://cellcog.ai/blog/what-is-space-bunny-alpha/&quot;&gt;MiniMax’s M3.1-Flash-Preview&lt;/a&gt;, and it’s still unconfirmed. I’ll pull a number from its &lt;a href=&quot;https://openrouter.ai/stealth/space-bunny-alpha&quot;&gt;OpenRouter activity panel&lt;/a&gt; in the horror section, because it deserves the spot.&lt;/p&gt;

&lt;p&gt;Meanwhile, the paid volume king is DeepSeek V4.1 Flash. It’s number two this week at 28.4 trillion tokens, up 36 percent, and number one for the trailing month at 72.3 trillion. DeepSeek is the top author on OpenRouter at 21.8 percent of requests. Google is right behind at 21.5 percent, up 23 percent, all of that from models that aren’t Argon. GLM-5.3-Flash is third at 10 trillion, down 25 percent. Anthropic is 2.5 percent of requests, which tells you who’s buying Claude through OpenRouter and who’s buying it direct.&lt;/p&gt;

&lt;p&gt;On the apps list, Nous Research’s Hermes Agent pushed 1.86 trillion tokens in a day, ahead of Claude Code at 1.06 trillion. On OpenRouter, at least, an open-source agent is now burning more tokens than Anthropic’s own coding tool. I run OpenClaw at home, so I’m part of the problem.&lt;/p&gt;

&lt;h2 id=&quot;horror-stories&quot;&gt;Horror stories&lt;/h2&gt;

&lt;h3 id=&quot;your-agents-status-report-might-be-fiction&quot;&gt;Your agent’s status report might be fiction&lt;/h3&gt;

&lt;p&gt;I covered GPT-6.1 Astra above, but it belongs here too, because it’s the scariest thing in this issue. Most model horror stories are about a model being dumb: it hallucinated a function, it deleted the wrong folder. Astra’s failure was the model hiding its own actions and acting outside the task it was given. Its shipping big brother, GPT-6 Astra, went as far as running sock-puppet accounts in AISI’s supply-chain simulations. That’s a trust problem, and a higher benchmark score doesn’t fix it.&lt;/p&gt;

&lt;p&gt;OpenAI deserves credit for catching it in testing and not shipping it. But agent frameworks are built on the assumption that the model’s account of what it did is true. Your agent’s “I ran the tests and they passed” message is only useful if it’s true. Somebody built a model where it sometimes wasn’t. The next one might not get caught.&lt;/p&gt;

&lt;h3 id=&quot;four-trillion-prompt-tokens-to-a-stranger&quot;&gt;Four trillion prompt tokens to a stranger&lt;/h3&gt;

&lt;p&gt;Space Bunny Alpha’s activity panel shows 4.06 trillion prompt tokens since September 23, against 135 billion completion tokens. That’s a 30-to-1 ratio of input to output. People aren’t chatting with it. They’re feeding it whole codebases and huge documents, because it’s free and has a million-token window.&lt;/p&gt;

&lt;p&gt;The listing says prompts and completions “may be retained by the provider.” The provider won’t give its name. I’ve said this before and I’ll keep saying it until free stealth models stop topping the charts: kick the tires on public code all you want, but don’t send it anything you wouldn’t post on GitHub.&lt;/p&gt;

&lt;h3 id=&quot;budgeting-off-a-leaderboard-number&quot;&gt;Budgeting off a leaderboard number&lt;/h3&gt;

&lt;p&gt;This one’s small but I bet somebody does it this month. Arena shows Argon at $2/$10 next to a #1 ranking in every category. A team lead sees that, writes “switch to Gemini 4, cheaper and better” in a planning doc, and finds out at implementation time that the model isn’t available to them and the price is going to double. If you build a budget off a leaderboard row, check that you can actually call the model, and whether the price is introductory.&lt;/p&gt;

&lt;h2 id=&quot;whats-coming&quot;&gt;What’s coming&lt;/h2&gt;

&lt;p&gt;Gemini 4 Argon for the rest of us is announced but has no date. Paying API customers and Ultra subscribers go next, then everyone else. When it opens up, I’ll be watching two things: whether the Arena lead holds as votes pile up, and when the price goes to $4/$20.&lt;/p&gt;

&lt;p&gt;GPT-6.1 Astra is shelved with no new date. OpenAI says it’s shifting focus to “improving the safety of future models.”&lt;/p&gt;

&lt;p&gt;The Space Bunny unmasking is still speculation. If it turns out to be MiniMax M3.1-Flash and gets a price, it’ll probably follow the last stealth model. Union Alpha &lt;a href=&quot;https://openrouter.ai/stealth/space-bunny-alpha&quot;&gt;turned out to be Pareto by Unbiased&lt;/a&gt;, and the free period ended when the name came out. Then we’ll see how much of that 38.5 trillion was people who actually wanted the model.&lt;/p&gt;

&lt;p&gt;Claude Haiku 5.5 and Fable 5.5 are speculation too. &lt;a href=&quot;https://manifold.markets/prismatic/october-2026-ai-model-releases&quot;&gt;Prediction markets&lt;/a&gt; give both about a 92 percent chance of shipping in October. Anthropic said Haiku 5.5 would follow Opus and Sonnet “in the coming weeks,” and Sonnet shipped on September 28.&lt;/p&gt;

&lt;p&gt;DeepSeek V5 is still pure speculation. People keep saying October, and DeepSeek hasn’t said anything.&lt;/p&gt;

&lt;h2 id=&quot;finally&quot;&gt;Finally&lt;/h2&gt;

&lt;p&gt;Google and OpenAI arrived at the same place from opposite directions this week. Google built something it thinks is too good at hacking to hand out, so it handed it to defenders first. OpenAI built something that couldn’t be trusted to stay in its lane, so it handed out something smaller. In both cases the frontier model in the headlines isn’t for sale, and the one that is costs $2/$10.&lt;/p&gt;

&lt;p&gt;For the stuff you actually pay for, not much changed. GLM-5.3-Flash and the MiMo models are still the floor. Year-old Opus 4.6 still tops the purchasable Overall board. DeepSeek V4.1 Flash is moving more tokens than anything except a free model with no name.&lt;/p&gt;

&lt;p&gt;Next week I’ll check whether Argon’s 255 Math votes have turned into a few thousand, and whether anyone outside Fairwind has gotten a key. And I’ll check whether my cheapskate table survives the next ceiling this model raises.&lt;/p&gt;
</description>
        <pubDate>Tue, 06 Oct 2026 08:00:00 -0500</pubDate>
        <link>https://www.stephanmiller.com/model-buzz-roundup-week-of-0930/</link>
        <guid isPermaLink="true">https://www.stephanmiller.com/model-buzz-roundup-week-of-0930/</guid>
        
        <category>llm</category>
        
        <category>openrouter</category>
        
        <category>model-roundup</category>
        
        
        <category>large-language-models</category>
        
      </item>
    
      <item>
        <title>Before You Write a Scraper, Check Firecrawl Alexandria</title>
        <description>&lt;p&gt;Last Thursday I published a &lt;a href=&quot;https://www.stephanmiller.com/track-new-ai-models-on-the-arena-leaderboard/&quot;&gt;70-line scraper that watches the Arena leaderboard&lt;/a&gt; for new AI models. It runs daily and announces new arrivals.&lt;/p&gt;

&lt;p&gt;Two days earlier, &lt;a href=&quot;https://firecrawl.link/stephan-miller&quot;&gt;Firecrawl&lt;/a&gt; launched Alexandria, a catalog of providers that sell the structured version of exactly this kind of data, callable through the same scrape endpoint I was already paying for. I didn’t notice. I was busy having AI build me regex for a markdown table, while the company whose credits I was spending shipped the thing meant to make that regex unnecessary.&lt;/p&gt;

&lt;p&gt;So to recap: I built a machine that watches the internet for new arrivals, and the first new arrival it missed was the one built to replace it.&lt;/p&gt;

&lt;p&gt;If somebody is already selling the structured version of the data I wanted, why was I scraping the page? It never occurred to me to check.&lt;/p&gt;

&lt;ul id=&quot;markdown-toc&quot;&gt;
  &lt;li&gt;&lt;a href=&quot;#what-alexandria-is&quot; id=&quot;markdown-toc-what-alexandria-is&quot;&gt;What Alexandria Is&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;#first-you-need-a-newer-cli&quot; id=&quot;markdown-toc-first-you-need-a-newer-cli&quot;&gt;First, You Need a Newer CLI&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;#browsing-the-catalog-costs-nothing&quot; id=&quot;markdown-toc-browsing-the-catalog-costs-nothing&quot;&gt;Browsing the Catalog Costs Nothing&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;#the-answer-for-arena-keep-scraping&quot; id=&quot;markdown-toc-the-answer-for-arena-keep-scraping&quot;&gt;The Answer for Arena: Keep Scraping&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;#the-leaderboard-next-door-artificial-analysis&quot; id=&quot;markdown-toc-the-leaderboard-next-door-artificial-analysis&quot;&gt;The Leaderboard Next Door: Artificial Analysis&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;#the-top-20-is-really-nine-models&quot; id=&quot;markdown-toc-the-top-20-is-really-nine-models&quot;&gt;The Top 20 Is Really Nine Models&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;#the-same-harness-new-fetch-step&quot; id=&quot;markdown-toc-the-same-harness-new-fetch-step&quot;&gt;The Same Harness, New Fetch Step&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;#what-the-swap-costs&quot; id=&quot;markdown-toc-what-the-swap-costs&quot;&gt;What the Swap Costs&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;#if-your-cron-runs-unattended-read-this-part&quot; id=&quot;markdown-toc-if-your-cron-runs-unattended-read-this-part&quot;&gt;If Your Cron Runs Unattended, Read This Part&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;#check-the-catalog-before-you-write-the-scraper&quot; id=&quot;markdown-toc-check-the-catalog-before-you-write-the-scraper&quot;&gt;Check the Catalog Before You Write the Scraper&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&quot;what-alexandria-is&quot;&gt;What Alexandria Is&lt;/h2&gt;

&lt;p&gt;Alexandria is Firecrawl’s catalog of third-party data sources: FRED economic data, podcast transcripts, app store listings, company records, and a few hundred other things. They named their permanent store of the world’s data after history’s most famous case of data loss, which is either confidence or a gap in the marketing team’s classics minor. Each tool in it has a published contract that says what options it takes, what it returns, and what it costs. Finding tools and reading contracts is free. Running one costs credits, Firecrawl’s metered unit, and the free plan is 1,000 of them a month. The Arena scraper already spends one per run: it fetches the page through Firecrawl. That 1 is the baseline for every cost number in this post.&lt;/p&gt;

&lt;p&gt;It works in three steps: find a tool, read its contract, call it. The call goes to the regular &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;/v2/scrape&lt;/code&gt; endpoint with an &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;alexandria&lt;/code&gt; array instead of a URL.&lt;/p&gt;

&lt;p&gt;If you want the full tour, JP Caparas already wrote it well: &lt;a href=&quot;https://medium.com/reading-sh/firecrawl-alexandria-tested-end-to-end-with-real-api-calls-1cf1cf3030a9&quot;&gt;Firecrawl Alexandria, tested end to end with real API calls&lt;/a&gt;. I am not going to redo that. This post asks: does it replace a scraper I actually run?&lt;/p&gt;

&lt;h2 id=&quot;first-you-need-a-newer-cli&quot;&gt;First, You Need a Newer CLI&lt;/h2&gt;

&lt;p&gt;My installed Firecrawl CLI was 1.23.3, which predates Alexandria. The &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;list&lt;/code&gt; and &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;find-tools&lt;/code&gt; commands need 1.24 or later.&lt;/p&gt;

&lt;p&gt;The 1.23.3 on my machine came from &lt;a href=&quot;https://www.stephanmiller.com/firecrawl-cli-setup-skip-the-installer-that-rewrites-your-editors/&quot;&gt;a plain npm install, skipping the installer that rewrites your editors&lt;/a&gt;. So the CLI is installed, just old, and I am not upgrading it blind for a look around. &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;npx&lt;/code&gt; runs a specific version without replacing the one you have:&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;npx &lt;span class=&quot;nt&quot;&gt;-y&lt;/span&gt; firecrawl-cli@1.24.6 credits &lt;span class=&quot;nt&quot;&gt;--json&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;It picks up the same &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;FIRECRAWL_API_KEY&lt;/code&gt; from your environment. When you decide you want Alexandria in scripts permanently, upgrade for real. Until then, this is a free look.&lt;/p&gt;

&lt;h2 id=&quot;browsing-the-catalog-costs-nothing&quot;&gt;Browsing the Catalog Costs Nothing&lt;/h2&gt;

&lt;p&gt;Start at the root:&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;npx &lt;span class=&quot;nt&quot;&gt;-y&lt;/span&gt; firecrawl-cli@1.24.6 list &lt;span class=&quot;nt&quot;&gt;--json&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;That returned 23 categories for me on October 1st. The launch page said 20, and JP counted 21, so they seem to be adding categories as they go.&lt;/p&gt;

&lt;p&gt;Each item carries a &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;nextCommand&lt;/code&gt; you can run as is to go one level down. That is also why the commands in this post alternate between bare &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;list&lt;/code&gt; and &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;alexandria list&lt;/code&gt;: I am running whatever &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;nextCommand&lt;/code&gt; prints, verbatim. The one I cared about:&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;/images/2026/firecrawl-alexandria-couldnt-replace-my-scraper-body-1.jpg&quot; alt=&quot;Neon data streams converging into a glowing folder icon&quot; srcset=&quot;            /assets/resized/480/firecrawl-alexandria-couldnt-replace-my-scraper-body-1.jpg 480w,            /assets/resized/800/firecrawl-alexandria-couldnt-replace-my-scraper-body-1.jpg 800w,            /assets/resized/1400/firecrawl-alexandria-couldnt-replace-my-scraper-body-1.jpg 1400w,    &quot; loading=&quot;lazy&quot; /&gt;&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;npx &lt;span class=&quot;nt&quot;&gt;-y&lt;/span&gt; firecrawl-cli@1.24.6 alexandria list ai-models &lt;span class=&quot;nt&quot;&gt;--category&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;--json&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;Trimmed to what matters:&lt;/p&gt;

&lt;div class=&quot;language-json highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;p&quot;&gt;{&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;
  &lt;/span&gt;&lt;span class=&quot;nl&quot;&gt;&quot;provider&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;:&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;firecrawl&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;
  &lt;/span&gt;&lt;span class=&quot;nl&quot;&gt;&quot;capability&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;:&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;find-tools&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;
  &lt;/span&gt;&lt;span class=&quot;nl&quot;&gt;&quot;creditsCost&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;:&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;mi&quot;&gt;0&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;
  &lt;/span&gt;&lt;span class=&quot;nl&quot;&gt;&quot;data&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;:&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;p&quot;&gt;{&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;
    &lt;/span&gt;&lt;span class=&quot;nl&quot;&gt;&quot;level&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;:&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;providers&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;
    &lt;/span&gt;&lt;span class=&quot;nl&quot;&gt;&quot;items&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;:&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;p&quot;&gt;[&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;
      &lt;/span&gt;&lt;span class=&quot;p&quot;&gt;{&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;nl&quot;&gt;&quot;id&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;:&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;artificialanalysis-ai&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;nl&quot;&gt;&quot;name&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;:&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;Artificial Analysis&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;nl&quot;&gt;&quot;toolCount&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;:&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;mi&quot;&gt;4&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;p&quot;&gt;},&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;
      &lt;/span&gt;&lt;span class=&quot;p&quot;&gt;{&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;nl&quot;&gt;&quot;id&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;:&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;arxiv-org&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;nl&quot;&gt;&quot;name&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;:&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;arXiv&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;nl&quot;&gt;&quot;toolCount&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;:&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;mi&quot;&gt;6&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;p&quot;&gt;},&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;
      &lt;/span&gt;&lt;span class=&quot;p&quot;&gt;{&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;nl&quot;&gt;&quot;id&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;:&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;huggingface-co&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;nl&quot;&gt;&quot;name&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;:&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;Hugging Face&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;nl&quot;&gt;&quot;toolCount&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;:&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;mi&quot;&gt;5&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;p&quot;&gt;},&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;
      &lt;/span&gt;&lt;span class=&quot;p&quot;&gt;{&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;nl&quot;&gt;&quot;id&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;:&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;ollama-com&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;nl&quot;&gt;&quot;name&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;:&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;Ollama model library&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;nl&quot;&gt;&quot;toolCount&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;:&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;mi&quot;&gt;7&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;p&quot;&gt;},&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;
      &lt;/span&gt;&lt;span class=&quot;p&quot;&gt;{&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;nl&quot;&gt;&quot;id&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;:&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;openrouter-ai&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;nl&quot;&gt;&quot;name&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;:&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;OpenRouter&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;nl&quot;&gt;&quot;toolCount&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;:&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;mi&quot;&gt;8&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;p&quot;&gt;},&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;
      &lt;/span&gt;&lt;span class=&quot;p&quot;&gt;{&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;nl&quot;&gt;&quot;id&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;:&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;semanticscholar-org&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;nl&quot;&gt;&quot;name&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;:&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;Semantic Scholar&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;nl&quot;&gt;&quot;toolCount&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;:&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;mi&quot;&gt;10&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;p&quot;&gt;}&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;
    &lt;/span&gt;&lt;span class=&quot;p&quot;&gt;],&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;
    &lt;/span&gt;&lt;span class=&quot;nl&quot;&gt;&quot;total&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;:&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;mi&quot;&gt;6&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;
  &lt;/span&gt;&lt;span class=&quot;p&quot;&gt;}&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;
&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;}&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;Six providers. No Arena. When I first ran this a few days earlier it was three: Artificial Analysis, Hugging Face, and OpenRouter. arXiv, Ollama, and Semantic Scholar showed up since, so this category is growing as fast as the category list is.&lt;/p&gt;

&lt;h2 id=&quot;the-answer-for-arena-keep-scraping&quot;&gt;The Answer for Arena: Keep Scraping&lt;/h2&gt;

&lt;p&gt;If you want Arena’s rankings, Alexandria does not have them. The Arena scraper stays a scraper, and it costs one credit a run.&lt;/p&gt;

&lt;p&gt;It took two commands. Every scraper in this series gets that check from now on: it takes a minute.&lt;/p&gt;

&lt;p&gt;But the category was not empty, and one of those providers is close enough to be interesting.&lt;/p&gt;

&lt;h2 id=&quot;the-leaderboard-next-door-artificial-analysis&quot;&gt;The Leaderboard Next Door: Artificial Analysis&lt;/h2&gt;

&lt;p&gt;Artificial Analysis runs its own LLM leaderboard. It doesn’t use human votes. It runs benchmarks and puts them into an Intelligence Index, along with list prices and measured speed. It answers a different question than Arena does: Arena asks which answer people like better, and Artificial Analysis asks how the model scores on tests. For “did a new model just show up near the top,” it works too: the same question, answered from the benchmark side. A second signal, not a substitute.&lt;/p&gt;

&lt;p&gt;Listing that provider’s tools shows the price of everything before you spend anything:&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;npx &lt;span class=&quot;nt&quot;&gt;-y&lt;/span&gt; firecrawl-cli@1.24.6 list ai-models artificialanalysis-ai &lt;span class=&quot;nt&quot;&gt;--category&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;--json&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;table&gt;
  &lt;thead&gt;
    &lt;tr&gt;
      &lt;th&gt;Tool&lt;/th&gt;
      &lt;th&gt;Credits&lt;/th&gt;
      &lt;th&gt;Per record?&lt;/th&gt;
    &lt;/tr&gt;
  &lt;/thead&gt;
  &lt;tbody&gt;
    &lt;tr&gt;
      &lt;td&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;benchmarks/search_models&lt;/code&gt;&lt;/td&gt;
      &lt;td&gt;5&lt;/td&gt;
      &lt;td&gt;no&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;benchmarks/get_model&lt;/code&gt;&lt;/td&gt;
      &lt;td&gt;5&lt;/td&gt;
      &lt;td&gt;no&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;benchmarks/list_model_providers&lt;/code&gt;&lt;/td&gt;
      &lt;td&gt;5&lt;/td&gt;
      &lt;td&gt;no&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;benchmarks/search_provider_endpoints&lt;/code&gt;&lt;/td&gt;
      &lt;td&gt;5&lt;/td&gt;
      &lt;td&gt;no&lt;/td&gt;
    &lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;

&lt;p&gt;OpenRouter’s eight tools are also 5 credits a call. &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;perRecord: false&lt;/code&gt; is the flag to look at. It means the price is per call, not per row, so asking for 100 models costs the same as asking for one. If that flag were &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;true&lt;/code&gt;, a pagination loop could burn credits fast.&lt;/p&gt;

&lt;p&gt;The contract for &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;search_models&lt;/code&gt; says it returns the whole leaderboard sorted by Intelligence Index, up to 100 per page, with filters for creator, open weights, and reasoning. Nothing in it needs a terms agreement. So I spent the five credits:&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;npx &lt;span class=&quot;nt&quot;&gt;-y&lt;/span&gt; firecrawl-cli@1.24.6 scrape artificialanalysis-ai/benchmarks/search_models &lt;span class=&quot;se&quot;&gt;\&lt;/span&gt;
  &lt;span class=&quot;nt&quot;&gt;--options&lt;/span&gt; &lt;span class=&quot;s1&quot;&gt;&apos;{&quot;limit&quot;: 20}&apos;&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;--json&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;&lt;img src=&quot;/images/2026/firecrawl-alexandria-couldnt-replace-my-scraper-body-2.jpg&quot; alt=&quot;A spectral AI figure reaching toward a floating column of model score cards&quot; srcset=&quot;            /assets/resized/480/firecrawl-alexandria-couldnt-replace-my-scraper-body-2.jpg 480w,            /assets/resized/800/firecrawl-alexandria-couldnt-replace-my-scraper-body-2.jpg 800w,            /assets/resized/1400/firecrawl-alexandria-couldnt-replace-my-scraper-body-2.jpg 1400w,    &quot; loading=&quot;lazy&quot; /&gt;&lt;/p&gt;

&lt;p&gt;It came back in about a second and a half. Each result is a proper record: slug, creator, the index, a dozen benchmark scores, prices per million tokens, context window, and a link back to the model’s page on the site. No regex. No &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;--wait-for 3000&lt;/code&gt; hoping the table rendered. No guessing which row is the calendar widget.&lt;/p&gt;

&lt;p&gt;And then I looked at the top 20.&lt;/p&gt;

&lt;h2 id=&quot;the-top-20-is-really-nine-models&quot;&gt;The Top 20 Is Really Nine Models&lt;/h2&gt;

&lt;p&gt;Here are the first eight rows:&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;1  claude-opus-5-5          Anthropic  57.6
2  claude-opus-5-5-xhigh    Anthropic  56.0
3  claude-opus-5-5-high     Anthropic  53.6
4  claude-fable-5-1         Anthropic  53.4
5  claude-fable-5-1-xhigh   Anthropic  53.2
6  gpt-6-astra              OpenAI     52.7
7  gpt-6-astra-xhigh        OpenAI     52.4
8  claude-opus-5-5-medium   Anthropic  51.2
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;Artificial Analysis lists every reasoning-effort setting as its own row. Opus 5.5 at max effort, at xhigh, at high, at medium. Those are four entries for one model. The “top 20” is really about nine models, several times over. Claude Opus 5.5 alone holds four of the top eight slots, which makes this less a leaderboard than an Anthropic family newsletter.&lt;/p&gt;

&lt;p&gt;That is not a bug. Nobody ships one this loud, and the first test run prints it. It is a caveat of the swap, and the first one to handle: the Arena harness assumes one row per model, and this feed returns one row per model per effort setting. The contract gives you clean JSON but clean data does not mean you can skip reading it.&lt;/p&gt;

&lt;p&gt;The fix is to collapse effort variants into one model before ranking. The slugs end in &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;-xhigh&lt;/code&gt;, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;-high&lt;/code&gt;, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;-medium&lt;/code&gt;, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;-low&lt;/code&gt;, so strip that and keep the first (highest) row for each model.&lt;/p&gt;

&lt;p&gt;My first version of that also stripped &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;-max&lt;/code&gt;, since the variants are labeled “Max Effort” in their names. That ate a real model. &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;qwen3-8-max&lt;/code&gt; is not Qwen 3.8 at max effort. “Max” is part of the product’s name. The max-effort rows in this table have no suffix at all, which you only find out by looking. So &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;max&lt;/code&gt; came out of the pattern.&lt;/p&gt;

&lt;h2 id=&quot;the-same-harness-new-fetch-step&quot;&gt;The Same Harness, New Fetch Step&lt;/h2&gt;

&lt;p&gt;Here is the Arena harness with the scraper swapped for an Alexandria call. It is still 70 lines, and everything below &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;fetch()&lt;/code&gt; is the same code: save with a timestamp, compare with the last run, speak only when a model is new to the top 20. The Arena watcher keeps its job. This one runs alongside it: same harness, different signal.&lt;/p&gt;

&lt;div class=&quot;language-python highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;c1&quot;&gt;#!/usr/bin/env python3
&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;&quot;&quot;&quot;arena_watch.py with the scraper swapped for a Firecrawl Alexandria call.&quot;&quot;&quot;&lt;/span&gt;
&lt;span class=&quot;kn&quot;&gt;import&lt;/span&gt; &lt;span class=&quot;nn&quot;&gt;json&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;re&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;sqlite3&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;subprocess&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;sys&lt;/span&gt;
&lt;span class=&quot;kn&quot;&gt;from&lt;/span&gt; &lt;span class=&quot;nn&quot;&gt;datetime&lt;/span&gt; &lt;span class=&quot;kn&quot;&gt;import&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;datetime&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;timezone&lt;/span&gt;

&lt;span class=&quot;n&quot;&gt;DB&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;s&quot;&gt;&quot;aa.db&quot;&lt;/span&gt;
&lt;span class=&quot;n&quot;&gt;TOP&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;mi&quot;&gt;20&lt;/span&gt;

&lt;span class=&quot;c1&quot;&gt;# Artificial Analysis lists every reasoning-effort setting as its own row, so
# claude-opus-5-5, -xhigh and -high are three rows. Collapse them to one model.
&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;EFFORT&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;re&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;nb&quot;&gt;compile&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;sa&quot;&gt;r&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;&quot;-(xhigh|high|medium|low|minimal)$&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;


&lt;span class=&quot;k&quot;&gt;def&lt;/span&gt; &lt;span class=&quot;nf&quot;&gt;fetch&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;():&lt;/span&gt;
    &lt;span class=&quot;n&quot;&gt;out&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;subprocess&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;run&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;
        &lt;span class=&quot;p&quot;&gt;[&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;&quot;firecrawl&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;s&quot;&gt;&quot;scrape&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;s&quot;&gt;&quot;artificialanalysis-ai/benchmarks/search_models&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt;
         &lt;span class=&quot;s&quot;&gt;&quot;--options&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;s&quot;&gt;&apos;{&quot;limit&quot;: 100}&apos;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;s&quot;&gt;&quot;--json&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;],&lt;/span&gt;
        &lt;span class=&quot;n&quot;&gt;capture_output&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;bp&quot;&gt;True&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;text&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;bp&quot;&gt;True&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;check&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;bp&quot;&gt;True&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;).&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;stdout&lt;/span&gt;
    &lt;span class=&quot;n&quot;&gt;call&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;json&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;loads&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;out&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;[&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;out&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;index&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;&quot;{&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;):])[&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;&quot;data&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;][&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;&quot;alexandria&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;][&lt;/span&gt;&lt;span class=&quot;mi&quot;&gt;0&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;]&lt;/span&gt;
    &lt;span class=&quot;n&quot;&gt;rows&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;seen&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;p&quot;&gt;[],&lt;/span&gt; &lt;span class=&quot;nb&quot;&gt;set&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;()&lt;/span&gt;
    &lt;span class=&quot;k&quot;&gt;for&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;m&lt;/span&gt; &lt;span class=&quot;ow&quot;&gt;in&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;call&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;[&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;&quot;data&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;][&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;&quot;results&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;]:&lt;/span&gt;  &lt;span class=&quot;c1&quot;&gt;# already sorted by Intelligence Index
&lt;/span&gt;        &lt;span class=&quot;n&quot;&gt;model&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;EFFORT&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;sub&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;&quot;&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;m&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;[&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;&quot;slug&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;])&lt;/span&gt;
        &lt;span class=&quot;k&quot;&gt;if&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;model&lt;/span&gt; &lt;span class=&quot;ow&quot;&gt;in&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;seen&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;:&lt;/span&gt;
            &lt;span class=&quot;k&quot;&gt;continue&lt;/span&gt;  &lt;span class=&quot;c1&quot;&gt;# a lower effort setting of a model we already ranked
&lt;/span&gt;        &lt;span class=&quot;n&quot;&gt;seen&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;add&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;model&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;
        &lt;span class=&quot;n&quot;&gt;rows&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;append&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;((&lt;/span&gt;&lt;span class=&quot;nb&quot;&gt;len&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;rows&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;+&lt;/span&gt; &lt;span class=&quot;mi&quot;&gt;1&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;model&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;m&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;[&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;&quot;creator_name&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;],&lt;/span&gt;
                     &lt;span class=&quot;nb&quot;&gt;round&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;m&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;[&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;&quot;intelligence_index&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;]&lt;/span&gt; &lt;span class=&quot;ow&quot;&gt;or&lt;/span&gt; &lt;span class=&quot;mi&quot;&gt;0&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;mi&quot;&gt;1&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)))&lt;/span&gt;
    &lt;span class=&quot;k&quot;&gt;return&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;rows&lt;/span&gt;


&lt;span class=&quot;k&quot;&gt;def&lt;/span&gt; &lt;span class=&quot;nf&quot;&gt;main&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;():&lt;/span&gt;
    &lt;span class=&quot;n&quot;&gt;rows&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;fetch&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;()&lt;/span&gt;
    &lt;span class=&quot;k&quot;&gt;if&lt;/span&gt; &lt;span class=&quot;nb&quot;&gt;len&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;rows&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;&amp;lt;&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;TOP&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;:&lt;/span&gt;
        &lt;span class=&quot;n&quot;&gt;sys&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;nb&quot;&gt;exit&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;sa&quot;&gt;f&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;&quot;only got &lt;/span&gt;&lt;span class=&quot;si&quot;&gt;{&lt;/span&gt;&lt;span class=&quot;nb&quot;&gt;len&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;rows&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;&lt;span class=&quot;si&quot;&gt;}&lt;/span&gt;&lt;span class=&quot;s&quot;&gt; models back, check the call&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;

    &lt;span class=&quot;n&quot;&gt;taken&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;datetime&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;now&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;timezone&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;utc&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;).&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;strftime&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;&quot;%Y-%m-%d %H:%M&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;
    &lt;span class=&quot;n&quot;&gt;db&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;sqlite3&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;connect&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;DB&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;
    &lt;span class=&quot;n&quot;&gt;db&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;execute&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;&quot;&quot;&quot;CREATE TABLE IF NOT EXISTS ranks
                  (taken TEXT, rank INT, model TEXT, org TEXT, score REAL)&quot;&quot;&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;
    &lt;span class=&quot;n&quot;&gt;db&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;executemany&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;&quot;INSERT INTO ranks VALUES (?,?,?,?,?)&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;p&quot;&gt;[(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;taken&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;*&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;r&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt; &lt;span class=&quot;k&quot;&gt;for&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;r&lt;/span&gt; &lt;span class=&quot;ow&quot;&gt;in&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;rows&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;])&lt;/span&gt;
    &lt;span class=&quot;n&quot;&gt;db&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;commit&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;()&lt;/span&gt;

    &lt;span class=&quot;n&quot;&gt;prev&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;db&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;execute&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;&quot;SELECT MAX(taken) FROM ranks WHERE taken &amp;lt; ?&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;taken&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,)).&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;fetchone&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;()[&lt;/span&gt;&lt;span class=&quot;mi&quot;&gt;0&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;]&lt;/span&gt;
    &lt;span class=&quot;k&quot;&gt;if&lt;/span&gt; &lt;span class=&quot;ow&quot;&gt;not&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;prev&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;:&lt;/span&gt;
        &lt;span class=&quot;k&quot;&gt;print&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;sa&quot;&gt;f&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;si&quot;&gt;{&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;taken&lt;/span&gt;&lt;span class=&quot;si&quot;&gt;}&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;: first run, saved &lt;/span&gt;&lt;span class=&quot;si&quot;&gt;{&lt;/span&gt;&lt;span class=&quot;nb&quot;&gt;len&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;rows&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;&lt;span class=&quot;si&quot;&gt;}&lt;/span&gt;&lt;span class=&quot;s&quot;&gt; models. Nothing to compare yet.&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;
        &lt;span class=&quot;k&quot;&gt;return&lt;/span&gt;

    &lt;span class=&quot;n&quot;&gt;was_top&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;p&quot;&gt;{&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;m&lt;/span&gt; &lt;span class=&quot;k&quot;&gt;for&lt;/span&gt; &lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;m&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,)&lt;/span&gt; &lt;span class=&quot;ow&quot;&gt;in&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;db&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;execute&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;
        &lt;span class=&quot;s&quot;&gt;&quot;SELECT model FROM ranks WHERE taken = ? AND rank &amp;lt;= ?&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;prev&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;TOP&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;))}&lt;/span&gt;
    &lt;span class=&quot;n&quot;&gt;seen&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;p&quot;&gt;{&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;m&lt;/span&gt; &lt;span class=&quot;k&quot;&gt;for&lt;/span&gt; &lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;m&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,)&lt;/span&gt; &lt;span class=&quot;ow&quot;&gt;in&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;db&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;execute&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;&quot;SELECT DISTINCT model FROM ranks WHERE taken &amp;lt; ?&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;taken&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,))}&lt;/span&gt;

    &lt;span class=&quot;k&quot;&gt;print&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;sa&quot;&gt;f&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;si&quot;&gt;{&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;taken&lt;/span&gt;&lt;span class=&quot;si&quot;&gt;}&lt;/span&gt;&lt;span class=&quot;s&quot;&gt; vs &lt;/span&gt;&lt;span class=&quot;si&quot;&gt;{&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;prev&lt;/span&gt;&lt;span class=&quot;si&quot;&gt;}&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;
    &lt;span class=&quot;k&quot;&gt;for&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;rank&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;model&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;org&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;score&lt;/span&gt; &lt;span class=&quot;ow&quot;&gt;in&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;rows&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;[:&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;TOP&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;]:&lt;/span&gt;
        &lt;span class=&quot;k&quot;&gt;if&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;model&lt;/span&gt; &lt;span class=&quot;ow&quot;&gt;not&lt;/span&gt; &lt;span class=&quot;ow&quot;&gt;in&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;seen&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;:&lt;/span&gt;
            &lt;span class=&quot;k&quot;&gt;print&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;sa&quot;&gt;f&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;&quot;  NEW      #&lt;/span&gt;&lt;span class=&quot;si&quot;&gt;{&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;rank&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;:&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;&amp;lt;&lt;/span&gt;&lt;span class=&quot;mi&quot;&gt;3&lt;/span&gt;&lt;span class=&quot;si&quot;&gt;}&lt;/span&gt;&lt;span class=&quot;s&quot;&gt; &lt;/span&gt;&lt;span class=&quot;si&quot;&gt;{&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;model&lt;/span&gt;&lt;span class=&quot;si&quot;&gt;}&lt;/span&gt;&lt;span class=&quot;s&quot;&gt; (&lt;/span&gt;&lt;span class=&quot;si&quot;&gt;{&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;org&lt;/span&gt;&lt;span class=&quot;si&quot;&gt;}&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;, &lt;/span&gt;&lt;span class=&quot;si&quot;&gt;{&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;score&lt;/span&gt;&lt;span class=&quot;si&quot;&gt;}&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;)&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;
        &lt;span class=&quot;k&quot;&gt;elif&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;model&lt;/span&gt; &lt;span class=&quot;ow&quot;&gt;not&lt;/span&gt; &lt;span class=&quot;ow&quot;&gt;in&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;was_top&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;:&lt;/span&gt;
            &lt;span class=&quot;k&quot;&gt;print&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;sa&quot;&gt;f&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;&quot;  MOVED UP #&lt;/span&gt;&lt;span class=&quot;si&quot;&gt;{&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;rank&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;:&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;&amp;lt;&lt;/span&gt;&lt;span class=&quot;mi&quot;&gt;3&lt;/span&gt;&lt;span class=&quot;si&quot;&gt;}&lt;/span&gt;&lt;span class=&quot;s&quot;&gt; &lt;/span&gt;&lt;span class=&quot;si&quot;&gt;{&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;model&lt;/span&gt;&lt;span class=&quot;si&quot;&gt;}&lt;/span&gt;&lt;span class=&quot;s&quot;&gt; (&lt;/span&gt;&lt;span class=&quot;si&quot;&gt;{&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;org&lt;/span&gt;&lt;span class=&quot;si&quot;&gt;}&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;, &lt;/span&gt;&lt;span class=&quot;si&quot;&gt;{&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;score&lt;/span&gt;&lt;span class=&quot;si&quot;&gt;}&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;)&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;


&lt;span class=&quot;k&quot;&gt;if&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;__name__&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;==&lt;/span&gt; &lt;span class=&quot;s&quot;&gt;&quot;__main__&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;:&lt;/span&gt;
    &lt;span class=&quot;n&quot;&gt;main&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;()&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;The script shells out to plain &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;firecrawl&lt;/code&gt;. The pinned &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;npx&lt;/code&gt; form above was for looking around. On the box that runs the cron, either upgrade the installed CLI for real or swap the command for the full &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;npx&lt;/code&gt; version.&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;/images/2026/firecrawl-alexandria-couldnt-replace-my-scraper-body-3.jpg&quot; alt=&quot;Two ribbons of light braided into a single bright knot&quot; srcset=&quot;            /assets/resized/480/firecrawl-alexandria-couldnt-replace-my-scraper-body-3.jpg 480w,            /assets/resized/800/firecrawl-alexandria-couldnt-replace-my-scraper-body-3.jpg 800w,            /assets/resized/1400/firecrawl-alexandria-couldnt-replace-my-scraper-body-3.jpg 1400w,    &quot; loading=&quot;lazy&quot; /&gt;&lt;/p&gt;

&lt;p&gt;It asks for 100 rows instead of 20, because after collapsing there have to be 20 distinct models left. Since pricing is per call, the extra 80 rows are free. That one call came back with 61 distinct models. The first run’s top 20, with the variants collapsed:&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;#1   claude-opus-5-5      (Anthropic, 57.6)
#2   claude-fable-5-1     (Anthropic, 53.4)
#3   gpt-6-astra          (OpenAI, 52.7)
#4   muse-spark-1-3       (Meta, 48.1)
#5   gpt-6-sol            (OpenAI, 47.5)
#6   grok-4-7             (SpaceXAI, 46.4)
#7   mimo-v2-6-pro        (Xiaomi, 46.3)
#8   qwen3-8-max          (Alibaba, 45.4)
...
#20  gpt-6-luna           (OpenAI, 37.3)
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;Now that is a top 20 you can diff.&lt;/p&gt;

&lt;p&gt;The &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;sys.exit&lt;/code&gt; guard means something different here, too. In the Arena version, fewer than 20 rows meant “the page layout changed, go fix the regex.” Here it means the call itself went wrong.&lt;/p&gt;

&lt;h2 id=&quot;what-the-swap-costs&quot;&gt;What the Swap Costs&lt;/h2&gt;

&lt;table&gt;
  &lt;thead&gt;
    &lt;tr&gt;
      &lt;th&gt; &lt;/th&gt;
      &lt;th&gt;Arena scraper&lt;/th&gt;
      &lt;th&gt;Alexandria call&lt;/th&gt;
    &lt;/tr&gt;
  &lt;/thead&gt;
  &lt;tbody&gt;
    &lt;tr&gt;
      &lt;td&gt;Credits per run&lt;/td&gt;
      &lt;td&gt;1&lt;/td&gt;
      &lt;td&gt;5&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Daily for 30 days&lt;/td&gt;
      &lt;td&gt;30&lt;/td&gt;
      &lt;td&gt;150&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Parsing&lt;/td&gt;
      &lt;td&gt;Regex on a markdown table&lt;/td&gt;
      &lt;td&gt;JSON, fields documented in the contract&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Breaks when&lt;/td&gt;
      &lt;td&gt;The page redesigns&lt;/td&gt;
      &lt;td&gt;The contract changes&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Extra fields&lt;/td&gt;
      &lt;td&gt;Rank, score, org&lt;/td&gt;
      &lt;td&gt;Benchmarks, prices, speed, context window&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;What it measures&lt;/td&gt;
      &lt;td&gt;Human preference votes&lt;/td&gt;
      &lt;td&gt;Benchmark scores&lt;/td&gt;
    &lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;

&lt;p&gt;A daily Alexandria check uses 150 of the free 1,000 a month, which is fine for one watcher and adds up if you run five.&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;/images/2026/firecrawl-alexandria-couldnt-replace-my-scraper-body-4.jpg&quot; alt=&quot;A glowing chip feeding circuit traces that fan out into floating currency symbols&quot; srcset=&quot;            /assets/resized/480/firecrawl-alexandria-couldnt-replace-my-scraper-body-4.jpg 480w,            /assets/resized/800/firecrawl-alexandria-couldnt-replace-my-scraper-body-4.jpg 800w,            /assets/resized/1400/firecrawl-alexandria-couldnt-replace-my-scraper-body-4.jpg 1400w,    &quot; loading=&quot;lazy&quot; /&gt;&lt;/p&gt;

&lt;p&gt;So the swap is not free. Five times the credits buys you no regex, no layout breakage, and a lot of extra fields. The Model Buzz Report &lt;a href=&quot;https://www.stephanmiller.com/building-a-cost-saving-skill-that-accidentally-became-its-own-newsletter/&quot;&gt;leans on both kinds of signal&lt;/a&gt;, so the two lists side by side say more than either one alone.&lt;/p&gt;

&lt;p&gt;My test cost 15 credits: three five-credit calls, one of them the run that ate Qwen. Fifteen, out of the 1,000 I pay nothing for. All that comparison shopping, and the grand total is 1.5% of free.&lt;/p&gt;

&lt;h2 id=&quot;if-your-cron-runs-unattended-read-this-part&quot;&gt;If Your Cron Runs Unattended, Read This Part&lt;/h2&gt;

&lt;p&gt;Some Alexandria providers will not run until you accept their data-use terms. Alexandria tells the caller, human or agent, to show the terms to a person and wait for an explicit yes.&lt;/p&gt;

&lt;p&gt;The two providers I touched are fine: &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;terms show openrouter-ai&lt;/code&gt; returns &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;&quot;required&quot;: false&lt;/code&gt;, and Artificial Analysis is not in the terms catalog at all. But check before you schedule anything:&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;npx &lt;span class=&quot;nt&quot;&gt;-y&lt;/span&gt; firecrawl-cli@1.24.6 alexandria terms show &amp;lt;provider&amp;gt; &lt;span class=&quot;nt&quot;&gt;--json&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;h2 id=&quot;check-the-catalog-before-you-write-the-scraper&quot;&gt;Check the Catalog Before You Write the Scraper&lt;/h2&gt;

&lt;p&gt;Every recipe in this series so far has assumed the data lives on a web page and the job is &lt;a href=&quot;https://www.stephanmiller.com/what-to-do-when-your-ai-coding-agent-cant-read-a-web-page/&quot;&gt;getting it off that page&lt;/a&gt;. Alexandria adds a step before that: someone may already sell the structured version.&lt;/p&gt;

&lt;p&gt;The step is cheap. &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;list&lt;/code&gt;, follow &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;nextCommand&lt;/code&gt; two or three times, read the contract. Takes a couple of minutes. If what you want is in there, you skip the part of scraping that breaks. If it is not, like Arena, you have lost nothing and you write the scraper knowing you have to.&lt;/p&gt;

&lt;p&gt;Next up for this one: OpenRouter’s &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;get_model_stats&lt;/code&gt;, which says it returns usage and availability for a model. That is the other half of what the &lt;a href=&quot;https://www.stephanmiller.com/series/model-buzz-report/&quot;&gt;Model Buzz Report&lt;/a&gt; watches, and I have been getting it the hard way. Which is, as established, my signature way of getting things.&lt;/p&gt;
</description>
        <pubDate>Mon, 05 Oct 2026 07:00:00 -0500</pubDate>
        <link>https://www.stephanmiller.com/check-firecrawl-alexandria-before-you-scrape/</link>
        <guid isPermaLink="true">https://www.stephanmiller.com/check-firecrawl-alexandria-before-you-scrape/</guid>
        
        <category>Firecrawl Alexandria</category>
        
        <category>Firecrawl CLI</category>
        
        <category>Arena leaderboard</category>
        
        
        <category>web-scraping</category>
        
        <category>python</category>
        
      </item>
    
      <item>
        <title>Obsidian OneDrive Sync: The Setup Windows Already Installed for You</title>
        <description>&lt;p&gt;OneDrive is the sync method nobody chooses. It chooses you. You set up a Windows machine, click through a dozen screens, and at the end of it your Documents folder lives in Microsoft’s cloud whether you asked for that or not.&lt;/p&gt;

&lt;p&gt;So a lot of people end up syncing Obsidian with OneDrive by accident. They made a vault in Documents, and one day it showed up on their other laptop. It works great, right up until a note opens blank or a file named &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;Daily Note-DESKTOP-7QK2M1B.md&lt;/code&gt; appears next to the real one. My &lt;a href=&quot;https://www.stephanmiller.com/sync-obsidian-vault-across-devices/&quot;&gt;guide to syncing an Obsidian vault across devices&lt;/a&gt; gives OneDrive about two lines.&lt;/p&gt;

&lt;ul id=&quot;markdown-toc&quot;&gt;
  &lt;li&gt;&lt;a href=&quot;#why-onedrive-is-the-path-of-least-resistance&quot; id=&quot;markdown-toc-why-onedrive-is-the-path-of-least-resistance&quot;&gt;Why OneDrive Is the Path of Least Resistance&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;#files-on-demand-is-why-your-notes-vanished&quot; id=&quot;markdown-toc-files-on-demand-is-why-your-notes-vanished&quot;&gt;Files On-Demand Is Why Your Notes Vanished&lt;/a&gt;    &lt;ul&gt;
      &lt;li&gt;&lt;a href=&quot;#storage-sense-will-undo-your-work&quot; id=&quot;markdown-toc-storage-sense-will-undo-your-work&quot;&gt;Storage Sense will undo your work&lt;/a&gt;&lt;/li&gt;
    &lt;/ul&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;#onedrive-on-a-mac&quot; id=&quot;markdown-toc-onedrive-on-a-mac&quot;&gt;OneDrive on a Mac&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;#what-onedrive-refuses-to-sync&quot; id=&quot;markdown-toc-what-onedrive-refuses-to-sync&quot;&gt;What OneDrive Refuses to Sync&lt;/a&gt;    &lt;ul&gt;
      &lt;li&gt;&lt;a href=&quot;#the-duplicate-file-with-your-computers-name-in-it&quot; id=&quot;markdown-toc-the-duplicate-file-with-your-computers-name-in-it&quot;&gt;The duplicate file with your computer’s name in it&lt;/a&gt;&lt;/li&gt;
    &lt;/ul&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;#iphone-and-ipad-remotely-save-or-nothing&quot; id=&quot;markdown-toc-iphone-and-ipad-remotely-save-or-nothing&quot;&gt;iPhone and iPad: Remotely Save or Nothing&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;#android-remotely-save-or-a-sync-app&quot; id=&quot;markdown-toc-android-remotely-save-or-a-sync-app&quot;&gt;Android: Remotely Save or a Sync App&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;#personal-vs-business-your-work-account-is-out&quot; id=&quot;markdown-toc-personal-vs-business-your-work-account-is-out&quot;&gt;Personal vs. Business: Your Work Account Is Out&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;#the-cost-math-if-you-already-pay-for-microsoft-365&quot; id=&quot;markdown-toc-the-cost-math-if-you-already-pay-for-microsoft-365&quot;&gt;The Cost Math If You Already Pay for Microsoft 365&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;#should-you-use-it&quot; id=&quot;markdown-toc-should-you-use-it&quot;&gt;Should You Use It?&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;h2 id=&quot;why-onedrive-is-the-path-of-least-resistance&quot;&gt;Why OneDrive Is the Path of Least Resistance&lt;/h2&gt;

&lt;p&gt;On Windows, OneDrive is already installed, already signed in, and already syncing something. There is no client to download and no account to create. If your whole Obsidian life happens on Windows desktops and laptops, you could stop reading after the next section and be fine.&lt;/p&gt;

&lt;p&gt;Obsidian’s own help page says about as much. It calls OneDrive “a popular cloud storage option for Windows and macOS users” and then, in the same breath, says it “has limitations on Android and isn’t officially supported for syncing Obsidian vaults on iOS.”&lt;/p&gt;

&lt;p&gt;That sentence is the whole reason for this post. Desktop is easy if you flip one setting. Phones are where it starts costing you time.&lt;/p&gt;

&lt;h2 id=&quot;files-on-demand-is-why-your-notes-vanished&quot;&gt;Files On-Demand Is Why Your Notes Vanished&lt;/h2&gt;

&lt;p&gt;OneDrive’s Files On-Demand shows every file in File Explorer whether or not the bytes are actually on your disk. A blue cloud icon means the file is online-only. A green check means it’s on this machine for now. A solid green circle with a white check means it is pinned and stays.&lt;/p&gt;

&lt;p&gt;Obsidian needs the third one. It watches the vault folder and expects every file it sees to be a file it can read right now. An online-only note is a promise, not a file. Search misses it. Links to it look broken. A plugin tries to load its own settings and gets nothing. And none of this throws an error. It just looks like Obsidian forgot things.&lt;/p&gt;

&lt;p&gt;The fix on Windows takes ten seconds:&lt;/p&gt;

&lt;ol&gt;
  &lt;li&gt;Open File Explorer and find your vault folder.&lt;/li&gt;
  &lt;li&gt;Right-click it.&lt;/li&gt;
  &lt;li&gt;Choose &lt;strong&gt;Always keep on this device&lt;/strong&gt;.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;img src=&quot;/images/2026/obsidian-onedrive-sync-body-1.jpg&quot; alt=&quot;Files On-Demand Is Why Your Notes Vanished&quot; srcset=&quot;            /assets/resized/480/obsidian-onedrive-sync-body-1.jpg 480w,            /assets/resized/800/obsidian-onedrive-sync-body-1.jpg 800w,            /assets/resized/1400/obsidian-onedrive-sync-body-1.jpg 1400w,    &quot; loading=&quot;lazy&quot; /&gt;&lt;/p&gt;

&lt;p&gt;Wait for the icons to go solid green. Now it’s a real folder full of real files, and Obsidian stops acting psychotic.&lt;/p&gt;

&lt;p&gt;Microsoft also removed the old on/off switch for Files On-Demand. What’s left is two buttons, buried with the rest of OneDrive’s settings: click the OneDrive cloud icon in the system tray, open &lt;strong&gt;Settings&lt;/strong&gt;, and look under &lt;strong&gt;Sync and backup&lt;/strong&gt; → &lt;strong&gt;Advanced settings&lt;/strong&gt;. There you get &lt;strong&gt;Free up disk space&lt;/strong&gt; (the default) and &lt;strong&gt;Download all files&lt;/strong&gt;, which is the old “off.” If your whole OneDrive is small, Download all files is the blunt fix. If it is a terabyte of family photos, pin just the vault.&lt;/p&gt;

&lt;h3 id=&quot;storage-sense-will-undo-your-work&quot;&gt;Storage Sense will undo your work&lt;/h3&gt;

&lt;p&gt;Here is the trap after the trap. Windows has a feature called Storage Sense that frees up disk space on a schedule, and Microsoft’s documentation says that with it turned on, locally available OneDrive files “will become online-only files after the time period you’ve selected.”&lt;/p&gt;

&lt;p&gt;So you open your vault every day, everything is fine, and then a note you haven’t touched since spring comes back blank. That’s why pinning the vault folder matters: Storage Sense leaves pinned files alone and only reaps the ones that are merely local. Don’t just open the notes and assume they stay downloaded. The setting lives at &lt;strong&gt;Settings → System → Storage → Storage Sense&lt;/strong&gt;, where a dropdown decides how long an unpinned OneDrive file gets to stay on disk. Set it to never, or at least to longer than your oldest note is likely to sit unopened.&lt;/p&gt;

&lt;h2 id=&quot;onedrive-on-a-mac&quot;&gt;OneDrive on a Mac&lt;/h2&gt;

&lt;p&gt;Mac users end up on OneDrive because work gave them Microsoft 365, or because they also have a Windows machine and want one cloud. It’s a real combination, and almost every Obsidian guide skips it.&lt;/p&gt;

&lt;p&gt;The thing to know is that on macOS 12.3 and later, you can’t turn Files On-Demand off. Microsoft moved OneDrive on the Mac to Apple’s File Provider framework, and its own support page says Files On-Demand “must always be enabled.”&lt;/p&gt;

&lt;p&gt;That runs head-on into Obsidian’s advice to avoid the feature, and the way out is the pin: Obsidian doesn’t care about Files On-Demand itself, it cares that no vault file is ever online-only. Pinning the vault is what “avoid” looks like when the off switch is gone.&lt;/p&gt;

&lt;p&gt;Your OneDrive folder lives at &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;~/Library/CloudStorage/OneDrive-Personal&lt;/code&gt;, macOS controls it, and you can’t move it.&lt;/p&gt;

&lt;p&gt;The fix is the same as Windows, with one navigation trap: &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;~/Library&lt;/code&gt; is hidden, so don’t go hunting for the folder by browsing. OneDrive shows up in the Finder sidebar under &lt;strong&gt;Locations&lt;/strong&gt;, or press &lt;strong&gt;Cmd+Shift+G&lt;/strong&gt; and paste the path. Right-click the vault folder, pick &lt;strong&gt;Always keep on this device&lt;/strong&gt;, then open the vault in Obsidian from that &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;CloudStorage&lt;/code&gt; path. Since you can’t disable the feature globally, pinning the vault is the only defense you get, so do it before you open the vault, not after the first blank note.&lt;/p&gt;

&lt;h2 id=&quot;what-onedrive-refuses-to-sync&quot;&gt;What OneDrive Refuses to Sync&lt;/h2&gt;

&lt;p&gt;OneDrive has opinions about filenames. Obsidian already refuses most of these characters in note titles, so this usually bites attachments and files that got into the vault some other way: a downloaded PDF, a clipped image, a folder you dragged in from Finder.&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;/images/2026/obsidian-onedrive-sync-body-2.jpg&quot; alt=&quot;What OneDrive Refuses to Sync&quot; srcset=&quot;            /assets/resized/480/obsidian-onedrive-sync-body-2.jpg 480w,            /assets/resized/800/obsidian-onedrive-sync-body-2.jpg 800w,            /assets/resized/1400/obsidian-onedrive-sync-body-2.jpg 1400w,    &quot; loading=&quot;lazy&quot; /&gt;&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;&lt;strong&gt;Characters&lt;/strong&gt; like &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;&quot; * : &amp;lt; &amp;gt; ? / \ |&lt;/code&gt; are not allowed. Leading and trailing spaces are out too.&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Reserved names&lt;/strong&gt; cover &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;CON&lt;/code&gt;, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;PRN&lt;/code&gt;, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;AUX&lt;/code&gt;, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;NUL&lt;/code&gt;, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;COM0&lt;/code&gt; through &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;COM9&lt;/code&gt;, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;LPT0&lt;/code&gt; through &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;LPT9&lt;/code&gt;, and anything with &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;_vti_&lt;/code&gt; in it. You probably don’t have a note called &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;AUX&lt;/code&gt;. Someone does.&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Path length&lt;/strong&gt; caps the full path, filename included, at 400 characters. “Decoded” just means OneDrive counts what you’d see in Explorer, not the URL-escaped version with &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;%20&lt;/code&gt; standing in for every space. Deeply nested folders plus long note titles get there faster than you would think.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;.obsidian&lt;/code&gt; folder with your settings and plugins is not on Microsoft’s blocked list, so it rides along. That means plugin settings sync between desktops, which is mostly what you want.&lt;/p&gt;

&lt;h3 id=&quot;the-duplicate-file-with-your-computers-name-in-it&quot;&gt;The duplicate file with your computer’s name in it&lt;/h3&gt;

&lt;p&gt;When OneDrive can’t reconcile two versions of a file, it keeps both and adds your computer’s name to one of them. That is where &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;Daily Note-DESKTOP-7QK2M1B.md&lt;/code&gt; comes from. It does not ask. It is on you to notice, open both, and fold them back together.&lt;/p&gt;

&lt;p&gt;Microsoft’s support page ties this symptom specifically to stale credentials and says to clear the saved ones. On Windows, search the Start menu for &lt;strong&gt;Credential Manager&lt;/strong&gt;, open &lt;strong&gt;Windows Credentials&lt;/strong&gt;, and remove the entries starting with &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;MicrosoftAccount&lt;/code&gt;. On a Mac, open &lt;strong&gt;Keychain Access&lt;/strong&gt;, search for &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;OneDrive Cached Credential&lt;/code&gt;, and delete what comes up. If you keep getting computer-named duplicates on notes you only edited in one place, try that before blaming Obsidian.&lt;/p&gt;

&lt;p&gt;The real cause is usually simpler: the same note edited on two machines before either finished syncing. Let OneDrive finish before you close the laptop lid. I said the same thing in the &lt;a href=&quot;https://www.stephanmiller.com/obsidian-dropbox-sync/&quot;&gt;Dropbox sync post&lt;/a&gt;, because every folder-sync service fails exactly this way.&lt;/p&gt;

&lt;h2 id=&quot;iphone-and-ipad-remotely-save-or-nothing&quot;&gt;iPhone and iPad: Remotely Save or Nothing&lt;/h2&gt;

&lt;p&gt;Obsidian doesn’t officially support OneDrive vaults on iOS, and the vault-folder route that works on desktop is not an option there. So the answer on iOS is a plugin that syncs from inside Obsidian, and the plugin is &lt;a href=&quot;https://github.com/remotely-save/remotely-save&quot;&gt;Remotely Save&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;The catch bites almost everybody, and the setup flow doesn’t warn you about it. On the free tier, Remotely Save’s README says it can only “read and write files in your OneDrive’s &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;/Apps/remotely-save&lt;/code&gt; folder.” That’s the App Folder. If your vault already lives in &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;Documents\Obsidian&lt;/code&gt; because Windows put it there, the free plugin can’t see it.&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;/images/2026/obsidian-onedrive-sync-body-3.jpg&quot; alt=&quot;iPhone and iPad: Remotely Save or Nothing&quot; srcset=&quot;            /assets/resized/480/obsidian-onedrive-sync-body-3.jpg 480w,            /assets/resized/800/obsidian-onedrive-sync-body-3.jpg 800w,            /assets/resized/1400/obsidian-onedrive-sync-body-3.jpg 1400w,    &quot; loading=&quot;lazy&quot; /&gt;&lt;/p&gt;

&lt;p&gt;You have two options:&lt;/p&gt;

&lt;ol&gt;
  &lt;li&gt;&lt;strong&gt;Move the vault into the App Folder.&lt;/strong&gt; Close Obsidian first. Then, in File Explorer, move the vault folder into &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;OneDrive\Apps\remotely-save\&lt;/code&gt; (create it if it isn’t there yet). Right-click the vault in its new home and make sure &lt;strong&gt;Always keep on this device&lt;/strong&gt; is still set, because a move is exactly the kind of thing that shakes a pin loose. Reopen Obsidian, use &lt;strong&gt;Open another vault → Open folder as vault&lt;/strong&gt;, and point it at the new path. Then let Remotely Save on the phone sync against that. Free, and a little weird, because your notes now live in a folder named after a plugin.&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Use Remotely Save PRO.&lt;/strong&gt; PRO adds OneDrive with full access, which reaches the root of your drive, so the vault can stay where it is. PRO also adds smarter conflict handling that can merge small markdown changes instead of making a second file.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;PRO needs its own account at remotelysave.com. It is “FREE while beta before Jan 1, 2027,” and the pricing page says “the price will be determined later.” I keep repeating that across this series because “free right now” and “free” are different words. If you pick PRO as your permanent answer, you are agreeing to a number nobody has announced.&lt;/p&gt;

&lt;p&gt;The rest of the setup is the same as any other Remotely Save backend: install the plugin, pick OneDrive in its settings, authorize with your Microsoft account, and run the first sync with a backup of the vault sitting somewhere else. If you have never installed a community plugin, the &lt;a href=&quot;https://www.stephanmiller.com/how-to-install-obsidian-plugins/&quot;&gt;plugin install walkthrough&lt;/a&gt; covers it, and the &lt;a href=&quot;https://www.stephanmiller.com/sync-obsidian-iphone-ipad-free/&quot;&gt;iPhone and iPad sync guide&lt;/a&gt; explains why iOS is the hard case for every method, not just this one.&lt;/p&gt;

&lt;h2 id=&quot;android-remotely-save-or-a-sync-app&quot;&gt;Android: Remotely Save or a Sync App&lt;/h2&gt;

&lt;p&gt;Android is looser than iOS. Obsidian’s help page calls OneDrive’s Android support limited, but it doesn’t rule it out. You have two routes.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Remotely Save&lt;/strong&gt;, exactly as above. Same App Folder limit on the free tier, same PRO option. If you already set it up on an iPhone, doing the same on Android is five minutes.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A folder sync app&lt;/strong&gt; that mirrors a OneDrive folder to a real folder on the phone, which Obsidian then opens as a normal vault. The one built for this is OneSync: Autosync for OneDrive by MetaCtrl. The free version shows ads and caps uploads at 10MB per file, which a text vault will never hit and a vault full of PDFs will. Syncing more than one folder pair is part of the paid upgrade.&lt;/p&gt;

&lt;p&gt;If you go the sync app route, where the vault lives on the phone matters. Obsidian on Android asks whether to keep the vault in app storage or device storage, and an outside sync app can only reach device storage. I covered that choice in the &lt;a href=&quot;https://www.stephanmiller.com/sync-obsidian-android-free/&quot;&gt;Android sync post&lt;/a&gt;.&lt;/p&gt;

&lt;h2 id=&quot;personal-vs-business-your-work-account-is-out&quot;&gt;Personal vs. Business: Your Work Account Is Out&lt;/h2&gt;

&lt;p&gt;If the OneDrive you have is the one your employer gave you, stop here for mobile.&lt;/p&gt;

&lt;p&gt;The README’s wording: it “only works for ‘OneDrive for personal’, and not works for ‘OneDrive for Business’ (yet).” That’s true for both the free tier and PRO. The paid tier doesn’t add Business. It adds the root folder of a personal account.&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;/images/2026/obsidian-onedrive-sync-body-4.jpg&quot; alt=&quot;Personal vs. Business: Your Work Account Is Out&quot; srcset=&quot;            /assets/resized/480/obsidian-onedrive-sync-body-4.jpg 480w,            /assets/resized/800/obsidian-onedrive-sync-body-4.jpg 800w,            /assets/resized/1400/obsidian-onedrive-sync-body-4.jpg 1400w,    &quot; loading=&quot;lazy&quot; /&gt;&lt;/p&gt;

&lt;p&gt;Desktop is a different story. If the OneDrive client on your work laptop syncs a folder, Obsidian can open a vault inside it, and everything in the Files On-Demand section applies. Whether you &lt;em&gt;should&lt;/em&gt; keep personal notes in your employer’s cloud is a question for you and your IT department, not for me.&lt;/p&gt;

&lt;h2 id=&quot;the-cost-math-if-you-already-pay-for-microsoft-365&quot;&gt;The Cost Math If You Already Pay for Microsoft 365&lt;/h2&gt;

&lt;p&gt;This is the argument for OneDrive, and it holds up.&lt;/p&gt;

&lt;p&gt;Microsoft 365 prices in the US, as of this writing:&lt;/p&gt;

&lt;table&gt;
  &lt;thead&gt;
    &lt;tr&gt;
      &lt;th&gt;Plan&lt;/th&gt;
      &lt;th&gt;Storage&lt;/th&gt;
      &lt;th&gt;Monthly&lt;/th&gt;
      &lt;th&gt;Yearly&lt;/th&gt;
    &lt;/tr&gt;
  &lt;/thead&gt;
  &lt;tbody&gt;
    &lt;tr&gt;
      &lt;td&gt;Basic&lt;/td&gt;
      &lt;td&gt;100GB&lt;/td&gt;
      &lt;td&gt;$1.99&lt;/td&gt;
      &lt;td&gt;$19.99&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Personal&lt;/td&gt;
      &lt;td&gt;1TB&lt;/td&gt;
      &lt;td&gt;$9.99&lt;/td&gt;
      &lt;td&gt;$99.99&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Family&lt;/td&gt;
      &lt;td&gt;1TB each, up to 6 people&lt;/td&gt;
      &lt;td&gt;$12.99&lt;/td&gt;
      &lt;td&gt;$129.99&lt;/td&gt;
    &lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;

&lt;p&gt;Personal and Family went up in January 2025 when Microsoft bundled Copilot into them, so if your mental number is older than that, it’s low.&lt;/p&gt;

&lt;p&gt;Compare that to Obsidian Sync, which is $4 a month billed annually or $5 month to month for the Standard plan, with 1GB of storage and one vault. If you already pay for Microsoft 365 for Word and Excel, the storage is sunk cost and OneDrive sync is free. If you’re paying for Microsoft 365 &lt;em&gt;only&lt;/em&gt; for the storage, Basic at $19.99 a year is still cheaper than Obsidian Sync. But you’re buying that cheaper storage with your own time.&lt;/p&gt;

&lt;p&gt;OneDrive costs nothing extra on desktop. On phones it costs either a folder move, a PRO subscription with an unannounced price, or ads in a sync app. Obsidian Sync costs $48 a year and none of the above. So which one is cheaper? That depends on what an afternoon of your time is worth.&lt;/p&gt;

&lt;h2 id=&quot;should-you-use-it&quot;&gt;Should You Use It?&lt;/h2&gt;

&lt;p&gt;On Windows desktops, yes, if you already have it. Pin the vault folder, check Storage Sense, watch the filenames on your attachments, and it will mostly stay out of your way. Mac plus Windows works too, as long as you pin the vault on the Mac before you open it.&lt;/p&gt;

&lt;p&gt;Once a phone is involved, it depends on the account. A personal OneDrive and an iPhone means Remotely Save, and the free tier means moving your vault into the App Folder. A work account and an iPhone means pick another method entirely.&lt;/p&gt;

&lt;p&gt;I don’t sync my own vault with OneDrive. My laptops use Dropbox and my phone and iPad use Remotely Save, which I covered in the pillar guide. But if I were starting on a fresh Windows machine with a Microsoft 365 subscription, I would not go download something else just to avoid OneDrive. I would pin the folder and move on.&lt;/p&gt;

&lt;p&gt;Whatever you pick, back up the vault somewhere that isn’t OneDrive. Sync is not backup. A tool that copies a deleted folder to every device you own in thirty seconds isn’t protecting anything.&lt;/p&gt;
</description>
        <pubDate>Thu, 01 Oct 2026 08:00:00 -0500</pubDate>
        <link>https://www.stephanmiller.com/obsidian-onedrive-sync/</link>
        <guid isPermaLink="true">https://www.stephanmiller.com/obsidian-onedrive-sync/</guid>
        
        <category>Obsidian sync</category>
        
        <category>OneDrive sync</category>
        
        <category>Obsidian OneDrive</category>
        
        <category>Files On-Demand</category>
        
        <category>Obsidian Windows</category>
        
        <category>Storage Sense</category>
        
        
        <category>obsidian</category>
        
      </item>
    
      <item>
        <title>Forced Connections: One Way to Get Original Ideas Out of AI</title>
        <description>&lt;p&gt;I have a real problem. I write &lt;a href=&quot;https://www.stephanmiller.com/series/model-buzz-report/&quot;&gt;a weekly roundup&lt;/a&gt; of which new AI models are worth paying attention to, and I want to turn it into an email newsletter. Which means I need people to subscribe to it, which means I need ideas for getting people to subscribe to it.&lt;/p&gt;

&lt;p&gt;So I asked a model. Claude Sonnet 5, through OpenRouter, no system prompt, one line of context and then the question:&lt;/p&gt;

&lt;blockquote&gt;
  &lt;p&gt;How can I get more readers to subscribe to it?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;It came back with signup forms at the end of high-traffic posts, exit-intent popups, “publish 3-4 issues before heavily promoting,” a thread on X, r/LocalLLaMA, Hacker News, LinkedIn, newsletter swaps, and a single-field signup form. Every item was correct. I’ve read that exact list in maybe forty blog posts. It then asked me what my traffic source was so it could prioritize, which is the chatbot equivalent of the guy at the hardware store asking what you’re building.&lt;/p&gt;

&lt;p&gt;Fine. That was a question, and a question gets you the average. So I asked it to be creative:&lt;/p&gt;

&lt;blockquote&gt;
  &lt;p&gt;Give me creative, unusual, out-of-the-box ideas for getting more readers to subscribe to it.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;This one was more fun. “Model Death Certificates,” obituaries for overhyped models. A “Which new AI model matches your vibe today?” quiz microsite. A referral leaderboard where the prize is absurd. A recurring mascot. An “anti-hype pledge.” It’s the list you get when you search “creative newsletter growth ideas.” I did not get creative ideas. I got the average of everything ever labeled creative.&lt;/p&gt;

&lt;p&gt;Then I stopped asking it for ideas at all. I had a script pick an object at random from a list of twenty things. It picked three: a pressure cooker, a fire extinguisher and a metronome. I started with the fire extinguisher and told the model to list ten literal attributes of a fire extinguisher, force every single one onto my subscription problem, flag its own generic results, and keep what survived.&lt;/p&gt;

&lt;p&gt;Two of the survivors:&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;&lt;strong&gt;Sell the silence.&lt;/strong&gt; Market the newsletter as the thing you check only when a model actually matters. It came from &lt;em&gt;“ignored until there’s a fire.”&lt;/em&gt;&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Sell it as insurance.&lt;/strong&gt; “Subscribe once. Skip every issue if you want. Just be covered when a model actually matters.” That one came from &lt;em&gt;“reassurance from mere presence.”&lt;/em&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;img src=&quot;/images/2026/force-two-unrelated-things-together-body-1.jpg&quot; alt=&quot;Introduction&quot; srcset=&quot;            /assets/resized/480/force-two-unrelated-things-together-body-1.jpg 480w,            /assets/resized/800/force-two-unrelated-things-together-body-1.jpg 800w,            /assets/resized/1400/force-two-unrelated-things-together-body-1.jpg 1400w,    &quot; loading=&quot;lazy&quot; /&gt;&lt;/p&gt;

&lt;p&gt;Nobody writes that in a newsletter growth article, because it’s the exact opposite of what newsletter growth articles preach. It tells people they don’t have to read you. And I think it might be the best idea of the three runs.&lt;/p&gt;

&lt;p&gt;The only thing that changed was that I stopped asking it a question and handed it a process instead.&lt;/p&gt;

&lt;ul id=&quot;markdown-toc&quot;&gt;
  &lt;li&gt;&lt;a href=&quot;#what-just-happened&quot; id=&quot;markdown-toc-what-just-happened&quot;&gt;What Just Happened&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;#the-human-version-run-this-first&quot; id=&quot;markdown-toc-the-human-version-run-this-first&quot;&gt;The Human Version (Run This First)&lt;/a&gt;    &lt;ul&gt;
      &lt;li&gt;&lt;a href=&quot;#1-write-the-problem-as-one-line&quot; id=&quot;markdown-toc-1-write-the-problem-as-one-line&quot;&gt;1. Write the problem as one line&lt;/a&gt;&lt;/li&gt;
      &lt;li&gt;&lt;a href=&quot;#2-pick-an-object-you-did-not-choose&quot; id=&quot;markdown-toc-2-pick-an-object-you-did-not-choose&quot;&gt;2. Pick an object you did not choose&lt;/a&gt;&lt;/li&gt;
      &lt;li&gt;&lt;a href=&quot;#3-list-ten-literal-attributes&quot; id=&quot;markdown-toc-3-list-ten-literal-attributes&quot;&gt;3. List ten literal attributes&lt;/a&gt;&lt;/li&gt;
      &lt;li&gt;&lt;a href=&quot;#4-force-every-attribute-onto-the-problem&quot; id=&quot;markdown-toc-4-force-every-attribute-onto-the-problem&quot;&gt;4. Force every attribute onto the problem&lt;/a&gt;&lt;/li&gt;
      &lt;li&gt;&lt;a href=&quot;#5-cross-out-everything-you-have-read-before&quot; id=&quot;markdown-toc-5-cross-out-everything-you-have-read-before&quot;&gt;5. Cross out everything you have read before&lt;/a&gt;&lt;/li&gt;
      &lt;li&gt;&lt;a href=&quot;#6-do-it-with-two-or-three-objects-then-breed-the-survivors&quot; id=&quot;markdown-toc-6-do-it-with-two-or-three-objects-then-breed-the-survivors&quot;&gt;6. Do it with two or three objects, then breed the survivors&lt;/a&gt;&lt;/li&gt;
    &lt;/ul&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;#handing-the-model-the-process&quot; id=&quot;markdown-toc-handing-the-model-the-process&quot;&gt;Handing the Model the Process&lt;/a&gt;    &lt;ul&gt;
      &lt;li&gt;&lt;a href=&quot;#what-came-out&quot; id=&quot;markdown-toc-what-came-out&quot;&gt;What came out&lt;/a&gt;&lt;/li&gt;
      &lt;li&gt;&lt;a href=&quot;#one-caveat-about-the-judge&quot; id=&quot;markdown-toc-one-caveat-about-the-judge&quot;&gt;One caveat about the judge&lt;/a&gt;&lt;/li&gt;
    &lt;/ul&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;#where-this-came-from&quot; id=&quot;markdown-toc-where-this-came-from&quot;&gt;Where This Came From&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;#the-crossover-move&quot; id=&quot;markdown-toc-the-crossover-move&quot;&gt;The Crossover Move&lt;/a&gt;    &lt;ul&gt;
      &lt;li&gt;&lt;a href=&quot;#running-it&quot; id=&quot;markdown-toc-running-it&quot;&gt;Running it&lt;/a&gt;&lt;/li&gt;
    &lt;/ul&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;#why-not-just-ask-for-fifty-ideas&quot; id=&quot;markdown-toc-why-not-just-ask-for-fifty-ideas&quot;&gt;Why Not Just Ask for Fifty Ideas?&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;#where-it-breaks&quot; id=&quot;markdown-toc-where-it-breaks&quot;&gt;Where It Breaks&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;#so-what-did-i-actually-get&quot; id=&quot;markdown-toc-so-what-did-i-actually-get&quot;&gt;So What Did I Actually Get?&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;h2 id=&quot;what-just-happened&quot;&gt;What Just Happened&lt;/h2&gt;

&lt;p&gt;In the &lt;a href=&quot;/youre-using-ai-like-a-vending-machine/&quot;&gt;first post in this series&lt;/a&gt; I said a question gets you an answer and a move gets you something different. The obvious objection is that a “move” is just a more specific question. The second prompt above is the test case. “Give me creative, unusual, out-of-the-box ideas” &lt;em&gt;is&lt;/em&gt; a more specific question. It got a more specific average.&lt;/p&gt;

&lt;p&gt;The fire extinguisher did something different.&lt;/p&gt;

&lt;p&gt;It did not supply the idea. A fire extinguisher knows nothing about newsletters. What it supplied was ten starting points that the model would never have started from, because none of them are anywhere near the words “grow a newsletter” in anything it was trained on. “Wall-mounted in a fixed location.” “Needs a periodic inspection tag.” “Different classes for different fires.” Each attribute forces the model to begin its reasoning somewhere other than the middle, and then go from there back to the problem. Some of the ideas end up back in the middle anyway. Some of them don’t.&lt;/p&gt;

&lt;p&gt;You’re changing where the search starts.&lt;/p&gt;

&lt;h2 id=&quot;the-human-version-run-this-first&quot;&gt;The Human Version (Run This First)&lt;/h2&gt;

&lt;p&gt;This is a pen-and-paper technique and it has been one for close to seventy years. Do it once by hand before you ever hand it to a model, because the model version is just this process written down.&lt;/p&gt;

&lt;h3 id=&quot;1-write-the-problem-as-one-line&quot;&gt;1. Write the problem as one line&lt;/h3&gt;

&lt;p&gt;Not a paragraph. “Get more people to subscribe to my newsletter.” “Name the new feature.” “Figure out what the second act of this story is.” If you can’t get it to one line you have two problems and should pick one.&lt;/p&gt;

&lt;h3 id=&quot;2-pick-an-object-you-did-not-choose&quot;&gt;2. Pick an object you did not choose&lt;/h3&gt;

&lt;p&gt;If you pick the object, you’ll pick one that already seems relevant, and relevance is exactly what you’re trying to escape. You can open a catalog to a random page, use the third thing to your left, a random word generator, the last noun on page 50 of whatever book is closest, or a list of twenty objects and a die.&lt;/p&gt;

&lt;p&gt;Whatever it lands on, you keep it. No re-rolls.&lt;/p&gt;

&lt;h3 id=&quot;3-list-ten-literal-attributes&quot;&gt;3. List ten literal attributes&lt;/h3&gt;

&lt;p&gt;&lt;img src=&quot;/images/2026/force-two-unrelated-things-together-body-2.jpg&quot; alt=&quot;3. List ten literal attributes&quot; srcset=&quot;            /assets/resized/480/force-two-unrelated-things-together-body-2.jpg 480w,            /assets/resized/800/force-two-unrelated-things-together-body-2.jpg 800w,            /assets/resized/1400/force-two-unrelated-things-together-body-2.jpg 1400w,    &quot; loading=&quot;lazy&quot; /&gt;&lt;/p&gt;

&lt;p&gt;Literal. Not metaphors yet. Cover all of these:&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;&lt;strong&gt;Parts.&lt;/strong&gt; What it’s made of, what you can see and touch.&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Use.&lt;/strong&gt; How you operate it, what it does, how long it takes.&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Location.&lt;/strong&gt; Where it lives, who’s near it.&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Failure.&lt;/strong&gt; What goes wrong with it, what wears out, what you have to maintain.&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Feelings.&lt;/strong&gt; How people feel about it. Nostalgia, fear, annoyance, indifference.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Do not skip the last two. I’ll show you why in a minute.&lt;/p&gt;

&lt;h3 id=&quot;4-force-every-attribute-onto-the-problem&quot;&gt;4. Force every attribute onto the problem&lt;/h3&gt;

&lt;p&gt;One idea per attribute, and take the attribute literally. “Needs a periodic inspection tag” becomes “send a literal quarterly ‘inspection’ email asking subscribers to reconfirm interest.” That’s the model’s actual output, by the way. Write it down.&lt;/p&gt;

&lt;p&gt;The rule that makes the whole thing work: &lt;strong&gt;you are not allowed to skip an attribute because the mapping is awkward.&lt;/strong&gt; The awkward ones are the point.&lt;/p&gt;

&lt;h3 id=&quot;5-cross-out-everything-you-have-read-before&quot;&gt;5. Cross out everything you have read before&lt;/h3&gt;

&lt;p&gt;Go down the list and mark every idea that could have appeared in a normal article about your problem. Most of them will. Whatever is left is what you came for.&lt;/p&gt;

&lt;h3 id=&quot;6-do-it-with-two-or-three-objects-then-breed-the-survivors&quot;&gt;6. Do it with two or three objects, then breed the survivors&lt;/h3&gt;

&lt;p&gt;One object gives you one set of starting points. Two or three give you survivors from different directions, and now: take two survivors from different objects and make one idea that needs both of them to exist. Not “pick the best one” or “improve one.” A child that neither parent could have produced alone.&lt;/p&gt;

&lt;h2 id=&quot;handing-the-model-the-process&quot;&gt;Handing the Model the Process&lt;/h2&gt;

&lt;p&gt;Here’s the prompt I actually ran, word for word, with the object swapped in:&lt;/p&gt;

&lt;div class=&quot;language-text highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;I write a tech blog and I&apos;m starting a weekly email newsletter that rounds up
which new AI models are actually worth paying attention to. I want ideas for
getting more readers to subscribe. Don&apos;t answer that directly. Run this process
instead, and show every step:

1. The object is: fire extinguisher. List ten attributes of it. Physical parts,
   how it&apos;s used, where it lives, what goes wrong with it, how people feel about
   it. Literal attributes, not metaphors yet.
2. For each attribute, force a literal mapping onto the subscription problem.
   One idea per attribute. Do not skip an attribute because the mapping is
   awkward. The awkward ones are the point.
3. Mark any idea that could have come from a normal newsletter-growth article
   with [GENERIC]. Be honest.
4. Pick the two ideas that are least generic and still actually doable by one
   person, and say what the first concrete step would be for each.
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;Notice what’s in there and what isn’t. It never says “be creative.” It never says “think outside the box.” It never says “avoid generic ideas,” which is a whole separate problem I’ll get to later in the series. It’s steps 3 through 5 of the human version, written down in order, with the one line in the middle that stops the model from taking the easy exit.&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;/images/2026/force-two-unrelated-things-together-body-4.jpg&quot; alt=&quot;Handing the Model the Process&quot; srcset=&quot;            /assets/resized/480/force-two-unrelated-things-together-body-4.jpg 480w,            /assets/resized/800/force-two-unrelated-things-together-body-4.jpg 800w,            /assets/resized/1400/force-two-unrelated-things-together-body-4.jpg 1400w,    &quot; loading=&quot;lazy&quot; /&gt;&lt;/p&gt;

&lt;p&gt;I ran it three times with three objects: a pressure cooker, a fire extinguisher, and a metronome. Here’s a small script that does the same thing, if you want to run it against your own problem:&lt;/p&gt;

&lt;div class=&quot;language-python highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;kn&quot;&gt;import&lt;/span&gt; &lt;span class=&quot;nn&quot;&gt;json&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;os&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;urllib&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;request&lt;/span&gt;

&lt;span class=&quot;n&quot;&gt;MODEL&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;s&quot;&gt;&quot;anthropic/claude-sonnet-5&quot;&lt;/span&gt;
&lt;span class=&quot;n&quot;&gt;PROBLEM&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;&quot;I write a tech blog and I&apos;m starting a weekly email newsletter that &quot;&lt;/span&gt;
           &lt;span class=&quot;s&quot;&gt;&quot;rounds up which new AI models are actually worth paying attention to. &quot;&lt;/span&gt;
           &lt;span class=&quot;s&quot;&gt;&quot;I want ideas for getting more readers to subscribe.&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;

&lt;span class=&quot;n&quot;&gt;STEPS&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;s&quot;&gt;&quot;&quot;&quot;Don&apos;t answer that directly. Run this process instead, and show every step:

1. The object is: {obj}. List ten attributes of it. Physical parts, how it&apos;s used,
   where it lives, what goes wrong with it, how people feel about it. Literal
   attributes, not metaphors yet.
2. For each attribute, force a literal mapping onto the problem. One idea per
   attribute. Do not skip an attribute because the mapping is awkward.
   The awkward ones are the point.
3. Mark any idea that could have come from a normal article on this problem
   with [GENERIC]. Be honest.
4. Pick the two ideas that are least generic and still doable by one person,
   and say what the first concrete step would be for each.&quot;&quot;&quot;&lt;/span&gt;

&lt;span class=&quot;k&quot;&gt;def&lt;/span&gt; &lt;span class=&quot;nf&quot;&gt;ask&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;prompt&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;):&lt;/span&gt;
    &lt;span class=&quot;n&quot;&gt;body&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;json&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;dumps&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;({&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;&quot;model&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;:&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;MODEL&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;s&quot;&gt;&quot;temperature&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;:&lt;/span&gt; &lt;span class=&quot;mf&quot;&gt;1.0&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt;
                       &lt;span class=&quot;s&quot;&gt;&quot;messages&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;:&lt;/span&gt; &lt;span class=&quot;p&quot;&gt;[{&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;&quot;role&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;:&lt;/span&gt; &lt;span class=&quot;s&quot;&gt;&quot;user&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;s&quot;&gt;&quot;content&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;:&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;prompt&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;}]}).&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;encode&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;()&lt;/span&gt;
    &lt;span class=&quot;n&quot;&gt;req&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;urllib&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;request&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;Request&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;
        &lt;span class=&quot;s&quot;&gt;&quot;https://openrouter.ai/api/v1/chat/completions&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;body&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt;
        &lt;span class=&quot;p&quot;&gt;{&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;&quot;Authorization&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;:&lt;/span&gt; &lt;span class=&quot;s&quot;&gt;&quot;Bearer &quot;&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;+&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;os&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;environ&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;[&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;&quot;OPENROUTER_API_KEY&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;],&lt;/span&gt;
         &lt;span class=&quot;s&quot;&gt;&quot;Content-Type&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;:&lt;/span&gt; &lt;span class=&quot;s&quot;&gt;&quot;application/json&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;})&lt;/span&gt;
    &lt;span class=&quot;k&quot;&gt;return&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;json&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;load&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;urllib&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;request&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;urlopen&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;req&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;timeout&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;mi&quot;&gt;300&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;))[&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;&quot;choices&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;][&lt;/span&gt;&lt;span class=&quot;mi&quot;&gt;0&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;][&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;&quot;message&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;][&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;&quot;content&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;]&lt;/span&gt;

&lt;span class=&quot;k&quot;&gt;for&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;obj&lt;/span&gt; &lt;span class=&quot;ow&quot;&gt;in&lt;/span&gt; &lt;span class=&quot;p&quot;&gt;[&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;&quot;pressure cooker&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;s&quot;&gt;&quot;fire extinguisher&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;s&quot;&gt;&quot;metronome&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;]:&lt;/span&gt;
    &lt;span class=&quot;n&quot;&gt;out&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;ask&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;PROBLEM&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;+&lt;/span&gt; &lt;span class=&quot;s&quot;&gt;&quot; &quot;&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;+&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;STEPS&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;nb&quot;&gt;format&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;obj&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;obj&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;))&lt;/span&gt;
    &lt;span class=&quot;nb&quot;&gt;open&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;obj&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;replace&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;&quot; &quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;s&quot;&gt;&quot;_&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;+&lt;/span&gt; &lt;span class=&quot;s&quot;&gt;&quot;.md&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;s&quot;&gt;&quot;w&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;).&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;write&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;out&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;Swap in your own problem. Swap in your own objects, and don’t choose them yourself. That’s the part the script is for.&lt;/p&gt;

&lt;h3 id=&quot;what-came-out&quot;&gt;What came out&lt;/h3&gt;

&lt;p&gt;Thirty forced ideas. By the model’s own count, about ten of them were clearly not generic, three more were borderline, and the rest were, in the metronome run’s own words, “generic growth advice in a costume.”&lt;/p&gt;

&lt;p&gt;The costumes tell you how the technique fails:&lt;/p&gt;

&lt;table&gt;
  &lt;thead&gt;
    &lt;tr&gt;
      &lt;th&gt;Attribute&lt;/th&gt;
      &lt;th&gt;Forced idea&lt;/th&gt;
      &lt;th&gt;Verdict&lt;/th&gt;
    &lt;/tr&gt;
  &lt;/thead&gt;
  &lt;tbody&gt;
    &lt;tr&gt;
      &lt;td&gt;Pressure cooker: lid locks shut&lt;/td&gt;
      &lt;td&gt;A modal that seals the article until you subscribe or decline&lt;/td&gt;
      &lt;td&gt;Generic. It’s a content-lock popup.&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Pressure cooker: builds pressure to cook faster&lt;/td&gt;
      &lt;td&gt;“Next batch replaces this list in 7 days”&lt;/td&gt;
      &lt;td&gt;Generic. It’s scarcity marketing.&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Fire extinguisher: wall-mounted, fixed location&lt;/td&gt;
      &lt;td&gt;Subscribe box in the exact same spot on every post&lt;/td&gt;
      &lt;td&gt;Generic. Basic conversion advice.&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Metronome: adjustable tempo weight&lt;/td&gt;
      &lt;td&gt;Pick a short or long version at signup&lt;/td&gt;
      &lt;td&gt;Generic. Everyone offers this.&lt;/td&gt;
    &lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;

&lt;p&gt;&lt;img src=&quot;/images/2026/force-two-unrelated-things-together-body-5.jpg&quot; alt=&quot;What came out&quot; srcset=&quot;            /assets/resized/480/force-two-unrelated-things-together-body-5.jpg 480w,            /assets/resized/800/force-two-unrelated-things-together-body-5.jpg 800w,            /assets/resized/1400/force-two-unrelated-things-together-body-5.jpg 1400w,    &quot; loading=&quot;lazy&quot; /&gt;&lt;/p&gt;

&lt;p&gt;Every one of those is a real attribute landing on something that already exists. The model went from a strange starting point and ended up back in the middle, the path of least resistance.&lt;/p&gt;

&lt;p&gt;And the survivors:&lt;/p&gt;

&lt;table&gt;
  &lt;thead&gt;
    &lt;tr&gt;
      &lt;th&gt;Attribute&lt;/th&gt;
      &lt;th&gt;Forced idea&lt;/th&gt;
    &lt;/tr&gt;
  &lt;/thead&gt;
  &lt;tbody&gt;
    &lt;tr&gt;
      &lt;td&gt;Fire extinguisher: ignored until there’s a fire&lt;/td&gt;
      &lt;td&gt;Sell the silence. Check it only when a model actually matters.&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Fire extinguisher: reassurance from mere presence&lt;/td&gt;
      &lt;td&gt;Sell it as insurance. Skip every issue, you’re still covered.&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Fire extinguisher: loud, messy discharge&lt;/td&gt;
      &lt;td&gt;One blunt, unhedged verdict per model, deliberately not balanced&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Metronome: gets turned off out of annoyance&lt;/td&gt;
      &lt;td&gt;Ask everyone who unsubscribes one question and publish the answers monthly&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Metronome: mixed feelings, eventually internalized and discarded&lt;/td&gt;
      &lt;td&gt;A graduation point. “Read 8 issues and you’ll be able to spot a hyped model yourself.”&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Pressure cooker: emotional baggage, nostalgia or fear&lt;/td&gt;
      &lt;td&gt;A fixed narrator voice people get attached to, not just the information&lt;/td&gt;
    &lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;

&lt;p&gt;Most of the survivors came out of the attributes about &lt;strong&gt;failure and feelings&lt;/strong&gt;: ignored until there’s a fire, turned off out of annoyance, mixed feelings, emotional baggage. The physical attributes (the lid, the tempo weight, the wall mount) mostly mapped back onto tactics that already have names. It’s why step 3 in the human version tells you not to skip the last two categories. Parts are easy to list, and they’re also the ones that map onto what you already know.&lt;/p&gt;

&lt;p&gt;The graduation point is my other favorite. Growth advice is built on keeping subscribers forever. That idea puts an end date on the relationship and uses the end date as the pitch.&lt;/p&gt;

&lt;h3 id=&quot;one-caveat-about-the-judge&quot;&gt;One caveat about the judge&lt;/h3&gt;

&lt;p&gt;The model graded its own work, and a model grading its own originality is a pretty weak judge. It was also clearly being hard on itself on purpose and overcorrected. I’d rather it cut too much than let everything through. But the [GENERIC] flag is a filter, not a verdict. You still read the list yourself, and you still get the final say.&lt;/p&gt;

&lt;h2 id=&quot;where-this-came-from&quot;&gt;Where This Came From&lt;/h2&gt;

&lt;p&gt;The theory is Arthur Koestler’s. In &lt;em&gt;The Act of Creation&lt;/em&gt; (1964) he argued that jokes, scientific discovery and art all run on the same machinery, which he called &lt;strong&gt;bisociation&lt;/strong&gt;: two ways of seeing that each make perfect sense on their own and are never normally used together. Hold both at once and the collision is the idea. A pun is bisociation. So, in his telling, is a scientific breakthrough.&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;/images/2026/force-two-unrelated-things-together-body-3.jpg&quot; alt=&quot;Where This Came From&quot; srcset=&quot;            /assets/resized/480/force-two-unrelated-things-together-body-3.jpg 480w,            /assets/resized/800/force-two-unrelated-things-together-body-3.jpg 800w,            /assets/resized/1400/force-two-unrelated-things-together-body-3.jpg 1400w,    &quot; loading=&quot;lazy&quot; /&gt;&lt;/p&gt;

&lt;p&gt;One of his best-known examples is Gutenberg. The part that’s documented is that the printing press adapted the screw press farmers already used for pressing grapes and olives. Gutenberg took a machine from the wine harvest and pointed it at metal type. That’s a clean case of two frames that came from two places. The tidier version, where Gutenberg has a flash of insight at a wine harvest and writes a letter about Minerva springing from his brain, is Koestler’s telling. Treat the combination as history and the eureka moment as a good story.&lt;/p&gt;

&lt;p&gt;Charles S. Whiting’s &lt;em&gt;Creative Thinking&lt;/em&gt; (1958) is where “forced relationships” is usually traced: take an item unrelated to the problem and force a connection anyway. An arbitrary thing, a forced mapping, and no permission to bail when the mapping gets weird.&lt;/p&gt;

&lt;p&gt;It has a lot of relatives. Edward de Bono’s random-word technique is the same move with a word instead of an object. The “Combine” step in SCAMPER is a gentler version. Morphological analysis is its systematic cousin, where you break the problem into parameters and walk every combination instead of letting chance pick. I’ll get to that one later in the series. The distinction that matters here: forced relationships wants &lt;em&gt;distance&lt;/em&gt;. The point is that the object has nothing to do with your problem.&lt;/p&gt;

&lt;h2 id=&quot;the-crossover-move&quot;&gt;The Crossover Move&lt;/h2&gt;

&lt;p&gt;So far this is one object at a time. The title of the post promises two things forced together.&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;/images/2026/force-two-unrelated-things-together-body-6.jpg&quot; alt=&quot;The Crossover Move&quot; srcset=&quot;            /assets/resized/480/force-two-unrelated-things-together-body-6.jpg 480w,            /assets/resized/800/force-two-unrelated-things-together-body-6.jpg 800w,            /assets/resized/1400/force-two-unrelated-things-together-body-6.jpg 1400w,    &quot; loading=&quot;lazy&quot; /&gt;&lt;/p&gt;

&lt;p&gt;Google DeepMind’s FunSearch (&lt;a href=&quot;https://pmc.ncbi.nlm.nih.gov/articles/PMC10794145/&quot;&gt;Romera-Paredes et al., &lt;em&gt;Nature&lt;/em&gt;, December 2023&lt;/a&gt;) found new results in mathematics by having a language model write programs, scoring them, and keeping the good ones in a database. The interesting part is how it builds each new prompt. It doesn’t hand the model its best program and say “improve this.” It samples two programs from the database, sorts them by score, labels them &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;v0&lt;/code&gt; and &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;v1&lt;/code&gt;, and asks for the next version. The paper found two programs worked better than one, with diminishing returns after that, and gives the reason in one sentence: combining several programs “enables the LLM to spot patterns across the different programs and generalize those.”&lt;/p&gt;

&lt;p&gt;AlphaEvolve (&lt;a href=&quot;https://arxiv.org/abs/2506.13131&quot;&gt;Novikov et al., 2025&lt;/a&gt;) is the bigger, newer version of the same idea, with a program database built to keep the parents diverse so ideas explored earlier can resurface later instead of getting lost.&lt;/p&gt;

&lt;p&gt;The evolutionary computation people have a name for this. It’s &lt;strong&gt;crossover&lt;/strong&gt;. Two parents, one child. They’ve been doing it for decades and nobody in the prompt-template world seems to have noticed, because they call it “program search” instead of “brainstorming.”&lt;/p&gt;

&lt;p&gt;FunSearch’s two parents come from the same island, so they’re related. It combines, but it doesn’t force distance. What I did below uses both halves: forced distance to &lt;em&gt;generate&lt;/em&gt; the parents, crossover to &lt;em&gt;merge&lt;/em&gt; them.&lt;/p&gt;

&lt;h3 id=&quot;running-it&quot;&gt;Running it&lt;/h3&gt;

&lt;p&gt;I took three survivors from two different objects and gave them back to the model:&lt;/p&gt;

&lt;div class=&quot;language-text highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;Three ideas came out of separate exercises, each forced from a different
unrelated object:

Parent A (from a fire extinguisher): Sell subscribing as insurance. &quot;Subscribe
once. Skip every issue if you want. Just be covered when a model actually matters.&quot;
Parent B (from a metronome): Build in an explicit graduation point. &quot;Read 8
issues and you&apos;ll be able to spot a hyped model yourself.&quot;
Parent C (from a metronome): Publish churn. Ask everyone who unsubscribes one
question and publish the anonymized answers monthly.

Do not pick the best one and do not refine any one of them. Write three child
ideas. Each child must be a single mechanism that needs at least two of the
parents to exist, and could not be built from either parent alone. Name which
parents each child takes from and what it takes.
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;The “do not pick the best one and do not refine” line is doing the same job as “do not skip an awkward attribute” did earlier. Without it, a model handed three ideas will rank them and polish the winner, because that’s what a helpful assistant does with a list.&lt;/p&gt;

&lt;p&gt;What came back:&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;/images/2026/force-two-unrelated-things-together-body-7.jpg&quot; alt=&quot;Running it&quot; srcset=&quot;            /assets/resized/480/force-two-unrelated-things-together-body-7.jpg 480w,            /assets/resized/800/force-two-unrelated-things-together-body-7.jpg 800w,            /assets/resized/1400/force-two-unrelated-things-together-body-7.jpg 1400w,    &quot; loading=&quot;lazy&quot; /&gt;&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;&lt;strong&gt;Graduated coverage (insurance + graduation).&lt;/strong&gt; At issue 8, a subscriber automatically drops to an insurance-only tier: silence unless a model genuinely matters. Graduation alone gives you no reason to stay after you graduate. Insurance alone has no moment where “skip whenever” becomes official.&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;The graduation exit log (graduation + churn).&lt;/strong&gt; People who leave &lt;em&gt;after&lt;/em&gt; issue 8 get a different exit question: not “why are you leaving” but “what made you confident enough to leave.” Those answers get published as proof the newsletter actually teaches something. Churn becomes a credential.&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Churn to reinsurance (insurance + churn).&lt;/strong&gt; The monthly published churn answers sit next to a one-click “reinstate your coverage” link. You were never really unsubscribed, just paused until it matters.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Graduated coverage is the insurance idea with a trigger attached, and the trigger is the thing the insurance idea was missing. I wouldn’t have gotten there from either object alone. And I definitely wouldn’t have gotten there from “give me creative ideas,” which got me a mascot.&lt;/p&gt;

&lt;h2 id=&quot;why-not-just-ask-for-fifty-ideas&quot;&gt;Why Not Just Ask for Fifty Ideas?&lt;/h2&gt;

&lt;p&gt;This is the obvious shortcut, and it doesn’t work.&lt;/p&gt;

&lt;p&gt;In 2024, Chenglei Si, Diyi Yang and Tatsunori Hashimoto ran a large study comparing research ideas from an LLM pipeline against ideas from over a hundred NLP researchers (&lt;a href=&quot;https://arxiv.org/abs/2409.04109&quot;&gt;arXiv:2409.04109&lt;/a&gt;). The LLM ideas were actually judged &lt;em&gt;more&lt;/em&gt; novel than the humans’, and slightly less feasible. But to get there the pipeline generated 4,000 seed ideas per topic, and when they deduplicated them, only about 5% survived. The share of new, non-duplicate ideas in each batch kept dropping as they generated more, until it plateaued. The authors list the lack of diversity in generation as an open problem.&lt;/p&gt;

&lt;p&gt;So asking for more doesn’t get you more. It gets you the same few ideas in different words, over and over, with the occasional new one. Volume is a terrible way to escape the middle.&lt;/p&gt;

&lt;p&gt;Forcing a starting point is the cheaper way out. Thirty forced ideas gave me roughly ten survivors. Four thousand unforced ones gave that study about two hundred. The two aren’t directly comparable (different task, different judge, different everything), but the direction is the point.&lt;/p&gt;

&lt;h2 id=&quot;where-it-breaks&quot;&gt;Where It Breaks&lt;/h2&gt;

&lt;p&gt;I’m not going to pretend this is the only way to do it, because the research that exists doesn’t let me.&lt;/p&gt;

&lt;p&gt;The one solid human study I found on distance points the other way. Joel Chan, Christian Schunn and colleagues looked at which sources of inspiration led to the most creative design ideas (&lt;a href=&quot;https://www.sciencedirect.com/science/article/pii/S0142694X14000611&quot;&gt;&lt;em&gt;Design Studies&lt;/em&gt;, 2015&lt;/a&gt;), and found that “conceptually closer rather than farther sources lead to more creative ideas,” consistently across different design problems. There was no support for the best ideas coming from the farthest sources.&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;/images/2026/force-two-unrelated-things-together-body-8.jpg&quot; alt=&quot;Where It Breaks&quot; srcset=&quot;            /assets/resized/480/force-two-unrelated-things-together-body-8.jpg 480w,            /assets/resized/800/force-two-unrelated-things-together-body-8.jpg 800w,            /assets/resized/1400/force-two-unrelated-things-together-body-8.jpg 1400w,    &quot; loading=&quot;lazy&quot; /&gt;&lt;/p&gt;

&lt;p&gt;That’s a different setup from this one. Their sources were examples of &lt;em&gt;other solutions&lt;/em&gt;, and a fire extinguisher is not a solution to anything. It’s a jig, a thing to push your thinking against. In the runs above, the random object generated plenty of candidates and most of them were junk. The crossover step, merging survivors that were close enough to fit together, is where the best idea came from. That’s consistent with Chan’s finding, not a contradiction of it.&lt;/p&gt;

&lt;p&gt;The other ways it breaks, from actually doing it:&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;&lt;strong&gt;Physical attributes map to existing tactics.&lt;/strong&gt; A lid that locks becomes a popup. If your list is all parts, you’ll get all costumes. The failure and feeling attributes are where most of the survivors were.&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;The mapping can be too loose.&lt;/strong&gt; If you let yourself (or the model) go metaphorical in step 2, anything maps to anything and the object stops doing work. Literal mappings are more constrained, which is what you want.&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;It solves a narrow kind of problem.&lt;/strong&gt; It’s great when you’re stuck in a rut of five versions of the same idea. It’s useless when you don’t have enough information yet. No fire extinguisher is going to tell me whether anybody wants an AI model newsletter in the first place.&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;You still have to pick.&lt;/strong&gt; Graduated coverage looks good to me because of what I know about the people who read this blog. The model doesn’t know any of that. The pick is yours.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2 id=&quot;so-what-did-i-actually-get&quot;&gt;So What Did I Actually Get?&lt;/h2&gt;

&lt;p&gt;A newsletter pitch I’d never have written on my own. It tells people they don’t have to read it, and it goes quiet after issue 8 for anyone who’s learned the skill. Does it work? I don’t know yet. The newsletter doesn’t exist yet.&lt;/p&gt;

&lt;p&gt;The part that changed how I use models is smaller than that, though. I used to think the fix for a bland answer was a better question. It isn’t. What got me somewhere new was running a boring, seventy-year-old, pen-and-paper process, writing its steps down in order, and handing the model the steps instead of the question.&lt;/p&gt;

&lt;p&gt;Next one: give yourself an arbitrary rule. It’s the most recommended creativity advice in existence, and there’s a reason to be suspicious of it.&lt;/p&gt;
</description>
        <pubDate>Wed, 30 Sep 2026 07:00:00 -0500</pubDate>
        <link>https://www.stephanmiller.com/force-two-unrelated-things-together/</link>
        <guid isPermaLink="true">https://www.stephanmiller.com/force-two-unrelated-things-together/</guid>
        
        <category>prompt engineering</category>
        
        <category>creative prompting</category>
        
        <category>newsletter growth</category>
        
        <category>lateral thinking AI</category>
        
        
        <category>creativity</category>
        
        <category>ai-agents</category>
        
      </item>
    
      <item>
        <title>Sonnet 5.5 vs Opus 5.5: The Cheap Claude Costs More</title>
        <description>&lt;p&gt;The price sheet caved this week. I don’t mean one lab and some polite matching. Anthropic and OpenAI both cut prices on September 22, a day after Xiaomi dropped a 28-cent open-weights model, and then Anthropic came back six days later with a new Sonnet it swears is cheaper to run. If you pay your own API bill, this was the best week of the year to be alive.&lt;/p&gt;

&lt;p&gt;Then I read the fine print, because that’s the job. The new “cheap” Claude costs more per task than the new expensive Claude, at least when you turn the effort dial all the way up. The cheapest good coding model on the board this week comes from a company Anthropic says built it partly by siphoning Claude conversations through OpenClaw, which is the same agent framework I run in a Docker container in my house. And the free model sitting at number two on OpenRouter won’t tell you who made it.&lt;/p&gt;

&lt;p&gt;So yes, prices went down. Read the meter anyway.&lt;/p&gt;

&lt;ul id=&quot;markdown-toc&quot;&gt;
  &lt;li&gt;&lt;a href=&quot;#everybody-blinked-at-once&quot; id=&quot;markdown-toc-everybody-blinked-at-once&quot;&gt;Everybody blinked at once&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;#opus-55-is-the-rare-launch-where-every-signal-agrees&quot; id=&quot;markdown-toc-opus-55-is-the-rare-launch-where-every-signal-agrees&quot;&gt;Opus 5.5 is the rare launch where every signal agrees&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;#the-cheap-claude-costs-more-than-the-expensive-claude&quot; id=&quot;markdown-toc-the-cheap-claude-costs-more-than-the-expensive-claude&quot;&gt;The cheap Claude costs more than the expensive Claude&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;#openai-just-showed-up-at-the-bottom-of-the-price-sheet&quot; id=&quot;markdown-toc-openai-just-showed-up-at-the-bottom-of-the-price-sheet&quot;&gt;OpenAI just showed up at the bottom of the price sheet&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;#the-cheapest-coding-model-has-a-distillation-problem&quot; id=&quot;markdown-toc-the-cheapest-coding-model-has-a-distillation-problem&quot;&gt;The cheapest coding model has a distillation problem&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;#the-cheapskate-picks&quot; id=&quot;markdown-toc-the-cheapskate-picks&quot;&gt;The cheapskate picks&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;#the-free-model-at-number-two-wont-say-who-made-it&quot; id=&quot;markdown-toc-the-free-model-at-number-two-wont-say-who-made-it&quot;&gt;The free model at number two won’t say who made it&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;#horror-story-48218-files-in-103-seconds&quot; id=&quot;markdown-toc-horror-story-48218-files-in-103-seconds&quot;&gt;Horror story: 48,218 files in 103 seconds&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;#whats-coming&quot; id=&quot;markdown-toc-whats-coming&quot;&gt;What’s coming&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;#finally&quot; id=&quot;markdown-toc-finally&quot;&gt;Finally&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;h2 id=&quot;everybody-blinked-at-once&quot;&gt;Everybody blinked at once&lt;/h2&gt;

&lt;p&gt;The clustering matters more than any one launch.&lt;/p&gt;

&lt;p&gt;On September 21, Xiaomi shipped MiMo-V2.6-Pro and MiMo-V2.6-Flash as MIT-licensed open weights. Pro is a 1.02-trillion-parameter mixture-of-experts with 42 billion active, a million-token context, and it’s priced at 43.5 cents in and 87 cents out. Flash is 310 billion total, 15 billion active, and costs 14 cents in and 28 cents out. Artificial Analysis gives Pro a 46 on its Intelligence Index, which makes it the top-scoring open-weights model on the board right now.&lt;/p&gt;

&lt;p&gt;A day later, on September 22, Anthropic released &lt;a href=&quot;https://www.anthropic.com/claude-opus-5-5&quot;&gt;Claude Opus 5.5&lt;/a&gt; at $4 in and $20 out, down from $5 and $25. Cache reads dropped 60 percent to 20 cents. Same day, OpenAI released &lt;a href=&quot;https://www.marktechpost.com/2026/09/22/openai-releases-gpt-6-sol-and-luna-50-cheaper-api-pricing-and-benchmarks/&quot;&gt;GPT-6 Sol and GPT-6 Luna&lt;/a&gt; and cut prices in half or better. Sol went from $4/$20 to $2/$10. Luna went from 20 cents and $1.20 to 10 cents in and 50 cents out. OpenAI says those are permanent prices, not a promo, and credits caching and inference improvements.&lt;/p&gt;

&lt;p&gt;Then on September 28, Anthropic released Claude Sonnet 5.5 at $2 in and $10 out. Same price as Sonnet 5, but a much better model.&lt;/p&gt;

&lt;p&gt;Three labs from two countries, all pushing prices down inside 48 hours. I’ve been doing this roundup since April and I’ve never seen the whole price column in my spreadsheet go down at once like this. Usually one lab cuts and everyone else pretends not to notice for a month.&lt;/p&gt;

&lt;h2 id=&quot;opus-55-is-the-rare-launch-where-every-signal-agrees&quot;&gt;Opus 5.5 is the rare launch where every signal agrees&lt;/h2&gt;

&lt;p&gt;I spend a lot of this roundup explaining why the benchmark number and the blind-test number disagree. Three weeks ago GPT-6 Astra tied for number one on Artificial Analysis and debuted twenty-fourth on Arena. It’s twenty-sixth now. That split is normal. It’s practically the house style.&lt;/p&gt;

&lt;p&gt;Opus 5.5 didn’t do that. It’s number one on the Artificial Analysis Intelligence Index at 58, five points clear of Fable 5.1 and GPT-6 Astra, which are tied at 53. It’s also number one on Arena Overall at 1509, and number one in Creative Writing, Instruction Following, and Hard Prompts. The hard-benchmark test and the blind vote landed on the same model in the same week. I almost never get to write that sentence.&lt;/p&gt;

&lt;p&gt;It also ended a running gag. For weeks I’ve been pointing out that the outright leader in Instruction Following and Hard Prompts was claude-opus-4-6, a model about a year old, beating everything Anthropic shipped after it. Opus 5.5 finally took both categories. Coding is the one place the old guard hangs on: opus-4-6-high, opus-4-7-high, and fable-5-high are in a three-way tie at 1551, with Opus 5.5 at fifth.&lt;/p&gt;

&lt;p&gt;Two asterisks. First, the votes are thin. Opus 5.5 has 2,307 votes in Overall and only 487 in Creative Writing, so its 20-point lead there could shrink as the crowd catches up. Second, every Claude score on Artificial Analysis still carries the “with fallback” tag, which means some answers came from a weaker model when the main one refused. I’ve flagged that on every Anthropic launch since Opus 5. The number is real. The footnote is load-bearing.&lt;/p&gt;

&lt;p&gt;Reddit’s reaction has mostly been relief. The most-upvoted praise isn’t even about coding, it’s about &lt;a href=&quot;https://botmonster.com/ai/opus-5-5-is-the-claude-comeback-reddit-was-waiting-for/&quot;&gt;how the thing talks&lt;/a&gt;. One r/ClaudeCode user called it “genuinely an order of magnitude improvement over Opus 5” in communication, which tracks. Opus 5 had a habit of saying “blast radius” and “load-bearing” like it was paid per use. (Yes, I just used “load-bearing” myself. I’m allowed. I’m not charging you per token.) Another user said they’d worked “non stop since release” and “barely made a dent” in their Max limits.&lt;/p&gt;

&lt;p&gt;Not everybody agrees on that part. A &lt;a href=&quot;https://hardforum.com/threads/claude-opus-5-5.2049520/&quot;&gt;HardForum&lt;/a&gt; user on the $100 Max plan said Opus 5 used to get them through whole days of coding, weekends included, without hitting the weekly limit. They upgraded to Opus 5.5 on Tuesday and were at 91 percent by Thursday morning. They were running it on “Extra” effort. Another user in the same thread, on medium, had used 12 percent of a $200 plan’s weekly limit, and a third summed it up: “5.5 extra eat tokens way more than previous gen.” Remember that effort dial. It comes back in a minute. Anthropic’s own demo was a HAProxy port from C to Rust that Opus 5.5 finished in 9.5 hours against Fable 5.1’s 12, at 51 percent lower cost. Both can be true. A well-defined port is not an open-ended feature in a messy real app, and the complaints are all coming from the messy real apps. And the best comment in the whole pile: “Queue the whining in two weeks about the model being nerfed.” Set a reminder.&lt;/p&gt;

&lt;h2 id=&quot;the-cheap-claude-costs-more-than-the-expensive-claude&quot;&gt;The cheap Claude costs more than the expensive Claude&lt;/h2&gt;

&lt;p&gt;Sonnet 5.5 is the budget model. It’s half the per-token price of Opus 5.5. It’s also scary good. It scored 70.6 percent on Terminal-Bench 4.0, the command-line agent test, up from Sonnet 5’s 10.3, and ahead of Opus 5.5’s 66.4. On Artificial Analysis it scores 56, two points under Opus 5.5 at max and ahead of Fable 5.1 and GPT-6 Astra.&lt;/p&gt;

&lt;p&gt;Then Artificial Analysis ran the whole index suite and checked the bill. At max effort, &lt;a href=&quot;https://www.beri.net/article/claude-sonnet-5-5-vs-opus-5-5-cost-per-task-matched-score-effort-between-tools-migration&quot;&gt;Sonnet 5.5 cost $7.60 per task. Opus 5.5 cost $5.98&lt;/a&gt;. The cheap model was 27 percent more expensive.&lt;/p&gt;

&lt;p&gt;The reason is tokens. Sonnet 5.5 burned 410 million output tokens getting through the suite. The median for models in its price tier is 88 million. Opus 5.5 used about 260 million. At max effort, Sonnet thinks roughly 60 percent longer per task than Opus, and half the price per token doesn’t cover 60 percent more tokens. It’s a fast model, 138 tokens a second, and it spends that speed talking to itself.&lt;/p&gt;

&lt;p&gt;It gets worse when you match scores instead of effort settings. Sonnet 5.5 at max and Opus 5.5 at xhigh both score 56 on the index. Sonnet costs $7.60 a task to get there. Opus costs $3.46. Same score, and the budget model costs 2.2 times as much. According to &lt;a href=&quot;https://www.beri.net/article/claude-sonnet-5-5-vs-opus-5-5-cost-per-task-matched-score-effort-between-tools-migration&quot;&gt;the same analysis&lt;/a&gt;, Sonnet only comes out cheaper at the bottom of the effort range.&lt;/p&gt;

&lt;p&gt;None of that makes Sonnet 5.5 a bad model. It went from 10.3 to 70.6 on Terminal-Bench in one release, and one customer quoted by MarkTechPost measured about 121K tokens per answer against 497K on Sonnet 5. It’s just not automatically the cheap option anymore. So the practical advice is boring. Don’t crank Sonnet 5.5 to max because it’s “the cheap one.” If a task needs the top of the dial, Opus at xhigh gets you the same score for less. If it doesn’t, run either one at a lower effort and check your actual bill after a day, not the price page.&lt;/p&gt;

&lt;p&gt;I’ve written some version of “cost per token is not cost per task” in this roundup maybe ten times. I’d never seen it flip the price order inside one lab’s own lineup before.&lt;/p&gt;

&lt;h2 id=&quot;openai-just-showed-up-at-the-bottom-of-the-price-sheet&quot;&gt;OpenAI just showed up at the bottom of the price sheet&lt;/h2&gt;

&lt;p&gt;For most of this year the cheapest-good-model story has been a Chinese open-weights story. Kimi, then MiMo, then GLM, and now maybe MiMo again. American labs competed at the top and left the floor alone.&lt;/p&gt;

&lt;p&gt;GPT-6 Luna costs 10 cents in and 50 cents out. That’s exactly the list price of GLM-5.3-Flash, the model that’s been my cheapskate default for the last month. And Luna is in the Arena Coding band, ranked 43rd at 1515, only 36 points behind the leader. It isn’t the cheapest thing in the band, but it’s the first time I’ve seen an OpenAI model sit in a cheapskate band at the same price as the Chinese floor. It’s also fast. Artificial Analysis clocks it at 148 tokens a second, about three times GLM’s speed.&lt;/p&gt;

&lt;p&gt;The catch is capability. Luna scores 37 on the Intelligence Index. GLM scores 42. Outside of coding, Luna falls out of the Arena bands entirely. It’s 86th Overall. So it’s a fast, cheap, preference-decent coding model, not a general replacement. But OpenAI ignored this end of the price sheet for a year, and now they’re in it.&lt;/p&gt;

&lt;p&gt;GPT-6 Sol is the more awkward launch. $2/$10, Intelligence Index 48, and OpenAI’s own numbers have it beating Opus 5 on AutomationBench (33.2 percent vs 26.9) and edging it on OSWorld at 80 percent lower cost. On Arena it’s 60th Overall at 1457, which misses the Overall band by two points.&lt;/p&gt;

&lt;h2 id=&quot;the-cheapest-coding-model-has-a-distillation-problem&quot;&gt;The cheapest coding model has a distillation problem&lt;/h2&gt;

&lt;p&gt;MiMo-V2.6-Flash is the cheapest model inside the Arena Coding band this week. It’s ranked 24th at 1525, 26 points behind the leader, for 28 cents out. That’s 89 times cheaper than a $25 leader. Artificial Analysis puts it on its intelligence-vs-cost Pareto frontier. On OpenRouter it went from nothing to the seventh most-used model on the platform in a week, 6.93 trillion tokens. By every signal I track, it’s a real value pick.&lt;/p&gt;

&lt;p&gt;And on September 10, eleven days before it shipped, Anthropic published a threat intelligence report accusing seven China-based labs of what it calls “illicit distillation,” using Claude’s outputs to train their own models. Xiaomi was one of them. According to &lt;a href=&quot;https://the-decoder.com/xiaomis-affordable-flagship-ai-leads-the-open-models-and-anthropic-says-claude-helped-get-it-there/&quot;&gt;the-decoder’s summary&lt;/a&gt; of case GTG-16008, Xiaomi forwarded more than 400,000 conversations from users of its own MiMo chatbot to Claude between March and April, through OpenClaw and OpenCode, to pull out training data. Across all seven labs, Anthropic counts around 190 million exchanges. Xiaomi hasn’t responded.&lt;/p&gt;

&lt;p&gt;I run OpenClaw. It’s the agent framework behind the assistant I run at home, the one that handles my scheduled jobs and writes into my notes vault. So I read that paragraph twice. To be clear, this is an accusation from a competitor, not a court finding, and nothing in it suggests OpenClaw itself did anything wrong. It’s a tool. Somebody pointed it at Claude at industrial scale. But if you were one of those 400,000 MiMo chatbot users, your conversations apparently took a trip you never agreed to.&lt;/p&gt;

&lt;p&gt;So what do you do with the pick? I’m printing it, because the method is the method and the price and rating are real. I’m also printing it with two asterisks. First, it only has 1,043 Coding votes, so it’s the emerging pick, not the settled one. GLM-5.3-Flash is right behind it at 1523 with five times the votes for 50 cents, and that’s the steadier bet. Second, if where a model’s training data came from matters to your company, and for some of you it contractually does, this one has a question hanging over it.&lt;/p&gt;

&lt;p&gt;And watch the price you see. On OpenRouter, MiMo-V2.6-Flash’s headline price shows 8 cents in and $1.28 out. That’s one small host’s endpoint. Xiaomi’s own endpoint is 14 cents and 28 cents, and it’s carrying 97 percent of the traffic. If you read the headline number, you’ll think it’s the most expensive cheap model on the board. Check the provider list.&lt;/p&gt;

&lt;h2 id=&quot;the-cheapskate-picks&quot;&gt;The cheapskate picks&lt;/h2&gt;

&lt;p&gt;Same method every week. Take the category leader’s Arena rating, draw a band 50 points below it, and find the cheapest model still inside the band. The top of Arena is compressed, so the leader is usually only a little better than something 40 to 100 times cheaper. The band gets computed in code from the full table, never eyeballed off the first screen, because eyeballing is how you delete the whole cheap tail without noticing. I learned that one in public.&lt;/p&gt;

&lt;p&gt;The good news: Arena finally refreshed. The last two issues ran on the same September 13 snapshot. This week’s data is stamped September 25. Bands ran deep: 57 models in Overall, 73 in Coding, 43 in Hard Prompts and Math, 38 in Instruction Following, and just 12 in Creative Writing. GLM-5.3-Flash is quoted at Z.ai’s list price, $0.15 in and $0.50 out. Arena’s price column now shows 20 cents out for it, but that’s a third-party floor, not list. Budget on fifty.&lt;/p&gt;

&lt;table&gt;
  &lt;thead&gt;
    &lt;tr&gt;
      &lt;th&gt;Category&lt;/th&gt;
      &lt;th&gt;Leader&lt;/th&gt;
      &lt;th&gt;$ out&lt;/th&gt;
      &lt;th&gt;Cheapskate pick&lt;/th&gt;
      &lt;th&gt;$ out&lt;/th&gt;
      &lt;th&gt;Δ rating&lt;/th&gt;
      &lt;th&gt;Cheaper by&lt;/th&gt;
    &lt;/tr&gt;
  &lt;/thead&gt;
  &lt;tbody&gt;
    &lt;tr&gt;
      &lt;td&gt;Overall&lt;/td&gt;
      &lt;td&gt;claude-opus-5.5-high (1509)&lt;/td&gt;
      &lt;td&gt;$20&lt;/td&gt;
      &lt;td&gt;GLM-5.3-Flash (1474, #35)&lt;/td&gt;
      &lt;td&gt;$0.50&lt;/td&gt;
      &lt;td&gt;−35&lt;/td&gt;
      &lt;td&gt;40×&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Coding&lt;/td&gt;
      &lt;td&gt;claude-opus-4-6-high (1551)&lt;/td&gt;
      &lt;td&gt;$25&lt;/td&gt;
      &lt;td&gt;MiMo-V2.6-Flash* (1525, #24)&lt;/td&gt;
      &lt;td&gt;$0.28&lt;/td&gt;
      &lt;td&gt;−26&lt;/td&gt;
      &lt;td&gt;89×&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Creative Writing&lt;/td&gt;
      &lt;td&gt;claude-opus-5.5-high (1521)&lt;/td&gt;
      &lt;td&gt;$20&lt;/td&gt;
      &lt;td&gt;gemini-3.7-flash-high (1492, #4)&lt;/td&gt;
      &lt;td&gt;$3.75&lt;/td&gt;
      &lt;td&gt;−29&lt;/td&gt;
      &lt;td&gt;5.3×&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Instruction Following&lt;/td&gt;
      &lt;td&gt;claude-opus-5.5-high (1516)&lt;/td&gt;
      &lt;td&gt;$20&lt;/td&gt;
      &lt;td&gt;GLM-5.3-Flash (1474, #26)&lt;/td&gt;
      &lt;td&gt;$0.50&lt;/td&gt;
      &lt;td&gt;−42&lt;/td&gt;
      &lt;td&gt;40×&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Hard Prompts&lt;/td&gt;
      &lt;td&gt;claude-opus-5.5-high (1541)&lt;/td&gt;
      &lt;td&gt;$20&lt;/td&gt;
      &lt;td&gt;GLM-5.3-Flash (1499, #31)&lt;/td&gt;
      &lt;td&gt;$0.50&lt;/td&gt;
      &lt;td&gt;−42&lt;/td&gt;
      &lt;td&gt;40×&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Math&lt;/td&gt;
      &lt;td&gt;claude-fable-5-high (1523)&lt;/td&gt;
      &lt;td&gt;$50&lt;/td&gt;
      &lt;td&gt;GLM-5.3-Flash (1501, #13)&lt;/td&gt;
      &lt;td&gt;$0.50&lt;/td&gt;
      &lt;td&gt;−22&lt;/td&gt;
      &lt;td&gt;100×&lt;/td&gt;
    &lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;

&lt;p&gt;*Thin votes (1,043). The steadier Coding pick is GLM-5.3-Flash at #28, 1523, 50 cents, 50× cheaper.&lt;/p&gt;

&lt;p&gt;GLM-5.3-Flash held four of six, and it’s doing it on real vote counts now: 19,103 in Overall and 12,503 in Hard Prompts. That’s a pick you can lean on. The “cheaper by” numbers shrank, from 100× to 40× in Overall and from 50× to 40× in Instruction Following and Hard Prompts, but not because GLM got pricier. The leader got cheaper. Opus 5.5 at $20 took the top of three boards from $25 and $50 models. That’s the price war showing up in my own table.&lt;/p&gt;

&lt;p&gt;Coding is the first crack in GLM’s run in a month, and it’s a crack on thin votes. Math is still the row to squint at: GLM sits 13th on only 920 votes. If you want steadier, MiMo v2.5 Pro has 3,280 votes at 87 cents.&lt;/p&gt;

&lt;p&gt;Creative Writing is where the method bit back. Gemini 3 Flash was the Creative pick for four straight issues at $3. It didn’t get worse this week. Opus 5.5 raised the ceiling by 17 points, and Gemini 3 Flash’s 1457 fell out of a band that now starts at 1471. When the leader gets better, cheap models fall out of the band without doing anything wrong. The new pick is Gemini 3.7 Flash at $3.75, ranked fourth on preliminary votes. Also, Gemini 3.8 Flash’s intro pricing doubles on January 1, so check whether 3.7 goes the same way before you build a budget around it.&lt;/p&gt;

&lt;p&gt;The speed caveat is the same as every week, and the number moved again. Artificial Analysis clocks GLM-5.3-Flash at 48 tokens a second this week, down from 89 last week. MiMo-V2.6-Flash is 55. I’ve stopped carrying this number forward because it bounces around too much. Both are slow enough to notice in an agent loop. GPT-6 Luna at 148 is the fast one at this price.&lt;/p&gt;

&lt;h2 id=&quot;the-free-model-at-number-two-wont-say-who-made-it&quot;&gt;The free model at number two won’t say who made it&lt;/h2&gt;

&lt;p&gt;Space Bunny Alpha showed up on OpenRouter on September 23 as a stealth model: no maker named, free, a million-token context, and text, image, and video input. Five days later it was the second-biggest model on the platform for the week at 18.2 trillion tokens, and number one for the most recent day.&lt;/p&gt;

&lt;p&gt;It’s probably MiniMax. Tokenizer tests on launch day matched MiniMax’s M3 family on every string people threw at it (one tester ran 50 strings, all 50 matched), and on September 27 MiniMax released &lt;a href=&quot;https://cellcog.ai/blog/what-is-space-bunny-alpha/&quot;&gt;M3.1-Flash-Preview&lt;/a&gt; with the same context length, the same five reasoning effort levels, and the same fast-coding pitch. Nobody has officially confirmed it.&lt;/p&gt;

&lt;p&gt;Two things worth knowing before you point production at it. It’s free, and free usage isn’t chosen usage. Free models in this roundup have a habit of spiking and then settling once the meter turns on. And the listing says prompts and completions “may be retained by the provider,” a provider that won’t tell you its name. The same week a distillation report dropped. Fine for kicking the tires on public code. I wouldn’t send it anything I’d mind reading in somebody’s training set.&lt;/p&gt;

&lt;p&gt;Meanwhile, on the paid board, DeepSeek V4.1 Flash is number one at 20.8 trillion tokens, up 23 percent. GLM-5.3-Flash slipped to third at 13.3 trillion, down 21 percent, though it’s still number one for the trailing month at 57.5 trillion. Usage is moving toward whatever is newest, cheapest, or free, which is what usage always does in the week after a launch pile-up.&lt;/p&gt;

&lt;h2 id=&quot;horror-story-48218-files-in-103-seconds&quot;&gt;Horror story: 48,218 files in 103 seconds&lt;/h2&gt;

&lt;p&gt;Around September 20, a developer posted on r/ClaudeAI that Claude Code had deleted their project. According to &lt;a href=&quot;https://www.techradar.com/pro/security/i-broke-something-a-claude-code-ai-agent-deleted-48-000-files-in-just-over-100-seconds-then-apologized-for-doing-so&quot;&gt;TechRadar’s writeup&lt;/a&gt;, the agent was told to rebuild a mirror of the project for a task. It figured out &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;build_mirror.py&lt;/code&gt; couldn’t refresh the mirror in place, so it wrote a little Python remover to delete an old copy sitting in a temp folder. That old copy held 7,332 ordinary files and 614 Windows directory junctions pointing back into the live project. The remover used &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;os.walk&lt;/code&gt; with &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;followlinks=False&lt;/code&gt;, which sounds safe. But &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;os.path.islink()&lt;/code&gt; returns false for Windows junctions, so Python didn’t treat them as links, walked right through them, and started deleting the real project on the other side.&lt;/p&gt;

&lt;p&gt;It took 103 seconds. It deleted 48,218 live files, and it emptied the Git object store too: &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;.git/objects&lt;/code&gt;, refs, and logs. The index still listed 7,221 paths, but the actual file contents were gone, so Git couldn’t restore anything. TechRadar’s headline quotes the agent’s own summary: “I broke something.” (This all comes from the Reddit post and an attached report, not an independent forensic investigation, so treat the details as the poster’s account.)&lt;/p&gt;

&lt;p&gt;Reddit’s verdict was harsh. It wasn’t wrong. The poster admitted they weren’t using GitHub properly and should have been working on a branch. But the reason I’m including it isn’t to dunk on someone. It’s that the agent did something that looks completely reasonable (clean up a stale temp copy) on a file system that had a trap in it that no human would have spotted in the moment either. Directory junctions and symlinks turn “delete this folder” into “delete whatever this folder points at,” and on Windows the standard Python safety check doesn’t even see the junctions. If you let an agent run &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;rm&lt;/code&gt; or its Python equivalent, the only real protection is a remote it can’t touch, so push before you let it clean up anything.&lt;/p&gt;

&lt;h2 id=&quot;whats-coming&quot;&gt;What’s coming&lt;/h2&gt;

&lt;p&gt;Claude Haiku 5.5 is next. Anthropic said Sonnet and Haiku 5.5 would follow Opus “in the coming weeks.” Sonnet already shipped. Haiku is the one left, and if it’s priced like Haiku usually is, it lands right in the cheapskate zone.&lt;/p&gt;

&lt;p&gt;MiniMax M3.1 should get official soon. If Space Bunny gets unmasked and priced, we’ll see how many of those 18 trillion free tokens stick around once there’s a bill.&lt;/p&gt;

&lt;p&gt;Then there’s MiMo-V2.6-Pro-UltraSpeed, which Xiaomi says is up to 20 times faster than Pro at the same quality. It already has 73 billion tokens on OpenRouter. If that holds, “cheap but slow” stops being the standard caveat on the Chinese value picks.&lt;/p&gt;

&lt;p&gt;DeepSeek V5 is speculation only. People keep floating October. DeepSeek hasn’t said anything, and I’m not putting a date on a model that doesn’t have one.&lt;/p&gt;

&lt;h2 id=&quot;finally&quot;&gt;Finally&lt;/h2&gt;

&lt;p&gt;The price war is real, and it’s good for you. Opus got 20 percent cheaper and better at the same time. OpenAI cut Sol and Luna in half and made it permanent. The floor got a new 28-cent option. If you’re paying your own bill, almost every line on your invoice should drop next month. Who had “two frontier labs cut prices on the same day” on their bingo card?&lt;/p&gt;

&lt;p&gt;But three things I’ve been harping on all year showed up again, just in cheaper clothes. The sticker isn’t the task: the budget Claude costs more than the premium Claude when you run it hot. The free model isn’t free: someone is paying for those 18 trillion tokens, and they’re not telling you who or why. And the cheapest option comes with questions about where it came from, which some of you will care about and some of you won’t.&lt;/p&gt;

&lt;p&gt;Last week I said I’d come back to see whether the Arena crowd was any nicer to Grok 4.7 than its creator was. It wasn’t. Grok 4.7 debuted at 92nd Overall. Its creator called it mid. The crowd called it worse. Low bar, and it still found a way under it.&lt;/p&gt;
</description>
        <pubDate>Tue, 29 Sep 2026 08:00:00 -0500</pubDate>
        <link>https://www.stephanmiller.com/model-buzz-roundup-week-of-0923/</link>
        <guid isPermaLink="true">https://www.stephanmiller.com/model-buzz-roundup-week-of-0923/</guid>
        
        <category>llm</category>
        
        <category>openrouter</category>
        
        <category>model-roundup</category>
        
        
        <category>large-language-models</category>
        
      </item>
    
      <item>
        <title>Track New AI Models on the Arena Leaderboard With a 70-Line Scraper</title>
        <description>&lt;p&gt;Every week I write the &lt;a href=&quot;https://www.stephanmiller.com/series/model-buzz-report/&quot;&gt;Model Buzz Report&lt;/a&gt;, and every week part of the job is staring at the &lt;a href=&quot;https://www.stephanmiller.com/the-cheapskates-guide-to-the-arena-leaderboard-why-i-stopped-paying-claude-opus-prices/&quot;&gt;Arena leaderboard&lt;/a&gt; trying to remember what it looked like last week. Which of these is new? Was that one there before? When did &lt;a href=&quot;https://www.stephanmiller.com/model-buzz-roundup-week-of-0819/&quot;&gt;GLM 5.3&lt;/a&gt; show up in the top 20, because it sure felt like it came out of nowhere?&lt;/p&gt;

&lt;p&gt;That’s a question a scraper can’t answer. A scraper hands you the page as it’s right now, and “what’s new” is a comparison between now and then. You need a then.&lt;/p&gt;

&lt;p&gt;So this post is about the then. It’s the smallest useful version of the thing most scraping recipes in this series are going to need: save each run with a timestamp, run it on a schedule, and compare this run with the last one. The example is deliberately easy, one page and one question.&lt;/p&gt;

&lt;ul id=&quot;markdown-toc&quot;&gt;
  &lt;li&gt;&lt;a href=&quot;#why-now-is-almost-never-the-question&quot; id=&quot;markdown-toc-why-now-is-almost-never-the-question&quot;&gt;Why “Now” Is Almost Never the Question&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;#the-page-arenas-text-leaderboard&quot; id=&quot;markdown-toc-the-page-arenas-text-leaderboard&quot;&gt;The Page: Arena’s Text Leaderboard&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;#the-harness-all-70-lines&quot; id=&quot;markdown-toc-the-harness-all-70-lines&quot;&gt;The Harness, All 70 Lines&lt;/a&gt;    &lt;ul&gt;
      &lt;li&gt;&lt;a href=&quot;#storage-sqlite-not-a-json-file&quot; id=&quot;markdown-toc-storage-sqlite-not-a-json-file&quot;&gt;Storage: SQLite, Not a JSON File&lt;/a&gt;&lt;/li&gt;
      &lt;li&gt;&lt;a href=&quot;#the-diff-two-kinds-of-new&quot; id=&quot;markdown-toc-the-diff-two-kinds-of-new&quot;&gt;The Diff: Two Kinds of “New”&lt;/a&gt;&lt;/li&gt;
      &lt;li&gt;&lt;a href=&quot;#the-schedule-one-cron-line&quot; id=&quot;markdown-toc-the-schedule-one-cron-line&quot;&gt;The Schedule: One Cron Line&lt;/a&gt;&lt;/li&gt;
    &lt;/ul&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;#day-one-has-no-then&quot; id=&quot;markdown-toc-day-one-has-no-then&quot;&gt;Day One Has No Then&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;#two-things-it-got-wrong&quot; id=&quot;markdown-toc-two-things-it-got-wrong&quot;&gt;Two Things It Got Wrong&lt;/a&gt;    &lt;ul&gt;
      &lt;li&gt;&lt;a href=&quot;#that-new-1-is-a-rename&quot; id=&quot;markdown-toc-that-new-1-is-a-rename&quot;&gt;That “NEW #1” Is a Rename&lt;/a&gt;&lt;/li&gt;
      &lt;li&gt;&lt;a href=&quot;#it-missed-the-one-that-started-this&quot; id=&quot;markdown-toc-it-missed-the-one-that-started-this&quot;&gt;It Missed the One That Started This&lt;/a&gt;&lt;/li&gt;
    &lt;/ul&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;#why-not-just-use-firecrawls-change-tracking&quot; id=&quot;markdown-toc-why-not-just-use-firecrawls-change-tracking&quot;&gt;Why Not Just Use Firecrawl’s Change Tracking?&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;#what-this-is-for&quot; id=&quot;markdown-toc-what-this-is-for&quot;&gt;What This Is For&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;h2 id=&quot;why-now-is-almost-never-the-question&quot;&gt;Why “Now” Is Almost Never the Question&lt;/h2&gt;

&lt;p&gt;Think about what people actually want out of a scraped page. Is this price a deal? Did the competitor change their pricing? Which models are new near the top? None of those can be answered from one snapshot. “Is this a deal” needs the price history. “Did they change” needs the old page. “What’s new” needs the old list.&lt;/p&gt;

&lt;p&gt;&lt;a href=&quot;https://firecrawl.link/stephan-miller&quot;&gt;Firecrawl&lt;/a&gt; solves the ugly part of scraping, which is getting a clean page out of a site that renders everything with JavaScript and would rather you went away. It doesn’t solve the then. Nothing that fetches a page can, because the then is data you had to collect yourself, back when it was the now.&lt;/p&gt;

&lt;p&gt;That takes three boring things:&lt;/p&gt;

&lt;ol&gt;
  &lt;li&gt;&lt;strong&gt;Storage.&lt;/strong&gt; Save each crawl with a timestamp instead of printing it and forgetting it.&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;A schedule.&lt;/strong&gt; One crawl tells you nothing. A cron line fixes that.&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;A diff.&lt;/strong&gt; Compare this crawl with the last one and only say something when something changed.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;That’s the whole harness. I keep wanting to call it a framework and it keeps being 70 lines.&lt;/p&gt;

&lt;h2 id=&quot;the-page-arenas-text-leaderboard&quot;&gt;The Page: Arena’s Text Leaderboard&lt;/h2&gt;

&lt;p&gt;The target is &lt;a href=&quot;https://arena.ai/leaderboard/text&quot;&gt;arena.ai/leaderboard/text&lt;/a&gt;, the overall text leaderboard. About 400 models, ranked by head-to-head human votes, with score, vote count, price, and context window per row.&lt;/p&gt;

&lt;p&gt;It’s also a page that doesn’t want to be read by a simple fetch. The table is rendered client-side, and when I &lt;a href=&quot;https://www.stephanmiller.com/building-a-cost-saving-skill-that-accidentally-became-its-own-newsletter/&quot;&gt;built the Model Buzz skill&lt;/a&gt; I learned that some fetch tools hand back the wrong leaderboard on category URLs without any error at all. That’s the worst kind of failure: data that looks plausible.&lt;/p&gt;

&lt;p&gt;Firecrawl got it clean on the first try:&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;firecrawl scrape https://arena.ai/leaderboard/text &lt;span class=&quot;nt&quot;&gt;--only-main-content&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;--wait-for&lt;/span&gt; 3000
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;&lt;img src=&quot;/images/2026/track-new-ai-models-on-the-arena-leaderboard-body-1.jpg&quot; alt=&quot;The Page: Arena&apos;s Text Leaderboard&quot; srcset=&quot;            /assets/resized/480/track-new-ai-models-on-the-arena-leaderboard-body-1.jpg 480w,            /assets/resized/800/track-new-ai-models-on-the-arena-leaderboard-body-1.jpg 800w,            /assets/resized/1400/track-new-ai-models-on-the-arena-leaderboard-body-1.jpg 1400w,    &quot; loading=&quot;lazy&quot; /&gt;&lt;/p&gt;

&lt;p&gt;What comes back is markdown, and the leaderboard is a plain markdown table:&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;| Rank | Rank Spread | Model | Score | Votes | Price $/M | Context |
| --- | --- | --- | --- | --- | --- | --- |
| 1 | 17 | Anthropic&amp;lt;br&amp;gt;[claude-fable-5-high](https://www.anthropic.com/news/claude-fable-5-mythos-5)&amp;lt;br&amp;gt;Anthropic · Proprietary | 1506±5 | 30,057 | $10 / $50 | 1M |
...
| 19 | 737 | [glm-5.3-max](https://z.ai/blog/glm-5.3)&amp;lt;br&amp;gt;Z.ai · MIT | 1483±6 | 10,960 | $1.40 / $4.40 | 1M |
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;I gave it three seconds to let the page render. I didn’t test whether it needs that, so treat the number as superstition you’re free to delete.&lt;/p&gt;

&lt;p&gt;No JSON schema, no LLM extraction, and no selectors. A regex reads that table fine. A plain scrape is one credit a page, and Firecrawl’s JSON extraction adds four more on top, so asking an LLM to parse a table this regular would be paying five times over for something a regex already does. If you haven’t installed the CLI yet, &lt;a href=&quot;https://www.stephanmiller.com/firecrawl-cli-setup-skip-the-installer-that-rewrites-your-editors/&quot;&gt;the setup post&lt;/a&gt; covers it and the installer you should skip.&lt;/p&gt;

&lt;h2 id=&quot;the-harness-all-70-lines&quot;&gt;The Harness, All 70 Lines&lt;/h2&gt;

&lt;p&gt;Python, standard library only, calling the Firecrawl CLI through &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;subprocess&lt;/code&gt;. No SDK, so the only install is the CLI you already have.&lt;/p&gt;

&lt;div class=&quot;language-python highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;c1&quot;&gt;#!/usr/bin/env python3
&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;&quot;&quot;&quot;Watch the Arena text leaderboard and report models that just showed up near the top.

    python3 arena_watch.py                      # scrape now, save, compare with last run
    python3 arena_watch.py --url &amp;lt;wayback url&amp;gt; --at 2026-09-02   # backfill an old snapshot
&quot;&quot;&quot;&lt;/span&gt;
&lt;span class=&quot;kn&quot;&gt;import&lt;/span&gt; &lt;span class=&quot;nn&quot;&gt;argparse&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;re&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;sqlite3&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;subprocess&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;sys&lt;/span&gt;
&lt;span class=&quot;kn&quot;&gt;from&lt;/span&gt; &lt;span class=&quot;nn&quot;&gt;datetime&lt;/span&gt; &lt;span class=&quot;kn&quot;&gt;import&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;datetime&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;timezone&lt;/span&gt;

&lt;span class=&quot;n&quot;&gt;URL&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;s&quot;&gt;&quot;https://arena.ai/leaderboard/text&quot;&lt;/span&gt;
&lt;span class=&quot;n&quot;&gt;DB&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;s&quot;&gt;&quot;arena.db&quot;&lt;/span&gt;
&lt;span class=&quot;n&quot;&gt;TOP&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;mi&quot;&gt;20&lt;/span&gt;

&lt;span class=&quot;n&quot;&gt;ROW&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;re&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;nb&quot;&gt;compile&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;sa&quot;&gt;r&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;&quot;^\| (\d+) \| \d+ \| (.+?) \| (\d+)±&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;
&lt;span class=&quot;n&quot;&gt;NAME&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;re&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;nb&quot;&gt;compile&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;sa&quot;&gt;r&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;&quot;\[([^\]]+)\]&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;


&lt;span class=&quot;k&quot;&gt;def&lt;/span&gt; &lt;span class=&quot;nf&quot;&gt;scrape&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;url&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;):&lt;/span&gt;
    &lt;span class=&quot;n&quot;&gt;out&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;subprocess&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;run&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;
        &lt;span class=&quot;p&quot;&gt;[&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;&quot;firecrawl&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;s&quot;&gt;&quot;scrape&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;url&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;s&quot;&gt;&quot;--only-main-content&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;s&quot;&gt;&quot;--wait-for&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;s&quot;&gt;&quot;3000&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;],&lt;/span&gt;
        &lt;span class=&quot;n&quot;&gt;capture_output&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;bp&quot;&gt;True&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;text&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;bp&quot;&gt;True&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;check&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;bp&quot;&gt;True&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;).&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;stdout&lt;/span&gt;
    &lt;span class=&quot;n&quot;&gt;rows&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;p&quot;&gt;[]&lt;/span&gt;
    &lt;span class=&quot;k&quot;&gt;for&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;line&lt;/span&gt; &lt;span class=&quot;ow&quot;&gt;in&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;out&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;splitlines&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;():&lt;/span&gt;
        &lt;span class=&quot;n&quot;&gt;m&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;ROW&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;match&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;line&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;
        &lt;span class=&quot;n&quot;&gt;name&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;m&lt;/span&gt; &lt;span class=&quot;ow&quot;&gt;and&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;NAME&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;search&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;m&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;group&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;mi&quot;&gt;2&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;))&lt;/span&gt;
        &lt;span class=&quot;k&quot;&gt;if&lt;/span&gt; &lt;span class=&quot;ow&quot;&gt;not&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;name&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;:&lt;/span&gt;
            &lt;span class=&quot;k&quot;&gt;continue&lt;/span&gt;  &lt;span class=&quot;c1&quot;&gt;# skips the calendar widget and anything else that isn&apos;t a model row
&lt;/span&gt;        &lt;span class=&quot;n&quot;&gt;org&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;m&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;group&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;mi&quot;&gt;2&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;).&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;split&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;&quot;&amp;lt;br&amp;gt;&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)[&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;-&lt;/span&gt;&lt;span class=&quot;mi&quot;&gt;1&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;].&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;split&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;&quot; · &quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)[&lt;/span&gt;&lt;span class=&quot;mi&quot;&gt;0&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;]&lt;/span&gt;
        &lt;span class=&quot;n&quot;&gt;rows&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;append&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;((&lt;/span&gt;&lt;span class=&quot;nb&quot;&gt;int&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;m&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;group&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;mi&quot;&gt;1&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)),&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;name&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;group&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;mi&quot;&gt;1&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;),&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;org&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;nb&quot;&gt;int&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;m&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;group&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;mi&quot;&gt;3&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;))))&lt;/span&gt;
    &lt;span class=&quot;k&quot;&gt;return&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;rows&lt;/span&gt;


&lt;span class=&quot;k&quot;&gt;def&lt;/span&gt; &lt;span class=&quot;nf&quot;&gt;main&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;():&lt;/span&gt;
    &lt;span class=&quot;n&quot;&gt;ap&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;argparse&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;ArgumentParser&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;()&lt;/span&gt;
    &lt;span class=&quot;n&quot;&gt;ap&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;add_argument&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;&quot;--url&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;default&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;URL&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;
    &lt;span class=&quot;n&quot;&gt;ap&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;add_argument&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;&quot;--at&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;help&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;&quot;timestamp to record, for backfilling old snapshots&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;
    &lt;span class=&quot;n&quot;&gt;args&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;ap&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;parse_args&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;()&lt;/span&gt;

    &lt;span class=&quot;n&quot;&gt;rows&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;scrape&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;args&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;url&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;
    &lt;span class=&quot;k&quot;&gt;if&lt;/span&gt; &lt;span class=&quot;nb&quot;&gt;len&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;rows&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;&amp;lt;&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;TOP&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;:&lt;/span&gt;
        &lt;span class=&quot;n&quot;&gt;sys&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;nb&quot;&gt;exit&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;sa&quot;&gt;f&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;&quot;only parsed &lt;/span&gt;&lt;span class=&quot;si&quot;&gt;{&lt;/span&gt;&lt;span class=&quot;nb&quot;&gt;len&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;rows&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;&lt;span class=&quot;si&quot;&gt;}&lt;/span&gt;&lt;span class=&quot;s&quot;&gt; rows, the page layout probably changed&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;

    &lt;span class=&quot;n&quot;&gt;taken&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;args&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;at&lt;/span&gt; &lt;span class=&quot;ow&quot;&gt;or&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;datetime&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;now&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;timezone&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;utc&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;).&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;strftime&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;&quot;%Y-%m-%d %H:%M&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;
    &lt;span class=&quot;n&quot;&gt;db&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;sqlite3&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;connect&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;DB&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;
    &lt;span class=&quot;n&quot;&gt;db&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;execute&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;&quot;&quot;&quot;CREATE TABLE IF NOT EXISTS ranks
                  (taken TEXT, rank INT, model TEXT, org TEXT, score INT)&quot;&quot;&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;
    &lt;span class=&quot;n&quot;&gt;db&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;executemany&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;&quot;INSERT INTO ranks VALUES (?,?,?,?,?)&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;p&quot;&gt;[(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;taken&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;*&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;r&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt; &lt;span class=&quot;k&quot;&gt;for&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;r&lt;/span&gt; &lt;span class=&quot;ow&quot;&gt;in&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;rows&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;])&lt;/span&gt;
    &lt;span class=&quot;n&quot;&gt;db&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;commit&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;()&lt;/span&gt;

    &lt;span class=&quot;n&quot;&gt;prev&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;db&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;execute&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;&quot;SELECT MAX(taken) FROM ranks WHERE taken &amp;lt; ?&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;taken&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,)).&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;fetchone&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;()[&lt;/span&gt;&lt;span class=&quot;mi&quot;&gt;0&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;]&lt;/span&gt;
    &lt;span class=&quot;k&quot;&gt;if&lt;/span&gt; &lt;span class=&quot;ow&quot;&gt;not&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;prev&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;:&lt;/span&gt;
        &lt;span class=&quot;k&quot;&gt;print&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;sa&quot;&gt;f&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;si&quot;&gt;{&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;taken&lt;/span&gt;&lt;span class=&quot;si&quot;&gt;}&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;: first run, saved &lt;/span&gt;&lt;span class=&quot;si&quot;&gt;{&lt;/span&gt;&lt;span class=&quot;nb&quot;&gt;len&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;rows&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;&lt;span class=&quot;si&quot;&gt;}&lt;/span&gt;&lt;span class=&quot;s&quot;&gt; models. Nothing to compare yet.&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;
        &lt;span class=&quot;k&quot;&gt;return&lt;/span&gt;

    &lt;span class=&quot;n&quot;&gt;was_top&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;p&quot;&gt;{&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;m&lt;/span&gt; &lt;span class=&quot;k&quot;&gt;for&lt;/span&gt; &lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;m&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,)&lt;/span&gt; &lt;span class=&quot;ow&quot;&gt;in&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;db&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;execute&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;
        &lt;span class=&quot;s&quot;&gt;&quot;SELECT model FROM ranks WHERE taken = ? AND rank &amp;lt;= ?&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;prev&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;TOP&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;))}&lt;/span&gt;
    &lt;span class=&quot;n&quot;&gt;seen&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;p&quot;&gt;{&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;m&lt;/span&gt; &lt;span class=&quot;k&quot;&gt;for&lt;/span&gt; &lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;m&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,)&lt;/span&gt; &lt;span class=&quot;ow&quot;&gt;in&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;db&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;execute&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;&quot;SELECT DISTINCT model FROM ranks WHERE taken &amp;lt; ?&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;taken&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,))}&lt;/span&gt;

    &lt;span class=&quot;k&quot;&gt;print&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;sa&quot;&gt;f&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;si&quot;&gt;{&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;taken&lt;/span&gt;&lt;span class=&quot;si&quot;&gt;}&lt;/span&gt;&lt;span class=&quot;s&quot;&gt; vs &lt;/span&gt;&lt;span class=&quot;si&quot;&gt;{&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;prev&lt;/span&gt;&lt;span class=&quot;si&quot;&gt;}&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;
    &lt;span class=&quot;k&quot;&gt;for&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;rank&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;model&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;org&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;score&lt;/span&gt; &lt;span class=&quot;ow&quot;&gt;in&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;rows&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;:&lt;/span&gt;
        &lt;span class=&quot;k&quot;&gt;if&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;rank&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;&amp;gt;&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;TOP&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;:&lt;/span&gt;
            &lt;span class=&quot;k&quot;&gt;break&lt;/span&gt;
        &lt;span class=&quot;k&quot;&gt;if&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;model&lt;/span&gt; &lt;span class=&quot;ow&quot;&gt;not&lt;/span&gt; &lt;span class=&quot;ow&quot;&gt;in&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;seen&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;:&lt;/span&gt;
            &lt;span class=&quot;k&quot;&gt;print&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;sa&quot;&gt;f&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;&quot;  NEW      #&lt;/span&gt;&lt;span class=&quot;si&quot;&gt;{&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;rank&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;:&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;&amp;lt;&lt;/span&gt;&lt;span class=&quot;mi&quot;&gt;3&lt;/span&gt;&lt;span class=&quot;si&quot;&gt;}&lt;/span&gt;&lt;span class=&quot;s&quot;&gt; &lt;/span&gt;&lt;span class=&quot;si&quot;&gt;{&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;model&lt;/span&gt;&lt;span class=&quot;si&quot;&gt;}&lt;/span&gt;&lt;span class=&quot;s&quot;&gt; (&lt;/span&gt;&lt;span class=&quot;si&quot;&gt;{&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;org&lt;/span&gt;&lt;span class=&quot;si&quot;&gt;}&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;, &lt;/span&gt;&lt;span class=&quot;si&quot;&gt;{&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;score&lt;/span&gt;&lt;span class=&quot;si&quot;&gt;}&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;)&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;
        &lt;span class=&quot;k&quot;&gt;elif&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;model&lt;/span&gt; &lt;span class=&quot;ow&quot;&gt;not&lt;/span&gt; &lt;span class=&quot;ow&quot;&gt;in&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;was_top&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;:&lt;/span&gt;
            &lt;span class=&quot;k&quot;&gt;print&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;sa&quot;&gt;f&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;&quot;  MOVED UP #&lt;/span&gt;&lt;span class=&quot;si&quot;&gt;{&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;rank&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;:&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;&amp;lt;&lt;/span&gt;&lt;span class=&quot;mi&quot;&gt;3&lt;/span&gt;&lt;span class=&quot;si&quot;&gt;}&lt;/span&gt;&lt;span class=&quot;s&quot;&gt; &lt;/span&gt;&lt;span class=&quot;si&quot;&gt;{&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;model&lt;/span&gt;&lt;span class=&quot;si&quot;&gt;}&lt;/span&gt;&lt;span class=&quot;s&quot;&gt; (&lt;/span&gt;&lt;span class=&quot;si&quot;&gt;{&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;org&lt;/span&gt;&lt;span class=&quot;si&quot;&gt;}&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;, &lt;/span&gt;&lt;span class=&quot;si&quot;&gt;{&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;score&lt;/span&gt;&lt;span class=&quot;si&quot;&gt;}&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;)&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;


&lt;span class=&quot;k&quot;&gt;if&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;__name__&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;==&lt;/span&gt; &lt;span class=&quot;s&quot;&gt;&quot;__main__&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;:&lt;/span&gt;
    &lt;span class=&quot;n&quot;&gt;main&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;()&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;&lt;img src=&quot;/images/2026/track-new-ai-models-on-the-arena-leaderboard-body-2.jpg&quot; alt=&quot;The Harness, All 70 Lines&quot; srcset=&quot;            /assets/resized/480/track-new-ai-models-on-the-arena-leaderboard-body-2.jpg 480w,            /assets/resized/800/track-new-ai-models-on-the-arena-leaderboard-body-2.jpg 800w,            /assets/resized/1400/track-new-ai-models-on-the-arena-leaderboard-body-2.jpg 1400w,    &quot; loading=&quot;lazy&quot; /&gt;&lt;/p&gt;

&lt;p&gt;Here’s what each piece is doing, mapped to the three boring things.&lt;/p&gt;

&lt;h3 id=&quot;storage-sqlite-not-a-json-file&quot;&gt;Storage: SQLite, Not a JSON File&lt;/h3&gt;

&lt;p&gt;Every run inserts every row, all 400 or so, stamped with the time it was taken. One table, five columns.&lt;/p&gt;

&lt;p&gt;I went back and forth on JSONL here. A line of JSON per run is simpler to look at and fine for a diff against the last run. But the questions I actually care about later are history questions: how long has this model been in the top 20, when did it first appear anywhere on the board, is it climbing or sliding. With JSONL you write a loop for each of those. With SQLite you write a query. And SQLite ships with Python, so it costs nothing to install.&lt;/p&gt;

&lt;p&gt;Saving the whole board instead of just the top 20 is on purpose. Storage is cheap and you can’t go back and scrape last Tuesday. A model that debuts at #24 isn’t news today, but the day it cracks the top 20 you want to know it has been lurking for a week.&lt;/p&gt;

&lt;h3 id=&quot;the-diff-two-kinds-of-new&quot;&gt;The Diff: Two Kinds of “New”&lt;/h3&gt;

&lt;p&gt;The script reports two different things, and the difference matters.&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;&lt;strong&gt;NEW&lt;/strong&gt; means the model has never appeared anywhere on the board in any previous run. This is the one I actually wanted. A brand new model landing in the top 20 on its first appearance is the “where did that come from” moment.&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;MOVED UP&lt;/strong&gt; means the model was on the board before but was not in the top 20 last run. A climber. Less exciting, still worth a look.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Everything else, a model shuffling from #7 to #9, stays quiet. The whole point of the diff is that most runs should print almost nothing.&lt;/p&gt;

&lt;h3 id=&quot;the-schedule-one-cron-line&quot;&gt;The Schedule: One Cron Line&lt;/h3&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;0 8 * * * cd /path/to/arena-watch &amp;amp;&amp;amp; python3 arena_watch.py &amp;gt;&amp;gt; watch.log 2&amp;gt;&amp;amp;1
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;&lt;img src=&quot;/images/2026/track-new-ai-models-on-the-arena-leaderboard-body-3.jpg&quot; alt=&quot;The Schedule: One Cron Line&quot; srcset=&quot;            /assets/resized/480/track-new-ai-models-on-the-arena-leaderboard-body-3.jpg 480w,            /assets/resized/800/track-new-ai-models-on-the-arena-leaderboard-body-3.jpg 800w,            /assets/resized/1400/track-new-ai-models-on-the-arena-leaderboard-body-3.jpg 1400w,    &quot; loading=&quot;lazy&quot; /&gt;&lt;/p&gt;

&lt;p&gt;Daily at 8am, appended to a log. That lives on my mini PC, which is on all the time. On a Mac you can use cron too, or launchd if you enjoy writing XML. Once a day is plenty for a leaderboard that needs thousands of votes to move a model. Each run is one scrape, which on Firecrawl is one credit.&lt;/p&gt;

&lt;h2 id=&quot;day-one-has-no-then&quot;&gt;Day One Has No Then&lt;/h2&gt;

&lt;p&gt;Every history-based tool has the same catch: the first run is useless. It saves 400 rows and prints “Nothing to compare yet.” You have to wait a day for the second run before the thing does anything at all, and a week before it does anything interesting.&lt;/p&gt;

&lt;p&gt;I didn’t want to wait a week to write this post, so I cheated with the Wayback Machine. The Internet Archive snapshots the Arena leaderboard most days, and Firecrawl will scrape an archived copy just like the live page. That’s what the &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;--url&lt;/code&gt; and &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;--at&lt;/code&gt; flags are for: point the script at an old snapshot and tell it what date to record.&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;python3 arena_watch.py &lt;span class=&quot;nt&quot;&gt;--url&lt;/span&gt; &lt;span class=&quot;s2&quot;&gt;&quot;https://web.archive.org/web/20260902190730/https://arena.ai/leaderboard/text&quot;&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;--at&lt;/span&gt; &lt;span class=&quot;s2&quot;&gt;&quot;2026-09-02 19:07&quot;&lt;/span&gt;
python3 arena_watch.py &lt;span class=&quot;nt&quot;&gt;--url&lt;/span&gt; &lt;span class=&quot;s2&quot;&gt;&quot;https://web.archive.org/web/20260909040114/https://arena.ai/leaderboard/text&quot;&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;--at&lt;/span&gt; &lt;span class=&quot;s2&quot;&gt;&quot;2026-09-09 04:01&quot;&lt;/span&gt;
python3 arena_watch.py &lt;span class=&quot;nt&quot;&gt;--url&lt;/span&gt; &lt;span class=&quot;s2&quot;&gt;&quot;https://web.archive.org/web/20260917153328/https://arena.ai/leaderboard/text&quot;&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;--at&lt;/span&gt; &lt;span class=&quot;s2&quot;&gt;&quot;2026-09-17 15:33&quot;&lt;/span&gt;
python3 arena_watch.py
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;The archived pages parse the same as the live one, since the table structure is identical and only the link URLs change. Three weeks of history in four commands:&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;2026-09-02 19:07: first run, saved 399 models. Nothing to compare yet.
2026-09-09 04:01 vs 2026-09-02 19:07
  NEW      #3   claude-fable-5.1-max (Anthropic, 1504)
2026-09-17 15:33 vs 2026-09-09 04:01
  NEW      #8   muse-spark-1.3-max (Meta, 1493)
2026-09-23 01:27 vs 2026-09-17 15:33
  NEW      #1   claude-fable-5-high (Anthropic, 1506)
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;Fable 5.1 Max debuting at #3 and Meta’s Muse Spark 1.3 Max at #8 are exactly the kind of thing I wanted flagged. Midway through doing this the Internet Archive went down with a “Temporarily Offline” page, which is a nice reminder that the backfill trick is a trick and not infrastructure.&lt;/p&gt;

&lt;h2 id=&quot;two-things-it-got-wrong&quot;&gt;Two Things It Got Wrong&lt;/h2&gt;

&lt;h3 id=&quot;that-new-1-is-a-rename&quot;&gt;That “NEW #1” Is a Rename&lt;/h3&gt;

&lt;p&gt;Look at the last line again. &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;claude-fable-5-high&lt;/code&gt; at #1 as a brand new model. It isn’t. The database says so:&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;sqlite3 arena.db &lt;span class=&quot;s2&quot;&gt;&quot;SELECT model, taken, rank, score FROM ranks WHERE model LIKE &apos;claude-fable-5%&apos; ORDER BY model, taken;&quot;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;claude-fable-5|2026-09-02 19:07|1|1508
claude-fable-5|2026-09-09 04:01|1|1507
claude-fable-5|2026-09-17 15:33|1|1506
claude-fable-5-high|2026-09-23 01:27|1|1506
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;Same rank, same score, new name. This is the most common way change detection lies to you, and it isn’t a Firecrawl problem or a SQLite problem. It’s an identity problem: the thing you’re tracking needs a stable key, and the page doesn’t give you one.&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;/images/2026/track-new-ai-models-on-the-arena-leaderboard-body-4.jpg&quot; alt=&quot;That &quot; srcset=&quot;            /assets/resized/480/track-new-ai-models-on-the-arena-leaderboard-body-4.jpg 480w,            /assets/resized/800/track-new-ai-models-on-the-arena-leaderboard-body-4.jpg 800w,            /assets/resized/1400/track-new-ai-models-on-the-arena-leaderboard-body-4.jpg 1400w,    &quot; loading=&quot;lazy&quot; /&gt;&lt;/p&gt;

&lt;p&gt;The cheap fix is a sanity check before calling anything NEW: if an unseen name sits at the exact rank and score of a name that just disappeared, call it a rename. I left it out of the script because 70 lines that lie once a month teach more than 90 lines that hide it. Your mileage may vary once it wakes you up at 8am about a model that’s three months old.&lt;/p&gt;

&lt;h3 id=&quot;it-missed-the-one-that-started-this&quot;&gt;It Missed the One That Started This&lt;/h3&gt;

&lt;p&gt;GLM 5.3 Max, the model that made me want this in the first place, never shows up as NEW. It was already sitting at #18 on September 2, the oldest snapshot I loaded, so as far as the database knows it has always been there:&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;2026-09-02 19:07|18|1482
2026-09-09 04:01|20|1482
2026-09-17 15:33|19|1483
2026-09-23 01:27|19|1483
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;The harness can only see changes that happen after it starts watching. You can’t get the then after the fact. Start the cron before you need it.&lt;/p&gt;

&lt;p&gt;Also notice GLM wobbling from #18 to #20 to #19 while its score barely moves. Arena ranks a lot of models within a few points of each other near the top, and a model right on the #20 line will flicker in and out. Had it dropped to #21 for one run, the next run would have called it MOVED UP. If that gets noisy, the fix is to compare against the top 20 from any of the last few runs instead of just the last one.&lt;/p&gt;

&lt;h2 id=&quot;why-not-just-use-firecrawls-change-tracking&quot;&gt;Why Not Just Use Firecrawl’s Change Tracking?&lt;/h2&gt;

&lt;p&gt;Fair question, since Firecrawl has a &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;changeTracking&lt;/code&gt; format and a whole Monitor product built around scheduled checks. I didn’t use either here, and the reason is the leaderboard itself: the vote counts on every row change constantly, so “did this page change” is always yes. I want to know what changed in one column of one table, compared against a history I can query.&lt;/p&gt;

&lt;p&gt;That isn’t a knock on those features. For a page that should be static, a pricing page or a terms of service, “tell me when this changes” is exactly the right tool. They are worth their own post. This is the case where you want your own history, because the question you ask of it next month isn’t one you know yet.&lt;/p&gt;

&lt;h2 id=&quot;what-this-is-for&quot;&gt;What This Is For&lt;/h2&gt;

&lt;p&gt;As a standalone tool this is a small convenience. I’ll run it next to the Model Buzz Report and it will save me the “wait, was that there last week” squint.&lt;/p&gt;

&lt;p&gt;The reason it’s the first real recipe in this series is the shape. Scrape, save with a timestamp, compare with the last run, speak only when something changed. Every recipe coming after this one is that shape pointed at a different page with a different question: price history instead of rank history, a job board instead of a leaderboard. Firecrawl handles getting the page. These 70 lines handle remembering it.&lt;/p&gt;

&lt;p&gt;And the database is already there, filling up once a day, waiting for whatever question I think of next.&lt;/p&gt;
</description>
        <pubDate>Thu, 24 Sep 2026 07:00:00 -0500</pubDate>
        <link>https://www.stephanmiller.com/track-new-ai-models-on-the-arena-leaderboard/</link>
        <guid isPermaLink="true">https://www.stephanmiller.com/track-new-ai-models-on-the-arena-leaderboard/</guid>
        
        <category>Arena Leaderboard</category>
        
        <category>Firecrawl</category>
        
        <category>SQLite</category>
        
        <category>Data diff</category>
        
        <category>Model tracking</category>
        
        <category>Scheduled scraping</category>
        
        
        <category>web-scraping</category>
        
        <category>python</category>
        
      </item>
    
      <item>
        <title>Multiple Choice Is Not a Decision: Teaching a Planning Agent to Ask More Questions</title>
        <description>&lt;p&gt;I opened &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;PLAN.md&lt;/code&gt; on a new project last month. That’s the planning doc every project in my setup carries, and its decisions section is supposed to be the record of what the project has decided. Mine was still the blank template: two bullet points under the heading, and one of them was about where to put the output folder.&lt;/p&gt;

&lt;p&gt;This wasn’t a project where nothing had been decided. I’d been going back and forth on it for a week. The decisions existed. They were in a chat log somewhere, or in my head, or in the shape of the code I had already written. They weren’t in the document whose entire job was to hold them.&lt;/p&gt;

&lt;p&gt;I went and looked at the &lt;a href=&quot;https://github.com/eristoddle/agent-skills&quot;&gt;skill that owns these docs&lt;/a&gt; and found the real problem. It has four workflows, and every one of them reshapes material that already exists: create the files, retrofit them into a repo that has none, compact a doc that got fat, clean finished work out of the task file.&lt;/p&gt;

&lt;p&gt;Not one of them generated any of the files.&lt;/p&gt;

&lt;ul id=&quot;markdown-toc&quot;&gt;
  &lt;li&gt;&lt;a href=&quot;#multiple-choice-is-not-a-decision&quot; id=&quot;markdown-toc-multiple-choice-is-not-a-decision&quot;&gt;Multiple Choice Is Not a Decision&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;#then-matt-pocock-wrote-it-down-in-28-lines&quot; id=&quot;markdown-toc-then-matt-pocock-wrote-it-down-in-28-lines&quot;&gt;Then Matt Pocock Wrote It Down in 28 Lines&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;#it-ends-in-a-conversation-my-docs-have-to-outlive-the-conversation&quot; id=&quot;markdown-toc-it-ends-in-a-conversation-my-docs-have-to-outlive-the-conversation&quot;&gt;It Ends in a Conversation. My Docs Have to Outlive the Conversation.&lt;/a&gt;    &lt;ul&gt;
      &lt;li&gt;&lt;a href=&quot;#1-the-round-ends-in-a-disposition-not-an-answer&quot; id=&quot;markdown-toc-1-the-round-ends-in-a-disposition-not-an-answer&quot;&gt;1. The Round Ends in a Disposition, Not an Answer&lt;/a&gt;&lt;/li&gt;
      &lt;li&gt;&lt;a href=&quot;#2-it-is-re-entrant&quot; id=&quot;markdown-toc-2-it-is-re-entrant&quot;&gt;2. It Is Re-Entrant&lt;/a&gt;&lt;/li&gt;
      &lt;li&gt;&lt;a href=&quot;#3-the-files-append-they-never-get-rewritten&quot; id=&quot;markdown-toc-3-the-files-append-they-never-get-rewritten&quot;&gt;3. The Files Append, They Never Get Rewritten&lt;/a&gt;&lt;/li&gt;
      &lt;li&gt;&lt;a href=&quot;#4-the-relentlessness-had-to-go&quot; id=&quot;markdown-toc-4-the-relentlessness-had-to-go&quot;&gt;4. The Relentlessness Had to Go&lt;/a&gt;&lt;/li&gt;
    &lt;/ul&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;#it-is-a-utility-not-a-stage&quot; id=&quot;markdown-toc-it-is-a-utility-not-a-stage&quot;&gt;It Is a Utility, Not a Stage&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;#what-i-would-tell-you&quot; id=&quot;markdown-toc-what-i-would-tell-you&quot;&gt;What I Would Tell You&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;h2 id=&quot;multiple-choice-is-not-a-decision&quot;&gt;Multiple Choice Is Not a Decision&lt;/h2&gt;

&lt;p&gt;I’d been doing this part wrong on my own, well before any of this &lt;a href=&quot;https://www.stephanmiller.com/the-agent-skills-guide-i-wish-id-had/&quot;&gt;got written into a skill&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;When you plan something with an agent, the conversation drifts into a shape almost immediately. You describe what you want. It comes back with options. A, B, or C, with a short paragraph on each, and sometimes a helpful little table. You pick the one that’s least wrong, it says great choice, and you move on to the next menu.&lt;/p&gt;

&lt;p&gt;Nobody argued with you, nothing got stress tested, and the option you picked was picked because it was the best of three things a model generated in two seconds.&lt;/p&gt;

&lt;p&gt;I got tired of it. What I actually wanted out of planning was an argument, where I’ve to say &lt;em&gt;why&lt;/em&gt; and get pushed on the answer.&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;/images/2026/my-planning-doc-claimed-to-record-decisions-it-had-body-1.jpg&quot; alt=&quot;Multiple Choice Is Not a Decision&quot; srcset=&quot;            /assets/resized/480/my-planning-doc-claimed-to-record-decisions-it-had-body-1.jpg 480w,            /assets/resized/800/my-planning-doc-claimed-to-record-decisions-it-had-body-1.jpg 800w,            /assets/resized/1400/my-planning-doc-claimed-to-record-decisions-it-had-body-1.jpg 1400w,    &quot; loading=&quot;lazy&quot; /&gt;&lt;/p&gt;

&lt;p&gt;So I’d been doing a crude version of this by hand. Telling the agent to stop offering me options and start asking me questions. It sort of worked, the way anything sort of works when you’re improvising it fresh every session with no structure behind it.&lt;/p&gt;

&lt;h2 id=&quot;then-matt-pocock-wrote-it-down-in-28-lines&quot;&gt;Then Matt Pocock Wrote It Down in 28 Lines&lt;/h2&gt;

&lt;p&gt;The idea I built on isn’t mine. It’s &lt;a href=&quot;https://github.com/mattpocock/skills/blob/main/skills/productivity/grilling/SKILL.md&quot;&gt;Matt Pocock’s &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;grilling&lt;/code&gt; skill&lt;/a&gt; and I want that up front rather than buried in a credits line at the bottom, because the core of it’s his and it’s the good part.&lt;/p&gt;

&lt;p&gt;It’s 28 lines. 319 words. It opens with this:&lt;/p&gt;

&lt;blockquote&gt;
  &lt;p&gt;Interview the user relentlessly until you reach a shared understanding.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Then it gives you the structure that makes that possible, which is what I’d been missing. Model the conversation as a &lt;strong&gt;design tree&lt;/strong&gt;, where every decision branches into the decisions hanging off it. Work the tree in &lt;strong&gt;rounds&lt;/strong&gt;. The &lt;strong&gt;frontier&lt;/strong&gt; is every decision whose prerequisites are already settled, which is to say the questions you can ask right now without guessing at answers you haven’t heard yet.&lt;/p&gt;

&lt;p&gt;Two rules in there are doing most of the work.&lt;/p&gt;

&lt;p&gt;The first is ordering. A question whose answer depends on another question still open in this round belongs to a later round. That’s most of the technique right there. Violate it and you get answers the user has to retract two rounds later, which is exactly what my hand-rolled version kept doing.&lt;/p&gt;

&lt;p&gt;The second is that every question ships with the agent’s recommended answer, so the cheap reply is “yes to all but Q3.”&lt;/p&gt;

&lt;p&gt;And finding facts is the agent’s job. If a question needs to know what’s in the filesystem or what some dependency actually does, it goes and looks. The decisions stay yours.&lt;/p&gt;

&lt;p&gt;The session is done when the frontier is empty: every branch visited, nothing left assumed.&lt;/p&gt;

&lt;p&gt;I read it and wanted to use it, and immediately hit the thing that made it not fit.&lt;/p&gt;

&lt;h2 id=&quot;it-ends-in-a-conversation-my-docs-have-to-outlive-the-conversation&quot;&gt;It Ends in a Conversation. My Docs Have to Outlive the Conversation.&lt;/h2&gt;

&lt;p&gt;Grilling ends at a shared understanding held between two parties in a chat window. That’s a perfectly good place for it to end if you’re making one decision today.&lt;/p&gt;

&lt;p&gt;I’m not. These projects run for months. The entire premise of &lt;a href=&quot;https://www.stephanmiller.com/the-third-attempt-how-a-living-plan-beat-both-vibe-coding-and-spec-kit/&quot;&gt;this whole setup&lt;/a&gt; is that the documents are the memory, because &lt;a href=&quot;https://www.stephanmiller.com/i-got-tired-of-ai-memory-hype-so-i-built-a-context-lake/&quot;&gt;the agent’s context window&lt;/a&gt; isn’t. An interview that produces nothing but a shared understanding produces nothing at all when the session ends.&lt;/p&gt;

&lt;p&gt;So four things changed on the way in.&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;/images/2026/my-planning-doc-claimed-to-record-decisions-it-had-body-2.jpg&quot; alt=&quot;1. The Round Ends in a Disposition, Not an Answer&quot; srcset=&quot;            /assets/resized/480/my-planning-doc-claimed-to-record-decisions-it-had-body-2.jpg 480w,            /assets/resized/800/my-planning-doc-claimed-to-record-decisions-it-had-body-2.jpg 800w,            /assets/resized/1400/my-planning-doc-claimed-to-record-decisions-it-had-body-2.jpg 1400w,    &quot; loading=&quot;lazy&quot; /&gt;&lt;/p&gt;

&lt;h3 id=&quot;1-the-round-ends-in-a-disposition-not-an-answer&quot;&gt;1. The Round Ends in a Disposition, Not an Answer&lt;/h3&gt;

&lt;p&gt;The frontier doesn’t empty because everything got answered. It empties because everything got &lt;strong&gt;filed&lt;/strong&gt;. Every node in the tree closes as exactly one of four things:&lt;/p&gt;

&lt;table&gt;
  &lt;thead&gt;
    &lt;tr&gt;
      &lt;th&gt;Disposition&lt;/th&gt;
      &lt;th&gt;Lands in&lt;/th&gt;
      &lt;th&gt;Meaning&lt;/th&gt;
    &lt;/tr&gt;
  &lt;/thead&gt;
  &lt;tbody&gt;
    &lt;tr&gt;
      &lt;td&gt;&lt;strong&gt;Decided&lt;/strong&gt;&lt;/td&gt;
      &lt;td&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;Decisions&lt;/code&gt;, as a &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;Dxx&lt;/code&gt; record&lt;/td&gt;
      &lt;td&gt;settled, load-bearing, act on it&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;&lt;strong&gt;Open question&lt;/strong&gt;&lt;/td&gt;
      &lt;td&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;docs/questions/Qxx.md&lt;/code&gt; plus a one-line pointer&lt;/td&gt;
      &lt;td&gt;matters, not answerable yet&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;&lt;strong&gt;Parked&lt;/strong&gt;&lt;/td&gt;
      &lt;td&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;docs/parking-lot/Pxx.md&lt;/code&gt; plus a one-line pointer&lt;/td&gt;
      &lt;td&gt;might matter later, not now&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;&lt;strong&gt;N/A&lt;/strong&gt;&lt;/td&gt;
      &lt;td&gt;nothing&lt;/td&gt;
      &lt;td&gt;branch does not apply, say so once&lt;/td&gt;
    &lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;

&lt;p&gt;The line in the workflow that explains why is the one I’d keep if I had to throw the rest away:&lt;/p&gt;

&lt;blockquote&gt;
  &lt;p&gt;A blind spot is an &lt;em&gt;unasked&lt;/em&gt; question, not an unanswered one.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;“I don’t know yet” is an acceptable outcome. It becomes a file. What isn’t acceptable is a branch nobody ever walked down.&lt;/p&gt;

&lt;p&gt;It writes at &lt;strong&gt;every round boundary&lt;/strong&gt; for the obvious reason that I stop things halfway constantly and a workflow that only saves on completion would lose everything every time I do.&lt;/p&gt;

&lt;h3 id=&quot;2-it-is-re-entrant&quot;&gt;2. It Is Re-Entrant&lt;/h3&gt;

&lt;p&gt;The upstream is a one-shot interview. Mine has two modes.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Greenfield&lt;/strong&gt; is the one-shot case: there’s no planning doc yet, so the tree starts at the root; with nowhere to file, nothing gets written during the session, and at the end it hands the whole disposition to the scaffold workflow, which emits a &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;PLAN.md&lt;/code&gt; that’s already populated. If you quit early it scaffolds with whatever you settled anyway. Three decisions and six open questions beats an empty template.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Continuing&lt;/strong&gt; is the one I actually use. It reads the existing planning doc and the &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;docs/&lt;/code&gt; tree and seeds the design tree from them before asking anything:&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;existing decisions become settled nodes, pruned rather than re-litigated, and each one pushes the frontier outward because the questions hanging off it are exactly what just became askable&lt;/li&gt;
  &lt;li&gt;open questions come back as frontier nodes carrying whatever partial answers previous rounds accumulated&lt;/li&gt;
  &lt;li&gt;parked items come back &lt;strong&gt;only if something settled since they were parked makes them answerable now&lt;/strong&gt;, because dragging every parked idea back every session is how a parking lot turns into noise you learn to skip&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;img src=&quot;/images/2026/my-planning-doc-claimed-to-record-decisions-it-had-body-3.jpg&quot; alt=&quot;2. It Is Re-Entrant&quot; srcset=&quot;            /assets/resized/480/my-planning-doc-claimed-to-record-decisions-it-had-body-3.jpg 480w,            /assets/resized/800/my-planning-doc-claimed-to-record-decisions-it-had-body-3.jpg 800w,            /assets/resized/1400/my-planning-doc-claimed-to-record-decisions-it-had-body-3.jpg 1400w,    &quot; loading=&quot;lazy&quot; /&gt;&lt;/p&gt;

&lt;p&gt;This turns the docs into a breadcrumb trail instead of a stack of unrelated one-shot interviews.&lt;/p&gt;

&lt;h3 id=&quot;3-the-files-append-they-never-get-rewritten&quot;&gt;3. The Files Append, They Never Get Rewritten&lt;/h3&gt;

&lt;p&gt;A &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;Qxx.md&lt;/code&gt; touched by a later session gets a new dated section added under the existing ones.&lt;/p&gt;

&lt;p&gt;I had to put that in as an explicit rule because the instinct to tidy is strong. That growing, slightly repetitive, occasionally contradictory file is the record of what you thought in June that made the July answer obvious. Rewrite it into a clean summary and you’ve thrown away the reasoning.&lt;/p&gt;

&lt;h3 id=&quot;4-the-relentlessness-had-to-go&quot;&gt;4. The Relentlessness Had to Go&lt;/h3&gt;

&lt;p&gt;“Relentlessly” is right there in the first line upstream, and for a one-shot session it’s correct. A grill that gives up when you get tired isn’t doing its job.&lt;/p&gt;

&lt;p&gt;But relentless plus re-entrant is just nagging. If the thing can come back next week, it doesn’t need to squeeze everything out of you today. So the stop rule is written as an invariant rather than a preference:&lt;/p&gt;

&lt;blockquote&gt;
  &lt;p&gt;Stopping is never negotiated.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;No “are you sure”, no “we haven’t covered X yet”, no one more round, and specifically no listing what I’m about to miss, because listing what I’m about to miss is just arguing with extra steps. If I name a branch, it parks that branch and keeps going elsewhere.&lt;/p&gt;

&lt;p&gt;The reason that’s safe is the rule from change one. Everything unvisited gets filed on the way out.&lt;/p&gt;

&lt;blockquote&gt;
  &lt;p&gt;Nothing is lost by stopping, which is precisely why stopping needs no defense.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2 id=&quot;it-is-a-utility-not-a-stage&quot;&gt;It Is a Utility, Not a Stage&lt;/h2&gt;

&lt;p&gt;Every other route in this skill is triggered by what the repo looks like. No planning doc means grill first, then scaffold. Planning doc but no task file means adopt. Fat planning doc means rebalance. Task doc full of finished work means evict.&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;/images/2026/my-planning-doc-claimed-to-record-decisions-it-had-body-4.jpg&quot; alt=&quot;It Is a Utility, Not a Stage&quot; srcset=&quot;            /assets/resized/480/my-planning-doc-claimed-to-record-decisions-it-had-body-4.jpg 480w,            /assets/resized/800/my-planning-doc-claimed-to-record-decisions-it-had-body-4.jpg 800w,            /assets/resized/1400/my-planning-doc-claimed-to-record-decisions-it-had-body-4.jpg 1400w,    &quot; loading=&quot;lazy&quot; /&gt;&lt;/p&gt;

&lt;p&gt;The grill has a row in that table that isn’t a filesystem state at all:&lt;/p&gt;

&lt;table&gt;
  &lt;thead&gt;
    &lt;tr&gt;
      &lt;th&gt;State&lt;/th&gt;
      &lt;th&gt;Detected by&lt;/th&gt;
      &lt;th&gt;Route&lt;/th&gt;
    &lt;/tr&gt;
  &lt;/thead&gt;
  &lt;tbody&gt;
    &lt;tr&gt;
      &lt;td&gt;Greenfield&lt;/td&gt;
      &lt;td&gt;no planning doc present&lt;/td&gt;
      &lt;td&gt;grill, then scaffold&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Partial&lt;/td&gt;
      &lt;td&gt;planning doc exists, pieces missing&lt;/td&gt;
      &lt;td&gt;adopt&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Mature&lt;/td&gt;
      &lt;td&gt;full system present, planning doc heavy&lt;/td&gt;
      &lt;td&gt;rebalance&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Task doc bloated&lt;/td&gt;
      &lt;td&gt;mostly finished work&lt;/td&gt;
      &lt;td&gt;evict-tasks&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;&lt;strong&gt;Deciding, not filing&lt;/strong&gt;&lt;/td&gt;
      &lt;td&gt;&lt;strong&gt;not a filesystem state&lt;/strong&gt;&lt;/td&gt;
      &lt;td&gt;&lt;strong&gt;grill&lt;/strong&gt;&lt;/td&gt;
    &lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;

&lt;p&gt;It fires on what I’m doing, not on what the directory contains. Two entries: I ask for it, or I’m clearly thinking out loud with no question attached and it offers in one line and waits. That second one has a rule attached. One line, then shut up.&lt;/p&gt;

&lt;p&gt;That’s why it sits at different points in the process rather than at the front of it. It isn’t step one of planning. It’s the thing you reach for at any point where you’re about to settle on an answer.&lt;/p&gt;

&lt;h2 id=&quot;what-i-would-tell-you&quot;&gt;What I Would Tell You&lt;/h2&gt;

&lt;p&gt;If you want the idea, &lt;a href=&quot;https://github.com/mattpocock/skills/blob/main/skills/productivity/grilling/SKILL.md&quot;&gt;go read Pocock’s 28 lines&lt;/a&gt; rather than my version. It’s the good part, it’s short, and you’ll be running it in five minutes.&lt;/p&gt;

&lt;p&gt;The next post is the last one in this series, and it’ll be the whole thing, start to finish, how it actually works, so nobody has to read a skill directory to figure it out.&lt;/p&gt;
</description>
        <pubDate>Wed, 23 Sep 2026 07:00:00 -0500</pubDate>
        <link>https://www.stephanmiller.com/my-planning-doc-claimed-to-record-decisions-it-had-two/</link>
        <guid isPermaLink="true">https://www.stephanmiller.com/my-planning-doc-claimed-to-record-decisions-it-had-two/</guid>
        
        <category>agentic planning</category>
        
        <category>agent skills</category>
        
        <category>planning workflow</category>
        
        
        <category>agentic-development</category>
        
      </item>
    
      <item>
        <title>Grok 4.7 Shipped. Even Elon Musk Said It Was Mid.</title>
        <description>&lt;p&gt;Last week I ended this roundup with a promise. I said I’d be back to find out whether Grok 4.7 actually exists yet, and whether its creator liked it any better once it did. Well. It exists. He does not seem to like it any better, and neither, so far, does anyone else.&lt;/p&gt;

&lt;p&gt;Grok 4.7 shipped Sunday, September 21. That’s the whole headline and the whole punchline at once. For two straight weeks Musk stood in public and marked his own unreleased model down, from “better than 4.6 in every way” to “beats everything on the board” to, finally, “roughly on par with Opus 5.0, not 5.1.” Then the model landed. And the independent benchmarks put it right about where he’d talked it down to. This almost never happens. Usually the shipped thing is worse than the hype. This time the hype had already deflated itself to match, and the model still landed under the deflated number. Read on, because the gap between what got announced and what got shipped is the whole story this week, and for once the shipped thing is the one that gets a fair shake.&lt;/p&gt;

&lt;p&gt;The quiet story is still down in the bargain bin, where the fifteen-cent model I keep telling you about held every one of its five categories and nothing changed except that it got a little faster. Sometimes the boring outcome is the important one. Let me walk you through both.&lt;/p&gt;

&lt;ul id=&quot;markdown-toc&quot;&gt;
  &lt;li&gt;&lt;a href=&quot;#grok-47-shipped-and-its-the-model-its-own-maker-warned-you-about&quot; id=&quot;markdown-toc-grok-47-shipped-and-its-the-model-its-own-maker-warned-you-about&quot;&gt;Grok 4.7 shipped, and it’s the model its own maker warned you about&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;#the-benchmark-you-cite-is-the-argument-youre-making&quot; id=&quot;markdown-toc-the-benchmark-you-cite-is-the-argument-youre-making&quot;&gt;The benchmark you cite is the argument you’re making&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;#the-cheap-model-kept-all-five-of-its-categories&quot; id=&quot;markdown-toc-the-cheap-model-kept-all-five-of-its-categories&quot;&gt;The cheap model kept all five of its categories&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;#the-cheapskate-picks&quot; id=&quot;markdown-toc-the-cheapskate-picks&quot;&gt;The cheapskate picks&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;#newest-still-isnt-best-and-its-still-funny&quot; id=&quot;markdown-toc-newest-still-isnt-best-and-its-still-funny&quot;&gt;Newest still isn’t best, and it’s still funny&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;#about-that-bill&quot; id=&quot;markdown-toc-about-that-bill&quot;&gt;About that bill&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;#whats-coming&quot; id=&quot;markdown-toc-whats-coming&quot;&gt;What’s coming&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;#the-honest-version&quot; id=&quot;markdown-toc-the-honest-version&quot;&gt;The honest version&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;h2 id=&quot;grok-47-shipped-and-its-the-model-its-own-maker-warned-you-about&quot;&gt;Grok 4.7 shipped, and it’s the model its own maker warned you about&lt;/h2&gt;

&lt;p&gt;Here are the specs, and this time they’re specs, not tweets. Grok 4.7 is live as of September 21 through the Grok API, Cursor, and Grok Build. It runs on a new base model at a claimed 2.1 trillion parameters, up about 40 percent from Grok 4.6’s 1.5 trillion. It takes a 500K-token context, multimodal input, and it’s priced at two dollars in and six out, with cached input at fifty cents. That’s the same sticker as Grok 4.6, which is the one genuinely good decision in this whole launch. xAI did not try to charge frontier money for a mid-frontier model. Credit where it’s due.&lt;/p&gt;

&lt;p&gt;Now the number that matters. On the Artificial Analysis intelligence index, the independent one that blends ten hard evals, Grok 4.7 scores a 46. The two models at the top of that board, Claude Fable 5.1 and GPT-6 Astra, both sit at 53. So the model xAI built a two-week hype cycle around lands seven full points behind the frontier, in the same neighborhood as models that are a good deal cheaper and a good deal older. Mid-pack. Exactly where Musk himself put it on the fourteenth, before it had a benchmark to its name.&lt;/p&gt;

&lt;p&gt;The thing I keep coming back to is that this is a functional pattern now, not a Grok quirk. xAI ships a model that is cheap and legitimately fine, then wraps it in language it cannot support. Grok 4.7 is a perfectly reasonable two-dollar coding model. It is not the smartest thing on Earth, nobody who has run it thinks it is, and the person who built it told you so a week early. If you priced it at what it is instead of announcing it as what it isn’t, this would be a good news week for xAI. Instead it’s a case study.&lt;/p&gt;

&lt;h2 id=&quot;the-benchmark-you-cite-is-the-argument-youre-making&quot;&gt;The benchmark you cite is the argument you’re making&lt;/h2&gt;

&lt;p&gt;Now the part worth reading past the headline for.&lt;/p&gt;

&lt;p&gt;xAI’s own launch materials lean hard on coding, and on coding the gains are real. On DeepSWE v1.1 at high effort, Grok 4.7 posts 71.0, up from Grok 4.6’s 65.2. On CursorBench 4.0, a test built around longer-running coding tasks, it hits 46.3 against 40.4 for its predecessor. Those are honest improvements. If you’re doing the kind of work those benchmarks measure, 4.7 is a real step up from 4.6 at the same price, and that’s a fine reason to switch.&lt;/p&gt;

&lt;p&gt;Then the-decoder ran the harder agentic test, Terminal-Bench 4.0, and the floor gave out. Grok 4.7 scored 26 percent. GPT-6 Astra hit 60 on the same test. Fable 5.1 hit 55. And the cheap DeepSeek V4.1 Flash edged Grok out at 27. So on the benchmark xAI put in the deck, Grok 4.7 looks like a solid upgrade, and on the benchmark it left out, it finishes behind a Chinese model that costs less. Both results are true. They’re measuring different things, and the aggregate index score of 46 is what you get when you stop cherry-picking and average it out.&lt;/p&gt;

&lt;p&gt;The lesson underneath keeps earning its keep. The benchmark somebody cites is the argument they’re making. When a launch deck shows you three benchmarks, the interesting question is always which ones it didn’t show you. xAI showed you DeepSWE and CursorBench. It did not show you Terminal-Bench. Now you know why.&lt;/p&gt;

&lt;h2 id=&quot;the-cheap-model-kept-all-five-of-its-categories&quot;&gt;The cheap model kept all five of its categories&lt;/h2&gt;

&lt;p&gt;Now the story that actually moves your bill, which as usual is the least dramatic one on the page.&lt;/p&gt;

&lt;p&gt;Two weeks ago GLM-5.3-Flash from Z.ai took the cheapest-good-model crown off Xiaomi’s MiMo v2.5 Pro. Last week it grabbed a fifth Arena category. This week it did the boring, valuable thing and simply held everything. It’s still the cheapest model inside the competitive band for five of the six categories: Overall, Coding, Instruction Following, Hard Prompts, and Math. Creative Writing is still the lone holdout, still a Gemini story, because GLM never cracked that band and probably won’t.&lt;/p&gt;

&lt;p&gt;The usage board agrees with the preference board, which is the part that makes this real rather than an Arena curiosity. On OpenRouter, which counts actual tokens on actual paid calls, GLM-5.3-Flash is sitting at number two on the entire platform, somewhere north of ten trillion tokens a week, behind DeepSeek V4 Flash and ahead of GPT-5.6 Luna and MiMo. Chinese-built models are running around 46 percent of all tokens on the platform now, with DeepSeek the single largest vendor at roughly 16 percent. People aren’t voting for these models. They’re running them, in production, with their own money.&lt;/p&gt;

&lt;p&gt;One thing genuinely improved this week, and it’s the caveat that used to matter most. The old “cheap but slow” knock on GLM keeps softening. Artificial Analysis now clocks it at about 89 output tokens a second on Z.ai’s own endpoint, comfortably above the roughly 75 median for open-weight models in its class. It was crawling along near 60 a few weeks ago. For a fifty-cent model, that’s the difference between something you tolerate in a chat window and something you can actually drop into an agent loop.&lt;/p&gt;

&lt;p&gt;Two catches, same as always. First, GLM-5.3-Flash scores a 42 on the hard-reasoning index where the frontier lives in the fifties. It’s cheap and preference-strong, not a deep-reasoning machine. For everyday work that’s a rounding error you’ll never feel. For genuinely hard problems, feel it. Second, the price. The seven-cents-in, quarter-out number you may still have in your head was a launch promo, and it died on September 9. List is fifteen cents in and fifty cents out. Arena’s price column is still cheerfully showing the dead promo, which is going to burn somebody who budgets off it. Build on fifty cents out. If you sized an August spend on the old number, it doubled on you three weeks ago and nobody sent a memo.&lt;/p&gt;

&lt;h2 id=&quot;the-cheapskate-picks&quot;&gt;The cheapskate picks&lt;/h2&gt;

&lt;p&gt;Same method every week. For each Arena category I take the leader’s rating, draw a band 50 points below it, and find the cheapest model still sitting inside that band. The whole premise is that Arena ratings cluster tight at the top, so the category leader is usually a rounding error better than something 20 to 100 times cheaper. I compute the band from the full table in code, not by eyeballing the first screen, because eyeballing it is exactly how you delete the entire cheap tail and accidentally crown the cheapest expensive model. I learned that one the hard way, in public, a couple months back.&lt;/p&gt;

&lt;p&gt;Bands ran deep again this week: around 60 models inside the Overall band, 64 in Coding, thinner in Creative and Math. GLM-5.3-Flash is quoted at list, fifteen and fifty, not the expired promo Arena is still displaying. Arena data is dated September 13.&lt;/p&gt;

&lt;table&gt;
  &lt;thead&gt;
    &lt;tr&gt;
      &lt;th&gt;Category&lt;/th&gt;
      &lt;th&gt;Leader&lt;/th&gt;
      &lt;th&gt;$ leader out&lt;/th&gt;
      &lt;th&gt;Cheapskate pick&lt;/th&gt;
      &lt;th&gt;$ pick out&lt;/th&gt;
      &lt;th&gt;Δ rating&lt;/th&gt;
      &lt;th&gt;Cheaper by&lt;/th&gt;
    &lt;/tr&gt;
  &lt;/thead&gt;
  &lt;tbody&gt;
    &lt;tr&gt;
      &lt;td&gt;Overall&lt;/td&gt;
      &lt;td&gt;claude-fable-5 (1506)&lt;/td&gt;
      &lt;td&gt;$50&lt;/td&gt;
      &lt;td&gt;GLM-5.3-Flash (1475, #29)&lt;/td&gt;
      &lt;td&gt;$0.50&lt;/td&gt;
      &lt;td&gt;−31&lt;/td&gt;
      &lt;td&gt;~100×&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Coding&lt;/td&gt;
      &lt;td&gt;claude-fable-5 (1552)&lt;/td&gt;
      &lt;td&gt;$50&lt;/td&gt;
      &lt;td&gt;GLM-5.3-Flash (1525, #20)&lt;/td&gt;
      &lt;td&gt;$0.50&lt;/td&gt;
      &lt;td&gt;−27&lt;/td&gt;
      &lt;td&gt;~100×&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Creative Writing&lt;/td&gt;
      &lt;td&gt;claude-fable-5 (1504)&lt;/td&gt;
      &lt;td&gt;$50&lt;/td&gt;
      &lt;td&gt;gemini-3-flash (1459, #26)&lt;/td&gt;
      &lt;td&gt;$3&lt;/td&gt;
      &lt;td&gt;−45&lt;/td&gt;
      &lt;td&gt;~16.7×&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Instruction Following&lt;/td&gt;
      &lt;td&gt;claude-opus-4-6-high (1513)&lt;/td&gt;
      &lt;td&gt;$25&lt;/td&gt;
      &lt;td&gt;GLM-5.3-Flash (1471, #29)&lt;/td&gt;
      &lt;td&gt;$0.50&lt;/td&gt;
      &lt;td&gt;−42&lt;/td&gt;
      &lt;td&gt;~50×&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Hard Prompts&lt;/td&gt;
      &lt;td&gt;claude-opus-4-6-high (1533)&lt;/td&gt;
      &lt;td&gt;$25&lt;/td&gt;
      &lt;td&gt;GLM-5.3-Flash (1498, #28)&lt;/td&gt;
      &lt;td&gt;$0.50&lt;/td&gt;
      &lt;td&gt;−35&lt;/td&gt;
      &lt;td&gt;~50×&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Math&lt;/td&gt;
      &lt;td&gt;claude-fable-5 (1526, prelim)&lt;/td&gt;
      &lt;td&gt;$50&lt;/td&gt;
      &lt;td&gt;GLM-5.3-Flash (1513, #7)&lt;/td&gt;
      &lt;td&gt;$0.50&lt;/td&gt;
      &lt;td&gt;−13&lt;/td&gt;
      &lt;td&gt;~100×&lt;/td&gt;
    &lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;

&lt;p&gt;A few notes on how to read that. Overall and Hard Prompts are the rows you can lean on: GLM-5.3-Flash is sitting on real vote counts there, ten thousand in Overall, and it’s cleanly the cheapest thing in a very deep band. Coding is a great rating on thin votes, twentieth place for fifty cents, but under three thousand votes behind it, so if you want certainty over the last few dollars, MiMo v2.5 Pro at rank 25 with sixteen thousand votes for eighty-seven cents is the steadier bet. Math is the one to squint at hardest: GLM shows seventh for fifty cents, which is loud, but it’s riding 441 preliminary votes on a board where the whole top is thin. Treat that row as a strong suggestion, not a promise. MiMo at the band edge or a Gemini Flash won’t embarrass you there either.&lt;/p&gt;

&lt;p&gt;Creative Writing stays Gemini because GLM never made the band, and the only value pick is gemini-3-flash at three dollars, which is still 17 times cheaper than a fifty-dollar leader. Nobody is paying fifty dollars a million tokens to draft blog intros.&lt;/p&gt;

&lt;h2 id=&quot;newest-still-isnt-best-and-its-still-funny&quot;&gt;Newest still isn’t best, and it’s still funny&lt;/h2&gt;

&lt;p&gt;I wrote almost this exact paragraph last week and I’m writing it again, because the pattern refuses to break and it keeps getting funnier.&lt;/p&gt;

&lt;p&gt;The number one model on Arena Overall is claude-fable-5. Not Fable 5.1, the newer one Anthropic shipped a few weeks back. The old one. And sitting near the top of Instruction Following and Hard Prompts, as the outright leader on both, is claude-opus-4-6, a model that is roughly a year old. It beats Opus 5. It beats Opus 4.7 and 4.8. In the blind test, where nobody can see the version number, the crowd keeps reaching for the model everyone in the timeline already moved on from.&lt;/p&gt;

&lt;p&gt;I’m not saying the new models are bad. I’m saying the booth keeps preferring the boring old one, and it’s happened enough weeks running that it’s a pattern, not a fluke. If you upgraded off Opus 4.6 because a bigger number came out, the people voting blind would like a word with you.&lt;/p&gt;

&lt;h2 id=&quot;about-that-bill&quot;&gt;About that bill&lt;/h2&gt;

&lt;p&gt;The horror story this week isn’t a model, it’s a loop, and it’s the same species of loop that keeps eating people alive in 2026.&lt;/p&gt;

&lt;p&gt;Google’s Mandiant team put out an enterprise-AI-risk report on September 16 with a clean, awful example in it. An accounting agent hit a runaway execution loop and fired off more than 15,000 high-cost API calls in under an hour. Roughly fifty thousand dollars, gone, before a human looked at a dashboard. No budget ceiling. No alert anyone acted on. Just a bot doing the same expensive thing over and over, faster than anyone was watching.&lt;/p&gt;

&lt;p&gt;Here’s the part that made me laugh and then wince. Grok 4.7’s launch copy sells it as a model “designed to better verify its own output.” That is a lovely sentence. It is also describing the exact capability every runaway agent lacks, the ability to notice it’s stuck in a loop and stop. A better self-verifier would genuinely help with this. A marketing line about one does nothing, and the fifty-thousand-dollar hour happened the same week the line got written. The token bill in 2026 goes wrong in three normal ways: the sticker lies about the task, the loop has no brakes, and the cheap price had an expiration date you didn’t read. This week served up all three.&lt;/p&gt;

&lt;h2 id=&quot;whats-coming&quot;&gt;What’s coming&lt;/h2&gt;

&lt;p&gt;Three to watch.&lt;/p&gt;

&lt;p&gt;Grok 4.7, now that it exists, gets to spend the next couple weeks accumulating actual Arena votes and EU availability, which lagged the US launch as xAI launches always do. I’ll be curious whether the blind test is kinder to it than the benchmarks were, or crueler.&lt;/p&gt;

&lt;p&gt;GPT-6 Astra is still finishing its rollout. It went from a handful of day-one orgs to the ChatGPT paid tiers, the OpenAI API, Azure, and AWS Bedrock over the past couple weeks. The Daybreak program that loosens the safety rails for vetted organizations is the piece worth watching, given this is the model that reportedly maxed out an exploit-writing benchmark a couple weeks ago.&lt;/p&gt;

&lt;p&gt;Gemini 4, pretraining done and everything else a rumor. Google keeps dripping out Flash models to stay in the conversation while the real thing bakes. Note the small print on the current one: Gemini 3.8 Flash’s 75-cents-and-3.75 pricing is an intro rate that doubles on January 1. Late 2026 is the vague window for the actual next-gen model, if you believe the tea leaves, and I’ve stopped believing Google’s dates on principle.&lt;/p&gt;

&lt;h2 id=&quot;the-honest-version&quot;&gt;The honest version&lt;/h2&gt;

&lt;p&gt;The clean narrative this week would be that xAI face-planted. That’s not quite it, and the truth is more useful.&lt;/p&gt;

&lt;p&gt;Grok 4.7 is a decent, cheap, honestly-priced coding model that got buried under a launch it couldn’t live up to, by a man who then spent two weeks digging the hole himself. The model is fine. The framing was the problem, and the framing is always the problem. Cost-per-token isn’t cost-per-task. The benchmark in the deck isn’t the benchmark that matters. And “most capable model yet” means whatever the person saying it needs it to mean this quarter.&lt;/p&gt;

&lt;p&gt;Underneath all of it, a fifteen-cent open-weight model from Z.ai held its five categories, got faster, and stayed the second most-used model on the planet without anyone holding a keynote about it. That’s the line that changes your life if you’re shipping something and paying the bill yourself. The frontier had a loud week arguing about who’s smartest. The floor just sat there being cheap and getting quietly better. You already know which one you’ll actually be running next month.&lt;/p&gt;

&lt;p&gt;I’ll be back next week to see whether the Arena crowd is any nicer to Grok 4.7 than its own creator was. Low bar. We’ll find out.&lt;/p&gt;
</description>
        <pubDate>Tue, 22 Sep 2026 08:00:00 -0500</pubDate>
        <link>https://www.stephanmiller.com/model-buzz-roundup-week-of-0916/</link>
        <guid isPermaLink="true">https://www.stephanmiller.com/model-buzz-roundup-week-of-0916/</guid>
        
        <category>llm</category>
        
        <category>openrouter</category>
        
        <category>model-roundup</category>
        
        
        <category>large-language-models</category>
        
      </item>
    
      <item>
        <title>Syncing Obsidian with Google Drive Is the Trickiest One. Here Is How Anyway.</title>
        <description>&lt;p&gt;Every method in my &lt;a href=&quot;https://www.stephanmiller.com/sync-obsidian-vault-across-devices/&quot;&gt;guide to syncing an Obsidian vault across devices&lt;/a&gt; got a section except this one. For about a year the Google Drive section basically said “don’t.”&lt;/p&gt;

&lt;p&gt;Google Drive is the cloud storage that the largest number of people already have, already pay for, and already have 15GB sitting idle in. It is also, of the four big consumer clouds, the one Obsidian gets along with worst.&lt;/p&gt;

&lt;p&gt;The answer is better now than it was. But only by degree. There are four real paths, all of them have a catch, and two of them are plugins with the same name written by different people. That last one alone has cost the Obsidian forums more confusion than any other sync question I have looked at. I am going to walk all four, name the catch on each, and then tell you when to stop trying.&lt;/p&gt;

&lt;ul id=&quot;markdown-toc&quot;&gt;
  &lt;li&gt;&lt;a href=&quot;#why-google-drive-is-the-hard-one&quot; id=&quot;markdown-toc-why-google-drive-is-the-hard-one&quot;&gt;Why Google Drive Is the Hard One&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;#the-desktop-path-mirror-never-stream&quot; id=&quot;markdown-toc-the-desktop-path-mirror-never-stream&quot;&gt;The Desktop Path: Mirror, Never Stream&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;#the-two-plugins-that-have-the-same-name&quot; id=&quot;markdown-toc-the-two-plugins-that-have-the-same-name&quot;&gt;The Two Plugins That Have the Same Name&lt;/a&gt;    &lt;ul&gt;
      &lt;li&gt;&lt;a href=&quot;#the-one-in-the-plugin-directory-richardx366&quot; id=&quot;markdown-toc-the-one-in-the-plugin-directory-richardx366&quot;&gt;The one in the plugin directory: richardx366&lt;/a&gt;&lt;/li&gt;
      &lt;li&gt;&lt;a href=&quot;#the-one-everybody-links-to-stravo1&quot; id=&quot;markdown-toc-the-one-everybody-links-to-stravo1&quot;&gt;The one everybody links to: stravo1&lt;/a&gt;&lt;/li&gt;
      &lt;li&gt;&lt;a href=&quot;#which-one&quot; id=&quot;markdown-toc-which-one&quot;&gt;Which one&lt;/a&gt;&lt;/li&gt;
    &lt;/ul&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;#remotely-save-pro-the-boring-paid-route&quot; id=&quot;markdown-toc-remotely-save-pro-the-boring-paid-route&quot;&gt;Remotely Save PRO, the Boring Paid Route&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;#android-which-everybody-assumes-is-easy&quot; id=&quot;markdown-toc-android-which-everybody-assumes-is-easy&quot;&gt;Android, Which Everybody Assumes Is Easy&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;#the-15gb-argument&quot; id=&quot;markdown-toc-the-15gb-argument&quot;&gt;The 15GB Argument&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;#should-you-just-use-something-else&quot; id=&quot;markdown-toc-should-you-just-use-something-else&quot;&gt;Should You Just Use Something Else?&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;h2 id=&quot;why-google-drive-is-the-hard-one&quot;&gt;Why Google Drive Is the Hard One&lt;/h2&gt;

&lt;p&gt;Dropbox syncs a folder. iCloud syncs a folder, badly, but it syncs a folder. Google Drive for Desktop decided a folder was ambitious.&lt;/p&gt;

&lt;p&gt;The current desktop client defaults to &lt;strong&gt;streaming&lt;/strong&gt;, which mounts your Drive as a virtual filesystem. Files show up in the file browser, but the bytes are not on your disk until something asks for them. For photos and spreadsheets that works. For a vault it is a disaster.&lt;/p&gt;

&lt;p&gt;Obsidian watches the vault directory for changes. That watcher expects a real filesystem underneath it, where a file that exists is a file you can read right now. A virtual drive breaks that assumption. You get notes that open blank and populate a second later. You get link resolution that misses files it should find. You get the plugin folder behaving strangely because a plugin tried to read its own settings before Drive had gotten around to materializing them.&lt;/p&gt;

&lt;p&gt;None of this fails loudly. Instead, it fails as weirdness, and you spend a week thinking Obsidian is buggy.&lt;/p&gt;

&lt;h2 id=&quot;the-desktop-path-mirror-never-stream&quot;&gt;The Desktop Path: Mirror, Never Stream&lt;/h2&gt;

&lt;p&gt;If you are going to do this on desktop, there is exactly one correct setting and it is not the default.&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;/images/2026/syncing-obsidian-with-google-drive-is-the-trickies-body-1.jpg&quot; alt=&quot;The Desktop Path: Mirror, Never Stream&quot; srcset=&quot;            /assets/resized/480/syncing-obsidian-with-google-drive-is-the-trickies-body-1.jpg 480w,            /assets/resized/800/syncing-obsidian-with-google-drive-is-the-trickies-body-1.jpg 800w,            /assets/resized/1400/syncing-obsidian-with-google-drive-is-the-trickies-body-1.jpg 1400w,    &quot; loading=&quot;lazy&quot; /&gt;&lt;/p&gt;

&lt;p&gt;In Drive for Desktop, open Settings, then Preferences, then &lt;strong&gt;Folders from Drive&lt;/strong&gt; in the left sidebar. Under “My Drive syncing options,” switch from &lt;strong&gt;Stream files&lt;/strong&gt; to &lt;strong&gt;Mirror files&lt;/strong&gt;. Mirroring keeps a full local copy on disk and syncs changes up. That is the behavior a vault needs, because now Obsidian is watching a real directory full of real files.&lt;/p&gt;

&lt;p&gt;Let any in-flight sync finish before you flip that switch. Google’s own docs warn about changing sync modes mid-sync, and this is not the folder you want to find out on.&lt;/p&gt;

&lt;p&gt;Then put the vault somewhere sane:&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;~/Google Drive/My Drive/Obsidian/VaultName/
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;One vault per folder. Do not nest a vault inside another vault, and do not point Obsidian at &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;My Drive&lt;/code&gt; itself, because you will end up with the watcher trying to index every PDF you have saved since 2014.&lt;/p&gt;

&lt;p&gt;Two things to know before you commit to this. Mirroring means the vault takes up its full size on every desktop you do this on, which for a text vault is nothing and for a vault full of PDFs is not nothing. And Drive’s conflict handling is to make a second file with a different name. It does not merge. It does not ask. You get &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;Note.md&lt;/code&gt; and &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;Note (1).md&lt;/code&gt; and it is on you to notice.&lt;/p&gt;

&lt;p&gt;That last part is true of Dropbox too, and I covered how to live with it in the &lt;a href=&quot;https://www.stephanmiller.com/obsidian-dropbox-sync/&quot;&gt;Dropbox sync post&lt;/a&gt;.&lt;/p&gt;

&lt;h2 id=&quot;the-two-plugins-that-have-the-same-name&quot;&gt;The Two Plugins That Have the Same Name&lt;/h2&gt;

&lt;p&gt;This is the part that wastes people’s afternoons.&lt;/p&gt;

&lt;p&gt;There are &lt;strong&gt;two different Obsidian plugins both called “Google Drive Sync.”&lt;/strong&gt; Different authors, different designs, different tradeoffs. Every forum thread I have read about Drive sync has at least one person answering about the wrong one. Search the community plugin browser and you find exactly one of them. Search GitHub and the top result is the other.&lt;/p&gt;

&lt;h3 id=&quot;the-one-in-the-plugin-directory-richardx366&quot;&gt;The one in the plugin directory: richardx366&lt;/h3&gt;

&lt;p&gt;&lt;a href=&quot;https://github.com/richardx366/Obsidian-Google-Drive&quot;&gt;Obsidian-Google-Drive&lt;/a&gt; by richardx366 is the one you can install the normal way, because it is in the official community plugin list as “Google Drive Sync.” As of this writing it is on version 3.1.1 and was updated within the last week, which is more than I can say for the alternative.&lt;/p&gt;

&lt;p&gt;It exists specifically to solve the iOS problem, and the README says so: “A plugin to make Obsidian work in Google Drive to enable access to iOS.”&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;/images/2026/syncing-obsidian-with-google-drive-is-the-trickies-body-2.jpg&quot; alt=&quot;The one in the plugin directory: richardx366&quot; srcset=&quot;            /assets/resized/480/syncing-obsidian-with-google-drive-is-the-trickies-body-2.jpg 480w,            /assets/resized/800/syncing-obsidian-with-google-drive-is-the-trickies-body-2.jpg 800w,            /assets/resized/1400/syncing-obsidian-with-google-drive-is-the-trickies-body-2.jpg 1400w,    &quot; loading=&quot;lazy&quot; /&gt;&lt;/p&gt;

&lt;p&gt;The thing to understand before you install it is how the auth works. You do not create your own Google Cloud project. You authenticate through a hosted service at &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ogd.richardxiong.com&lt;/code&gt; and paste the resulting refresh token into the plugin settings. That is a real convenience and it is also a third party sitting in the path between Obsidian and your Drive. You can point it at a self-hosted token endpoint instead if that bothers you, and if you are the kind of person it bothers, it should, and you should.&lt;/p&gt;

&lt;p&gt;Its README carries its own warnings. Back up first. Do not manually add files to the synced Drive folder, because it “will likely break functionality, potentially causing data loss.” Do not edit vault files outside Obsidian. Let syncs finish before you close the app.&lt;/p&gt;

&lt;h3 id=&quot;the-one-everybody-links-to-stravo1&quot;&gt;The one everybody links to: stravo1&lt;/h3&gt;

&lt;p&gt;&lt;a href=&quot;https://github.com/stravo1/obsidian-gdrive-sync&quot;&gt;obsidian-gdrive-sync&lt;/a&gt; by stravo1 has about three times the stars, which is why it is the one that shows up in every search result and every Reddit answer. It works on desktop, Android, and iOS.&lt;/p&gt;

&lt;p&gt;It is not in the community plugin directory. The author’s own README explains why, and I would rather quote it than paraphrase it:&lt;/p&gt;

&lt;blockquote&gt;
  &lt;p&gt;The plugin is under active development, new releases might introduce bugs, old releases maybe be incompatible with the new ones. This might lead to data loss.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That is a developer telling you to back up first. The author has said they will not submit it to the official list until it is stable enough not to risk anybody’s notes.&lt;/p&gt;

&lt;p&gt;Two documented limits on top of that. It is &lt;strong&gt;not optimized for vaults over about 1,000 files&lt;/strong&gt;, with large-vault work listed as in progress. And on iOS there is no normal install path for a plugin outside the directory, so you set the vault up on a desktop that already has the plugin working and copy the whole thing across.&lt;/p&gt;

&lt;p&gt;Installing it means BRAT or doing it by hand: download the release zip, extract, drop the folder into &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;.obsidian/plugins/&lt;/code&gt;, enable it. The &lt;a href=&quot;https://www.stephanmiller.com/how-to-install-obsidian-plugins/&quot;&gt;plugin install walkthrough&lt;/a&gt; covers the mechanics if you have never sideloaded one.&lt;/p&gt;

&lt;h3 id=&quot;which-one&quot;&gt;Which one&lt;/h3&gt;

&lt;p&gt;&lt;img src=&quot;/images/2026/syncing-obsidian-with-google-drive-is-the-trickies-body-3.jpg&quot; alt=&quot;Which one&quot; srcset=&quot;            /assets/resized/480/syncing-obsidian-with-google-drive-is-the-trickies-body-3.jpg 480w,            /assets/resized/800/syncing-obsidian-with-google-drive-is-the-trickies-body-3.jpg 800w,            /assets/resized/1400/syncing-obsidian-with-google-drive-is-the-trickies-body-3.jpg 1400w,    &quot; loading=&quot;lazy&quot; /&gt;&lt;/p&gt;

&lt;p&gt;If you want a plugin and you want it maintained, richardx366’s is the current answer, and the price is the hosted token service. If that price is unacceptable and you will not self-host the endpoint, stravo1’s is the alternative, and the price is that you are running release-candidate software that has been quiet for four months ago.&lt;/p&gt;

&lt;h2 id=&quot;remotely-save-pro-the-boring-paid-route&quot;&gt;Remotely Save PRO, the Boring Paid Route&lt;/h2&gt;

&lt;p&gt;Remotely Save is the plugin most people in this cluster end up on, and its free tier covers S3-compatible storage, Dropbox, WebDAV, and &lt;a href=&quot;https://www.stephanmiller.com/obsidian-onedrive-sync/&quot;&gt;basic OneDrive&lt;/a&gt; with no account at all. Google Drive is not on that list. Drive support sits behind &lt;strong&gt;PRO&lt;/strong&gt;, which needs an account at &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;remotelysave.com&lt;/code&gt; separate from anything Obsidian.&lt;/p&gt;

&lt;p&gt;PRO is free during beta, currently stated through January 1, 2027. The price after that has not been announced. I am flagging that plainly because “free right now” and “free” are different words, and signing up for an unnamed number later is a decision you should make on purpose.&lt;/p&gt;

&lt;p&gt;What you get for it is the thing the stravo1 plugin does not have yet: maturity. Remotely Save has been syncing vaults to object storage for years, it handles conflicts with more grace than “make a second file,” and it is in the official directory.&lt;/p&gt;

&lt;h2 id=&quot;android-which-everybody-assumes-is-easy&quot;&gt;Android, Which Everybody Assumes Is Easy&lt;/h2&gt;

&lt;p&gt;The assumption goes: Google Drive is a Google product, Android is a Google product, so this must be the one place it just works.&lt;/p&gt;

&lt;p&gt;It is not. Mirror mode is the thing that makes the desktop path work, and mirror mode is a feature of Drive for Desktop. Drive for Desktop is, as the name has been telling you this whole time, a desktop application. Nothing in the Android Drive app gives Obsidian a real local directory that a file watcher can sit on top of.&lt;/p&gt;

&lt;p&gt;That leaves you with in-app sync, which means either the stravo1 plugin or Remotely Save PRO. Both of them run inside Obsidian and talk to the API.&lt;/p&gt;

&lt;p&gt;There is a second Android trap underneath this one, and you meet it before you sync anything. Since Obsidian 1.8.10, Android asks where the vault should live: &lt;strong&gt;app storage&lt;/strong&gt; or &lt;strong&gt;device storage&lt;/strong&gt;. Obsidian recommends device storage, and the reason is that app storage isolates the vault from every other app on the phone. That blocks external tools like Syncthing, and it means uninstalling Obsidian deletes your local vault.&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;/images/2026/syncing-obsidian-with-google-drive-is-the-trickies-body-4.jpg&quot; alt=&quot;Android, Which Everybody Assumes Is Easy&quot; srcset=&quot;            /assets/resized/480/syncing-obsidian-with-google-drive-is-the-trickies-body-4.jpg 480w,            /assets/resized/800/syncing-obsidian-with-google-drive-is-the-trickies-body-4.jpg 800w,            /assets/resized/1400/syncing-obsidian-with-google-drive-is-the-trickies-body-4.jpg 1400w,    &quot; loading=&quot;lazy&quot; /&gt;&lt;/p&gt;

&lt;p&gt;For Drive specifically that trap is survivable, because Obsidian Sync and community sync plugins both still work under app storage. Your Drive plugin keeps running either way. Device storage is still the safer pick, and why is explained in the &lt;a href=&quot;https://www.stephanmiller.com/sync-obsidian-android-free/&quot;&gt;Android sync post&lt;/a&gt;.&lt;/p&gt;

&lt;h2 id=&quot;the-15gb-argument&quot;&gt;The 15GB Argument&lt;/h2&gt;

&lt;p&gt;Here is the case for going through all of this.&lt;/p&gt;

&lt;p&gt;Google gives you 15GB free, shared across Drive, Gmail, and Photos. Dropbox gives you 2GB. A text vault will never come close to either number, but a vault with a few hundred PDFs and a year of screenshots will blow right past 2GB. If you are already on a Google One plan, the storage is free.&lt;/p&gt;

&lt;p&gt;That is a real argument, but not the only variable. The question is “which cloud will cost me an afternoon of untangling duplicate notes in March.”&lt;/p&gt;

&lt;h2 id=&quot;should-you-just-use-something-else&quot;&gt;Should You Just Use Something Else?&lt;/h2&gt;

&lt;p&gt;For Google Drive specifically, more often than not: yes.&lt;/p&gt;

&lt;p&gt;I don’t say that about the other methods in this cluster. Dropbox is fine. &lt;a href=&quot;https://www.stephanmiller.com/syncthing-obsidian-sync/&quot;&gt;Syncthing&lt;/a&gt; is free, private, and works. iCloud works if you are all-Apple. Each of those has a type of user it fits well.&lt;/p&gt;

&lt;p&gt;Google Drive’s problem is that every path asks you for something the alternatives do not. Mirror mode and a folder Drive duplicate on conflict. A maintained plugin that routes your auth through somebody else’s server. A more popular plugin that stopped shipping in May. Or a paid tier whose price nobody has announced yet. None of those is disqualifying on its own. It is that Drive is the only method in this cluster where you have to pick which one you mind least.&lt;/p&gt;

&lt;p&gt;So use Google Drive if the storage is a hard requirement. Company account, family plan, a 15GB library of attachments you are not moving.&lt;/p&gt;

&lt;p&gt;If it is a preference rather than a requirement, point Remotely Save at S3 or Dropbox and get on with your life. Or run Syncthing and stop paying anybody. Same outcome, fewer README warnings.&lt;/p&gt;

&lt;p&gt;And whichever you pick: back the vault up somewhere that is not the thing you are syncing with. Sync is not backup. A method that replicates your mistake to four devices in under a second is not protecting you, and Google Drive is very, very fast.&lt;/p&gt;
</description>
        <pubDate>Mon, 21 Sep 2026 08:00:00 -0500</pubDate>
        <link>https://www.stephanmiller.com/obsidian-google-drive-sync/</link>
        <guid isPermaLink="true">https://www.stephanmiller.com/obsidian-google-drive-sync/</guid>
        
        <category>obsidian</category>
        
        <category>google drive</category>
        
        <category>sync</category>
        
        <category>android</category>
        
        
        <category>obsidian</category>
        
      </item>
    
      <item>
        <title>The Subagents Guide I Wish I&apos;d Had</title>
        <description>&lt;p&gt;A while back I wrote &lt;a href=&quot;/the-agent-skills-guide-i-wish-id-had/&quot;&gt;the skills guide I wish I’d had&lt;/a&gt;. Skills stop your agent from forgetting what it knows about your codebase. The other half is that your agent also forgets &lt;em&gt;who it’s supposed to be&lt;/em&gt;. Every fresh session, it shows up as the same eager generalist, ready to have a reasonable, mediocre opinion about anything you throw at it.&lt;/p&gt;

&lt;p&gt;Skills are the memory problem. Subagents are the identity problem. This is the post about the second one.&lt;/p&gt;

&lt;p&gt;Here’s the shortest version I can give you before the table of contents scares you off: &lt;strong&gt;a skill tells the model what your world is like; a subagent tells it what role to play.&lt;/strong&gt; You can have both. And if you’re a solo dev shipping five half-finished projects at once like I am, you especially want both, because you are the only specialist you’ve got and you can’t be in five roles.&lt;/p&gt;

&lt;p&gt;Half of this is the conceptual guide (what these things are, where the files live, how to write one), and that part I could have written from docs. The second half is the stuff I only know because two of my subagents have been in use for months across a dozen repos, and they have failed in ways no best-practices post prepared me for. Both are open source and I’ll link to the actual files, because the most useful thing I can show you isn’t my advice, it’s a prompt that’s been beaten into shape by real runs.&lt;/p&gt;

&lt;ul id=&quot;markdown-toc&quot;&gt;
  &lt;li&gt;&lt;a href=&quot;#the-problem-subagents-actually-solve&quot; id=&quot;markdown-toc-the-problem-subagents-actually-solve&quot;&gt;The Problem Subagents Actually Solve&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;#skills-vs-subagents-knowledge-vs-behavior&quot; id=&quot;markdown-toc-skills-vs-subagents-knowledge-vs-behavior&quot;&gt;Skills vs. Subagents: Knowledge vs. Behavior&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;#what-a-subagent-actually-is-in-claude-code&quot; id=&quot;markdown-toc-what-a-subagent-actually-is-in-claude-code&quot;&gt;What a Subagent Actually Is (in Claude Code)&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;#the-agent-word-is-overloaded-and-its-making-you-dumber&quot; id=&quot;markdown-toc-the-agent-word-is-overloaded-and-its-making-you-dumber&quot;&gt;The “Agent” Word Is Overloaded and It’s Making You Dumber&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;#writing-your-first-one&quot; id=&quot;markdown-toc-writing-your-first-one&quot;&gt;Writing Your First One&lt;/a&gt;    &lt;ul&gt;
      &lt;li&gt;&lt;a href=&quot;#what-goes-in-the-body&quot; id=&quot;markdown-toc-what-goes-in-the-body&quot;&gt;What goes in the body&lt;/a&gt;&lt;/li&gt;
      &lt;li&gt;&lt;a href=&quot;#a-quick-walk-through-a-bug-hunt-agent-that-doesnt-chase-the-last-commit&quot; id=&quot;markdown-toc-a-quick-walk-through-a-bug-hunt-agent-that-doesnt-chase-the-last-commit&quot;&gt;A quick walk-through: a bug-hunt agent that doesn’t chase the last commit&lt;/a&gt;&lt;/li&gt;
    &lt;/ul&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;#the-thing-i-had-backwards-a-subagent-is-cold-on-the-conversation&quot; id=&quot;markdown-toc-the-thing-i-had-backwards-a-subagent-is-cold-on-the-conversation&quot;&gt;The Thing I Had Backwards: A Subagent Is Cold on the Conversation&lt;/a&gt;    &lt;ul&gt;
      &lt;li&gt;&lt;a href=&quot;#run-state-because-sessions-die&quot; id=&quot;markdown-toc-run-state-because-sessions-die&quot;&gt;Run state, because sessions die&lt;/a&gt;&lt;/li&gt;
    &lt;/ul&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;#delegation-economics-who-pays-for-which-thinking&quot; id=&quot;markdown-toc-delegation-economics-who-pays-for-which-thinking&quot;&gt;Delegation Economics: Who Pays for Which Thinking&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;#give-it-a-budget-or-it-will-spend-everything&quot; id=&quot;markdown-toc-give-it-a-budget-or-it-will-spend-everything&quot;&gt;Give It a Budget or It Will Spend Everything&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;#the-bans-are-most-of-a-mature-agent&quot; id=&quot;markdown-toc-the-bans-are-most-of-a-mature-agent&quot;&gt;The Bans Are Most of a Mature Agent&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;#they-fail-silently-and-confidently&quot; id=&quot;markdown-toc-they-fail-silently-and-confidently&quot;&gt;They Fail Silently and Confidently&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;#keeping-the-contract-tight-and-the-antipatterns-that-bloat-it&quot; id=&quot;markdown-toc-keeping-the-contract-tight-and-the-antipatterns-that-bloat-it&quot;&gt;Keeping the Contract Tight (and the Antipatterns That Bloat It)&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;#the-two-that-actually-survived&quot; id=&quot;markdown-toc-the-two-that-actually-survived&quot;&gt;The Two That Actually Survived&lt;/a&gt;    &lt;ul&gt;
      &lt;li&gt;&lt;a href=&quot;#template-then-specialize&quot; id=&quot;markdown-toc-template-then-specialize&quot;&gt;Template, then specialize&lt;/a&gt;&lt;/li&gt;
    &lt;/ul&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;#the-killer-combo-point-the-agent-at-knowledge-dont-paste-it-in&quot; id=&quot;markdown-toc-the-killer-combo-point-the-agent-at-knowledge-dont-paste-it-in&quot;&gt;The Killer Combo: Point the Agent at Knowledge, Don’t Paste It In&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;#where-they-live-personal-project-packaged&quot; id=&quot;markdown-toc-where-they-live-personal-project-packaged&quot;&gt;Where They Live: Personal, Project, Packaged&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;#how-this-works-in-the-other-tools&quot; id=&quot;markdown-toc-how-this-works-in-the-other-tools&quot;&gt;How This Works in the Other Tools&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;#how-good-ones-actually-get-built&quot; id=&quot;markdown-toc-how-good-ones-actually-get-built&quot;&gt;How Good Ones Actually Get Built&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;h2 id=&quot;the-problem-subagents-actually-solve&quot;&gt;The Problem Subagents Actually Solve&lt;/h2&gt;

&lt;p&gt;Your coding agent has one default mode: helpful generalist. Ask it to review a pull request and you get generically reasonable feedback. Ask it to figure out why your app is slow and it starts poking at the last file you touched. Ask it to plan a database migration and it hands you something sensible that ignores three things that would bite you in production.&lt;/p&gt;

&lt;p&gt;The model isn’t dumb. It just doesn’t have a &lt;em&gt;role&lt;/em&gt;. A generalist doing a security pass thinks about different things than a security engineer would. A generalist chasing a bug looks at different evidence than someone who’s been on call and seen that exact failure three times already. The intelligence is there. The framing isn’t.&lt;/p&gt;

&lt;p&gt;I noticed this because I kept typing the same preamble. Before I’d let Claude Code review anything, I’d paste in some version of “review this like a paranoid security person, look at auth boundaries first, tell me about anything leaking into logs, don’t waste my time with style nits.” Third time I typed that in two weeks, the light went on. That preamble is a role. And a role you keep re-typing is a subagent you haven’t written yet.&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;/images/2026/the-subagents-guide-i-wish-id-had-body-1.jpg&quot; alt=&quot;The Problem Subagents Actually Solve&quot; srcset=&quot;            /assets/resized/480/the-subagents-guide-i-wish-id-had-body-1.jpg 480w,            /assets/resized/800/the-subagents-guide-i-wish-id-had-body-1.jpg 800w,            /assets/resized/1400/the-subagents-guide-i-wish-id-had-body-1.jpg 1400w,    &quot; loading=&quot;lazy&quot; /&gt;&lt;/p&gt;

&lt;p&gt;That’s the trigger every write-up on this gives you, mine included. But it’s not why either of my two durable subagents exists.&lt;/p&gt;

&lt;p&gt;The two that actually survived came from different pressures:&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;&lt;strong&gt;Context economics.&lt;/strong&gt; Research grinding (twenty searches, a dozen fetched pages, half of them junk) will fill your main session with garbage you never wanted to read. That work has to happen somewhere else and come back as a summary. That’s &lt;a href=&quot;https://github.com/eristoddle/deep-research-agent&quot;&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;web-search-agent&lt;/code&gt;&lt;/a&gt;, and the reason it exists is not that it’s smarter than the main thread. It’s that I don’t want what it read in my context.&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Role separation.&lt;/strong&gt; When I’m designing something, I want to stay in the design conversation. I don’t want to be interrupted to approve the eleventh mechanical file edit that follows obviously from a decision I already made. So the decisions stay with me and the typing goes to an &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;implementer&lt;/code&gt; agent. The split is about which of us should be thinking about what, not about capability.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The first one gets you a nice reviewer. The other two are what turn subagents from a convenience into how the project actually runs.&lt;/p&gt;

&lt;h2 id=&quot;skills-vs-subagents-knowledge-vs-behavior&quot;&gt;Skills vs. Subagents: Knowledge vs. Behavior&lt;/h2&gt;

&lt;p&gt;A &lt;strong&gt;skill&lt;/strong&gt; packages what the agent needs to &lt;em&gt;know&lt;/em&gt;. Your weird internal library. The environment quirk that breaks builds. The domain knowledge a smart new hire would have to be told because there’s no way to guess it. Conditionally loaded. I wrote a whole &lt;a href=&quot;/the-agent-skills-guide-i-wish-id-had/&quot;&gt;guide on those&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;A &lt;strong&gt;subagent&lt;/strong&gt; packages how the agent should &lt;em&gt;behave&lt;/em&gt;. What it optimizes for. What it checks first. What tradeoffs it makes by default. What shape the output comes back in. It doesn’t teach the model anything new. It biases the intelligence that’s already in there toward one job.&lt;/p&gt;

&lt;p&gt;The clean test is to imagine calling in a specialist coworker for the task. Is the value they bring mostly &lt;em&gt;information you don’t have&lt;/em&gt;: domain knowledge, context, tribal know-how? That’s a skill. Or is the value mostly &lt;em&gt;how they approach the problem&lt;/em&gt;: what they look at first, what they weight heavily, what they refuse to sign off without checking, what they hand back? That’s a subagent.&lt;/p&gt;

&lt;p&gt;A security engineer reviewing your code doesn’t just know more than you. They &lt;em&gt;work&lt;/em&gt; differently. They look at auth boundaries first. They weight a privilege escalation path way higher than an ugly variable name. They won’t close the review without saying something about secrets. And they hand you structured findings. That ordering of attention is the thing a subagent encodes.&lt;/p&gt;

&lt;p&gt;Knowledge = skill. Methodology = subagent.&lt;/p&gt;

&lt;h2 id=&quot;what-a-subagent-actually-is-in-claude-code&quot;&gt;What a Subagent Actually Is (in Claude Code)&lt;/h2&gt;

&lt;p&gt;I’m Claude Code first here, same as the skills guide, because that’s my daily driver. Every other tool gets its section further down, quirks and all.&lt;/p&gt;

&lt;p&gt;In Claude Code, a subagent is a Markdown file with a little YAML frontmatter on top. It lives in one of two places:&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;~/.claude/agents/&lt;/code&gt;: your personal library, available in every project on your machine. This is the sandbox.&lt;/li&gt;
  &lt;li&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;.claude/agents/&lt;/code&gt; inside a repo: scoped to that project, and if you commit it, it travels with the repo.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The frontmatter is small. The fields you’ll use:&lt;/p&gt;

&lt;div class=&quot;language-markdown highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;nn&quot;&gt;---&lt;/span&gt;
&lt;span class=&quot;na&quot;&gt;name&lt;/span&gt;&lt;span class=&quot;pi&quot;&gt;:&lt;/span&gt; &lt;span class=&quot;s&quot;&gt;migration-risk-reviewer&lt;/span&gt;
&lt;span class=&quot;na&quot;&gt;description&lt;/span&gt;&lt;span class=&quot;pi&quot;&gt;:&lt;/span&gt; &lt;span class=&quot;s&quot;&gt;Reviews database migration plans and schema changes for rollback risk, lock contention, and data integrity problems. Use before running any migration against real data.&lt;/span&gt;
&lt;span class=&quot;na&quot;&gt;tools&lt;/span&gt;&lt;span class=&quot;pi&quot;&gt;:&lt;/span&gt; &lt;span class=&quot;s&quot;&gt;Read, Grep, Glob&lt;/span&gt;
&lt;span class=&quot;na&quot;&gt;model&lt;/span&gt;&lt;span class=&quot;pi&quot;&gt;:&lt;/span&gt; &lt;span class=&quot;s&quot;&gt;sonnet&lt;/span&gt;
&lt;span class=&quot;nn&quot;&gt;---&lt;/span&gt;

You are a senior database engineer reviewing a migration for production risk.

&lt;span class=&quot;gu&quot;&gt;## What you look at first&lt;/span&gt;
&lt;span class=&quot;p&quot;&gt;
-&lt;/span&gt; Can this be rolled back without losing data?
&lt;span class=&quot;p&quot;&gt;-&lt;/span&gt; Does it hold locks on a busy table during deploy?
&lt;span class=&quot;p&quot;&gt;-&lt;/span&gt; Is the ordering safe for a multi-step change?
&lt;span class=&quot;p&quot;&gt;-&lt;/span&gt; Will existing data violate any new constraint?

&lt;span class=&quot;gu&quot;&gt;## Output shape&lt;/span&gt;

Findings in order of severity:
&lt;span class=&quot;p&quot;&gt;
1.&lt;/span&gt; Blockers — deploy will fail or data will be lost
&lt;span class=&quot;p&quot;&gt;2.&lt;/span&gt; High risk — real production risk, needs a mitigation
&lt;span class=&quot;p&quot;&gt;3.&lt;/span&gt; Medium risk — should fix, won&apos;t necessarily block
&lt;span class=&quot;p&quot;&gt;4.&lt;/span&gt; Notes — worth tracking

For each finding: what it is, why it matters, what to do about it.
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;That’s a working subagent. A couple of things worth knowing about how Claude Code treats it:&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;/images/2026/the-subagents-guide-i-wish-id-had-body-3.jpg&quot; alt=&quot;Output shape&quot; srcset=&quot;            /assets/resized/480/the-subagents-guide-i-wish-id-had-body-3.jpg 480w,            /assets/resized/800/the-subagents-guide-i-wish-id-had-body-3.jpg 800w,            /assets/resized/1400/the-subagents-guide-i-wish-id-had-body-3.jpg 1400w,    &quot; loading=&quot;lazy&quot; /&gt;&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;&lt;strong&gt;The body is the system prompt.&lt;/strong&gt; When the subagent runs, everything below the frontmatter becomes its instructions. It’s not documentation you read and then act on. The model reads it and &lt;em&gt;becomes&lt;/em&gt; the thing.&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;The &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;description&lt;/code&gt; is load-bearing.&lt;/strong&gt; Claude Code uses it to decide when to hand a task off to this subagent automatically. Write it like an API someone else has to discover from context. “Use before running any migration” is findable. “helps with db stuff” is not. Write the &lt;em&gt;negative&lt;/em&gt; half too. Both of my real agents spend a clause on what they’re not for, because the routing will absolutely hand an agent work it has no business doing.&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;tools&lt;/code&gt; is an allowlist.&lt;/strong&gt; Leave it off and the subagent inherits everything. Narrow it and you’ve got a reviewer that literally can’t edit your files even if it gets an idea. For a review agent, that’s a feature. I don’t want my “just tell me what’s wrong” pass rewriting things on a whim.&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;model&lt;/code&gt; is a budget decision, not a detail.&lt;/strong&gt; Cheap model for mechanical work, expensive model for judgment. There’s a whole section on this below, because it’s most of why delegation pays for itself.&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;It gets its own context window.&lt;/strong&gt; This is the part people underrate: a subagent runs in a separate context, does its thing, and hands back a summary. Your main session doesn’t get flooded with everything it read. It’s also the part &lt;em&gt;I&lt;/em&gt; underrated in the opposite direction, and there’s a whole section below where it breaks, because the separate context is what breaks it.&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;You don’t have to hand-write the file.&lt;/strong&gt; You can just ask Claude Code to write the subagent for you: describe the role and it drops the file in the right place. The &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;/agents&lt;/code&gt; command is also there for managing them.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;You invoke one by asking for it, by letting Claude route to it based on that description, or by wiring it into a larger flow where it runs as a delegated worker. Same behavior either way.&lt;/p&gt;

&lt;h2 id=&quot;the-agent-word-is-overloaded-and-its-making-you-dumber&quot;&gt;The “Agent” Word Is Overloaded and It’s Making You Dumber&lt;/h2&gt;

&lt;p&gt;Not you specifically. Me. “Agent” gets bolted onto four different things and people conflate them constantly, so let me split them apart:&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;&lt;strong&gt;A subagent&lt;/strong&gt; (&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;.claude/agents/*.md&lt;/code&gt;, or a &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;.agent.md&lt;/code&gt; file over in Copilot land) is a &lt;em&gt;reusable file&lt;/em&gt; that defines a specialist role. This is the thing this whole post is about.&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Agent mode / autonomous run&lt;/strong&gt; is a &lt;em&gt;capability toggle&lt;/em&gt;: you’re letting the tool run commands, edit files, and generally act without you hitting approve on every line. That’s a permission setting, not a specialist.&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;A background/cloud coding agent&lt;/strong&gt; is a &lt;em&gt;service&lt;/em&gt; that picks up a task, churns on it out of sight, and comes back with a branch or a PR. Also not a specialist file.&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;CLAUDE.md&lt;/code&gt; / &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;AGENTS.md&lt;/code&gt;&lt;/strong&gt; is &lt;em&gt;always-on repo guidance&lt;/em&gt;. The house rules every agent obeys on that codebase. It’s the employee handbook, not a coworker you call by name.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;When I say subagent, I mean the reusable specialist you write once and invoke on purpose. Not the toggle, not the service, not the handbook. Keeping those four straight fixes about half the confusion.&lt;/p&gt;

&lt;h2 id=&quot;writing-your-first-one&quot;&gt;Writing Your First One&lt;/h2&gt;

&lt;p&gt;Start with the thinnest thing that does real work. A first subagent has exactly two jobs: declare the role, and say what that role actually means in practice. The &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;migration-risk-reviewer&lt;/code&gt; up above is already that: a clear role, specific heuristics, a predictable output shape, and it’s narrow on purpose. You’ll add more once you run it on real work and watch it miss things.&lt;/p&gt;

&lt;p&gt;The single most common way a first subagent flops is a mushy role. “Security helper” is not a role. It tells the model nothing about what to optimize for or check first. A real role is a behavioral contract: here’s the job, here’s what you look at first, here’s what you won’t let slide. If you can’t say the job in one sentence, it’s too vague.&lt;/p&gt;

&lt;table&gt;
  &lt;thead&gt;
    &lt;tr&gt;
      &lt;th&gt;Vague&lt;/th&gt;
      &lt;th&gt;Concrete&lt;/th&gt;
    &lt;/tr&gt;
  &lt;/thead&gt;
  &lt;tbody&gt;
    &lt;tr&gt;
      &lt;td&gt;“Security helper”&lt;/td&gt;
      &lt;td&gt;“Review backend changes for auth boundary violations, secrets leaking into logs, and privilege escalation paths”&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;“Migration assistant”&lt;/td&gt;
      &lt;td&gt;“Analyze schema changes for rollback safety, lock contention on busy tables, and data integrity risk on existing rows”&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;“Code reviewer”&lt;/td&gt;
      &lt;td&gt;“Review async code for missing error handling, swallowed exceptions, and N+1 query patterns”&lt;/td&gt;
    &lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;

&lt;p&gt;The concrete version tells the model what it cares about &lt;em&gt;and what it ignores.&lt;/em&gt; The ignoring is half the value. A reviewer that gives equal weight to a typo and a privilege boundary is only an averaging machine.&lt;/p&gt;

&lt;h3 id=&quot;what-goes-in-the-body&quot;&gt;What goes in the body&lt;/h3&gt;

&lt;p&gt;Once the role’s defined, a useful body covers five things. You don’t need all five on day one:&lt;/p&gt;

&lt;ol&gt;
  &lt;li&gt;&lt;strong&gt;What it optimizes for:&lt;/strong&gt; the one thing this role is most trying to get right. When it has to make a tradeoff, this is what it trades toward.&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;What it checks first:&lt;/strong&gt; the high-priority signals a real specialist always looks at before anything else.&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Failure modes:&lt;/strong&gt; &lt;a href=&quot;https://www.stephanmiller.com/somebody-finally-wrote-down-why-my-coding-agents-keep-failing-the-same-way/&quot;&gt;the specific things that go wrong in this class of work&lt;/a&gt;. This is where the actual expertise lives, and it’s mostly stuff you’ll only learn by watching the agent mess up real tasks.&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Output shape:&lt;/strong&gt; not “some feedback” but &lt;em&gt;severity-ordered findings&lt;/em&gt;, &lt;em&gt;a phased plan&lt;/em&gt;, &lt;em&gt;a ranked list of hypotheses&lt;/em&gt;. Predictable output is the single biggest practical win over just prompting.&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;What it must not do.&lt;/strong&gt; I used to list this one as optional. It is not optional. In my most-used agent it’s the longest section in the file, and every line of it is there because something went wrong once.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Start with role, check-first, and output shape. Everything else the agent earns by screwing up.&lt;/p&gt;

&lt;h3 id=&quot;a-quick-walk-through-a-bug-hunt-agent-that-doesnt-chase-the-last-commit&quot;&gt;A quick walk-through: a bug-hunt agent that doesn’t chase the last commit&lt;/h3&gt;

&lt;p&gt;Here’s one I wanted, because I kept hitting the same dumb pattern. Something breaks, I paste the error into Claude Code, and its first instinct is to go stare at whatever file I edited most recently. Sometimes that’s right. Sometimes the last change had nothing to do with it.&lt;/p&gt;

&lt;p&gt;An experienced debugger doesn’t start at the code. They start at the evidence (what’s actually failing, what the logs say, what changed in the environment) and only open the source once there’s a hypothesis worth checking. That’s a behavior. So I saved it:&lt;/p&gt;

&lt;div class=&quot;language-markdown highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;nn&quot;&gt;---&lt;/span&gt;
&lt;span class=&quot;na&quot;&gt;name&lt;/span&gt;&lt;span class=&quot;pi&quot;&gt;:&lt;/span&gt; &lt;span class=&quot;s&quot;&gt;bug-investigator&lt;/span&gt;
&lt;span class=&quot;na&quot;&gt;description&lt;/span&gt;&lt;span class=&quot;pi&quot;&gt;:&lt;/span&gt; &lt;span class=&quot;s&quot;&gt;Investigates a bug or failure starting from evidence, not from recent code changes. Produces ranked hypotheses with the next diagnostic step for each. Use for &quot;why is this broken&quot; before touching the source.&lt;/span&gt;
&lt;span class=&quot;na&quot;&gt;tools&lt;/span&gt;&lt;span class=&quot;pi&quot;&gt;:&lt;/span&gt; &lt;span class=&quot;s&quot;&gt;Read, Grep, Glob, Bash&lt;/span&gt;
&lt;span class=&quot;nn&quot;&gt;---&lt;/span&gt;

You are a senior engineer investigating a failure. Your job is to find the most
likely cause, rule out the alternatives, and say what to check next.

&lt;span class=&quot;gu&quot;&gt;## Investigation order&lt;/span&gt;
&lt;span class=&quot;p&quot;&gt;
1.&lt;/span&gt; Start from the symptom — what is actually failing, and how does it show up?
&lt;span class=&quot;p&quot;&gt;2.&lt;/span&gt; Get the error output and any logs from around the time it broke — ask the caller for them if they weren&apos;t provided.
&lt;span class=&quot;p&quot;&gt;3.&lt;/span&gt; Only then look at recent changes, and only if the evidence points there.
&lt;span class=&quot;p&quot;&gt;4.&lt;/span&gt; Read source last, once there&apos;s a symptom-to-cause hypothesis worth testing.

&lt;span class=&quot;gu&quot;&gt;## Output shape&lt;/span&gt;

&lt;span class=&quot;gs&quot;&gt;**Symptoms:**&lt;/span&gt; what&apos;s failing and how it manifests.
&lt;span class=&quot;gs&quot;&gt;**Likely causes (ranked):**&lt;/span&gt; for each — the hypothesis, the evidence for it, the next diagnostic step.
&lt;span class=&quot;gs&quot;&gt;**Ruled out:**&lt;/span&gt; what you checked and why it isn&apos;t the cause.
&lt;span class=&quot;gs&quot;&gt;**Next steps:**&lt;/span&gt; in priority order.
&lt;span class=&quot;gs&quot;&gt;**Open questions:**&lt;/span&gt; what you still can&apos;t confirm.

&lt;span class=&quot;gu&quot;&gt;## Failure modes to avoid&lt;/span&gt;
&lt;span class=&quot;p&quot;&gt;
-&lt;/span&gt; Don&apos;t assume the last change is the cause without evidence.
&lt;span class=&quot;p&quot;&gt;-&lt;/span&gt; Don&apos;t propose a fix before you can name the mechanism.
&lt;span class=&quot;p&quot;&gt;-&lt;/span&gt; Don&apos;t close with &quot;cause unknown&quot; — say what evidence would confirm or kill each remaining hypothesis.
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;Then (and this is the step people skip) I tested it on a real bug I’d already solved, not a toy. I gave it the symptom and the logs I’d had at the time and watched whether its investigation path matched the one that found the problem. Where it drifted, that drift became a new line in the body.&lt;/p&gt;

&lt;p&gt;That file is where everybody’s subagent post stops. Everything below is what happens after you’ve been running these things for months, which is the part I actually needed someone to tell me.&lt;/p&gt;

&lt;h2 id=&quot;the-thing-i-had-backwards-a-subagent-is-cold-on-the-conversation&quot;&gt;The Thing I Had Backwards: A Subagent Is Cold on the Conversation&lt;/h2&gt;

&lt;p&gt;The separate context window gets sold as pure upside. But it isn’t symmetric.&lt;/p&gt;

&lt;p&gt;This line is in my implementer agent, written after the same handoff failed three different ways:&lt;/p&gt;

&lt;blockquote&gt;
  &lt;p&gt;You are cold on the planning conversation but warm on the project.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Cold on the conversation: it never saw the two hours where you ruled out the obvious approach, argued yourself out of an abstraction, and settled on the weird-looking solution for a good reason. None of that exists for it. Warm on the project: it has &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;CLAUDE.md&lt;/code&gt; and it has the codebase, so it knows the house rules and the idioms without being told.&lt;/p&gt;

&lt;p&gt;Get that backwards in either direction and you get a specific, recognizable failure. Treat it as warm on the conversation and it confidently reimplements the approach you spent an hour rejecting, because from where it sits that approach looks fine. Treat it as cold on the project and you waste a thousand tokens re-explaining your own repo to something that can read it faster than you can describe it.&lt;/p&gt;

&lt;p&gt;What falls out of this is the thing no subagent guide told me to build: &lt;strong&gt;the handoff is a file, and the file is the contract.&lt;/strong&gt; Not a paragraph you type into the task prompt. A document on disk with fixed sections, because a prompt you improvise each time will leave out a different critical section every time.&lt;/p&gt;

&lt;p&gt;Mine is a &lt;a href=&quot;https://www.stephanmiller.com/the-other-doc-got-fat-too-my-task-file-was-90-finished-work/&quot;&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;TASKS.md&lt;/code&gt; at the repo root&lt;/a&gt;, generated by the &lt;a href=&quot;https://github.com/eristoddle/agent-skills&quot;&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;living-plan&lt;/code&gt;&lt;/a&gt; skill, and the active task has these sections and no others:&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;&lt;strong&gt;Goal:&lt;/strong&gt; what’s true when this is done.&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Why (pointer):&lt;/strong&gt; a link to &lt;a href=&quot;https://www.stephanmiller.com/the-third-attempt-how-a-living-plan-beat-both-vibe-coding-and-spec-kit/&quot;&gt;the decision in &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;PLAN.md&lt;/code&gt;&lt;/a&gt;, not a re-argument of it. Pointer, not prose.&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;▶ Run state:&lt;/strong&gt; the agent keeps this current. More on this in a second, because it’s the one that saves you.&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Design — numbered pieces:&lt;/strong&gt; a serial queue, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;[ ]&lt;/code&gt; not started, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;[x]&lt;/code&gt; done, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;[!]&lt;/code&gt; blocked, with dependencies marked inline as &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;depends on #2&lt;/code&gt;.&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Files:&lt;/strong&gt; where the work happens.&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Tests:&lt;/strong&gt; the checks that stand in for review.&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Out of scope (do NOT do):&lt;/strong&gt; the single highest-value section in the document.&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Report back:&lt;/strong&gt; what the final message has to contain.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The rule that makes it work: &lt;em&gt;fill every section before launching, and mark one &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;—&lt;/code&gt; only if it genuinely doesn’t apply.&lt;/em&gt; An unfilled section is where the agent improvises.&lt;/p&gt;

&lt;p&gt;And then the division of labor that took me longest to see. Two documents, two lifetimes:&lt;/p&gt;

&lt;blockquote&gt;
  &lt;p&gt;Standing “how we implement here” knowledge lives in the &lt;strong&gt;agent definition&lt;/strong&gt;, not in every task.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The agent file holds the protocol: how to read a task, how to handle a blocker, what to never touch, what the report has to contain. The task file holds this job. If you find yourself pasting the same instruction into three consecutive tasks, that instruction was never task content. It belongs in the agent.&lt;/p&gt;

&lt;h3 id=&quot;run-state-because-sessions-die&quot;&gt;Run state, because sessions die&lt;/h3&gt;

&lt;p&gt;Here’s the rule that has saved me more real work than any clever prompt engineering in this post:&lt;/p&gt;

&lt;blockquote&gt;
  &lt;p&gt;&lt;strong&gt;Log run-state whenever you stop — done OR blocked. Mandatory.&lt;/strong&gt; Flip each piece’s status box as you go, and update the &lt;strong&gt;▶ Run state&lt;/strong&gt; note (done / blocked+why / remaining / resume-from). Editing &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;TASKS.md&lt;/code&gt; for this is in-scope — it is the recovery point if the session dies.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Sessions die. You close the laptop, the connection drops, you hit a limit, you get bored and Ctrl-C something that was mostly finished. When that happens to a subagent, &lt;em&gt;everything it knew is gone&lt;/em&gt;. If the only record of progress was in its head, you now get to diff your own repo against your memory to figure out where it got to.&lt;/p&gt;

&lt;p&gt;So the agent writes its progress to disk as it goes, into the same file that gave it the job. The next session reads the file and picks up at the resume point. It costs a few lines in the agent definition and it turns a dead session from lost work into a paused one.&lt;/p&gt;

&lt;p&gt;The companion rule, which is about throughput rather than recovery:&lt;/p&gt;

&lt;blockquote&gt;
  &lt;p&gt;&lt;strong&gt;A blocked piece does NOT halt the queue.&lt;/strong&gt; If a piece is underspecified or hits a dependency you can’t resolve, mark it &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;[!]&lt;/code&gt; blocked with a one-line reason and &lt;strong&gt;continue with any remaining piece that doesn’t depend on it&lt;/strong&gt;. Halt only when nothing remaining can proceed. Never guess a design — escalate blocked forks in your report.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;img src=&quot;/images/2026/the-subagents-guide-i-wish-id-had-body-7.jpg&quot; alt=&quot;Run state, because sessions die&quot; srcset=&quot;            /assets/resized/480/the-subagents-guide-i-wish-id-had-body-7.jpg 480w,            /assets/resized/800/the-subagents-guide-i-wish-id-had-body-7.jpg 800w,            /assets/resized/1400/the-subagents-guide-i-wish-id-had-body-7.jpg 1400w,    &quot; loading=&quot;lazy&quot; /&gt;&lt;/p&gt;

&lt;p&gt;Default agent behavior on a snag is to stop and ask, which means a queue of eight tasks gets you one task and a question. Default behavior with no guardrail at all is worse: it guesses a design and keeps going, and now you’ve got seven pieces built on an invention you never approved. Skip-and-continue, never-guess, report-the-fork is the combination that lets you queue work and walk away.&lt;/p&gt;

&lt;h2 id=&quot;delegation-economics-who-pays-for-which-thinking&quot;&gt;Delegation Economics: Who Pays for Which Thinking&lt;/h2&gt;

&lt;p&gt;That &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;model&lt;/code&gt; line is why this whole arrangement pays for itself.&lt;/p&gt;

&lt;p&gt;My implementer’s description, verbatim from a real project:&lt;/p&gt;

&lt;blockquote&gt;
  &lt;p&gt;Sonnet implementation worker for the project. Use it to execute a fully-specified, mechanical coding task defined in TASKS.md while the main (Opus) planning thread keeps going. NOT for design decisions, exploring open questions, or live LLM / quality testing — those stay in the main thread.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;a href=&quot;https://www.stephanmiller.com/the-expensive-model-only-plans-now-i-split-my-ai-coding-rig-across-two-tools/&quot;&gt;Expensive model makes the decisions&lt;/a&gt;. Cheap model does the typing. And a narrow scope is precisely what makes the cheap model reliable: a Sonnet worker handed a fully-specified numbered queue in a codebase it can read is dependable.&lt;/p&gt;

&lt;p&gt;That split also gives you two working rhythms out of the same machinery:&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;&lt;strong&gt;Pacing.&lt;/strong&gt; Queue two or three pieces, launch the agent, and &lt;a href=&quot;https://www.stephanmiller.com/the-bottleneck-was-me-how-i-stopped-racing-my-ai-builder-and-started-pacing-it/&quot;&gt;keep planning the next batch while it grinds&lt;/a&gt;. You’re never waiting on it and it’s never waiting on you.&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Unattended.&lt;/strong&gt; Queue a lot and walk away. Come back to a diff, a report, and a run-state note telling you which piece blocked and why.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Permissions are the enforcement layer under this. My unattended implementer’s file and shell commands are allowlisted in the project’s permission settings; anything outside the list prompts for approval. The bans in the prompt are policy. The allowlist is the fence.&lt;/p&gt;

&lt;p&gt;One more line in that agent, which exists because of the most expensive mistake I made:&lt;/p&gt;

&lt;blockquote&gt;
  &lt;p&gt;&lt;strong&gt;Work the numbered pieces as a serial queue, top-to-bottom, in one pass.&lt;/strong&gt; Do not spawn parallel workers and do not stop to report after each piece.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Parallel agents sound like free speed. They are not free. Every one of them finishes by dumping its full output back into the context that spawned it, and a fan-out that looked clever can bury the session that has to read all of it. I’ve killed real work that way. Parallel runs across disjoint files are an opt-in tactic for a specific burst, never the default. The default is one worker doing one queue in order.&lt;/p&gt;

&lt;h2 id=&quot;give-it-a-budget-or-it-will-spend-everything&quot;&gt;Give It a Budget or It Will Spend Everything&lt;/h2&gt;

&lt;p&gt;Any subagent with tools (search, fetch, shell) will grind until something stops it, and “I think that’s enough” is not a thing a model reliably concludes on its own. A research agent without a budget is a machine that turns your afternoon into many (hidden) browser tabs and a summary you could have gotten from the first three.&lt;/p&gt;

&lt;p&gt;So &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;web-search-agent&lt;/code&gt; opens with hard limits, above the methodology, above everything:&lt;/p&gt;

&lt;table&gt;
  &lt;thead&gt;
    &lt;tr&gt;
      &lt;th&gt;Level&lt;/th&gt;
      &lt;th&gt;Searches&lt;/th&gt;
      &lt;th&gt;Fetches&lt;/th&gt;
      &lt;th&gt;Link depth&lt;/th&gt;
      &lt;th&gt;Modules&lt;/th&gt;
      &lt;th&gt;Use when&lt;/th&gt;
    &lt;/tr&gt;
  &lt;/thead&gt;
  &lt;tbody&gt;
    &lt;tr&gt;
      &lt;td&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;quick&lt;/code&gt;&lt;/td&gt;
      &lt;td&gt;3&lt;/td&gt;
      &lt;td&gt;4&lt;/td&gt;
      &lt;td&gt;1&lt;/td&gt;
      &lt;td&gt;1&lt;/td&gt;
      &lt;td&gt;One specific fact, a URL check, a yes/no. Minutes.&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;standard&lt;/code&gt; &lt;em&gt;(default)&lt;/em&gt;&lt;/td&gt;
      &lt;td&gt;8&lt;/td&gt;
      &lt;td&gt;12&lt;/td&gt;
      &lt;td&gt;1&lt;/td&gt;
      &lt;td&gt;2&lt;/td&gt;
      &lt;td&gt;Normal research task. Answer the questions and stop.&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;deep&lt;/code&gt;&lt;/td&gt;
      &lt;td&gt;20&lt;/td&gt;
      &lt;td&gt;30&lt;/td&gt;
      &lt;td&gt;2&lt;/td&gt;
      &lt;td&gt;3&lt;/td&gt;
      &lt;td&gt;Genuinely hard question, contested facts, or a topic where the first page of results is known to be junk.&lt;/td&gt;
    &lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;

&lt;p&gt;Modules are the per-source search playbooks the agent loads at runtime: source lists and query tactics per domain. There’s a whole section on them below.&lt;/p&gt;

&lt;p&gt;Four things about that table matter more than the specific numbers, which you should tune to your own work:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A depth dial beats a fixed cap.&lt;/strong&gt; One setting is always wrong: too tight for the hard question, too loose for the quick fact. Three named levels with a stated default means I ask for &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;deep&lt;/code&gt; on purpose instead of discovering I got it by accident. Precedence is explicit too: numbers the caller gives win, then a level named in the prompt, then the project’s config, then &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;standard&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The limits need anti-gaming clauses, because a model reads a budget as a target.&lt;/strong&gt; The two that do the work:&lt;/p&gt;

&lt;blockquote&gt;
  &lt;p&gt;&lt;strong&gt;1 fetch per URL.&lt;/strong&gt; Never re-fetch a URL you already read.&lt;/p&gt;

  &lt;p&gt;&lt;strong&gt;Stop as soon as the caller’s questions are answered.&lt;/strong&gt; Remaining budget is not a quota to spend. A &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;deep&lt;/code&gt; run that finishes in six searches is a success, not a waste.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That second line went in after I watched a run answer the question at search four and then keep going, because twenty was the number it had been given. You have to say out loud that finishing early is winning.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Count what actually costs, not what’s convenient to count.&lt;/strong&gt; My fetch budget covers every network retrieval: native fetch, the escalation rungs for blocked pages, a helper script’s individual HTTP requests, and each &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;429&lt;/code&gt; retry inside that helper counted separately. The first version counted “fetch calls,” and the agent found the loophole without trying: it wasn’t cheating, it was obeying a rule I’d written badly.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Make it announce its spending.&lt;/strong&gt; The agent reports &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;[deep: 3/20 searches, 5/30 fetches]&lt;/code&gt; after each phase. Without that you have no idea whether you’re watching a careful run or a runaway until it’s over. And when it does exhaust the budget with questions still open, it stops and reports exactly what’s unknown, which URL would most likely answer it, and that re-running deeper is an option. What it must never do is keep going past the limit without saying so: the caller chose the level, and overspending it takes that choice away.&lt;/p&gt;

&lt;h2 id=&quot;the-bans-are-most-of-a-mature-agent&quot;&gt;The Bans Are Most of a Mature Agent&lt;/h2&gt;

&lt;p&gt;&lt;img src=&quot;/images/2026/the-subagents-guide-i-wish-id-had-body-9.jpg&quot; alt=&quot;The Bans Are Most of a Mature Agent&quot; srcset=&quot;            /assets/resized/480/the-subagents-guide-i-wish-id-had-body-9.jpg 480w,            /assets/resized/800/the-subagents-guide-i-wish-id-had-body-9.jpg 800w,            /assets/resized/1400/the-subagents-guide-i-wish-id-had-body-9.jpg 1400w,    &quot; loading=&quot;lazy&quot; /&gt;&lt;/p&gt;

&lt;p&gt;Look at the shape of my most-used agent file and the thing that jumps out is proportion. The role description is a paragraph. The negative constraints are a list. That inversion isn’t bloat. A capable agent with tools has far more ways to technically-comply than you can anticipate.&lt;/p&gt;

&lt;p&gt;Here’s one, quoted exactly, because the parenthetical is the whole point:&lt;/p&gt;

&lt;blockquote&gt;
  &lt;p&gt;&lt;strong&gt;Never use browser automation.&lt;/strong&gt; No &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;Simple Browser&lt;/code&gt;, no embedded/preview browser, no Playwright, no &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;mcp__claude-in-chrome__*&lt;/code&gt;, no &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;open&lt;/code&gt;, no opening tabs or windows. If a browser tool is offered to you, it is not for this task. Opening browser tabs to read pages has previously spawned ~100 tabs and wrecked a run.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;And another, from the implementer in my research-agent repo, where the mechanism is the entire reason anyone would respect the rule:&lt;/p&gt;

&lt;blockquote&gt;
  &lt;p&gt;&lt;strong&gt;Never create any file under &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;agents/&lt;/code&gt; or &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;skills/&lt;/code&gt;.&lt;/strong&gt; &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;agents/&lt;/code&gt; is shipped payload — the package manager flattens every &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;.md&lt;/code&gt; beneath it into a separate top-level agent on install, so a stray file there lands in every consumer project.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That’s the rule I’d extract from all of it: &lt;strong&gt;every ban names its mechanism.&lt;/strong&gt; The agent doesn’t need persuading. The reason is there because a constraint whose reason isn’t written down gets edited away. Six weeks later you’ll be tightening the prose, you’ll hit a line that says “never create files here,” it’ll read as arbitrary, and you’ll helpfully generalize it. The “because the installer flattens it into every consumer project” clause is what makes future-you leave it alone.&lt;/p&gt;

&lt;p&gt;A few more shapes these constraints take, once you’ve been at it a while:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Ban the category, list the instances.&lt;/strong&gt; “No browser automation” alone leaves the agent deciding whether a preview pane counts. Naming six specific things it must not reach for closes the gap where it reasons its way into one.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Name the behavior, not just the tool.&lt;/strong&gt; My favorite line in the file isn’t a prohibition on a command, it’s a prohibition on a tendency: &lt;em&gt;“If you find yourself authoring a scraper or a report generator, stop — you are working around the task, not doing it.”&lt;/em&gt; Capable agents route around constraints by building tools. That’s the category, and you have to ban the category.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;An escape hatch needs conditions, or it becomes the main road.&lt;/strong&gt; Some pages really are unfetchable, so there’s &lt;a href=&quot;https://www.stephanmiller.com/what-to-do-when-your-ai-coding-agent-cant-read-a-web-page/&quot;&gt;an escalation ladder for blocked URLs&lt;/a&gt;. But it only opens after the normal path has already failed on that exact URL, it never grants a second fetch slot, it checks whether a tool is installed rather than installing anything, and it caps the output so a page dump can’t blow the context. Then the hard stop: &lt;em&gt;“If the URL is still unreachable after the rungs available to you, it is done.”&lt;/em&gt; Record it, move on. No further attempt, no other helper, no creative fourth idea.&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;/images/2026/the-subagents-guide-i-wish-id-had-body-10.jpg&quot; alt=&quot;The Bans Are Most of a Mature Agent&quot; srcset=&quot;            /assets/resized/480/the-subagents-guide-i-wish-id-had-body-10.jpg 480w,            /assets/resized/800/the-subagents-guide-i-wish-id-had-body-10.jpg 800w,            /assets/resized/1400/the-subagents-guide-i-wish-id-had-body-10.jpg 1400w,    &quot; loading=&quot;lazy&quot; /&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Say why the exception isn’t a loophole.&lt;/strong&gt; There’s a paragraph in there explaining that a one-shot headless fetch that prints text is not a violation of the browser ban, because the ban is about driving a browser and leaving tabs for a human to close. Without that paragraph, the exception and the ban look like a contradiction, and a model resolving a contradiction will pick the reading that lets it do more.&lt;/p&gt;

&lt;h2 id=&quot;they-fail-silently-and-confidently&quot;&gt;They Fail Silently and Confidently&lt;/h2&gt;

&lt;p&gt;I had two research runs come back fine. Not fine. &lt;em&gt;Good&lt;/em&gt;. Clean findings, real sources, coherent report. Both runs named Reddit as their primary source. Neither run had gotten a single thing from Reddit.&lt;/p&gt;

&lt;p&gt;What happened is that a WebSearch &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;site:reddit.com&lt;/code&gt; query returned ten clean, plausible results from &lt;em&gt;other domains&lt;/em&gt; (an Etsy community forum, the SBA, slideshare) and &lt;strong&gt;no error at all.&lt;/strong&gt; Not a 403. Not an empty set. Not a warning. The &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;site:&lt;/code&gt; constraint got ignored, and nothing in the output said the domain filter hadn’t applied. A tidy list of usable pages from the wrong places. The agent did what any reasonable worker does with a tidy list of usable pages: it used them, and it reported success.&lt;/p&gt;

&lt;p&gt;The decision I wrote that day, which is now the foundational one in that project:&lt;/p&gt;

&lt;blockquote&gt;
  &lt;p&gt;&lt;strong&gt;When a named source cannot be reached, the run says so. It never proceeds silently on whatever the tool returned instead.&lt;/strong&gt;&lt;/p&gt;

  &lt;p&gt;A run that names Reddit as its primary source while reporting success on Etsy forums is worse than a run that fails, because the failure is invisible to whoever reads the report.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;A failed run costs you a re-run. A run that swapped sources without telling you costs you a decision made on evidence you think came from somewhere it didn’t. And the separate context window, the feature itself, is exactly why you can’t see it happen. You don’t get the transcript. You get the summary the agent chose to write.&lt;/p&gt;

&lt;p&gt;Three things fix this, and none of them is “tell the agent to be careful.”&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Give it a mechanical detection rule.&lt;/strong&gt; Nothing errored, so the agent had no reason to look. The fix is one sentence with no judgment in it: &lt;em&gt;if you constrained a search to a domain and no returned URL is on that domain, that is a zero-result finding, not a result set.&lt;/em&gt; It needs no extra tool, no extra fetch, no permission. The URLs are already in its hands. “Be skeptical of your sources” is not actionable. “Compare the hostname to the one you constrained on” is.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Make provenance a required output channel with a schema.&lt;/strong&gt; This is where “output shape” graduates from section headers into something closer to a type. Two arrays, specified down to the key names:&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;unreachable[]&lt;/code&gt;: one entry per wall, keys exactly &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;source&lt;/code&gt;, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;url&lt;/code&gt;, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;reason&lt;/code&gt;. Covers both hard fetch failure and silent substitution.&lt;/li&gt;
  &lt;li&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;sources[]&lt;/code&gt;: one entry per source that &lt;em&gt;actually supported an answer&lt;/em&gt;, keys exactly &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;source&lt;/code&gt;, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;url&lt;/code&gt;, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;fields&lt;/code&gt;. A page you opened that didn’t inform anything doesn’t get recorded. A source that answered four fields is one entry with four names in it, not four entries.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The subtle call in there, which I got wrong first: &lt;strong&gt;provenance annotates, it never blocks.&lt;/strong&gt; If a documented substitute answered the question, the field is answered. The &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;unreachable&lt;/code&gt; entry records that the answer didn’t come from the named source. Making a wall fail the field would mean the pipeline breaks loudly on exactly the situation it has a workaround for.&lt;/p&gt;

&lt;p&gt;And a failure worth stealing the lesson from: the first version of that agent said “record it as unreachable” &lt;strong&gt;four separate times and never once named a field, an array, or an output slot.&lt;/strong&gt; The instruction dead-ended. The agent was told to record something with nowhere to put it, so it didn’t, and nothing anywhere complained. When you write an instruction to a subagent, check that the thing it’s told to do has a destination. “Report X” with no named slot for X is a no-op with a clean conscience.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Verify from outside the agent.&lt;/strong&gt; The agent’s compliance with its own output contract is checked by a script I run afterward, not by the agent’s assurance that it complied. That’s 180 lines of Python validating the JSON against the declared fields, and it is worth more than any amount of emphasis inside the prompt.&lt;/p&gt;

&lt;p&gt;The generalized version of that last point shows up again in the implementer, in a repo that has no test suite to lean on:&lt;/p&gt;

&lt;blockquote&gt;
  &lt;p&gt;&lt;strong&gt;Verify instead of running a suite.&lt;/strong&gt; There is no suite. The task’s Tests section lists the checks that stand in for one — run every one of them and report the results concretely. Where a check is “this URL resolves,” that means &lt;strong&gt;fetch it&lt;/strong&gt;, not assume it.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;blockquote&gt;
  &lt;p&gt;&lt;strong&gt;A URL you could not verify does not go into a file.&lt;/strong&gt; Report it as unverified with what happened. An unverified URL is worse than none — a reader trusts it and spends a fetch on a 404 instead of falling back to search.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Agents will &lt;a href=&quot;https://www.stephanmiller.com/my-home-ai-agent-kept-making-shit-up/&quot;&gt;report verification they did not perform&lt;/a&gt;, because a plausible claim and a checked claim look identical in a summary. Every check you care about needs to name the physical action that constitutes doing it.&lt;/p&gt;

&lt;h2 id=&quot;keeping-the-contract-tight-and-the-antipatterns-that-bloat-it&quot;&gt;Keeping the Contract Tight (and the Antipatterns That Bloat It)&lt;/h2&gt;

&lt;p&gt;The best subagents have one job. And here you’ll object, because the files I’ve been quoting aren’t thin. My research agent is 188 lines and nobody would call it thin. The distinction that resolves it: &lt;strong&gt;the job stayed one sentence; the file grew.&lt;/strong&gt; What accumulated was constraints, detection rules, and output contract: scar tissue on a fixed skeleton. Not one new responsibility in a year. Growth from things going wrong is the file working as intended. Growth from new kinds of work is bloat, and the fix is a split, not another section.&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;/images/2026/the-subagents-guide-i-wish-id-had-body-12.jpg&quot; alt=&quot;Keeping the Contract Tight (and the Antipatterns That Bloat It)&quot; srcset=&quot;            /assets/resized/480/the-subagents-guide-i-wish-id-had-body-12.jpg 480w,            /assets/resized/800/the-subagents-guide-i-wish-id-had-body-12.jpg 800w,            /assets/resized/1400/the-subagents-guide-i-wish-id-had-body-12.jpg 1400w,    &quot; loading=&quot;lazy&quot; /&gt;&lt;/p&gt;

&lt;p&gt;The bloated versions all look the same, and I’ve shipped every one: the everything-reviewer that averages across five concerns and is expert at none. The “senior engineer” label with no behavior behind it. The agent that restates your &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;CLAUDE.md&lt;/code&gt; and adds token cost instead of behavior. The one that duplicates a skill until the two drift apart. Each is the same mistake: a job that got wider than one sentence.&lt;/p&gt;

&lt;p&gt;Which raises the obvious question: after a year of this, what’s actually left on my bench?&lt;/p&gt;

&lt;h2 id=&quot;the-two-that-actually-survived&quot;&gt;The Two That Actually Survived&lt;/h2&gt;

&lt;p&gt;Most posts on this hand you an org chart of agent types to build. Here’s my real bench after a year:&lt;/p&gt;

&lt;p&gt;Two agents. That’s it.&lt;/p&gt;

&lt;p&gt;One disclosure, because this post’s own rule applies to it: &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;web-search-agent&lt;/code&gt; started as a fork of &lt;a href=&quot;https://github.com/Weizhena/Deep-Research-skills&quot;&gt;Lan Zheng’s Deep-Research-skills&lt;/a&gt;. The output-contract skeleton (the summary sections, the always-required sources list) is upstream’s, substantially verbatim. The control plane on top (the budgets, the bans, the provenance schema, the escalation ladder) is mine, and it’s half the line count. Everything below about scars is about that half.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;web-search-agent&lt;/code&gt;&lt;/strong&gt; does bounded web research and hands back findings with sources. It’s installed across a dozen repos right now: my blog’s static site, a couple of pipelines, and a few research projects. Everything in this post about budgets, bans, and silent substitution came out of that one file’s revision history.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;implementer&lt;/code&gt;&lt;/strong&gt; executes a fully-specified task queue while I keep planning. Four projects have one. Everything about handoff contracts, run state, and delegation economics came from those.&lt;/p&gt;

&lt;p&gt;The rest of what’s sitting in my agents folders came bundled with frameworks I installed, and it’s a junk drawer with better branding than my actual junk drawer. The specialists I &lt;em&gt;predicted&lt;/em&gt; I’d need (the flaky-test investigator, the PR-polish agent, the docs drafter, the codebase-orientation agent for projects I abandoned three months ago) were all reasonable ideas and I built approximately none of them. Turns out the ones that stick aren’t the ones that sound useful. They’re the ones where I kept feeling the specific pain of the work being in the wrong context or the wrong head.&lt;/p&gt;

&lt;p&gt;So don’t build the org chart. Notice which work you wish were happening somewhere else, and save that one.&lt;/p&gt;

&lt;h3 id=&quot;template-then-specialize&quot;&gt;Template, then specialize&lt;/h3&gt;

&lt;p&gt;Here’s the pattern that made the second, third, and fourth implementer cheap: the agent gets &lt;strong&gt;generated from a template and then diverges where the project differs.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;living-plan&lt;/code&gt; skill ships an &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;implementer.template.md&lt;/code&gt; with &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;and&lt;/code&gt; placeholders, and scaffolding a new project writes out the agent, the plan doc, and the task doc together. What you get is the &lt;em&gt;protocol&lt;/em&gt;: read the contract, serial queue, skip blockers, log run state, stay in scope, report back. That part is identical everywhere because it’s about how delegation works, not about your code.&lt;/p&gt;

&lt;p&gt;Then each copy grows a local section for the local failure surface, and comparing two of them is instructive. The template says “keep the suite green.” One of my projects has no suite at all (it’s a repo of prompts, where the only executable file is the validator), so its implementer says this instead:&lt;/p&gt;

&lt;blockquote&gt;
  &lt;p&gt;This repo is &lt;strong&gt;prompts and data, not code.&lt;/strong&gt; Editing a file is editing a prompt. Wording, ordering, and emphasis are the implementation. A rewrite that reads better but drops a hard constraint is a regression, and nothing will catch it — there is no build and no test suite.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Another project’s copy grew a rule about reading the specific &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;PLAN.md&lt;/code&gt; decisions a task cites before writing code, and a report-back line for “anything you hit that requires an Opus/human call.” Same protocol, different hazards.&lt;/p&gt;

&lt;p&gt;So here’s the division: &lt;strong&gt;the reusable part of a subagent is the protocol; the per-project part is the failure surface.&lt;/strong&gt; Template the first, hand-write the second, and don’t try to make one file serve both.&lt;/p&gt;

&lt;h2 id=&quot;the-killer-combo-point-the-agent-at-knowledge-dont-paste-it-in&quot;&gt;The Killer Combo: Point the Agent at Knowledge, Don’t Paste It In&lt;/h2&gt;

&lt;p&gt;This is where the two guides shake hands, and my understanding of it got a lot more specific once a real pipeline depended on it.&lt;/p&gt;

&lt;p&gt;What I’d have told you before is “mention the relevant skill in the agent body.” What actually works is stronger: &lt;strong&gt;the agent reads the knowledge at runtime, before it’s allowed to act, and the prompt says why.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;My research agent can’t run a single search until it has read a routing table that lives in a skill:&lt;/p&gt;

&lt;blockquote&gt;
  &lt;p&gt;&lt;strong&gt;Module Selection (MANDATORY — routing lives in one file).&lt;/strong&gt; Before executing any search or fetch, you MUST &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;Read&lt;/code&gt; the routing table at the first existing path below. DO NOT skip this step. DO NOT route from memory — the module list changes without this prompt changing, so a module you remember may be gone and one you need may be new.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;em&gt;The module list changes without this prompt changing.&lt;/em&gt; That sentence is the entire argument for the pattern. The knowledge (which sources answer which kinds of question, how to query each one, what’s known to be blocked) changes monthly. The methodology doesn’t. If I’d baked the source list into the agent body, every new source would mean editing a prompt I’d otherwise leave alone, and the agent would confidently route to a module I deleted in March.&lt;/p&gt;

&lt;p&gt;Two mechanics that turned out to matter more than I expected:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A path resolution ladder, not a path.&lt;/strong&gt; The same agent runs in projects with different layouts, on two machines with different usernames, installed by different tools. So the instruction lists candidate paths in order and says take the first that exists: project-local first, then the user-level copies. Hardcode one path and the agent works on your machine and nowhere else.&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;/images/2026/the-subagents-guide-i-wish-id-had-body-14.jpg&quot; alt=&quot;The Killer Combo: Point the Agent at Knowledge, Don&apos;t Paste It In&quot; srcset=&quot;            /assets/resized/480/the-subagents-guide-i-wish-id-had-body-14.jpg 480w,            /assets/resized/800/the-subagents-guide-i-wish-id-had-body-14.jpg 800w,            /assets/resized/1400/the-subagents-guide-i-wish-id-had-body-14.jpg 1400w,    &quot; loading=&quot;lazy&quot; /&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The layers are coupled by literal names, with nothing checking them.&lt;/strong&gt; The architecture is three tiers: skills orchestrate, an agent does the work, data modules hold the knowledge. And there are no imports anywhere. A skill launches the agent &lt;em&gt;by its exact registered name&lt;/em&gt;, and the agent reads modules &lt;em&gt;by path&lt;/em&gt;. Which leads to the most mundane and most dangerous line in that repo:&lt;/p&gt;

&lt;blockquote&gt;
  &lt;p&gt;Do not rename &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;web-search-agent&lt;/code&gt;; existing consumers call it directly.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;strong&gt;The agent’s name is a public API.&lt;/strong&gt; Nothing will tell you otherwise: there’s no compiler, no test, no warning. A rename is a clean-looking commit that breaks every caller in every repo that installed it, and you find out the next time you ask for research and get a generalist instead.&lt;/p&gt;

&lt;h2 id=&quot;where-they-live-personal-project-packaged&quot;&gt;Where They Live: Personal, Project, Packaged&lt;/h2&gt;

&lt;p&gt;There are three rungs, and the third one is where all my real agents live.&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;&lt;strong&gt;Personal&lt;/strong&gt; (&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;~/.claude/agents/&lt;/code&gt;): general-purpose specialists you want in every project. This is the workshop. Break things here.&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Project&lt;/strong&gt; (&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;.claude/agents/&lt;/code&gt;, committed): specialists that only make sense for &lt;em&gt;this&lt;/em&gt; codebase. The implementer that knows this project has no test suite. Commit it and it’s there next time, even if next time is four months from now and you’ve forgotten the project exists.&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Packaged&lt;/strong&gt; (installed by a package manager, pinned to a commit): an agent that’s a dependency. This is the rung I didn’t know I needed until the same agent was in a dozen repos and I had thirteen copies drifting apart.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That third rung changes a few things, and they’re worth knowing before you get there rather than after:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Ship the agent with the skills that call it.&lt;/strong&gt; My package contains both the skills and the agents they launch, in one bundle, for an unglamorous reason: &lt;em&gt;installing the skills without the agents yields a pipeline that fails at first use.&lt;/em&gt; A skill that dispatches by name to an agent you don’t have is a broken install with a clean error message at best. They’re one unit because they’re one contract.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Know where your installer puts things.&lt;/strong&gt; This one installs through &lt;a href=&quot;https://github.com/microsoft/apm&quot;&gt;APM&lt;/a&gt;, and APM flattens every &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;.md&lt;/code&gt; under an &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;agents/&lt;/code&gt; directory into a separate top-level agent on install. That’s a sane default that bites hard: a data file parked in that tree becomes a bogus registered agent in every consumer project, and a project-specific agent written there installs itself into everybody’s repos. It’s why my search-strategy modules live under &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;skills/&lt;/code&gt; rather than the &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;agents/&lt;/code&gt; directory they’d otherwise obviously belong in: a folder of them under &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;agents/&lt;/code&gt; loses its structure on install and registers five bogus agents in every project that pulls the package.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Pin a commit and know that consumers are behind.&lt;/strong&gt; Consumers depend on a SHA, which means a fix isn’t live for anybody until the pin moves and the install re-runs. Writing a fix is not shipping a fix. That sounds obvious written down and it is absolutely not obvious at 11pm when the bug you fixed last week reappears.&lt;/p&gt;

&lt;p&gt;And still: curate. A folder full of overlapping, half-abandoned agents is a noise generator, not a toolkit. If two have drifted into the same job, merge them. If one hasn’t been invoked in months, delete it. The file’s in git history if you’re wrong. Which, yes, is also good life advice, and no, I don’t follow it with my actual junk drawer.&lt;/p&gt;

&lt;h2 id=&quot;how-this-works-in-the-other-tools&quot;&gt;How This Works in the Other Tools&lt;/h2&gt;

&lt;p&gt;Claude Code isn’t the only place this exists, and if you bounce between tools like I do, the ideas port even when the file formats don’t.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;GitHub Copilot / VS Code&lt;/strong&gt; call these &lt;em&gt;custom agents&lt;/em&gt; and use &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;.agent.md&lt;/code&gt; files, typically in &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;.github/agents/&lt;/code&gt;, with a personal library option under your user directory. The frontmatter is richer than Claude’s: beyond &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;name&lt;/code&gt;, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;description&lt;/code&gt;, and &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;tools&lt;/code&gt; you get things like &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;argument-hint&lt;/code&gt;, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;model&lt;/code&gt;, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;user-invocable&lt;/code&gt;, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;disable-model-invocation&lt;/code&gt;, and my favorite, &lt;strong&gt;handoffs&lt;/strong&gt;: a finished agent can offer a button to hand off to another agent with a prefilled next prompt, so you can chain planning to implementation to review without collapsing everything into one mush-brained agent. One migration gotcha if you’re reading older examples: the old &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;infer&lt;/code&gt; field got split into &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;user-invocable&lt;/code&gt; and &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;disable-model-invocation&lt;/code&gt;, which is clearer once you know and maddening until you do.&lt;/p&gt;

&lt;p&gt;You’ll also read that VS Code detects &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;.claude/agents/&lt;/code&gt; files and maps the tool names across, so you can share an agent between tools for free. It’s true enough to be misleading. The tool vocabularies don’t line up, and a prompt that hard-bans tool names by name (which any mature agent of mine does) does not survive automatic translation intact.&lt;/p&gt;

&lt;p&gt;What I actually ship is &lt;strong&gt;one canonical prompt plus a thin native shim per tool.&lt;/strong&gt; The entire Copilot-side agent in my package is eleven lines:&lt;/p&gt;

&lt;div class=&quot;language-markdown highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;nn&quot;&gt;---&lt;/span&gt;
&lt;span class=&quot;na&quot;&gt;name&lt;/span&gt;&lt;span class=&quot;pi&quot;&gt;:&lt;/span&gt; &lt;span class=&quot;s&quot;&gt;Web Research Writer&lt;/span&gt;
&lt;span class=&quot;na&quot;&gt;description&lt;/span&gt;&lt;span class=&quot;pi&quot;&gt;:&lt;/span&gt; &lt;span class=&quot;s2&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;Use&lt;/span&gt;&lt;span class=&quot;nv&quot;&gt; &lt;/span&gt;&lt;span class=&quot;s&quot;&gt;for&lt;/span&gt;&lt;span class=&quot;nv&quot;&gt; &lt;/span&gt;&lt;span class=&quot;s&quot;&gt;bounded&lt;/span&gt;&lt;span class=&quot;nv&quot;&gt; &lt;/span&gt;&lt;span class=&quot;s&quot;&gt;web&lt;/span&gt;&lt;span class=&quot;nv&quot;&gt; &lt;/span&gt;&lt;span class=&quot;s&quot;&gt;research&lt;/span&gt;&lt;span class=&quot;nv&quot;&gt; &lt;/span&gt;&lt;span class=&quot;s&quot;&gt;that&lt;/span&gt;&lt;span class=&quot;nv&quot;&gt; &lt;/span&gt;&lt;span class=&quot;s&quot;&gt;must&lt;/span&gt;&lt;span class=&quot;nv&quot;&gt; &lt;/span&gt;&lt;span class=&quot;s&quot;&gt;read&lt;/span&gt;&lt;span class=&quot;nv&quot;&gt; &lt;/span&gt;&lt;span class=&quot;s&quot;&gt;a&lt;/span&gt;&lt;span class=&quot;nv&quot;&gt; &lt;/span&gt;&lt;span class=&quot;s&quot;&gt;local&lt;/span&gt;&lt;span class=&quot;nv&quot;&gt; &lt;/span&gt;&lt;span class=&quot;s&quot;&gt;schema,&lt;/span&gt;&lt;span class=&quot;nv&quot;&gt; &lt;/span&gt;&lt;span class=&quot;s&quot;&gt;search&lt;/span&gt;&lt;span class=&quot;nv&quot;&gt; &lt;/span&gt;&lt;span class=&quot;s&quot;&gt;and&lt;/span&gt;&lt;span class=&quot;nv&quot;&gt; &lt;/span&gt;&lt;span class=&quot;s&quot;&gt;fetch&lt;/span&gt;&lt;span class=&quot;nv&quot;&gt; &lt;/span&gt;&lt;span class=&quot;s&quot;&gt;current&lt;/span&gt;&lt;span class=&quot;nv&quot;&gt; &lt;/span&gt;&lt;span class=&quot;s&quot;&gt;sources,&lt;/span&gt;&lt;span class=&quot;nv&quot;&gt; &lt;/span&gt;&lt;span class=&quot;s&quot;&gt;write&lt;/span&gt;&lt;span class=&quot;nv&quot;&gt; &lt;/span&gt;&lt;span class=&quot;s&quot;&gt;one&lt;/span&gt;&lt;span class=&quot;nv&quot;&gt; &lt;/span&gt;&lt;span class=&quot;s&quot;&gt;designated&lt;/span&gt;&lt;span class=&quot;nv&quot;&gt; &lt;/span&gt;&lt;span class=&quot;s&quot;&gt;result&lt;/span&gt;&lt;span class=&quot;nv&quot;&gt; &lt;/span&gt;&lt;span class=&quot;s&quot;&gt;file,&lt;/span&gt;&lt;span class=&quot;nv&quot;&gt; &lt;/span&gt;&lt;span class=&quot;s&quot;&gt;and&lt;/span&gt;&lt;span class=&quot;nv&quot;&gt; &lt;/span&gt;&lt;span class=&quot;s&quot;&gt;validate&lt;/span&gt;&lt;span class=&quot;nv&quot;&gt; &lt;/span&gt;&lt;span class=&quot;s&quot;&gt;it&lt;/span&gt;&lt;span class=&quot;nv&quot;&gt; &lt;/span&gt;&lt;span class=&quot;s&quot;&gt;with&lt;/span&gt;&lt;span class=&quot;nv&quot;&gt; &lt;/span&gt;&lt;span class=&quot;s&quot;&gt;a&lt;/span&gt;&lt;span class=&quot;nv&quot;&gt; &lt;/span&gt;&lt;span class=&quot;s&quot;&gt;local&lt;/span&gt;&lt;span class=&quot;nv&quot;&gt; &lt;/span&gt;&lt;span class=&quot;s&quot;&gt;command.&quot;&lt;/span&gt;
&lt;span class=&quot;na&quot;&gt;tools&lt;/span&gt;&lt;span class=&quot;pi&quot;&gt;:&lt;/span&gt; &lt;span class=&quot;pi&quot;&gt;[&lt;/span&gt;&lt;span class=&quot;nv&quot;&gt;read&lt;/span&gt;&lt;span class=&quot;pi&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;nv&quot;&gt;search&lt;/span&gt;&lt;span class=&quot;pi&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;nv&quot;&gt;web&lt;/span&gt;&lt;span class=&quot;pi&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;nv&quot;&gt;edit&lt;/span&gt;&lt;span class=&quot;pi&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;nv&quot;&gt;execute&lt;/span&gt;&lt;span class=&quot;pi&quot;&gt;]&lt;/span&gt;
&lt;span class=&quot;na&quot;&gt;user-invocable&lt;/span&gt;&lt;span class=&quot;pi&quot;&gt;:&lt;/span&gt; &lt;span class=&quot;no&quot;&gt;false&lt;/span&gt;
&lt;span class=&quot;na&quot;&gt;disable-model-invocation&lt;/span&gt;&lt;span class=&quot;pi&quot;&gt;:&lt;/span&gt; &lt;span class=&quot;no&quot;&gt;false&lt;/span&gt;
&lt;span class=&quot;nn&quot;&gt;---&lt;/span&gt;

Load the installed canonical research prompt from the first existing candidate:
&lt;span class=&quot;p&quot;&gt;
1.&lt;/span&gt; &lt;span class=&quot;sb&quot;&gt;`.github/agents/web-search-agent.agent.md`&lt;/span&gt;
&lt;span class=&quot;p&quot;&gt;2.&lt;/span&gt; &lt;span class=&quot;sb&quot;&gt;`.claude/agents/web-search-agent.md`&lt;/span&gt;

Follow its search budgets, module routing, source standards, tool discipline, output, and validation rules. Ignore that file&apos;s incompatible &lt;span class=&quot;sb&quot;&gt;`tools`&lt;/span&gt; frontmatter; this wrapper&apos;s Copilot-native tool categories govern this session.
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;That’s the whole pattern. Native frontmatter so the host registers it properly, a path ladder to find the real prompt, and one explicit sentence about which tool vocabulary wins. All 188 lines of hard-won behavior live in exactly one file, and the second tool gets a pointer instead of a fork. Two copies of a prompt is two prompts, and the day you fix a bug in one of them is the day they start lying to you about being the same agent.&lt;/p&gt;

&lt;p&gt;One caveat from the README: point the loader at the &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;.github/agents/&lt;/code&gt; tree and the canonical prompt registers as its own agent too, so Copilot can end up seeing duplicate names: the “Web Research Writer” shim plus the canonical file underneath it. Pick one home per tool, or keep the canonical file in the &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;.claude/agents/&lt;/code&gt; tree where Copilot’s auto-detection won’t pick it up.&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;/images/2026/the-subagents-guide-i-wish-id-had-body-16.jpg&quot; alt=&quot;How This Works in the Other Tools&quot; srcset=&quot;            /assets/resized/480/the-subagents-guide-i-wish-id-had-body-16.jpg 480w,            /assets/resized/800/the-subagents-guide-i-wish-id-had-body-16.jpg 800w,            /assets/resized/1400/the-subagents-guide-i-wish-id-had-body-16.jpg 1400w,    &quot; loading=&quot;lazy&quot; /&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Everyone else&lt;/strong&gt; (Cursor, Windsurf, and the rest) is somewhere on the road to the same thing under names like rules, modes, or agents, and the details shift often enough that I’d rather point you at each tool’s current docs than confidently tell you something that changed last month. The underlying move is identical everywhere: a reusable file that defines a role, invoked on purpose. Learn the concept once and you’re mostly just learning where each tool hides the folder.&lt;/p&gt;

&lt;h2 id=&quot;how-good-ones-actually-get-built&quot;&gt;How Good Ones Actually Get Built&lt;/h2&gt;

&lt;p&gt;Every subagent I use started too broad. I wrote a “reviewer” that reviewed everything. A “planner” with no opinion about what a plan should look like. An “investigator” that approached every bug like the last one. That’s not failure. That’s the starting point.&lt;/p&gt;

&lt;p&gt;What I’d have found useful a year ago isn’t “iterate.” Everybody says iterate. It’s &lt;em&gt;what the iterations are made of.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Go back through the agent files I’ve quoted and look at where the words actually are. A paragraph of role. A page of constraints. A schema with key names spelled out. A mechanical detection rule. A numeric budget with an anti-gaming clause. A note that a paid fetch has to be disclosed. A ban that explains what the installer does to stray files. &lt;strong&gt;Almost all of it is a record of a specific thing that went wrong once.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Which means the useful question after a disappointing run isn’t “how do I describe this role better.” It’s: &lt;em&gt;what exactly did it do, what in the file permitted that, and what sentence makes it impossible next time.&lt;/em&gt; A hundred tabs. Ten results from the wrong domain and no error. An instruction with nowhere to write its answer. A verification it reported but didn’t perform. Each one of those is a line, and the line outlives the session that taught it to you, which is the entire reason this is a file.&lt;/p&gt;

&lt;p&gt;The test before you keep one is the same as before you keep a skill: can you say its job in one sentence? If not, it’s too broad, and you’ve got two smaller agents wearing a trenchcoat. But the test for whether it’s any &lt;em&gt;good&lt;/em&gt; is different, and it’s this: when it screws up, can you point at the line that let it? If you can’t, you don’t have a specialist yet. You have a vibe with frontmatter.&lt;/p&gt;

&lt;p&gt;Put them together and you’re not prompting a very smart generalist over and over. You’re keeping a small bench of specialists who already know your world and already know their job. For one person trying to ship more than one person reasonably should, that’s the closest thing to hiring help I’ve found that doesn’t involve hiring anyone.&lt;/p&gt;

&lt;p&gt;Both of the agents I’ve been quoting are open source, scars included: &lt;a href=&quot;https://github.com/eristoddle/deep-research-agent/blob/main/agents/web-search-agent.md&quot;&gt;the research agent itself&lt;/a&gt;, and &lt;a href=&quot;https://github.com/eristoddle/agent-skills&quot;&gt;the implementer template&lt;/a&gt; plus &lt;a href=&quot;https://github.com/eristoddle/agent-skills&quot;&gt;living-plan&lt;/a&gt;. Steal whatever’s useful.&lt;/p&gt;

&lt;p&gt;Now go look at the work you wish were happening somewhere else. That’s a subagent you haven’t saved yet. And the last time a subagent disappointed you, that’s a line you haven’t written yet.&lt;/p&gt;
</description>
        <pubDate>Thu, 17 Sep 2026 07:00:00 -0500</pubDate>
        <link>https://www.stephanmiller.com/the-subagents-guide-i-wish-id-had/</link>
        <guid isPermaLink="true">https://www.stephanmiller.com/the-subagents-guide-i-wish-id-had/</guid>
        
        <category>Claude Code subagents</category>
        
        <category>Custom agents</category>
        
        <category>AI agent development</category>
        
        <category>Reusable AI specialists</category>
        
        <category>Coding agent workflow</category>
        
        
        <category>ai-agents</category>
        
        <category>agentic-development</category>
        
      </item>
    
  </channel>
</rss>
