Stephan Miller
Before You Write a Scraper, Check Firecrawl Alexandria

Before You Write a Scraper, Check Firecrawl Alexandria

Last Thursday I published a 70-line scraper that watches the Arena leaderboard for new AI models. It runs daily and announces new arrivals.

Two days earlier, Firecrawl launched Alexandria, a catalog of providers that sell the structured version of exactly this kind of data, callable through the same scrape endpoint I was already paying for. I didn’t notice. I was busy having AI build me regex for a markdown table, while the company whose credits I was spending shipped the thing meant to make that regex unnecessary.

So to recap: I built a machine that watches the internet for new arrivals, and the first new arrival it missed was the one built to replace it.

If somebody is already selling the structured version of the data I wanted, why was I scraping the page? It never occurred to me to check.

What Alexandria Is

Alexandria is Firecrawl’s catalog of third-party data sources: FRED economic data, podcast transcripts, app store listings, company records, and a few hundred other things. They named their permanent store of the world’s data after history’s most famous case of data loss, which is either confidence or a gap in the marketing team’s classics minor. Each tool in it has a published contract that says what options it takes, what it returns, and what it costs. Finding tools and reading contracts is free. Running one costs credits, Firecrawl’s metered unit, and the free plan is 1,000 of them a month. The Arena scraper already spends one per run: it fetches the page through Firecrawl. That 1 is the baseline for every cost number in this post.

It works in three steps: find a tool, read its contract, call it. The call goes to the regular /v2/scrape endpoint with an alexandria array instead of a URL.

If you want the full tour, JP Caparas already wrote it well: Firecrawl Alexandria, tested end to end with real API calls. I am not going to redo that. This post asks: does it replace a scraper I actually run?

First, You Need a Newer CLI

My installed Firecrawl CLI was 1.23.3, which predates Alexandria. The list and find-tools commands need 1.24 or later.

The 1.23.3 on my machine came from a plain npm install, skipping the installer that rewrites your editors. So the CLI is installed, just old, and I am not upgrading it blind for a look around. npx runs a specific version without replacing the one you have:

npx -y firecrawl-cli@1.24.6 credits --json

It picks up the same FIRECRAWL_API_KEY from your environment. When you decide you want Alexandria in scripts permanently, upgrade for real. Until then, this is a free look.

Browsing the Catalog Costs Nothing

Start at the root:

npx -y firecrawl-cli@1.24.6 list --json

That returned 23 categories for me on October 1st. The launch page said 20, and JP counted 21, so they seem to be adding categories as they go.

Each item carries a nextCommand you can run as is to go one level down. That is also why the commands in this post alternate between bare list and alexandria list: I am running whatever nextCommand prints, verbatim. The one I cared about:

Neon data streams converging into a glowing folder icon

npx -y firecrawl-cli@1.24.6 alexandria list ai-models --category --json

Trimmed to what matters:

{
  "provider": "firecrawl",
  "capability": "find-tools",
  "creditsCost": 0,
  "data": {
    "level": "providers",
    "items": [
      { "id": "artificialanalysis-ai", "name": "Artificial Analysis", "toolCount": 4 },
      { "id": "arxiv-org", "name": "arXiv", "toolCount": 6 },
      { "id": "huggingface-co", "name": "Hugging Face", "toolCount": 5 },
      { "id": "ollama-com", "name": "Ollama model library", "toolCount": 7 },
      { "id": "openrouter-ai", "name": "OpenRouter", "toolCount": 8 },
      { "id": "semanticscholar-org", "name": "Semantic Scholar", "toolCount": 10 }
    ],
    "total": 6
  }
}

Six providers. No Arena. When I first ran this a few days earlier it was three: Artificial Analysis, Hugging Face, and OpenRouter. arXiv, Ollama, and Semantic Scholar showed up since, so this category is growing as fast as the category list is.

The Answer for Arena: Keep Scraping

If you want Arena’s rankings, Alexandria does not have them. The Arena scraper stays a scraper, and it costs one credit a run.

It took two commands. Every scraper in this series gets that check from now on: it takes a minute.

But the category was not empty, and one of those providers is close enough to be interesting.

The Leaderboard Next Door: Artificial Analysis

Artificial Analysis runs its own LLM leaderboard. It doesn’t use human votes. It runs benchmarks and puts them into an Intelligence Index, along with list prices and measured speed. It answers a different question than Arena does: Arena asks which answer people like better, and Artificial Analysis asks how the model scores on tests. For “did a new model just show up near the top,” it works too: the same question, answered from the benchmark side. A second signal, not a substitute.

Listing that provider’s tools shows the price of everything before you spend anything:

npx -y firecrawl-cli@1.24.6 list ai-models artificialanalysis-ai --category --json
ToolCreditsPer record?
benchmarks/search_models5no
benchmarks/get_model5no
benchmarks/list_model_providers5no
benchmarks/search_provider_endpoints5no

OpenRouter’s eight tools are also 5 credits a call. perRecord: false is the flag to look at. It means the price is per call, not per row, so asking for 100 models costs the same as asking for one. If that flag were true, a pagination loop could burn credits fast.

The contract for search_models says it returns the whole leaderboard sorted by Intelligence Index, up to 100 per page, with filters for creator, open weights, and reasoning. Nothing in it needs a terms agreement. So I spent the five credits:

npx -y firecrawl-cli@1.24.6 scrape artificialanalysis-ai/benchmarks/search_models \
  --options '{"limit": 20}' --json

A spectral AI figure reaching toward a floating column of model score cards

It came back in about a second and a half. Each result is a proper record: slug, creator, the index, a dozen benchmark scores, prices per million tokens, context window, and a link back to the model’s page on the site. No regex. No --wait-for 3000 hoping the table rendered. No guessing which row is the calendar widget.

And then I looked at the top 20.

The Top 20 Is Really Nine Models

Here are the first eight rows:

1  claude-opus-5-5          Anthropic  57.6
2  claude-opus-5-5-xhigh    Anthropic  56.0
3  claude-opus-5-5-high     Anthropic  53.6
4  claude-fable-5-1         Anthropic  53.4
5  claude-fable-5-1-xhigh   Anthropic  53.2
6  gpt-6-astra              OpenAI     52.7
7  gpt-6-astra-xhigh        OpenAI     52.4
8  claude-opus-5-5-medium   Anthropic  51.2

Artificial Analysis lists every reasoning-effort setting as its own row. Opus 5.5 at max effort, at xhigh, at high, at medium. Those are four entries for one model. The “top 20” is really about nine models, several times over. Claude Opus 5.5 alone holds four of the top eight slots, which makes this less a leaderboard than an Anthropic family newsletter.

That is not a bug. Nobody ships one this loud, and the first test run prints it. It is a caveat of the swap, and the first one to handle: the Arena harness assumes one row per model, and this feed returns one row per model per effort setting. The contract gives you clean JSON but clean data does not mean you can skip reading it.

The fix is to collapse effort variants into one model before ranking. The slugs end in -xhigh, -high, -medium, -low, so strip that and keep the first (highest) row for each model.

My first version of that also stripped -max, since the variants are labeled “Max Effort” in their names. That ate a real model. qwen3-8-max is not Qwen 3.8 at max effort. “Max” is part of the product’s name. The max-effort rows in this table have no suffix at all, which you only find out by looking. So max came out of the pattern.

The Same Harness, New Fetch Step

Here is the Arena harness with the scraper swapped for an Alexandria call. It is still 70 lines, and everything below fetch() is the same code: save with a timestamp, compare with the last run, speak only when a model is new to the top 20. The Arena watcher keeps its job. This one runs alongside it: same harness, different signal.

#!/usr/bin/env python3
"""arena_watch.py with the scraper swapped for a Firecrawl Alexandria call."""
import json, re, sqlite3, subprocess, sys
from datetime import datetime, timezone

DB = "aa.db"
TOP = 20

# Artificial Analysis lists every reasoning-effort setting as its own row, so
# claude-opus-5-5, -xhigh and -high are three rows. Collapse them to one model.
EFFORT = re.compile(r"-(xhigh|high|medium|low|minimal)$")


def fetch():
    out = subprocess.run(
        ["firecrawl", "scrape", "artificialanalysis-ai/benchmarks/search_models",
         "--options", '{"limit": 100}', "--json"],
        capture_output=True, text=True, check=True).stdout
    call = json.loads(out[out.index("{"):])["data"]["alexandria"][0]
    rows, seen = [], set()
    for m in call["data"]["results"]:  # already sorted by Intelligence Index
        model = EFFORT.sub("", m["slug"])
        if model in seen:
            continue  # a lower effort setting of a model we already ranked
        seen.add(model)
        rows.append((len(rows) + 1, model, m["creator_name"],
                     round(m["intelligence_index"] or 0, 1)))
    return rows


def main():
    rows = fetch()
    if len(rows) < TOP:
        sys.exit(f"only got {len(rows)} models back, check the call")

    taken = datetime.now(timezone.utc).strftime("%Y-%m-%d %H:%M")
    db = sqlite3.connect(DB)
    db.execute("""CREATE TABLE IF NOT EXISTS ranks
                  (taken TEXT, rank INT, model TEXT, org TEXT, score REAL)""")
    db.executemany("INSERT INTO ranks VALUES (?,?,?,?,?)", [(taken, *r) for r in rows])
    db.commit()

    prev = db.execute("SELECT MAX(taken) FROM ranks WHERE taken < ?", (taken,)).fetchone()[0]
    if not prev:
        print(f"{taken}: first run, saved {len(rows)} models. Nothing to compare yet.")
        return

    was_top = {m for (m,) in db.execute(
        "SELECT model FROM ranks WHERE taken = ? AND rank <= ?", (prev, TOP))}
    seen = {m for (m,) in db.execute("SELECT DISTINCT model FROM ranks WHERE taken < ?", (taken,))}

    print(f"{taken} vs {prev}")
    for rank, model, org, score in rows[:TOP]:
        if model not in seen:
            print(f"  NEW      #{rank:<3} {model} ({org}, {score})")
        elif model not in was_top:
            print(f"  MOVED UP #{rank:<3} {model} ({org}, {score})")


if __name__ == "__main__":
    main()

The script shells out to plain firecrawl. The pinned npx form above was for looking around. On the box that runs the cron, either upgrade the installed CLI for real or swap the command for the full npx version.

Two ribbons of light braided into a single bright knot

It asks for 100 rows instead of 20, because after collapsing there have to be 20 distinct models left. Since pricing is per call, the extra 80 rows are free. That one call came back with 61 distinct models. The first run’s top 20, with the variants collapsed:

#1   claude-opus-5-5      (Anthropic, 57.6)
#2   claude-fable-5-1     (Anthropic, 53.4)
#3   gpt-6-astra          (OpenAI, 52.7)
#4   muse-spark-1-3       (Meta, 48.1)
#5   gpt-6-sol            (OpenAI, 47.5)
#6   grok-4-7             (SpaceXAI, 46.4)
#7   mimo-v2-6-pro        (Xiaomi, 46.3)
#8   qwen3-8-max          (Alibaba, 45.4)
...
#20  gpt-6-luna           (OpenAI, 37.3)

Now that is a top 20 you can diff.

The sys.exit guard means something different here, too. In the Arena version, fewer than 20 rows meant “the page layout changed, go fix the regex.” Here it means the call itself went wrong.

What the Swap Costs

 Arena scraperAlexandria call
Credits per run15
Daily for 30 days30150
ParsingRegex on a markdown tableJSON, fields documented in the contract
Breaks whenThe page redesignsThe contract changes
Extra fieldsRank, score, orgBenchmarks, prices, speed, context window
What it measuresHuman preference votesBenchmark scores

A daily Alexandria check uses 150 of the free 1,000 a month, which is fine for one watcher and adds up if you run five.

A glowing chip feeding circuit traces that fan out into floating currency symbols

So the swap is not free. Five times the credits buys you no regex, no layout breakage, and a lot of extra fields. The Model Buzz Report leans on both kinds of signal, so the two lists side by side say more than either one alone.

My test cost 15 credits: three five-credit calls, one of them the run that ate Qwen. Fifteen, out of the 1,000 I pay nothing for. All that comparison shopping, and the grand total is 1.5% of free.

If Your Cron Runs Unattended, Read This Part

Some Alexandria providers will not run until you accept their data-use terms. Alexandria tells the caller, human or agent, to show the terms to a person and wait for an explicit yes.

The two providers I touched are fine: terms show openrouter-ai returns "required": false, and Artificial Analysis is not in the terms catalog at all. But check before you schedule anything:

npx -y firecrawl-cli@1.24.6 alexandria terms show <provider> --json

Check the Catalog Before You Write the Scraper

Every recipe in this series so far has assumed the data lives on a web page and the job is getting it off that page. Alexandria adds a step before that: someone may already sell the structured version.

The step is cheap. list, follow nextCommand two or three times, read the contract. Takes a couple of minutes. If what you want is in there, you skip the part of scraping that breaks. If it is not, like Arena, you have lost nothing and you write the scraper knowing you have to.

Next up for this one: OpenRouter’s get_model_stats, which says it returns usage and availability for a model. That is the other half of what the Model Buzz Report watches, and I have been getting it the hard way. Which is, as established, my signature way of getting things.

Stephan Miller

Written by

Kansas City Software Engineer and Author

Twitter | Github | LinkedIn

Updated

* This website contains affiliate links. This means that if you click on a link and purchase a product or service, I may receive a small commission at no extra cost to you. Please note that I only recommend products and services that I believe in and that will add value to my readers. Not all links on this website are affiliate links. Learn more.