Before You Write a Scraper, Check Firecrawl Alexandria
Last Thursday I published a 70-line scraper that watches the Arena leaderboard for new AI models. It runs daily and announces new arrivals.
Two days earlier, Firecrawl launched Alexandria, a catalog of providers that sell the structured version of exactly this kind of data, callable through the same scrape endpoint I was already paying for. I didn’t notice. I was busy having AI build me regex for a markdown table, while the company whose credits I was spending shipped the thing meant to make that regex unnecessary.
So to recap: I built a machine that watches the internet for new arrivals, and the first new arrival it missed was the one built to replace it.
If somebody is already selling the structured version of the data I wanted, why was I scraping the page? It never occurred to me to check.
- What Alexandria Is
- First, You Need a Newer CLI
- Browsing the Catalog Costs Nothing
- The Answer for Arena: Keep Scraping
- The Leaderboard Next Door: Artificial Analysis
- The Top 20 Is Really Nine Models
- The Same Harness, New Fetch Step
- What the Swap Costs
- If Your Cron Runs Unattended, Read This Part
- Check the Catalog Before You Write the Scraper
What Alexandria Is
Alexandria is Firecrawl’s catalog of third-party data sources: FRED economic data, podcast transcripts, app store listings, company records, and a few hundred other things. They named their permanent store of the world’s data after history’s most famous case of data loss, which is either confidence or a gap in the marketing team’s classics minor. Each tool in it has a published contract that says what options it takes, what it returns, and what it costs. Finding tools and reading contracts is free. Running one costs credits, Firecrawl’s metered unit, and the free plan is 1,000 of them a month. The Arena scraper already spends one per run: it fetches the page through Firecrawl. That 1 is the baseline for every cost number in this post.
It works in three steps: find a tool, read its contract, call it. The call goes to the regular /v2/scrape endpoint with an alexandria array instead of a URL.
If you want the full tour, JP Caparas already wrote it well: Firecrawl Alexandria, tested end to end with real API calls. I am not going to redo that. This post asks: does it replace a scraper I actually run?
First, You Need a Newer CLI
My installed Firecrawl CLI was 1.23.3, which predates Alexandria. The list and find-tools commands need 1.24 or later.
The 1.23.3 on my machine came from a plain npm install, skipping the installer that rewrites your editors. So the CLI is installed, just old, and I am not upgrading it blind for a look around. npx runs a specific version without replacing the one you have:
npx -y firecrawl-cli@1.24.6 credits --json
It picks up the same FIRECRAWL_API_KEY from your environment. When you decide you want Alexandria in scripts permanently, upgrade for real. Until then, this is a free look.
Browsing the Catalog Costs Nothing
Start at the root:
npx -y firecrawl-cli@1.24.6 list --json
That returned 23 categories for me on October 1st. The launch page said 20, and JP counted 21, so they seem to be adding categories as they go.
Each item carries a nextCommand you can run as is to go one level down. That is also why the commands in this post alternate between bare list and alexandria list: I am running whatever nextCommand prints, verbatim. The one I cared about:

npx -y firecrawl-cli@1.24.6 alexandria list ai-models --category --json
Trimmed to what matters:
{
"provider": "firecrawl",
"capability": "find-tools",
"creditsCost": 0,
"data": {
"level": "providers",
"items": [
{ "id": "artificialanalysis-ai", "name": "Artificial Analysis", "toolCount": 4 },
{ "id": "arxiv-org", "name": "arXiv", "toolCount": 6 },
{ "id": "huggingface-co", "name": "Hugging Face", "toolCount": 5 },
{ "id": "ollama-com", "name": "Ollama model library", "toolCount": 7 },
{ "id": "openrouter-ai", "name": "OpenRouter", "toolCount": 8 },
{ "id": "semanticscholar-org", "name": "Semantic Scholar", "toolCount": 10 }
],
"total": 6
}
}
Six providers. No Arena. When I first ran this a few days earlier it was three: Artificial Analysis, Hugging Face, and OpenRouter. arXiv, Ollama, and Semantic Scholar showed up since, so this category is growing as fast as the category list is.
The Answer for Arena: Keep Scraping
If you want Arena’s rankings, Alexandria does not have them. The Arena scraper stays a scraper, and it costs one credit a run.
It took two commands. Every scraper in this series gets that check from now on: it takes a minute.
But the category was not empty, and one of those providers is close enough to be interesting.
The Leaderboard Next Door: Artificial Analysis
Artificial Analysis runs its own LLM leaderboard. It doesn’t use human votes. It runs benchmarks and puts them into an Intelligence Index, along with list prices and measured speed. It answers a different question than Arena does: Arena asks which answer people like better, and Artificial Analysis asks how the model scores on tests. For “did a new model just show up near the top,” it works too: the same question, answered from the benchmark side. A second signal, not a substitute.
Listing that provider’s tools shows the price of everything before you spend anything:
npx -y firecrawl-cli@1.24.6 list ai-models artificialanalysis-ai --category --json
| Tool | Credits | Per record? |
|---|---|---|
benchmarks/search_models | 5 | no |
benchmarks/get_model | 5 | no |
benchmarks/list_model_providers | 5 | no |
benchmarks/search_provider_endpoints | 5 | no |
OpenRouter’s eight tools are also 5 credits a call. perRecord: false is the flag to look at. It means the price is per call, not per row, so asking for 100 models costs the same as asking for one. If that flag were true, a pagination loop could burn credits fast.
The contract for search_models says it returns the whole leaderboard sorted by Intelligence Index, up to 100 per page, with filters for creator, open weights, and reasoning. Nothing in it needs a terms agreement. So I spent the five credits:
npx -y firecrawl-cli@1.24.6 scrape artificialanalysis-ai/benchmarks/search_models \
--options '{"limit": 20}' --json

It came back in about a second and a half. Each result is a proper record: slug, creator, the index, a dozen benchmark scores, prices per million tokens, context window, and a link back to the model’s page on the site. No regex. No --wait-for 3000 hoping the table rendered. No guessing which row is the calendar widget.
And then I looked at the top 20.
The Top 20 Is Really Nine Models
Here are the first eight rows:
1 claude-opus-5-5 Anthropic 57.6
2 claude-opus-5-5-xhigh Anthropic 56.0
3 claude-opus-5-5-high Anthropic 53.6
4 claude-fable-5-1 Anthropic 53.4
5 claude-fable-5-1-xhigh Anthropic 53.2
6 gpt-6-astra OpenAI 52.7
7 gpt-6-astra-xhigh OpenAI 52.4
8 claude-opus-5-5-medium Anthropic 51.2
Artificial Analysis lists every reasoning-effort setting as its own row. Opus 5.5 at max effort, at xhigh, at high, at medium. Those are four entries for one model. The “top 20” is really about nine models, several times over. Claude Opus 5.5 alone holds four of the top eight slots, which makes this less a leaderboard than an Anthropic family newsletter.
That is not a bug. Nobody ships one this loud, and the first test run prints it. It is a caveat of the swap, and the first one to handle: the Arena harness assumes one row per model, and this feed returns one row per model per effort setting. The contract gives you clean JSON but clean data does not mean you can skip reading it.
The fix is to collapse effort variants into one model before ranking. The slugs end in -xhigh, -high, -medium, -low, so strip that and keep the first (highest) row for each model.
My first version of that also stripped -max, since the variants are labeled “Max Effort” in their names. That ate a real model. qwen3-8-max is not Qwen 3.8 at max effort. “Max” is part of the product’s name. The max-effort rows in this table have no suffix at all, which you only find out by looking. So max came out of the pattern.
The Same Harness, New Fetch Step
Here is the Arena harness with the scraper swapped for an Alexandria call. It is still 70 lines, and everything below fetch() is the same code: save with a timestamp, compare with the last run, speak only when a model is new to the top 20. The Arena watcher keeps its job. This one runs alongside it: same harness, different signal.
#!/usr/bin/env python3
"""arena_watch.py with the scraper swapped for a Firecrawl Alexandria call."""
import json, re, sqlite3, subprocess, sys
from datetime import datetime, timezone
DB = "aa.db"
TOP = 20
# Artificial Analysis lists every reasoning-effort setting as its own row, so
# claude-opus-5-5, -xhigh and -high are three rows. Collapse them to one model.
EFFORT = re.compile(r"-(xhigh|high|medium|low|minimal)$")
def fetch():
out = subprocess.run(
["firecrawl", "scrape", "artificialanalysis-ai/benchmarks/search_models",
"--options", '{"limit": 100}', "--json"],
capture_output=True, text=True, check=True).stdout
call = json.loads(out[out.index("{"):])["data"]["alexandria"][0]
rows, seen = [], set()
for m in call["data"]["results"]: # already sorted by Intelligence Index
model = EFFORT.sub("", m["slug"])
if model in seen:
continue # a lower effort setting of a model we already ranked
seen.add(model)
rows.append((len(rows) + 1, model, m["creator_name"],
round(m["intelligence_index"] or 0, 1)))
return rows
def main():
rows = fetch()
if len(rows) < TOP:
sys.exit(f"only got {len(rows)} models back, check the call")
taken = datetime.now(timezone.utc).strftime("%Y-%m-%d %H:%M")
db = sqlite3.connect(DB)
db.execute("""CREATE TABLE IF NOT EXISTS ranks
(taken TEXT, rank INT, model TEXT, org TEXT, score REAL)""")
db.executemany("INSERT INTO ranks VALUES (?,?,?,?,?)", [(taken, *r) for r in rows])
db.commit()
prev = db.execute("SELECT MAX(taken) FROM ranks WHERE taken < ?", (taken,)).fetchone()[0]
if not prev:
print(f"{taken}: first run, saved {len(rows)} models. Nothing to compare yet.")
return
was_top = {m for (m,) in db.execute(
"SELECT model FROM ranks WHERE taken = ? AND rank <= ?", (prev, TOP))}
seen = {m for (m,) in db.execute("SELECT DISTINCT model FROM ranks WHERE taken < ?", (taken,))}
print(f"{taken} vs {prev}")
for rank, model, org, score in rows[:TOP]:
if model not in seen:
print(f" NEW #{rank:<3} {model} ({org}, {score})")
elif model not in was_top:
print(f" MOVED UP #{rank:<3} {model} ({org}, {score})")
if __name__ == "__main__":
main()
The script shells out to plain firecrawl. The pinned npx form above was for looking around. On the box that runs the cron, either upgrade the installed CLI for real or swap the command for the full npx version.

It asks for 100 rows instead of 20, because after collapsing there have to be 20 distinct models left. Since pricing is per call, the extra 80 rows are free. That one call came back with 61 distinct models. The first run’s top 20, with the variants collapsed:
#1 claude-opus-5-5 (Anthropic, 57.6)
#2 claude-fable-5-1 (Anthropic, 53.4)
#3 gpt-6-astra (OpenAI, 52.7)
#4 muse-spark-1-3 (Meta, 48.1)
#5 gpt-6-sol (OpenAI, 47.5)
#6 grok-4-7 (SpaceXAI, 46.4)
#7 mimo-v2-6-pro (Xiaomi, 46.3)
#8 qwen3-8-max (Alibaba, 45.4)
...
#20 gpt-6-luna (OpenAI, 37.3)
Now that is a top 20 you can diff.
The sys.exit guard means something different here, too. In the Arena version, fewer than 20 rows meant “the page layout changed, go fix the regex.” Here it means the call itself went wrong.
What the Swap Costs
| Arena scraper | Alexandria call | |
|---|---|---|
| Credits per run | 1 | 5 |
| Daily for 30 days | 30 | 150 |
| Parsing | Regex on a markdown table | JSON, fields documented in the contract |
| Breaks when | The page redesigns | The contract changes |
| Extra fields | Rank, score, org | Benchmarks, prices, speed, context window |
| What it measures | Human preference votes | Benchmark scores |
A daily Alexandria check uses 150 of the free 1,000 a month, which is fine for one watcher and adds up if you run five.

So the swap is not free. Five times the credits buys you no regex, no layout breakage, and a lot of extra fields. The Model Buzz Report leans on both kinds of signal, so the two lists side by side say more than either one alone.
My test cost 15 credits: three five-credit calls, one of them the run that ate Qwen. Fifteen, out of the 1,000 I pay nothing for. All that comparison shopping, and the grand total is 1.5% of free.
If Your Cron Runs Unattended, Read This Part
Some Alexandria providers will not run until you accept their data-use terms. Alexandria tells the caller, human or agent, to show the terms to a person and wait for an explicit yes.
The two providers I touched are fine: terms show openrouter-ai returns "required": false, and Artificial Analysis is not in the terms catalog at all. But check before you schedule anything:
npx -y firecrawl-cli@1.24.6 alexandria terms show <provider> --json
Check the Catalog Before You Write the Scraper
Every recipe in this series so far has assumed the data lives on a web page and the job is getting it off that page. Alexandria adds a step before that: someone may already sell the structured version.
The step is cheap. list, follow nextCommand two or three times, read the contract. Takes a couple of minutes. If what you want is in there, you skip the part of scraping that breaks. If it is not, like Arena, you have lost nothing and you write the scraper knowing you have to.
Next up for this one: OpenRouter’s get_model_stats, which says it returns usage and availability for a model. That is the other half of what the Model Buzz Report watches, and I have been getting it the hard way. Which is, as established, my signature way of getting things.
* This website contains affiliate links. This means that if you click on a link and purchase a product or service, I may receive a small commission at no extra cost to you. Please note that I only recommend products and services that I believe in and that will add value to my readers. Not all links on this website are affiliate links. Learn more.
