On small jobs, it didn't. On a real component, building it yourself cost about twice the tokens of reusing a library, and the scratch version still did less.
Build or reuse is one of the oldest decisions in system design. AI agents changed its cost model. Generating code is now cheap at the moment you ask for it, so the cost moved downstream: into the cases nobody tested, the fixes that follow, and the context every agent carries on every call. The question isn't "can the agent write it?" anymore. It's "what does it cost to own what it wrote?"
Since Friday I've run about 3.3 billion tokens through Claude Code, Codex and opencode. Before that sounds like bragging: 97% of it was the agents re-reading context they'd already seen. The new work was closer to 100 million. I'd also just started testing free models in a new harness, so some of this was me learning where the edges are.
My actual cash outlay for API calls was 21 cents. The rest ran on flat subscriptions, worth about $1,166 at API list prices. The surprise was a security-review hook I'd installed: it quietly called a frontier model 316 times on its own, about $97 of that. You don't see the bill for the things you didn't ask for. That's the carrying cost this piece is about.
I built RepoHunter to stop AI agents from rewriting code that already exists as maintained open source. The pitch writes itself: reuse instead of regenerate, save tokens. I had never measured it. So I did, and I'm publishing the results, including the parts that didn't go my way.
I ran out of money for AI subscriptions this year. One terminal running about 20 agents, each with around six subagents, emptied a $200 plan in roughly six hours (before weekly caps existed). Since then I've treated tokens like a budget, and this is what I've learned about where they actually go.
Four everyday jobs an agent gets asked to do: write a Word document, parse CSV files, generate a two-page PDF invoice, and render GitHub-flavored Markdown. Each job ran two ways:
Same worker model (GPT-5.6 Sol through the Codex CLI), same harness, and the same executed pass/fail test for both arms. Every one of the 23 runs passed its test on the first attempt. Token counts are what the CLI reported. MEASURED
For the three small jobs, reusing a library saved little or nothing. A frontier agent writes 100 lines of CSV parsing about as cheaply as it finds and wires up a library.
| Job | Built (each run) | Reused (each run) | Library chosen |
|---|---|---|---|
| Word document | 34,479 · 34,206 | 37,385 · 22,533 | docx |
| CSV parser | 18,722 · 29,969 | 7,955 · 34,191 | Papa Parse |
| PDF invoice | 19,651 · 35,161 · 19,662 | 27,277 · 22,935 · 21,431 | jsPDF, pdfkit, pdfkit |
If the job is small and you only need what you asked for, let the agent write it.
The Markdown renderer told a different story. Both arms passed the same feature test (tables, code blocks, task lists, footnotes) for similar token counts. Then I ran 16 everyday Markdown cases the agents had never seen.
No bold, no italics, no links, images or numbered lists. It did exactly what the test asked for.
All three runs, on the first pass. markdown-it had already done the work.
So I measured the second bill: what it costs to bring each scratch renderer up to the library's everyday quality, under the same rules. MEASURED
Reuse took about half the tokens. And even after the fix, the scratch renderers only handle what I tested. The library handles the cases nobody thought to ask for yet.
What I'm not claiming. This is a field log, not a benchmark: 2–3 runs per arm, one worker model, one machine, four jobs. Small jobs came out roughly even, and the 2× comes from one component type.
The build arm was not allowed packages, so this compares a library versus no library. It does not compare RepoHunter against an agent that picks a library on its own. I also chose the 16 everyday cases after reading one scratch renderer, and the fix runs were given those same cases.
Every token count and the method are in the raw data, so you can check my math.
The prompt you type is the smallest part of the bill. Every turn re-sends the conversation history, the system prompt, and the description of every tool, skill and MCP server your harness has loaded.
On my machine, 40 synced plugins added roughly 38,000 tokens of skill descriptions to every session. 39 of them had never been used. ESTIMATE · /doctor · SEP 29 Others have measured the same effect with a logging proxy: about 33k tokens before the first prompt in one harness versus about 7k in another. THIRD-PARTY
That fixed cost multiplies with agents. With 20 agents running six subagents each, the same catalog loads up to 120 times before any work starts. Caching helps, but it's a discount, not an exemption: cached tokens still count toward your limits.
Routers pick the cheapest chat model. They rarely ask whether the job needs a chat model at all.
| Job | Cheapest thing that works |
|---|---|
| Yes/no, classify, score, route, gate | A decision model, or a small model with structured output |
| A settled, repeatable transform | A script, with no model at all |
| A bounded code change with tests | A free or cheap coding model, in its own worktree |
| Design, ambiguous debugging, client-facing writing | A frontier model |
| A component that already exists as open source | A library, found and checked before your agent writes it |
Context: this was my first week running free models, in a harness I had only just started testing. These are first-week numbers, not a tuned setup.
I benchmarked free coding models with hidden tests on September 28. MEASURED The best free lanes passed every task, and most of them log your code or may train on it. For private code, my only free options were the zero-retention lanes, and those were the least reliable. Rate limits capped batches at about 6 to 9 tasks before cost ever did.
The hidden cost is review. In one batch of 10 free agents, 5 branches were ready to merge. The rest needed review, rewrote docs they were asked to extend, or never committed. Your hour reviewing them is the real bill.
The most expensive token is the one spent on a run you throw away. Each failure from that first week is now a rule the harness enforces, and none of them costs a token:
write a feature
install a plugin, skill or MCP server
add a dependency you'll ship
Does it already exist?
Maintained, or archived?
Can you ship its license?
Anything risky in it? (CLI)
Runs on your machine?
GO reuse it
MAYBE check first
SKIP build or pick another
RepoHunter is a reuse gate. Before your agent writes a component from scratch, it checks whether a maintained open-source repo already does the job, and flags what to check before you adopt it: real GitHub data, maintenance, license and resale risk, and (in the CLI) a pattern-based safety scan for prompt injection, piped installs and leaked secrets. You get GO, MAYBE or SKIP with the reasons.
It practices what this piece preaches. Fully loaded, its three tools and two skills add 1,038 tokens to a session, and Claude Code can defer them until a task needs them. MEASURED · OCT 4
In all 10 reuse runs, the library the agent chose came straight from RepoHunter's search results and was then vetted (2 GO, 8 MAYBE): docx, Papa Parse, pdfkit or jsPDF, and markdown-it. Search, vetting and integration together still came in at about half the tokens of building and fixing a real component.
Free and MIT-licensed. Install it in Claude Code:
Other agents: uvx --from git+https://github.com/meetziggy/[email protected] repohunter-mcp