Field Notes 01 · October 4, 2026 · Brian Gorzelic

I tried to prove my tool saves tokens.

On small jobs, it didn't. On a real component, building it yourself cost about twice the tokens of reusing a library, and the scratch version still did less.

Agent builds it56.0k tokens
  1. Writes a renderer from scratch
  2. Passes the feature test
  3. Unseen everyday cases: 5–6 / 16
  4. Pays again to fix it
build 33.8kfix 22.2k
Agent reuses a library28.6k tokens
  1. Asks RepoHunter what exists
  2. Vets and installs markdown-it
  3. Passes the feature test
  4. Unseen everyday cases: 16 / 16
reuse 28.6k
Markdown renderer · median of 3 runs per arm · same model, harness and tests · bars drawn to scale
My weekend in tokensFri–Sun · from my own session logs
3.28Btokens, all harnesses
$0.21actual pay-per-use API spend
≈$1,166the subscription usage at API list prices (not billed)
≈$97of that from one unattended hook
Where the tokens went
re-reading context they'd already seen · 96.9%
cache reads 96.9%new work 3.1% (≈100M tokens)
How it was paid for
flat subscriptions, frontier models · 92.3%
Claude Code + Codex subscriptions 92.3%free and local models 7.7%pay-per-use API 0.02%
Claude Code, Codex and opencode logs, deduplicated · list prices from OpenRouter, Oct 4 · chat apps not included

The architecture question

Build or reuse is one of the oldest decisions in system design. AI agents changed its cost model. Generating code is now cheap at the moment you ask for it, so the cost moved downstream: into the cases nobody tested, the fixes that follow, and the context every agent carries on every call. The question isn't "can the agent write it?" anymore. It's "what does it cost to own what it wrote?"

My weekend, by the numbers

Since Friday I've run about 3.3 billion tokens through Claude Code, Codex and opencode. Before that sounds like bragging: 97% of it was the agents re-reading context they'd already seen. The new work was closer to 100 million. I'd also just started testing free models in a new harness, so some of this was me learning where the edges are.

My actual cash outlay for API calls was 21 cents. The rest ran on flat subscriptions, worth about $1,166 at API list prices. The surprise was a security-review hook I'd installed: it quietly called a frontier model 316 times on its own, about $97 of that. You don't see the bill for the things you didn't ask for. That's the carrying cost this piece is about.

I built RepoHunter to stop AI agents from rewriting code that already exists as maintained open source. The pitch writes itself: reuse instead of regenerate, save tokens. I had never measured it. So I did, and I'm publishing the results, including the parts that didn't go my way.

I ran out of money for AI subscriptions this year. One terminal running about 20 agents, each with around six subagents, emptied a $200 plan in roughly six hours (before weekly caps existed). Since then I've treated tokens like a budget, and this is what I've learned about where they actually go.

The test

Four everyday jobs an agent gets asked to do: write a Word document, parse CSV files, generate a two-page PDF invoice, and render GitHub-flavored Markdown. Each job ran two ways:

Same worker model (GPT-5.6 Sol through the Codex CLI), same harness, and the same executed pass/fail test for both arms. Every one of the 23 runs passed its test on the first attempt. Token counts are what the CLI reported. MEASURED

Small jobs: a wash

For the three small jobs, reusing a library saved little or nothing. A frontier agent writes 100 lines of CSV parsing about as cheaply as it finds and wires up a library.

Median tokens to a passing build

2–3 runs per arm · scale 0–56k tokens
Agent built itAgent reused a library
Word document
34.3k
30.0k
CSV parser
24.3k
21.1k
PDF invoice
19.7k
22.9k
Show the numbers as a table
JobBuilt (each run)Reused (each run)Library chosen
Word document34,479 · 34,20637,385 · 22,533docx
CSV parser18,722 · 29,9697,955 · 34,191Papa Parse
PDF invoice19,651 · 35,161 · 19,66227,277 · 22,935 · 21,431jsPDF, pdfkit, pdfkit

If the job is small and you only need what you asked for, let the agent write it.

Real components: you pay twice

The Markdown renderer told a different story. Both arms passed the same feature test (tables, code blocks, task lists, footnotes) for similar token counts. Then I ran 16 everyday Markdown cases the agents had never seen.

Agent built it
5–6/16

No bold, no italics, no links, images or numbered lists. It did exactly what the test asked for.

Agent reused a library
16/16

All three runs, on the first pass. markdown-it had already done the work.

So I measured the second bill: what it costs to bring each scratch renderer up to the library's everyday quality, under the same rules. MEASURED

Markdown renderer: total tokens to everyday quality

Median of 3 runs per arm · scale 0–56k tokens
BuildFix it laterReuse
Agent built it
56.0k
Agent reused a library
28.6k
Built + fixed per run: 29,532 · 55,973 · 56,688. Ratio of medians: 1.96×.

Reuse took about half the tokens. And even after the fix, the scratch renderers only handle what I tested. The library handles the cases nobody thought to ask for yet.

What I'm not claiming. This is a field log, not a benchmark: 2–3 runs per arm, one worker model, one machine, four jobs. Small jobs came out roughly even, and the 2× comes from one component type.

The build arm was not allowed packages, so this compares a library versus no library. It does not compare RepoHunter against an agent that picks a library on its own. I also chose the 16 everyday cases after reading one scratch renderer, and the fix runs were given those same cases.

Every token count and the method are in the raw data, so you can check my math.

Where tokens actually go: fixed cost per turn

System prompt
Every loaded tool, skill and MCP description ≈ 38k tokens on my setup (estimate)
The conversation so far
Your prompt
× every turn, × every subagent
What one agent turn re-sends. Illustration; only the 38k estimate is measured.

The prompt you type is the smallest part of the bill. Every turn re-sends the conversation history, the system prompt, and the description of every tool, skill and MCP server your harness has loaded.

On my machine, 40 synced plugins added roughly 38,000 tokens of skill descriptions to every session. 39 of them had never been used. ESTIMATE · /doctor · SEP 29 Others have measured the same effect with a logging proxy: about 33k tokens before the first prompt in one harness versus about 7k in another. THIRD-PARTY

That fixed cost multiplies with agents. With 20 agents running six subagents each, the same catalog loads up to 120 times before any work starts. Caching helps, but it's a discount, not an exemption: cached tokens still count toward your limits.

Six rules I design agent setups around

  1. Start a new session when the job changes.Old history rides along on every turn of the new job.
  2. Carry the result, not the route.Feed the accepted output forward. Drop the drafts, corrections and dead ends.
  3. Grep, don't read.Pull the lines you need instead of loading a whole large file.
  4. Keep MCP tools deferred.Load tool schemas only when a task needs them, and turn off what you haven't used in a month.
  5. Script anything settled.Once a procedure is fixed, a script costs zero tokens per run.
  6. End every session with a handoff file.Starting from a short handoff instead of an old transcript is the cheapest context there is.

Model fit before model price

Routers pick the cheapest chat model. They rarely ask whether the job needs a chat model at all.

JobCheapest thing that works
Yes/no, classify, score, route, gateA decision model, or a small model with structured output
A settled, repeatable transformA script, with no model at all
A bounded code change with testsA free or cheap coding model, in its own worktree
Design, ambiguous debugging, client-facing writingA frontier model
A component that already exists as open sourceA library, found and checked before your agent writes it

Free tokens aren't free

Context: this was my first week running free models, in a harness I had only just started testing. These are first-week numbers, not a tuned setup.

I benchmarked free coding models with hidden tests on September 28. MEASURED The best free lanes passed every task, and most of them log your code or may train on it. For private code, my only free options were the zero-retention lanes, and those were the least reliable. Rate limits capped batches at about 6 to 9 tasks before cost ever did.

The hidden cost is review. In one batch of 10 free agents, 5 branches were ready to merge. The rest needed review, rewrote docs they were asked to extend, or never committed. Your hour reviewing them is the real bill.

Guardrails that save whole runs

The most expensive token is the one spent on a run you throw away. Each failure from that first week is now a rule the harness enforces, and none of them costs a token:

Where RepoHunter fits

Your agent is about to…

write a feature
install a plugin, skill or MCP server
add a dependency you'll ship

RepoHunter checks

Does it already exist?
Maintained, or archived?
Can you ship its license?
Anything risky in it? (CLI)
Runs on your machine?

You get

GO reuse it
MAYBE check first
SKIP build or pick another

RepoHunter is a reuse gate. Before your agent writes a component from scratch, it checks whether a maintained open-source repo already does the job, and flags what to check before you adopt it: real GitHub data, maintenance, license and resale risk, and (in the CLI) a pattern-based safety scan for prompt injection, piped installs and leaked secrets. You get GO, MAYBE or SKIP with the reasons.

It practices what this piece preaches. Fully loaded, its three tools and two skills add 1,038 tokens to a session, and Claude Code can defer them until a task needs them. MEASURED · OCT 4

In all 10 reuse runs, the library the agent chose came straight from RepoHunter's search results and was then vetted (2 GO, 8 MAYBE): docx, Papa Parse, pdfkit or jsPDF, and markdown-it. Search, vetting and integration together still came in at about half the tokens of building and fixing a real component.

Your agent's best line of code is import.

Free and MIT-licensed. Install it in Claude Code:

$ /plugin marketplace add meetziggy/repohunter
$ /plugin install repohunter

Other agents: uvx --from git+https://github.com/meetziggy/[email protected] repohunter-mcp

Get RepoHunter

Sources and data