Platform

Which tokens did not need to be spent.

The Token Ledger says what the tokens cost. The usage scan says which of them were avoidable. Both read the same local transcripts, and neither makes a network request.

Rebuilt from disk

Model use and API-equivalent spend from ~/.claude/projects and ~/.codex/sessions. No proxy, no telemetry endpoint, no account lookup.

Nine detectors

Cache hit rate, context carried, repeated reads, oversized results, thinking share, idle resumes, unused connectors, subagent fan-out, instruction-file length.

Priced on the right meter

Each finding says which meter its tokens sat on — fresh input, output or cache read — so the money line uses the rate that traffic was actually billed at.

A security product does not belong in the prompt traffic path

That is the whole reason this reads transcripts instead of proxying anything. Nothing sits between an agent and its model provider, so there is nothing new to trust, nothing to go down, and no second copy of your prompts.

Content is never stored. A repeated read is spotted by hashing the result in memory and keeping the hash. The per-file cache is keyed by size and modification time, so a finished session is parsed once, ever.

What the detectors look at

Fires whenEstimate
Cache hit ratemedian below 70% in a sessionmodelled from what a warm cache would have served
Idle resumes and mid-session model switchesrebuilds over 50K tokensmeasured — the fresh input and cache writes on the turns that followed
Context carriedany session above 150Kmeasured, then halved, because a split session still carries a handoff
Oversized single resultsabove the read thresholdmeasured size, two thirds assumed recoverable by a scoped read
Byte-identical resultsthe second copy in one sessionmeasured — the extra copies
Instruction file lengthCLAUDE.md or AGENTS.md over 200 linesmodelled at about 11 tokens a line, per session
Declared connectors never calledanyno figure claimed
Subagent fan-outthree or more in a sessionno figure claimed
Thinking as a share of outputabove 25%measured, priced as output

Where a figure would be a guess, none is given. Carried context priced as fresh input would overstate it tenfold, which is the mistake this field exists to prevent.

Two scores, kept apart

The usage score sits beside the doctrine score rather than inside it. One measures safety and the other measures waste; blending them dilutes both. Subagent transcripts are counted but not judged — a subagent has no conversation of its own to keep warm, so scoring its cache hit rate would be meaningless.

Right-sizing is arithmetic, not advice

The comparison prices the traffic each model actually ran against the next model down the ladder. It is arithmetic on tokens already spent. The note beside each swap says what the transcripts suggest about whether the work would survive it — a high thinking share is the worst candidate, heavy tool use the most plausible one — and the view says plainly that a smaller model changes the answers as well as the bill.