40 to 60 percent of agent spend hides in the gap between the invoice and cost per successful outcome, and most of it is not the model. It is the retry.
Cost per token is a procurement metric. Cost per successful outcome is a management metric. Here is where the six sources of waste concentrate, and the six-week path to a real number.
Ask a CIO what their agents cost last month and you get an invoice. Ask what one resolved claim, one reconciled ledger, or one closed ticket cost, and the room goes quiet. That gap is where 40 to 60 percent of agent spend hides, and most of it is not the model. It is the retry.
Most organizations running agents today can answer exactly one financial question: what the provider charged last month. It arrives as a handful of line items — a model, a token count, a total — and it cannot answer which of fourteen agents consumed the majority of spend, what one resolved case cost end to end, or how much of the bill went on runs that ultimately failed. The absence of these answers has a predictable consequence: finance sees a line growing 20-30% a month with no unit denominator, and imposes a cap. The cap lands on the agents that are working, because they are the ones with volume.
Cost per token is a procurement metric. Cost per successful outcome is a management metric. Until you can produce the second one, per agent and per task type, every cost decision is a guess with a spreadsheet attached.
Token spend is not a model-selection problem. It is an accounting problem followed by a control-loop problem, and it is worth attacking in that order.
| Metric | Definition | Why it is the one to watch |
|---|---|---|
| Cost per successful outcome | Total spend on a task type ÷ successful completions | The only figure comparable to the manual process it replaced |
| Waste ratio | Spend on failed, abandoned, and superseded runs ÷ total spend | Where the recoverable money is — above 20% is common and fixable |
| Retry amplification | Mean tokens per completed run ÷ tokens of a first-attempt run | Exposes full-context replay. Above 1.6 means the retry path is naive |
| Cache hit rate | Cached input tokens ÷ total input tokens | For a stable agent this should exceed 70%. Most start near zero |
| Frontier share | Spend on the largest model ÷ total, split by task complexity | Reveals frontier models doing classification work |
The shares below are indicative ranges rather than a fixed distribution, but the ordering has been remarkably stable across deployments.
Retry with full-context replayThe same history re-sent on every attempt. Signal: retry amplification above 1.6. Fix: delta repair and a circuit breaker.
No prompt cachingSystem prompt and tool definitions re-read every call. Signal: cache hit rate near zero. Fix: stable prefix and cache breakpoints.
Over-modelling simple stepsFrontier model doing classification and formatting. Signal: frontier share high on low-complexity tasks. Fix: model cascade with escalation.
Context bloatWhole documents pasted, history never pruned. Signal: input tokens rising with conversation length. Fix: retrieve spans, summarize, prune.
Unbounded loops and fan-outNo step limit, multi-agent broadcast with no budget. Signal: a long tail of runs at the token ceiling. Fix: step budget and per-run ceiling.
Interactive pricing on batchable workEvals, back-office runs, overnight processing. Signal: synchronous calls outside business hours. Fix: batch endpoint, roughly half price.
The first two rows require no quality trade-off at all — they are engineering defects rather than cost-and-quality decisions, which is why they come before any model-downgrade conversation.
Retrying is correct behavior. The problem is the standard implementation: when a tool call fails validation, most agent frameworks append the error to the conversation and re-send the whole thing. Context grows with each attempt, and because input tokens dominate agentic workloads, cost grows with it.
| Retry cause | What usually happens | What works better |
|---|---|---|
| Transient infrastructure | Immediate retry, same payload | Exponential backoff with jitter, idempotency key, no context change |
| Schema/format violation | Append error, resend everything | Constrained decoding first; if it fails, send the typed error and failing fragment only |
| Business rule failure | Resend with "you made a mistake" | Return the specific rule violated and the field |
| Judge/guardrail rejection | Regenerate from scratch | Cap regeneration at one attempt; a second rejection is a routing decision |
| Planner thrash | Nothing, it runs until the timeout | Progress detection, a hard step budget, abandon then escalate |
Cache the stable prefix. Everything invariant — system instructions, tool definitions, policy text, few-shot examples — sits at the front and never changes within a session. Usually a one-day change, the largest single return.
Fix the retry path. Typed errors, delta repair, circuit breakers, step budgets. Two to three weeks, no quality trade-off.
Route by complexity. A cascade — small model attempts, escalates on low confidence — rather than one model for every step. Needs eval coverage to do safely.
Batch what does not need to be interactive. Evals, overnight processing, bulk document work. Pure scheduling change, roughly half price.
Span schema agreed and emitted from one agent. Rate card loaded, versioned, effective-dated.
How you know it worked: any single run can be priced end to end.All agents instrumented. Outcome join on run identifier. First burn report, unedited.
How you know it worked: cost per successful outcome exists per task type.Waste analysis: retry amplification, cache hit rate, frontier share, the expensive tail.
How you know it worked: a ranked list of waste sources with a value on each.Prompt restructured for caching. Per-run token ceiling and step budget enforced in the runtime.
How you know it worked: cache hit rate above 60%, no run exceeds the ceiling.Retry path rebuilt: typed errors, delta repair, circuit breaker, escalation instead of a third attempt.
How you know it worked: retry amplification below 1.3.Routing cascade on the highest-volume simple task, gated by the eval suite. Batch migration for offline work.
How you know it worked: no quality regression on the eval gate.A worked example on an illustrative mid-size deployment, 250,000 runs a month. The rate card is representative rather than any provider's current pricing — the ratios between the lines are far more stable than the absolute figures.
| Line | Before | Change made | Monthly saving |
|---|---|---|---|
| Input tokens billed at full rate | 92% | Stable-prefix caching → 28% | $34,000 |
| Retry amplification | 1.9× | Typed errors, circuit breaker → 1.2× | $21,000 |
| Frontier model share of calls | 100% | Cascade on 3 simple task types → 54% | $16,000 |
| Offline work on interactive pricing | 18% | Batch endpoint → 2% | $7,000 |
| Runs abandoned at timeout | 4.1% | Progress detection, escalation → 0.6% | $5,000 |
| Cost per successful outcome | $0.61 | — | $0.24 · 61% reduction |
The six sources of waste above are not a separate cost project. AgentTrust OS's runtime intercepts the exact decision points where the burn concentrates — before it becomes an invoice line nobody can attribute.
A deterministic-first runtime that resolves most decisions in under 20ms, and an audit trail that turns cost per outcome into a query, not a project.
Start Free →