Free tool
AI cost calculator: what the four main levers are worth on your bill
Enter your volumes and the two rates from your own invoice. The model applies caching, batching, retry elimination and model routing in the order we would actually ship them — the ones that cannot change your output first.
Your workload
Your rates
Take these from your invoice, not from a pricing page. We do not hard-code provider prices here because they change and negotiated rates differ.
The levers
Current monthly spend
£6,989
£0.0175 per completed task. Of the total, £748.80 goes on attempts that had to be repeated.
No output change
No output change
No output change
Needs evaluation
Estimated after optimisation
£2,810−60%
£0.007 per task · £50,142 a year
An order-of-magnitude model, not a quote. It assumes your averages are representative and that the levers compose independently. Real estates are skewed — a small share of requests usually carries most of the cost — which is exactly what the audit measures.
Get the real numberHow the model works
Four levers, applied in the order we would ship them
Each lever operates on what is left after the previous one, which is why four levers that each save 30% do not add up to 120%. The three that cannot change your output run first.
1. Prompt caching
Applies the cache discount to the cacheable share of input tokens only — output is never cached. The cacheable share is whatever repeats byte-identically at the front of your prompt: system instructions, tool definitions, few-shot examples.
Prompt caching explained2. Batch endpoints
Applies the batch discount to the share of remaining spend that nobody is waiting on — nightly enrichment, backfills, evaluation runs, scheduled digests.
The twelve levers3. Retry elimination
Removes half of the extra attempts implied by your retry rate. Half rather than all, because some failures come from ambiguous input and provider errors you do not control.
Why failures belong in the numerator4. Model routing
Moves the routable share of what remains to a model priced at the percentage you specified. This is the only lever here that can change what a customer sees, which is why it is applied last and why it needs an evaluation harness.
Model routing explainedFAQ
About this calculator
How accurate is this AI cost calculator?
It is an order-of-magnitude model, not a quote. It assumes your averages are representative, that the levers compose independently, and that the rates you entered match your invoice. Real estates have skewed distributions where a small share of requests carries most of the cost, which this model cannot see. Treat the output as a reason to look properly, not as a number to put in a board pack.
Why do I have to enter my own prices?
Because provider list prices change frequently and negotiated rates differ from published ones. A hard-coded price table would silently make every result on this page wrong. Taking the two rates from your own invoice keeps the arithmetic honest and keeps this page correct for longer than any price table would be.
Why is the retry saving only half of my retry rate?
Because eliminating every retry is not realistic. Some failures come from genuinely ambiguous input, provider errors and timeouts you do not control. Recovering half of a retry rate through structured outputs, validation and better timeout handling is an achievable target; recovering all of it is not.
Does the order of the levers matter?
Yes, and this model applies them in the order we would ship them: the three that cannot change your output first, then routing, which can. Each lever operates on what is left after the previous one, which is why applying four levers that each save 30% does not save 120%.
An estimate is a reason to look. The audit is the look.
Two weeks, a fixed fee, and a cost per successful task computed from your own traces rather than your own averages.