Guides
The whole AI cost playbook, published in full
None of our method is proprietary. These guides cover every lever we use on client work, in the order we use it, including the ones that make an engagement unnecessary. If you can ship it yourself in a fortnight, you should.
How to reduce LLM API costs: twelve levers, ranked by risk
The complete playbook, in the order we actually apply it: the changes that cannot affect your output, then the ones that can — and how to tell the difference before you ship.
Read the guideCost per successful task
Per-token dashboards will tell you a change was a success while your retry rate doubles. Here is the denominator that catches it, and how to build it so no single service can quietly fake the number.
11 min readPrompt caching explained
Caching is prefix-based, which means one timestamp in the wrong place can cost you the entire discount. Here is how it actually works and how to structure prompts so it fires.
10 min readModel routing explained
Everyone knows a smaller model would do for some of the traffic. The hard part is proving which part, before a customer finds out you guessed wrong.
12 min readWhy AI agents cost more than you modelled
A chat turn costs what a chat turn costs. An agent run costs whatever it decides to spend. Here are the four structural reasons, and the controls that make agent spend predictable.
10 min readAzure OpenAI cost optimisation
Caching and routing apply everywhere. These are the decisions that only exist because you are on Azure — and the ones teams most often inherit from a proof of concept and never revisit.
11 min readAI FinOps metrics
Most AI dashboards show tokens over time, which is the one chart that never tells you what to do next. These seven do.
10 min readWhy we publish this
Because the method is not the moat.
Every technique on this site is documented somewhere in a provider's own guidance. The hard part was never knowing that prompt caching exists — it is finding, in your specific estate, which of a dozen levers is worth the effort, sizing each one honestly, and proving the risky ones before a customer is affected.
If reading these means you fix it yourself and never speak to us, that is a good outcome. If reading them means you would rather someone did it accurately in two weeks with the measurement left behind, that is what the audit is for.
Find out what your AI actually costs per completed task.
Two weeks, a fixed fee, and a ranked savings plan with the quality risk of every move stated up front. If the numbers say an audit is not worth it for you, we will say so on the first call.