Skip to content

    Start here

    AI spend audit: find out exactly where your LLM budget goes

    A two-week, fixed-fee forensic review of your AI bill. We instrument your calls, attribute every pound to a feature and an outcome, and return a ranked list of savings with the expected quality impact of each one.

    Outcome: A savings plan you can act on in the same quarter, with the risky moves clearly separated from the free ones.

    Talk about ai spend audit

    The invoice tells you the total. It does not tell you the truth.

    Almost every team we speak to can read their AI bill and almost none can explain it. The provider dashboard groups spend by model and by day. Your business runs on features, customers and outcomes. Nothing in the default reporting connects the two, so the conversation stalls at “it went up again”.

    The gap matters because the cheapest-looking change is often the most expensive one. Dropping to a smaller model looks like an instant saving per call — right up until the retry rate doubles, support tickets rise, and the same task now takes three calls instead of one. Per-token dashboards will report that as a win.

    What we actually do for two weeks

    Week one — instrument and attribute

    We put a trace on every model call: model, tokens in and out, cached tokens, latency, the feature that made the call, and the request it belonged to. Where you already have observability we read it; where you do not, we add a thin gateway that captures it without touching your application logic.

    Week one — join spend to outcomes

    Separately, your application tells us whether the customer was actually served: the answer was accepted, the lead was captured, the document was produced, the ticket was closed. Two independent writers, one join key. Neither side can quietly flatter the number.

    Week two — find the waste

    With both halves in place, the waste surfaces in hours rather than months: prompts that resend an entire conversation on every turn, system prompts that never change but are never cached, a top-tier model doing classification, silent retries billing twice for one answer, batch-eligible work running at interactive prices.

    Week two — size and rank

    Every finding gets three numbers: what it saves per month at your current volume, how long it takes to ship, and what it might cost you in quality. Then we sort by the ratio, not by the headline.

    What you get on the last day

    A document and a walkthrough, not a dashboard login you will never open. The document opens with the single number — your current cost per successful task, per journey — and then works down the ranked list of moves that change it.

    A large share of what we find is normally free: work that is provably identical in output and simply cheaper to execute. Caching a static system prompt does not change a single token of the answer. Moving overnight enrichment to a batch endpoint does not change the result, only the latency you were not using. We ship those first and pay for the audit out of them wherever we can.

    The rest involves a real trade. Those we do not guess at — we prove them with the evaluation harness described in the routing service, on your traffic, before anything reaches a customer.

    Deliverables

    What you have at the end

    • Cost decomposition: spend by model, feature, customer, prompt and time of day
    • Cost per successful task for your top three user journeys
    • The waste inventory — retries, abandoned sessions, over-long context, duplicated calls, wrong-tier models
    • A ranked savings plan: each move sized in £/month, effort, and quality risk
    • A quality baseline (golden set) so later changes can be proved, not hoped
    • One 90-minute walkthrough with your engineering and finance leads

    Everything on that list lives in your repositories and your cloud account. Ending an engagement does not take the capability with it.

    FAQ

    AI Spend Audit: common questions

    How much access do you need to our systems?

    Less than most teams expect. Read access to billing and model logs, and a short engineering session to add tracing to the call path. We can work entirely inside your cloud account and your network boundary. No customer data needs to leave your environment for the audit itself, and de-identified samples are used for any replay work.

    We already have observability. Is the audit still worth it?

    Often more so, because the raw material is already there. Most observability stacks capture tokens and latency but never the outcome, so they can tell you what a call cost and not whether it worked. Joining those two is usually the fastest part of the engagement when the tracing already exists.

    What if you don't find enough savings to justify the fee?

    Tell us your monthly AI spend on the intro call and we will tell you honestly whether an audit makes sense. Below roughly £3,000 a month the maths rarely works, and we will say so and point you at our free guides instead.

    Do you need to see our prompts?

    Yes, for the parts of the review that concern context size and caching — prompt structure is where a large share of avoidable spend lives. Prompts are treated as confidential under the engagement NDA and are never used to train anything.

    Find out what your AI actually costs per completed task.

    Two weeks, a fixed fee, and a ranked savings plan with the quality risk of every move stated up front. If the numbers say an audit is not worth it for you, we will say so on the first call.