Skip to content

    Services

    AI cost optimisation services, from first measurement to standing practice

    Six engagements that fit together in a sequence but are bought one at a time. Everything starts with a measurement, everything is judged on cost per successful task, and everything we build ends up in your repositories.

    Start here

    AI Spend Audit

    A two-week, fixed-fee forensic review of your AI bill. We instrument your calls, attribute every pound to a feature and an outcome, and return a ranked list of savings with the expected quality impact of each one.

    Typical duration
    2 weeks
    From
    £3,500 fixed fee

    You get

    • Cost decomposition: spend by model, feature, customer, prompt and time of day
    • Cost per successful task for your top three user journeys
    • The waste inventory — retries, abandoned sessions, over-long context, duplicated calls, wrong-tier models
    • A ranked savings plan: each move sized in £/month, effort, and quality risk
    AI Spend Audit in detail
    Foundation

    LLM Gateway Implementation

    The single OpenAI-compatible endpoint every AI call in your estate goes through. It routes, retries, caches, enforces budgets and records what everything cost — deployed in your own cloud, not ours.

    Typical duration
    3–5 weeks
    From
    From £11,000 fixed scope

    You get

    • An OpenAI-compatible gateway deployed in your cloud account and network
    • Per-project keys, budgets and hard spend caps with alerting before the cap
    • Provider failover and retry policy that does not silently double-bill
    • Response and prompt caching, with cache-hit reporting
    LLM Gateway Implementation in detail
    The big lever

    Model Routing & Evaluation

    The largest single lever in most AI bills is that one model serves every request. We segment your traffic by task, find the cheapest model that still passes, and prove it on your own traffic before it reaches a customer.

    Typical duration
    4–6 weeks
    From
    From £9,500 fixed scope

    You get

    • Traffic segmentation into task classes with volume and cost per class
    • A golden evaluation dataset per class, built from your real traffic
    • Shadow replay results: candidate models scored on quality, cost and latency
    • A versioned routing policy with per-class model assignment and fallbacks
    Model Routing & Evaluation in detail
    Free money first

    Prompt, Context & Caching Engineering

    The cheapest token is the one you never send. We restructure prompts for caching, trim context that adds cost but not accuracy, and move eligible work to batch — changes that are provably output-identical.

    Typical duration
    2–4 weeks
    From
    From £6,500 fixed scope

    You get

    • Token profile: where volume comes from, per feature and per turn
    • Cache-aware prompt restructuring, with before/after hit rates
    • Context and retrieval trimming, validated against your golden set
    • Conversation-history strategy for multi-turn products
    Prompt, Context & Caching Engineering in detail
    Platform specialists

    Azure AI Cost Optimisation

    Azure-specific AI cost work: choosing between provisioned throughput and pay-as-you-go, getting deployment types and regions right, and wiring model spend into the Azure cost tooling your finance team already uses.

    Typical duration
    3–4 weeks
    From
    From £8,500 fixed scope

    You get

    • Deployment-type review: standard, global, data-zone and provisioned, per workload
    • Provisioned throughput versus pay-as-you-go modelling against your real traffic curve
    • Region and quota strategy, including capacity and failover posture
    • Model spend joined to Azure Cost Management and your existing tagging scheme
    Azure AI Cost Optimisation in detail
    Keep it saved

    AI FinOps Retainer

    Cost reductions decay. Prompts grow, traffic shifts, models get deprecated and new features ship without a budget. The retainer keeps the measurement running and the backlog of savings moving.

    Typical duration
    Rolling, 3-month minimum
    From
    From £2,400 per month

    You get

    • Monthly unit economics pack: cost per successful task by feature, and the trend
    • Budget and anomaly alerting, tuned so the alerts stay worth reading
    • Golden set and evaluation gate maintenance as prompts and models change
    • Model deprecation and migration handling before the provider forces it
    AI FinOps Retainer in detail

    Sequence

    How the six fit together

    You do not need all of them, and most clients never buy all of them. This is the order that wastes the least effort.

    1. 01

      First

      AI spend audit

      Establishes the baseline, the attribution and the golden datasets that every later phase depends on. Two weeks, fixed fee, credited against what follows.

    2. 02

      Usually next

      Prompt, context and caching engineering

      The output-identical savings. Ships fastest, needs the least sign-off, and frequently pays for the audit and part of the next phase.

    3. 03

      In parallel, if the estate is fragmented

      LLM gateway implementation

      Needed once more than one team calls a provider directly. It is what makes routing a config change and cost attribution a property of the system rather than a project.

    4. 04

      Then

      Model routing and evaluation

      The largest single number and all of the quality risk. Comes after the golden sets exist and after the free savings have banked.

    5. 05

      Where relevant

      Azure AI cost optimisation

      Deployment types, provisioned throughput versus consumption, quota and region strategy. Sequenced after volume work, because committing to capacity locks in current inefficiency.

    6. 06

      Ongoing

      AI FinOps retainer

      Keeps the measurement honest and the backlog moving. Optional, and it has to keep earning renewal — everything it maintains already belongs to you.

    Frequently asked questions

    Which service should we start with?

    The AI spend audit, in almost every case. Without a baseline, an attribution model and a golden dataset, every other engagement spends its first fortnight building those anyway — at a worse rate than the fixed-fee audit charges. The audit is also credited in full against implementation work started within 60 days, so starting there costs nothing if you continue.

    Do we have to buy the whole programme?

    No. Implementation is phased and each phase is approved on its own. You can stop after any phase and keep everything built so far, because it all lives in your repositories and your cloud account rather than ours.

    How quickly can you start?

    Audits are usually scheduled within two to three weeks of the first call, and the audit itself is two weeks. Implementation phases follow directly from its findings, so the first output-identical savings often ship inside the first month of implementation.

    Do you work outside the UK?

    Yes. We are UK-based and most of our clients are in the UK and Europe, which suits the working hours and the data-residency conversations. The work itself is remote and the deliverables are the same wherever you are.

    Find out what your AI actually costs per completed task.

    Two weeks, a fixed fee, and a ranked savings plan with the quality risk of every move stated up front. If the numbers say an audit is not worth it for you, we will say so on the first call.