Services
AI cost optimisation services, from first measurement to standing practice
Six engagements that fit together in a sequence but are bought one at a time. Everything starts with a measurement, everything is judged on cost per successful task, and everything we build ends up in your repositories.
AI Spend Audit
A two-week, fixed-fee forensic review of your AI bill. We instrument your calls, attribute every pound to a feature and an outcome, and return a ranked list of savings with the expected quality impact of each one.
- Typical duration
- 2 weeks
- From
- £3,500 fixed fee
You get
- Cost decomposition: spend by model, feature, customer, prompt and time of day
- Cost per successful task for your top three user journeys
- The waste inventory — retries, abandoned sessions, over-long context, duplicated calls, wrong-tier models
- A ranked savings plan: each move sized in £/month, effort, and quality risk
LLM Gateway Implementation
The single OpenAI-compatible endpoint every AI call in your estate goes through. It routes, retries, caches, enforces budgets and records what everything cost — deployed in your own cloud, not ours.
- Typical duration
- 3–5 weeks
- From
- From £11,000 fixed scope
You get
- An OpenAI-compatible gateway deployed in your cloud account and network
- Per-project keys, budgets and hard spend caps with alerting before the cap
- Provider failover and retry policy that does not silently double-bill
- Response and prompt caching, with cache-hit reporting
Model Routing & Evaluation
The largest single lever in most AI bills is that one model serves every request. We segment your traffic by task, find the cheapest model that still passes, and prove it on your own traffic before it reaches a customer.
- Typical duration
- 4–6 weeks
- From
- From £9,500 fixed scope
You get
- Traffic segmentation into task classes with volume and cost per class
- A golden evaluation dataset per class, built from your real traffic
- Shadow replay results: candidate models scored on quality, cost and latency
- A versioned routing policy with per-class model assignment and fallbacks
Prompt, Context & Caching Engineering
The cheapest token is the one you never send. We restructure prompts for caching, trim context that adds cost but not accuracy, and move eligible work to batch — changes that are provably output-identical.
- Typical duration
- 2–4 weeks
- From
- From £6,500 fixed scope
You get
- Token profile: where volume comes from, per feature and per turn
- Cache-aware prompt restructuring, with before/after hit rates
- Context and retrieval trimming, validated against your golden set
- Conversation-history strategy for multi-turn products
Azure AI Cost Optimisation
Azure-specific AI cost work: choosing between provisioned throughput and pay-as-you-go, getting deployment types and regions right, and wiring model spend into the Azure cost tooling your finance team already uses.
- Typical duration
- 3–4 weeks
- From
- From £8,500 fixed scope
You get
- Deployment-type review: standard, global, data-zone and provisioned, per workload
- Provisioned throughput versus pay-as-you-go modelling against your real traffic curve
- Region and quota strategy, including capacity and failover posture
- Model spend joined to Azure Cost Management and your existing tagging scheme
AI FinOps Retainer
Cost reductions decay. Prompts grow, traffic shifts, models get deprecated and new features ship without a budget. The retainer keeps the measurement running and the backlog of savings moving.
- Typical duration
- Rolling, 3-month minimum
- From
- From £2,400 per month
You get
- Monthly unit economics pack: cost per successful task by feature, and the trend
- Budget and anomaly alerting, tuned so the alerts stay worth reading
- Golden set and evaluation gate maintenance as prompts and models change
- Model deprecation and migration handling before the provider forces it
Sequence
How the six fit together
You do not need all of them, and most clients never buy all of them. This is the order that wastes the least effort.
- 01
First
AI spend audit
Establishes the baseline, the attribution and the golden datasets that every later phase depends on. Two weeks, fixed fee, credited against what follows.
- 02
Usually next
Prompt, context and caching engineering
The output-identical savings. Ships fastest, needs the least sign-off, and frequently pays for the audit and part of the next phase.
- 03
In parallel, if the estate is fragmented
LLM gateway implementation
Needed once more than one team calls a provider directly. It is what makes routing a config change and cost attribution a property of the system rather than a project.
- 04
Then
Model routing and evaluation
The largest single number and all of the quality risk. Comes after the golden sets exist and after the free savings have banked.
- 05
Where relevant
Azure AI cost optimisation
Deployment types, provisioned throughput versus consumption, quota and region strategy. Sequenced after volume work, because committing to capacity locks in current inefficiency.
- 06
Ongoing
AI FinOps retainer
Keeps the measurement honest and the backlog moving. Optional, and it has to keep earning renewal — everything it maintains already belongs to you.
Frequently asked questions
Which service should we start with?
The AI spend audit, in almost every case. Without a baseline, an attribution model and a golden dataset, every other engagement spends its first fortnight building those anyway — at a worse rate than the fixed-fee audit charges. The audit is also credited in full against implementation work started within 60 days, so starting there costs nothing if you continue.
Do we have to buy the whole programme?
No. Implementation is phased and each phase is approved on its own. You can stop after any phase and keep everything built so far, because it all lives in your repositories and your cloud account rather than ours.
How quickly can you start?
Audits are usually scheduled within two to three weeks of the first call, and the audit itself is two weeks. Implementation phases follow directly from its findings, so the first output-identical savings often ship inside the first month of implementation.
Do you work outside the UK?
Yes. We are UK-based and most of our clients are in the UK and Europe, which suits the working hours and the data-residency conversations. The work itself is remote and the deliverables are the same wherever you are.
Find out what your AI actually costs per completed task.
Two weeks, a fixed fee, and a ranked savings plan with the quality risk of every move stated up front. If the numbers say an audit is not worth it for you, we will say so on the first call.