The bill has no owner
What we do. A tagging standard, allocation model and showback built on your own platform-native tooling, so every line has a team against it before anybody is asked to cut anything.
AISTORSServicesCloud cost and FinOps
Service 06Rightsize first. Commit second.
For engineering and finance leaders who cannot say which team caused the cloud bill. We give every line an owner, attribute AI token and GPU spend to the team that caused it, and change nothing without the person who owns that spend agreeing to it first.
Book a 30-minute call
How savings get agreed
Nothing is switched off until the person who owns that spend has agreed to it. That is why the reduction is still there three months later.
If the left column is not you, say so on the call and we will tell you what would help instead.
For a solo operator or a small business the whole engagement compresses into a single day, because the estate fits in one account and the answer is usually three changes rather than thirty.
Four things go wrong with cost programmes. Each has a specific answer, and none of them starts by switching something off.
What we do. A tagging standard, allocation model and showback built on your own platform-native tooling, so every line has a team against it before anybody is asked to cut anything.
What we do. Token consumption, request volume and GPU utilisation attributed to the team and the feature that caused them. Most tools report the bill well and attribute inference badly, and that gap is the work.
What we do. Nothing is switched off without the owner's agreement and a rollback path recorded first. Rightsizing comes before commitment, so you never commit three years of budget to waste.
What we do. A measured baseline first, then a ranked and costed waste report drawn from your own billing data, then the same four measures reported every month against that baseline.
Rightsizing and idle removal come before commitments, because a commitment locks waste in for a year.
Read-only until a change is approved. Nothing is switched off without a named owner agreeing.
Billing export, cost and usage data, tagging coverage and commitment position. We record the current run rate and the unit economics before proposing anything, so every later claim has a number to be compared against.
Each finding carries the estimated saving, the effort, the blast radius and the owner who has to approve it. Findings you decline stay on the register with the reason recorded, because next quarter someone will ask.
Executed against approved findings with a rollback path for each. This precedes commitment purchase deliberately, because buying a commitment against an oversized estate locks the waste in for the term.
Tagging and allocation land first so the estate is legible, then AI token and GPU attribution, then commitments sized against the cleaned estate. Cost gates go into the deployment pipeline so the next architecture is priced before it ships.
Run rate against baseline, unit cost per workload, commitment utilisation and waste percentage. The same format every month so the record stays comparable after we leave, and so your team can run it without us.
Delivered natively on whichever platform you are already on. All four, no reseller relationship on any.
Azure
Cost Management and Billing, reservations and savings plans, Azure Policy for tagging enforcement, AKS cost analysis, Azure OpenAI token metering.
AWS
Cost Explorer and CUR, Savings Plans and Reserved Instances, Compute Optimizer, Budgets and anomaly detection, Bedrock invocation logging, EKS split-cost allocation.
Google Cloud
Billing export to BigQuery, committed use discounts, Recommender, GKE cost allocation, Vertex AI request and token accounting.
DigitalOcean
Billing and usage reporting, project-level grouping, Droplet and DOKS rightsizing, bandwidth overage review against included transfer.
Where FinOps tooling is worth buying, we will name it and you will buy it directly. We take no commission and hold no reseller agreement, on any platform or any tool.
You do not need these numbers to recognise the problem. They are here because somebody in your approval chain will ask.
29%
of cloud spend is wasted, the first increase in five years, attributed to surging cloud-based AI workloads. Managing cloud spend remains a top challenge for 85 percent of respondents.
Source: Flexera, 2026 State of the Cloud Report; survey of more than 750 cloud decision-makers.
98%
of FinOps practitioners now manage AI spend, up from 31 percent two years earlier. It has gone from a side concern to the default.
Source: FinOps Foundation, State of FinOps 2026; 1,192 respondents representing more than 83 billion dollars in annual cloud spend.
5×
is the increase Gartner predicts in AI inference cost per agentic workflow through 2028. This is a structural cost trend, not a billing spike.
Source: Gartner, 17 August 2026.
The capability gap is documented, and it is precisely this. The FinOps Foundation asked practitioners which tooling capabilities they want but cannot buy. The top answer was granular monitoring of AI spend across tokens, LLM requests and GPU utilisation, followed by costing an architecture before deployment and a single view across all technology spend.
FinOps stopped being a cloud-bill discipline. The scope doubled in twelve months.
We will not give you a percentage before seeing the estate, because the honest answer depends entirely on tagging coverage, commitment position and how much of your spend is AI inference. What we will do is give you a ranked, costed waste report after the assessment, so the number comes from your own billing data rather than from an average.
Most tools are good at showing you the bill and poor at attributing AI inference. That is not a marketing claim, it is what practitioners reported to the FinOps Foundation: granular monitoring of AI spend across tokens, LLM requests and GPU utilisation is the number one capability they want and cannot buy. If your tool already does this well, you need engineering time to act on it, not another dashboard.
Not by default. Percentage-of-savings pricing rewards the volume of savings identified rather than the ones you can safely act on, and it puts our incentive behind aggressive changes to a production estate. Where an outcome-linked structure genuinely fits, it is agreed against the measured baseline recorded in the assessment and written into the contract.
Possibly not on compute, and we will say so. The Foundation's own research records mature practitioners reaching around 97 percent optimisation with the remainder intentionally unactioned. The work that is usually still open on a mature estate is AI spend attribution and pre-deployment cost design, because neither existed in the playbook the estate was optimised against.
The assessment, yes, entirely. It needs billing and usage data plus read-only inventory access. Execution needs change access, granted per action after the finding has been approved, with a rollback path recorded before anything is altered.
The tagging standard, the allocation model, the cost gates and the monthly report format are all yours, documented, and built with your own platform's native tooling rather than ours. Exit is designed in from the start, so ending the engagement is an administrative act rather than a migration.
Thirty minutes, no obligation. Bring a recent bill and the shape of your estate. You will get a straight answer on whether there is enough recoverable spend here to be worth an engagement, and if there is not, we will say so on the call.
Or write to [email protected].
Managed AI and cloud operations
The same four measures every month, plus drift detection and model upgrade handling. Where cost control becomes permanent rather than a project.
Service 01AI agents and assistants
Inside the system, not beside it. The work that creates the inference spend this page attributes.
DisclosureAI capability and data handling
Which model providers, on what terms, with training disabled and token cost instrumented per run.