AISTORSApplied AI and cloud engineeringBook a 30-minute call

AISTORSServicesCloud cost and FinOps

Service 06

Cloud cost reduction that holds.

Rightsize first. Commit second.

For engineering and finance leaders who cannot say which team caused the cloud bill. We give every line an owner, attribute AI token and GPU spend to the team that caused it, and change nothing without the person who owns that spend agreeing to it first.

Book a 30-minute call
Two colleagues in a bright meeting room reviewing a cloud cost dashboard on a large monitor, one seated and pointing at a usage chart, the other standing with an open notebook.

How savings get agreed

Nothing is switched off until the person who owns that spend has agreed to it. That is why the reduction is still there three months later.

01

Who this is for

If the left column is not you, say so on the call and we will tell you what would help instead.

This is for you if

  • The bill grew faster than the usage. Spend is up and nobody can point at the change that caused it, because there is no baseline to compare against.
  • No line has a team against it. Tagging coverage is partial and showback is a spreadsheet somebody rebuilds by hand each month.
  • AI inference is now a visible line. Token and GPU spend is climbing and your tooling reports it as one undifferentiated number nobody owns.
  • You are being asked to commit. Savings plans or reserved instances are on the table and you want rightsizing done before you lock in three years.
  • This sits with you in engineering. In the FinOps Foundation's 2026 research, 78 percent of practices report into the CTO or CIO organisation and 8 percent to the CFO. The work is scoped for that reader.

This is not for you yet if

  • Compute is genuinely already optimised. FinOps Foundation research records mature practices reaching around 97 percent optimisation, with the rest intentionally left alone. What is usually still open is AI attribution and cost design before deployment.
  • You want a percentage of savings deal. Not by default. That model rewards the volume of savings identified rather than the ones you can safely act on.
  • Nobody will grant billing access. The assessment needs billing and usage data plus read-only inventory. Without it there is nothing honest to find.
  • You only want a tool. Tools report the bill well. If nobody will act on the findings, the tool is simply the cheaper disappointment.
  • The monthly run rate is small. Below a certain level the engagement costs more than it recovers, and we would rather say so than take the work.

For a solo operator or a small business the whole engagement compresses into a single day, because the estate fits in one account and the answer is usually three changes rather than thirty.

02

What we fix, and how

Four things go wrong with cost programmes. Each has a specific answer, and none of them starts by switching something off.

PROBLEM 01

The bill has no owner

What we do. A tagging standard, allocation model and showback built on your own platform-native tooling, so every line has a team against it before anybody is asked to cut anything.

PROBLEM 02

AI spend is invisible

What we do. Token consumption, request volume and GPU utilisation attributed to the team and the feature that caused them. Most tools report the bill well and attribute inference badly, and that gap is the work.

PROBLEM 03

The last optimisation round did not stick

What we do. Nothing is switched off without the owner's agreement and a rollback path recorded first. Rightsizing comes before commitment, so you never commit three years of budget to waste.

PROBLEM 04

You cannot tell whether it worked

What we do. A measured baseline first, then a ranked and costed waste report drawn from your own billing data, then the same four measures reported every month against that baseline.

03

What you get

Rightsizing and idle removal come before commitments, because a commitment locks waste in for a year.

What we build

  • Cost assessment and waste report. Every account, subscription and project, with the findings ranked by recoverable amount and effort, not by how easy they are to describe.
  • Tagging, allocation, showback and chargeback. Untagged spend is the reason nobody owns the bill. This is the unglamorous work that makes every later number arguable.
  • Rightsizing and idle resource removal. Done first, and deliberately before any commitment purchase.
  • Commitment coverage. Savings plans, reserved instances and committed use discounts, sized against a cleaned estate rather than the current one.
  • Storage lifecycle and tiering. Including the orphaned snapshots and unattached volumes that survive every reorganisation.
  • Kubernetes cost control. Namespace and workload attribution, requests and limits reviewed against actual use.
  • AI token and GPU spend attribution. Inference cost broken down by model, by feature and by the team that triggered it, so the charge lands where the decision was made.
  • Cost gates in CI/CD. An architecture priced before it ships, which is the second capability practitioners told the FinOps Foundation they cannot buy.
  • Monthly FinOps review. Run rate against baseline, actions taken, actions declined and why, in the same format every month.

What this is not

  • Not a software licence. No platform of ours sits in your billing path. If a tool is worth buying we will say so and you will buy it directly.
  • Not a percentage-of-savings arrangement by default. That model rewards finding savings, not the ones you can safely act on.
  • Not a reseller relationship. We hold no reseller agreement on any of the four platforms, so there is no margin behind a recommendation to move or to commit.
  • Not a one-off report. A waste report with nobody accountable for acting on it regenerates the same waste inside two quarters.
  • Not unlimited. Mature estates hit diminishing returns. Practitioners in the Foundation's own research describe reaching around 97 percent optimisation with the remainder deliberately left alone. We will tell you when you are close to that line.
04

How it runs

Read-only until a change is approved. Nothing is switched off without a named owner agreeing.

01

Read-only access and a measured baseline

Billing export, cost and usage data, tagging coverage and commitment position. We record the current run rate and the unit economics before proposing anything, so every later claim has a number to be compared against.

Typically 3 to 5 days · No production change

02

Waste report, ranked by recoverable amount

Each finding carries the estimated saving, the effort, the blast radius and the owner who has to approve it. Findings you decline stay on the register with the reason recorded, because next quarter someone will ask.

Written deliverable · Yours to keep

03

Rightsizing and idle removal, in that order

Executed against approved findings with a rollback path for each. This precedes commitment purchase deliberately, because buying a commitment against an oversized estate locks the waste in for the term.

Approval gate on every consequential change

04

Attribution, then commitment coverage

Tagging and allocation land first so the estate is legible, then AI token and GPU attribution, then commitments sized against the cleaned estate. Cost gates go into the deployment pipeline so the next architecture is priced before it ships.

Coverage sized on cleaned run rate

05

Monthly review, same four measures

Run rate against baseline, unit cost per workload, commitment utilisation and waste percentage. The same format every month so the record stays comparable after we leave, and so your team can run it without us.

Ongoing · Cancellable

05

Platform native

Delivered natively on whichever platform you are already on. All four, no reseller relationship on any.

Azure

Cost Management and Billing, reservations and savings plans, Azure Policy for tagging enforcement, AKS cost analysis, Azure OpenAI token metering.

AWS

Cost Explorer and CUR, Savings Plans and Reserved Instances, Compute Optimizer, Budgets and anomaly detection, Bedrock invocation logging, EKS split-cost allocation.

Google Cloud

Billing export to BigQuery, committed use discounts, Recommender, GKE cost allocation, Vertex AI request and token accounting.

DigitalOcean

Billing and usage reporting, project-level grouping, Droplet and DOKS rightsizing, bandwidth overage review against included transfer.

Where FinOps tooling is worth buying, we will name it and you will buy it directly. We take no commission and hold no reseller agreement, on any platform or any tool.

06

The evidence, if you want it

You do not need these numbers to recognise the problem. They are here because somebody in your approval chain will ask.

29%

of cloud spend is wasted, the first increase in five years, attributed to surging cloud-based AI workloads. Managing cloud spend remains a top challenge for 85 percent of respondents.

Source: Flexera, 2026 State of the Cloud Report; survey of more than 750 cloud decision-makers.

98%

of FinOps practitioners now manage AI spend, up from 31 percent two years earlier. It has gone from a side concern to the default.

Source: FinOps Foundation, State of FinOps 2026; 1,192 respondents representing more than 83 billion dollars in annual cloud spend.

5×

is the increase Gartner predicts in AI inference cost per agentic workflow through 2028. This is a structural cost trend, not a billing spike.

Source: Gartner, 17 August 2026.

The capability gap is documented, and it is precisely this. The FinOps Foundation asked practitioners which tooling capabilities they want but cannot buy. The top answer was granular monitoring of AI spend across tokens, LLM requests and GPU utilisation, followed by costing an architecture before deployment and a single view across all technology spend.

07

What changed in one year

FinOps stopped being a cloud-bill discipline. The scope doubled in twelve months.

Artifact / What FinOps practices now manage Share of practitioners, 2025 to 2026
AI spend rose to 98 percent of practitioners in 2026 from 31 percent two years earlier. SaaS rose from 65 to 90 percent. Licensing rose from 49 to 64 percent. Private cloud rose from 39 to 57 percent. Data centre rose from 36 to 48 percent. Labour costs are newly tracked by 28 percent. CATEGORY SHARE OF FINOPS PRACTITIONERS MANAGING THIS SPEND 25% 50% 75% 100% AI spend SaaS Licensing Private cloud Data centre Labour costs 98% 90% 64% 57% 48% 28% 2026 SHARE 2025 SHARE. AI SPEND COMPARES TO 31% TWO YEARS EARLIER.
Source: FinOps Foundation, State of FinOps 2026, and the Foundation's own account of why it changed its mission from managing the value of cloud to managing the value of technology. Labour cost tracking is newly reported and has no prior-year comparison. The practical consequence for a buyer: a FinOps engagement scoped only around public cloud IaaS now covers less than half the spend the discipline is expected to govern.
08

Questions we are actually asked

How much can we expect to save?

We will not give you a percentage before seeing the estate, because the honest answer depends entirely on tagging coverage, commitment position and how much of your spend is AI inference. What we will do is give you a ranked, costed waste report after the assessment, so the number comes from your own billing data rather than from an average.

We already have a FinOps tool. Why would we need this?

Most tools are good at showing you the bill and poor at attributing AI inference. That is not a marketing claim, it is what practitioners reported to the FinOps Foundation: granular monitoring of AI spend across tokens, LLM requests and GPU utilisation is the number one capability they want and cannot buy. If your tool already does this well, you need engineering time to act on it, not another dashboard.

Will you take a percentage of the savings?

Not by default. Percentage-of-savings pricing rewards the volume of savings identified rather than the ones you can safely act on, and it puts our incentive behind aggressive changes to a production estate. Where an outcome-linked structure genuinely fits, it is agreed against the measured baseline recorded in the assessment and written into the contract.

Our estate has already been optimised. Is there anything left?

Possibly not on compute, and we will say so. The Foundation's own research records mature practitioners reaching around 97 percent optimisation with the remainder intentionally unactioned. The work that is usually still open on a mature estate is AI spend attribution and pre-deployment cost design, because neither existed in the playbook the estate was optimised against.

Can you do this without production access?

The assessment, yes, entirely. It needs billing and usage data plus read-only inventory access. Execution needs change access, granted per action after the finding has been approved, with a rollback path recorded before anything is altered.

What happens to the work when the engagement ends?

The tagging standard, the allocation model, the cost gates and the monthly report format are all yours, documented, and built with your own platform's native tooling rather than ours. Exit is designed in from the start, so ending the engagement is an administrative act rather than a migration.

09

Next step

Bring one month of billing data.

Thirty minutes, no obligation. Bring a recent bill and the shape of your estate. You will get a straight answer on whether there is enough recoverable spend here to be worth an engagement, and if there is not, we will say so on the call.

Assessment duration
Three to five days. One day for a solo operator or small business.
Investment
Scoped on the introductory call, and credited in full against the engagement that follows.
If you stop after the assessment
You keep the waste report, the tagging standard and the ranked findings. No further commitment.