AISTORSApplied AI and cloud engineeringBook a 30-minute call

AISTORSServicesManaged AI and cloud operations

Service 09

Managed operation, so it keeps working.

The same four measures, every month.

For teams with an AI system already live and nobody whose actual job is keeping it right. Most providers stop at handover. Models get retired, accuracy drifts, and the person who understood it moves team. We run it, report against the baseline recorded at build time, and give you a decision each month rather than another dashboard.

Book a 30-minute call
An ordinary two-monitor operations desk in morning daylight, one screen showing a monitoring dashboard and the other a plain table of rows, with an open notebook of hand-drawn diagrams beside the keyboard.
Someone has to be watching on the Tuesday after go-live. Monitoring is only the input; the output is a decision every month and a change every quarter.
01

Who this is for

If the left column is not you, say so on the call and we will tell you what would help instead.

This is for you if

  • You have something live. An agent, an automation or a retrieval system that people now depend on to get through the day.
  • Nobody owns it on a Tuesday. No named person whose job is accuracy, cost and drift once the build team has gone.
  • The model underneath will be retired. It will. Somebody has to track the notice and score the successor before cutover rather than after.
  • You have to report on it. To a regulator, a board or an internal risk function that will ask what changed and why.
  • You want to take it in-house later. That is the intended end state for most clients, and everything is documented for it as we go.

This is not for you yet if

  • Nothing is live yet. Start with the Diagnostic and the build. This is the layer that comes after, and buying it early wastes your money.
  • There is no baseline to report against. We can establish one in an onboarding assessment, but we cannot report against a measure nobody ever recorded.
  • You want round-the-clock cover by default. Not offered by default from a founder-led practice. Where it is genuinely needed it is scoped with named associates.
  • You only want monitoring. Tools do that more cheaply and we will name them. What we sell is the decision that follows the monitoring.
  • The system cannot be instrumented. If we cannot see what it did, we cannot be accountable for it, and we will tell you that before you sign.
02

What we fix, and how

Four things go wrong after go-live. Each has a named owner and a written threshold before we take the system on.

PROBLEM 01

The provider retires the model you are on

What we do. Tracked from the announcement rather than from the outage. The successor is scored against your fixed evaluation set, behavioural differences are reported, thresholds re-established, and cutover scheduled with a rollback path.

PROBLEM 02

Accuracy drifted and nobody noticed

What we do. Drift detection against the baseline recorded at build time, with thresholds agreed in writing and an escalation path that names a person rather than a shared mailbox.

PROBLEM 03

Three months of identical reports

What we do. Monitoring is the input, not the product. The output is a decision every month and a change every quarter. Unchanged reports on a live AI system mean nobody is looking hard enough.

PROBLEM 04

Somebody else built it, so nobody will run it

What we do. Usually we still can, after an onboarding assessment that establishes the measures nobody recorded, finds the gaps in logging and access control, and says plainly whether it can be operated or needs remediation first.

03

What you get

Reported against the baseline recorded in the Diagnostic, so the comparison stays honest.

What we build

  • Drift detection. Input distribution and output quality measured against what was recorded at build, not against a vendor's default.
  • Model upgrade handling. Deprecation notices tracked, successor models scored on your fixed evaluation set, cutover planned with a rollback path.
  • Accuracy scoring on a versioned evaluation set. Real cases labelled by your team, re-run on every change, so a regression is visible before a customer finds it.
  • Escalation framework. Confidence thresholds, the boundary of what the system may decide alone, and a named human on the other side of it.
  • Incident response. A written, tested playbook for AI-specific failure, which 72 percent of organisations in UST's survey do not have.
  • Cost and token monitoring. Cost per completed task tracked against baseline, with the attribution model from the FinOps engagement if you have one.
  • Compliance reporting. Audit trails, model documentation and the evidence a regulator or a customer's procurement team will eventually ask for.
  • Managed cloud and platform operations. Observability, incident response, backup and recovery verification, patching, to your chosen service level.
  • Monthly report and quarterly optimisation review. Same four measures every month. The quarterly session is where things get changed, not just reported.

What this is not

  • Not a 24/7 network operations centre. We are a senior practice, not a follow-the-sun desk. Coverage hours are stated in writing and we will not imply more.
  • Not a licence for monitoring software. Built on your platform's native observability and your own tooling. Nothing of ours has to stay running for the reports to work.
  • Not a guarantee the system will never be wrong. Any provider claiming that is selling something. What is guaranteed is that being wrong is detected, bounded and logged.
  • Not a lock-in. Runbooks, infrastructure as code and access records are written during delivery. You can move this to another provider or in-house without a migration project.
  • Not open-ended. Monthly, cancellable, with the exit handover documented before it is needed rather than after you ask.
04

How it runs

A measure below threshold is a reason to pause the system, and the report says so on the same line.

01

Baseline and thresholds agreed in writing

Task completion rate, human intervention rate, cost per completed task and accuracy against benchmark. Each gets a baseline value and a threshold, and the threshold is a contract term rather than a dashboard setting we can quietly widen later.

Set once · Changed only by agreement

02

Instrumentation and the evaluation set

Logging of inputs, outputs and decisions; cost attribution per feature; and a fixed, versioned set of real cases labelled by your team. That set is the only thing that makes a model change measurable rather than a matter of opinion.

Owned by you · Handed over

03

Continuous monitoring, monthly reporting

Drift, accuracy, intervention rate and cost run continuously with alerting. The report is monthly and always the same four measures in the same format, so twelve months of it is a comparable record rather than twelve different documents.

Monthly · Same format every time

04

Provider change handling

Deprecation notices are tracked as they are published. When a version is retired, the successor is scored against your evaluation set, differences are reported, thresholds are re-established and cutover is scheduled with a rollback path.

On provider announcement

05

Quarterly optimisation review

Where the system is over-escalating, where a cheaper model would score the same, where a threshold is set too loose, and what should be retired. This is the session that changes things rather than reporting them.

Quarterly · Decisions recorded

05

Platform native

Built on your platform's own observability, so nothing of ours has to keep running for the reports to work.

Azure

Azure Monitor and Application Insights, Log Analytics, Azure AI Foundry evaluations and content safety, Entra ID access reviews, Azure Policy.

AWS

CloudWatch and X-Ray, Bedrock model evaluation and invocation logging, Bedrock Guardrails, IAM Access Analyzer, CloudTrail, Config.

Google Cloud

Cloud Monitoring and Logging, Vertex AI Model Monitoring and evaluation, IAM Recommender, Cloud Audit Logs.

DigitalOcean

Monitoring and alert policies, managed database metrics, plus OpenTelemetry and Grafana where the native surface is thinner than the workload needs.

Where an open-source stack is the better answer, we build on OpenTelemetry and Grafana rather than on anything proprietary to us. Flexera's 2026 research notes that managed service providers remain critical for handling complexity while customers retain ownership of governance and cost accountability. That split is the correct one and it is how this engagement is structured.

06

The evidence, if you want it

You do not need these numbers to recognise the problem. They are here because somebody in your approval chain will ask.

150,000

agents or more will be in use at an average global Fortune 500 enterprise by 2028, up from fewer than 15 in 2025. Gartner published six steps for managing the resulting sprawl.

Source: Gartner, 28 April 2026.

28%

of organisations have an incident-response playbook for when an AI system fails. 23 percent conduct adversarial testing.

Source: UST, Enterprise AI at Scale, 2026; global survey of 510 senior leaders.

40%

of organisations deploying AI are forecast to use AI observability to monitor model performance by 2028. Which means most still will not, two years from now.

Source: Gartner, 12 May 2026.

The incidents are already happening. The OneTrust 2026 AI-Ready Governance Report found that 86 percent of organisations had experienced AI-related incidents, that 87 percent encourage AI agent use while only 47 percent have clear governance, oversight and controls in place, and that 28 percent had two or more incidents in the past year in which AI systems or agents took unapproved actions.

07

What actually changes after go-live

Four things move on their own. Each has a detection method, a cadence and a defined response.

Artifact / The four things that change without anyone touching the system 04 categories
What moves, how it is detected, and what happens when it trips
What changes How it is detected Cadence What happens when it trips
Model version Provider deprecation notices tracked against a pinned version register, plus a fixed evaluation set re-run on every provider change. On announcement, and monthly Successor model scored against the same evaluation set before cutover. Thresholds re-established rather than assumed. Migration is planned work, not an outage.
Data distribution Input drift monitoring against the distribution recorded at build, plus accuracy scored on a versioned set of real cases labelled by your team. Continuous, reported monthly Below threshold is a reason to pause the system, and the monthly report says so in the same line rather than in a footnote.
Cost per task Token, request and GPU cost attributed per feature and per team, indexed to the cost per completed task recorded at baseline. Monthly, with anomaly alerts Cost regression is treated as a defect. Gartner forecasts inference cost per agentic workflow rising more than fivefold through 2028, so this is a trend, not a spike.
Permissions and access Credential expiry register, scope review against the actions the system actually performs, and audit logs of every write. Quarterly, and on any scope change Unused permissions are withdrawn. Extension is a decision someone makes and signs, not a default that survives because nobody looked.
Model deprecation is not hypothetical and both examples come from the providers themselves. OpenAI announced the retirement of GPT-4o, GPT-4.1, GPT-4.1 mini and o4-mini from ChatGPT on 13 February 2026. Anthropic's published policy is at least 60 days notice before retiring a publicly released model. A system pinned to a version with no successor plan has an expiry date set by someone else's roadmap.
08

Questions we are actually asked

We did not build the system with you. Can you still run it?

Usually yes, and this is a common way engagements start. It requires an onboarding assessment first, because we cannot report against a baseline that was never recorded. That assessment establishes the current measures, finds the gaps in logging and access control, and tells you honestly whether the system is in a state that can be operated or needs remediation first.

What service levels do you offer?

Stated in writing per engagement, with coverage hours, response commitments and escalation paths named. We deliberately do not offer round-the-clock cover from a senior founder-led practice, because that would be a promise the structure cannot keep. Where you need it, it is scoped with the associate network and the individuals are disclosed to you by name before they touch your systems.

Who is accountable when the AI gets something wrong?

The boundary between what the system may decide alone and what always requires a person is agreed during the Diagnostic and written into the contract as a term. When the system operates inside that boundary and is wrong, the approval gate and the audit log are what contain it. When we have widened that boundary without agreement, that is our failure and it is a contractual matter, not a conversation.

What happens when a model provider retires the version we are on?

It is tracked from the announcement rather than discovered from an error. The successor is scored against your fixed evaluation set, the behavioural differences are reported to you, thresholds are re-established, and cutover is scheduled with a rollback path. Providers publish these dates: OpenAI announced retirements from ChatGPT effective 13 February 2026, and Anthropic commits to at least 60 days notice on publicly released models.

Can we take this in-house later?

That is the intended end state for most clients and it is designed for from the first month. The runbooks, evaluation set, infrastructure as code, monitoring configuration and report format are all yours and all documented as delivery proceeds. Handover is an administrative act rather than a project.

Is this just monitoring with a monthly PDF?

Monitoring is the input. The output is a decision each month and a change each quarter. If a report ever tells you everything is fine three months running with nothing changed, ask us what we are being paid for, because on a live AI system that is a sign nobody is looking hard enough.

09

Next step

Bring a system that is already live.

Thirty minutes, no obligation. Tell us what is running, what you measure today and what you would notice if it degraded. If the honest answer is that you do not need this yet, we will say so.

Onboarding
Assessment first where we did not build the system, so there is a baseline to report against.
Investment
Scoped on the introductory call. Monthly and cancellable.
Exit
Runbooks, evaluation set and configuration handed over documented. No migration project.