Call
You describe the process. We say whether it is a measurement problem, an integration problem or neither, and what the Diagnostic would look at. If it is not worth measuring we say that on the call.
AI, cloud and research brought together for the systems that move your organisation forward.
AISTORS designs, engineers and operates intelligent systems for real-world work. From AI automation and data to cloud foundations, governance and ongoing operations, we work inside your environment to turn complex technology into capable, dependable systems.
Book a 30-minute call
Enterprise AI adoption is accelerating faster than organisations’ ability to realise, govern and measure value.
0%
of companies reported minimal revenue and cost gains from AI despite substantial investment. Only 5 percent were achieving AI value at scale.
Source: BCG, The Widening AI Value Gap, September 2025; study of more than 1,250 firms worldwide.
0 in 3
of enterprises expect providers to build and operationalise their priority use cases. One third are already scaling agentic deployments.
Source: BCG, The $200 Billion AI Opportunity in Tech Services, 2026.
0 in 5
of cloud-based workloads and data were reported as repatriated by respondents, at 21 percent, even as public cloud adoption kept rising.
Source: Flexera, 2026 State of the Cloud Report.
0%
of 510 senior leaders in UST’s 2026 global survey said they had incident-response playbooks for AI failures. 23 percent said they conducted adversarial testing.
Source: UST, Enterprise AI at Scale, 2026.
Eight problems, every figure attributed, and not one of them solved by a better model.
The failure is rarely the model. BCG puts only 10 percent of AI value in the algorithm. The rest sits in data, technology and the operating model around it, and the same eight problems recur there whatever the sector.
That gap is expensive in a specific way. Teams cannot tell a working system from a demonstration, cannot defend the spend at renewal, and cannot decide what to scale. The second project is then approved on the same evidence as the first, which is none.
We start with the measurement. Cycle time, error rate, volume and hours on the target process, recorded before anything is built, in your systems, with the method written down so your team can repeat it after we leave.
A pilot is approved on a slide, shipped into a process nobody timed, and then judged by anecdote. When the invoice arrives there is no number to compare it against, so the honest answer to "did it help" is that nobody knows.
Evidence: BCG, Are You Generating Value from AI? The Widening Gap, 2025. 60 percent report minimal or no material value.
Data quality is the single most-cited barrier, ahead of budget, models and talent. It is also the least glamorous, which is why it stays unfixed while the pilot count goes up.
Evidence: UST, Enterprise AI at Scale, 2026. 44 percent name data quality the number one barrier, while 90 percent are piloting or scaling. PYMNTS Intelligence, 2026, puts it at 63 percent of executives. RSM US Middle Market AI Survey, 2026, at 34 percent.
The model can reason. It cannot reach the record. Where data is batch, siloed or unowned, an agent is a demonstration with a login.
Evidence: Confluent, 2026 Data Streaming Report. 72 percent of IT leaders cite insufficient infrastructure for real-time data processing, up from 61 percent the year before. 65 percent cite fragmented ownership of data.
A system that cannot read the order, write the note or close the ticket has not automated anything. Most of the work in a real deployment is plumbing into systems that were never built to be plumbed into.
Evidence: RSM US Middle Market AI Survey, 2026. Legacy systems integration cited by 28 percent, level with the talent gap. Integration readiness is assessed before the build because a model that cannot read the order, write the note or close the ticket has not automated the work.
Agents are being given permissions faster than anyone is writing down what happens when they misuse them. The gap is not model safety. It is that no one has drafted the runbook.
Evidence: UST, Enterprise AI at Scale, 2026, survey of 510 senior leaders. 28 percent have an AI incident-response playbook, 23 percent run adversarial testing. OneTrust 2026 AI-Ready Governance Report: 86 percent of organisations experienced AI-related incidents, 87 percent encourage agent use but only 47 percent have clear governance, oversight and controls in place.
Inference spend lands in one budget and is caused by another. Meanwhile the first vendor choice quietly becomes permanent, because nobody wrote down how to reverse it.
Evidence: Flexera, 2026 State of the Cloud Report. Wasted cloud spend rose to 29 percent, its first increase in five years, alongside 81 percent generative AI usage. IBM, The Calculus of AI Sovereignty, 2026: 71 percent say switching their primary AI vendor or model would be difficult.
Staff adopt tools faster than any policy can cover them, and company data leaves with them. An unapproved tool is not a training problem. It is an unlogged data path out of the business, and it does not appear in any architecture diagram.
Evidence: Verizon 2026 Data Breach Investigations Report: employee use of unapproved shadow AI tripled to 45 percent. PagerDuty, 2026: 66 percent of office professionals have used unauthorised AI tools at work. Deloitte UK: 31 percent of generative AI users do so without their employer knowing, and 46 percent use free-to-use tools at work.
Providers retire versions on their timetable, not yours. When the version a working system was built and approved on is withdrawn, every prompt, threshold and evaluation calibrated against it has to be re-established on a successor that behaves differently.
Evidence: OpenAI announced the retirement of GPT-4o, GPT-4.1, GPT-4.1 mini and o4-mini from ChatGPT on 13 February 2026. Anthropic's published policy is at least 60 days notice before retiring a publicly released model. Both are primary provider sources.
REVIEWED 17 SEPTEMBER 2026. EIGHT PROBLEMS, EACH WITH A NAMED AND DATED SOURCE, AND SURVEY SCOPE STATED WHERE IT CHANGES HOW A FIGURE SHOULD BE READ. THE LIST IS RE-DATED ANNUALLY. WHEN ONE OF THESE STOPS BEING TRUE WE WILL MARK IT RESOLVED AND SAY WHAT REPLACED IT, RATHER THAN DELETE IT QUIETLY.
We build it. We run it. We prove it.

Each stop produces something you keep: a written scope, a measured report, a working system inside your own environment, a monthly record of how it is behaving. Nothing here requires the next stage to have been worth doing.
You describe the process. We say whether it is a measurement problem, an integration problem or neither, and what the Diagnostic would look at. If it is not worth measuring we say that on the call.
A measured baseline of the target process, three prioritised opportunities, integration and data readiness findings, and an implementation plan with expected outcome ranges. Investment is credited against the build that follows.
Built against the plan, in your tenancy, on your accounts, with approval gates on consequential actions. Fixed scope agreed in writing before work starts. Changes are re-scoped rather than absorbed quietly.
Drift detection, model upgrade handling, escalation paths and a monthly report against the baseline. The same four measures every month, so the record stays comparable.
If the honest answer is that AI won't pay for itself here, we tell you that and you keep the report.
You keep the baseline, the plan and the reasoning, whichever way the recommendation goes.

Cycle time, error rate, volume and hours, recorded in your systems over a defined window, with the counting method written down.
Each scored on value and feasibility, with the reasoning shown, so the ranking can be argued with rather than accepted.
Which systems hold the data, which have usable interfaces, where the records disagree, and what has to be fixed before anything is built. Data readiness is the most commonly cited reason AI work stalls, so it is a named verdict here rather than a paragraph in an appendix.
Sequence, scope, dependencies and the range we expect on each measure, stated as a range because a single number would be a guess.
Where the process touches regulated data, where a human decision is required, and which obligations apply under your jurisdiction and sector.
Stated as a recommendation with the reasoning attached, including the case for not building anything. A no-go with a clear reason is a successful Diagnostic, and it is the cheapest possible outcome for you.
What the Diagnostic is not
If a group is not on this list, we do not do it.
Work is delivered inside your systems and your accounts. Nothing on this list requires a platform of ours to sit in the middle, and nothing on it is billed as a licence.
Inside the system, not beside it.
Assistants that run inside real systems rather than in a chat window beside them: intake and qualification, ticket triage, internal question answering over your own records, scheduling, escalation routing, drafting with a human approval step.
The procedure already exists. It just isn't running itself.
Document processing, CRM hygiene, quoting, reconciliation, ticketing. The work that is already written down as a procedure and is being done by hand anyway.
Answers from your own records.
Enterprise search, retrieval over your own documents, pipelines, private and self-hosted models where the data cannot leave your estate.
Only where the history is long enough to test against.
Forecasting, inventory, pricing, churn, fraud, predictive maintenance, computer vision. Scoped only where enough history exists to test against.
Built once, documented, handed over.
Landing zones, migration, Kubernetes, infrastructure as code, observability, site reliability engineering, cloud security.
Rightsize first. Commit second.
Rightsizing, commitment coverage, storage lifecycle, and AI token and GPU attribution so inference spend is charged back to the team that caused it.
Move the few that are cheaper. Leave the rest alone.
Placement analysis, selective repatriation, hybrid design. Moving the small number of workloads that are genuinely cheaper elsewhere, and leaving the rest alone.
Written down before anyone asks for it.
EU AI Act readiness, ISO 42001 readiness, HIPAA, SOC 2, GDPR and DPDP programme support, red teaming of deployed systems.
EU AI Act timetable, checked 17 September 2026. The Digital Omnibus on AI came into force on 27 July 2026 and moved Annex III standalone high-risk obligations to 2 December 2027 and Annex I embedded high-risk obligations to 2 August 2028. Transparency and watermarking obligations for AI-generated content apply from 2 December 2026. Readiness work is scoped against this timetable and re-checked at each review.
The same four measures, every month.
Drift detection, model upgrade handling, escalation, compliance reporting against the baseline recorded in the Diagnostic.
You should not need us in year two.
Runbooks, infrastructure as code handover, prompt and evaluation practice for your team, internal AI usage policy, working sessions with the people who will own the system. Scoped as delivery, not as a course.
The measure is named before the build, and it is the same measure afterwards.

Scheduling and intake, clinical documentation support with a clinician approval step, prior authorisation packaging, revenue cycle exception handling. Every path that touches patient data is designed with the audit record first.
The most common engagement here is a HIPAA and AI compliance retrofit: an organisation deployed an assistant before governance caught up, and now needs the access model, the logging, the vendor terms and the human approval boundary brought into line without switching the system off.
Reconciliation, claims and case triage, know your customer file assembly, fraud review queues, exception handling in the operations team rather than in the model. Decisions that carry a regulatory consequence stay with a named person, and the system's role is to prepare the file, not to sign it.
Model behaviour is recorded against a fixed evaluation set so a change in a provider's model can be detected in reporting rather than in a complaint.


Demand forecasting against real sell-through, replenishment, pricing support, catalogue and product data cleanup, returns and dispute handling, service assistants that can read order state rather than guess it.
The measurement here is usually simple and unforgiving: units available when a customer asks, and hours spent moving data between the store system and the finance system.
Quoting from a rules sheet nobody has updated in three years, order entry from emailed purchase orders and scanned documents, inventory placement across branches, route and load planning support, supplier invoice matching.
This sector tends to have the cleanest baseline available anywhere: the work is already counted in the warehouse management system, so the before and after argument is short.

Sector determines what the process looks like. It rarely determines whether the work is possible. Data readiness does, and that varies more between two companies in the same sector than it does between sectors.
For context on how uneven adoption still is: US Census Bureau Business Trends and Outlook Survey data reported 37 percent of firms with at least 250 employees using AI in their business operations, against 32 percent of firms with 100 to 249 employees, in the collection period ending 3 May 2026. Census and NBER researchers put firm-level use at 18 percent for November 2025 to January 2026, rising to 32 percent on an employment-weighted basis. Adoption is concentrated in larger firms, not evenly spread.
Proposal and bid assembly, timesheet and billing hygiene, research synthesis across your own past engagements, contract and scope review with a human sign-off.
The knowledge is already in the files nobody can search.
Measured onBillable hours recovered from non-billable work.
Route and load planning support, proof of delivery capture, exception handling on late and damaged consignments, freight invoice audit against the rate card.
Measured onExceptions cleared per person per day.
Predictive maintenance where sensor history exists, visual quality inspection, work order triage, supplier document matching.
Scoped only where enough failure history exists to test against.
Measured onUnplanned downtime hours and false-positive rate on inspection.
First notice of loss intake, claims triage and file assembly, policy document comparison, subrogation review. Deeper claims-specific work than the general financial services engagement above.
The system prepares the file. A named person decides.
Measured onCycle time from notification to decision-ready file.
Enrolment and enquiry handling, timetable and resource scheduling, marking support with an instructor approval step, accessibility remediation of existing material.
Measured onAdministrative hours per enrolled student.
Meter and billing exception handling, outage and fault triage, asset inspection from imagery, regulatory reporting assembly.
Measured onException backlog and reporting preparation time.
Case intake and eligibility packaging, records digitisation and search, grant and tender assembly, freedom of information response drafting with a reviewer.
Audit trail is the first requirement, not the last.
Measured onMedian time to first substantive response.
Tender and subcontractor bid comparison, drawing and specification search, site report and snag list capture, variation and claim substantiation.
Measured onHours spent locating information that already exists.
We advise on placement. You choose the platform. We build natively on it either way.
All four platforms can run almost everything described on this page. Treating the choice as a contest between three feature lists produces the wrong answer, because the feature lists converged years ago. So the consultation gives you a recommendation built on the constraints you already live with: where the data sits, what you have already committed to, which accelerators you can actually get, where your identity plane lives, what your regulator will sign off, and who you can hire.
Then you decide, and we build what you decided. If you have standardised on Azure, if your board has mandated AWS, if your data team wants Google Cloud, or if DigitalOcean is the right size for where you are, we deliver fully native on that platform. Four certifications and no reseller relationship on any of them means there is no commission pulling the recommendation one way, and no gap in what we can deliver once you have chosen.
On the framing: Technolynx, AWS vs Azure vs GCP for AI and Data Workloads, 2026, and CIO, Your AI cloud strategy isn't about cost, it's about gravity, 2026. On relative scale: Synergy Research Group put Q4 worldwide cloud infrastructure share at AWS 28 percent, Microsoft 21 percent and Google 14 to 15 percent, roughly two thirds of the market between them.
Single-platform native delivery, on whichever platform you have chosen. A constraints-based recommendation is what a consultation is for. It is not a condition of working with us.
Azure native
Entra ID, Azure AI Foundry, AKS, Bicep, Azure Policy, Microsoft Fabric, Purview governance.
AWS native
IAM Identity Center, Bedrock, SageMaker, EKS, CDK, Control Tower, Well-Architected review.
Google Cloud native
Vertex AI, BigQuery, GKE, Terraform, Dataplex, Organisation Policy, TPU access where justified.
DigitalOcean native
Droplets, DOKS, managed Postgres, Spaces, App Platform, GenAI Platform, predictable billing.
Where a single-platform requirement carries a cost, a limit or a compliance consequence, we put that in writing before the build starts rather than discovering it at renewal. You keep the requirement. You also keep the analysis.
| Constraint you already live with | What it actually decides | Why |
|---|---|---|
| Where the data already sits, and how much of it | Usually keeps the workload on the platform already holding the data | Egress is paid once on the way out and latency is paid forever after. Data gravity beats a feature comparison, and it is the constraint people discover last. |
| Commitments already signed | Narrows the field for the remaining term, whatever the technical answer is | An unused commitment is money you keep paying. Placement works around it until renewal, and renewal is the moment the decision reopens. |
| Where your identity and productivity plane lives | Anchors the governance-heavy workloads | If Entra ID and Microsoft 365 are already the source of truth, Azure removes a whole class of integration and audit work. That is an organisational fact, not a preference. |
| Accelerator availability in your region | Decides where training and heavy inference can physically run | Quota, not the price list, is the binding constraint on GPU and TPU capacity. A region on a map is not the same as capacity you can get this quarter. |
| Residency, sector and contractual obligations | Eliminates regions, and occasionally providers | A compliance team signs off a contract and an audit trail, not a marketing page. This is checked before architecture, not after. |
| Who you already employ, and who you can hire | Breaks ties, and it should | A platform nobody on the team knows becomes a single-person dependency. That is a bigger operational risk than a ten percent list-price gap. |
| Billing predictability at your size | Decides simple against broad | Below a certain size the hyperscaler console and account structure cost more attention than they return. Above it, the managed data services are the reason to pay the tax. |
| Platform | Where it genuinely pulls ahead | Where it costs you |
|---|---|---|
| Azure | Identity, governance and the Microsoft estate. If Entra ID, Microsoft 365 and existing licensing are already in place, a large amount of integration and audit work simply disappears. | The account and licensing model is complex, and the value is weakest if you are not already a Microsoft organisation. |
| AWS | Service breadth and the deepest hiring pool. If a managed service exists anywhere, it usually exists here, and someone you interview will have used it. | Breadth is also the problem: more ways to build the same thing, and list compute prices that generally sit above Google Cloud. |
| Google Cloud | Data and machine learning. BigQuery and TPU access are real differentiators when the workload is analytics-heavy or training-heavy. | The smallest service catalogue of the three, and a narrower pool of engineers who have run it in production. |
| DigitalOcean | Simplicity and predictable billing. A small surface one person can hold in their head, and enough managed database and object storage to ship a real product. | Fewer managed services and fewer regions, support is a paid tier, and bandwidth is charged past the included transfer. Right for a startup, usually wrong for a regulated enterprise. |
The honest weakness
Every platform in the table above has one. We name it before you ask, because the weakness you find in year two is the expensive one.
Integration is where the time goes, so it is scoped before the model is chosen.
These are the systems the work usually has to reach into. A model that cannot read the order, write the note or close the ticket is a demonstration.
CRM and sales
Salesforce
HubSpot
Zoho
Dynamics 365
Finance and ERP
QuickBooks
Xero
Tally
NetSuite
SAP Business One
Odoo
Commerce and point of sale
Shopify
WooCommerce
Magento
Square
Lightspeed
Field service
ServiceTitan
Jobber
Communications
Slack
Microsoft Teams
WhatsApp Business
Twilio
Support and service management
Zendesk
Freshdesk
Intercom
ServiceNow
Data platforms
Snowflake
BigQuery
Databricks
Postgres
Automation and orchestration
n8n
Make
Zapier
Power Automate
Airflow
Model providers
OpenAI
Anthropic
Google Gemini
AWS Bedrock
Azure AI Foundry
Self-hosted
Open weight models in your own tenancy
Where data cannot leave the estate
IF IT HAS AN API WE CAN ALMOST CERTAINLY WORK WITH IT. IF IT DOES NOT, WE BUILD BROWSER-LEVEL AUTOMATION AND TREAT IT AS A MAINTENANCE COMMITMENT RATHER THAN A ONE-OFF, BECAUSE THE VENDOR WILL CHANGE THE SCREEN EVENTUALLY.
Access is the whole security story here. Everything else follows from who holds which key, and for how long.
Procurement asks for this list eventually. Publishing it now saves a round of correspondence, and saying "not held" is cheaper than being found out in a questionnaire.
| Item | Status |
|---|---|
| SOC 2 Type II | Not held |
| ISO 27001 | Not held |
| ISO 42001 | Not held. Readiness work offered |
| Microsoft Azure certification | Held |
| Amazon Web Services certification | Held |
| Google Cloud certification | Held |
| DigitalOcean certification | Held |
| Mutual non-disclosure agreement | Always signed before scoping |
| Professional indemnity insurance | ‹FILL: policy status and limit› |
| Cyber liability insurance | ‹FILL: policy status and limit› |
Access is requested for the specific system and the specific action, not at account level because it is quicker.
Credentials carry an expiry from the day they are issued. Extension is a decision someone makes, not a default.
No shared accounts, on either side. Every action in your systems traces to one person.
Model access runs on enterprise or business terms with training on your data switched off, and the terms are shown to you.
Region is a requirement, not a preference. Where residency or sensitivity demands it, the models run fully self-hosted in your own tenancy.
Every third party in the path is listed before work starts, and you keep a right to refuse any of them.
New integrations begin with read access. Write access is granted per action after the behaviour has been reviewed against real records.
Runbooks, infrastructure as code and credentials handover are written during the build, so ending the engagement is an administrative act.
| Measure | Baseline | Threshold | Latest | How it is counted |
|---|---|---|---|---|
| Task completion rate | 0.0% | 55.0% | 61.4% | Cases closed by the system with no human edit, over all cases entering the queue. |
| Human intervention rate | 100.0% | 45.0% | 38.6% | Cases escalated to a person, whether by confidence threshold, policy rule or reviewer recall. |
| Cost per completed task | 1.00 | 0.60 | 0.41 | Indexed to the baseline. Includes model tokens, compute and the reviewer time still required. |
| Accuracy against benchmark | 94.2% | 97.0% | 98.1% | Scored against a fixed, versioned set of real cases labelled by your team, re-run on every change. |
AI systems make mistakes, and any provider claiming otherwise is selling something. The useful question is not whether a system will be wrong, but what happens in the minute after it is.
Every system ships with human approval gates on consequential actions, confidence thresholds that escalate to a person, full audit logging of inputs and outputs, and a rollback path that has been tested rather than assumed.
The boundary between what the system may decide alone and what always needs a human is agreed during the Diagnostic and written into the contract. It is a term, not a setting we can quietly widen later.
Capacity is published, because a founder-led practice can run out of it.
Deliberately empty. A generated face here would be the first dishonest thing on the page.
There is no account layer and no junior handoff. That is the constraint the practice is built around, and it is also the reason capacity is stated in writing rather than implied.
Parallel work is covered by a vetted associate network, and anyone new is disclosed to you by name before they touch your systems. Runbooks, infrastructure as code and access records are written during delivery rather than at the end, so no single person is a point of failure, including the founder.
Current available capacity: ‹FILL: engagements open this quarter›
| Platform | Certification | Credential ID | Issued / expires |
|---|---|---|---|
| Azure | ‹FILL: exact certification name› | ‹FILL: ID› | ‹FILL: dates› |
| AWS | ‹FILL: exact certification name› | ‹FILL: ID› | ‹FILL: dates› |
| Google Cloud | ‹FILL: exact certification name› | ‹FILL: ID› | ‹FILL: dates› |
| DigitalOcean | ‹FILL: exact certification name› | ‹FILL: ID› | ‹FILL: dates› |
Start with one operation. Find the intelligence worth applying.
Thirty minutes, no obligation. Bring one operation, system or decision that needs to move forward. We will give you a direct view of where applied AI, cloud engineering or research can create the next useful step.
Applied intelligence for what comes next.
Booking link: ‹FILL: scheduling URL›
Or write to [email protected]
