AISTORSApplied AI and cloud engineeringBook a 30-minute call

AISTORSServicesDocument and process automation

Service 02

Automation for the work nobody should be doing.

The procedure already exists. It just isn't running itself.

For finance, operations and back-office teams where documents arrive faster than people can key them in. The procedure is already written down, already followed by hand, and already counted somewhere. We automate the keying and the routing, keep a named person on the exceptions, and report against the count you started with.

Book a 30-minute call
A clean modern back-office document station in daylight, with a flatbed scanner, two trays of printed invoices and forms, and a monitor showing a plain table of rows.

Before anything is built

We count what the work costs you today: how many arrive, how long each takes, how often one comes back. Without that number there is nothing to prove afterwards.

01

Who this is for

If the left column is not you, say so on the call and we will tell you what would help instead.

This is for you if

  • The procedure is already written down. There is a policy, checklist or template covering what happens to each document, even if nobody follows it exactly.
  • Volume runs to hundreds or thousands a month. Enough that keying time is a real cost and accuracy can be scored properly rather than estimated.
  • The documents are broadly consistent. Invoices, claims, purchase orders, statements, forms. Variation is expected; total inconsistency is not.
  • Somebody can label a sample. A person who can mark fifty real documents right or wrong, so accuracy becomes a number instead of an impression.
  • You can live with an exception queue. There will always be one. The measure that matters is whether it shrinks, and that is reported monthly.

This is not for you yet if

  • Every document is genuinely unique. One-off contracts with no repeating structure suit a person with an assistant better than an extraction pipeline.
  • The volume is small. A few dozen a month rarely repays a build. We would rather say so than sell it to you.
  • Nobody owns the exceptions. Automation concentrates the hard cases. Without a named owner for them, the queue simply becomes the new problem.
  • The upstream process is broken. If the document arrives wrong, automating the keying only makes wrong data arrive faster. Fix the procedure first.
  • You need it to be perfect. Any provider quoting an accuracy figure before seeing your documents is guessing. We commit to a measured figure and a threshold.
02

What we fix, and how

Four things go wrong with document automation. Each has a specific answer, agreed before any build starts.

PROBLEM 01

People are keying data that already exists

What we do. We extract the fields, write them into the system of record through its own interface, and route the document. The keying stops rather than moving to a review screen that takes just as long.

PROBLEM 02

The last attempt created a worse queue

What we do. A confidence threshold is set before go-live. Below it, work goes to a named person rather than straight through. The exception queue has an owner from day one and its size is reported every month.

PROBLEM 03

Nobody can tell you the accuracy

What we do. Extraction is scored against a labelled sample of your own documents before it ships. You get the figure per field, and the confidence threshold that follows from it.

PROBLEM 04

The vendor will change the screen and it will break

What we do. Where a real interface exists we use it, because it is cheaper to keep working. Where only the screen exists we say so, price the maintenance honestly, and never present it as a one-off build.

03

What you get

Where an API exists we use it. Screen-level automation is a last resort with a maintenance cost attached.

What we build

  • Document processing. Invoices, contracts, purchase orders, claims and forms, extracted into the system of record with a confidence threshold and a human queue for anything below it.
  • Reconciliation and month-end close. Matching across systems, exception queues that shrink rather than grow, and an audit trail for every automated match.
  • Quoting from the rules that actually apply. Including the rules sheet nobody has updated in three years, which the assessment will find and name.
  • CRM hygiene and enrichment. Deduplication, normalisation and field completion, with the merge rules agreed before anything is merged.
  • Ticket and case handling. Classification, routing, enrichment from history, and drafting for a person to approve.
  • Order entry from unstructured input. Emailed purchase orders, scanned documents and attachments, into the order system with the exceptions surfaced.
  • Reporting and dashboards. Built once against the source of truth, rather than rebuilt monthly in a spreadsheet.
  • Prior authorisation and revenue cycle work. Packaging, submission and exception handling, designed against the CMS deadlines above with the audit record first.

What this is not

  • Not a robotic process automation licence. Where an API exists we use it. Screen-level automation is a last resort and is scoped as a maintenance commitment, because the vendor will change the screen.
  • Not automation of a broken procedure. Automating a process nobody has fixed makes the wrong output arrive faster. Where the procedure is the problem we will say so.
  • Not full autonomy on regulated decisions. The system prepares the file. A named person decides, and that boundary is a contract term.
  • Not scoped without a count. If nobody knows how many of these are processed a week, that is the first thing the Diagnostic establishes.
  • Not a promise of zero exceptions. There will be an exception queue. The measure that matters is whether it shrinks.
04

How it runs

The count comes before the build, because the work is usually already counted somewhere.

01

Count the work first

Volume, cycle time, error rate and hours on the target procedure. This is the easiest line in the catalogue to baseline, because the work is usually already counted in a system somewhere.

3 to 5 days. One day for small business

02

Read the procedure as it is actually performed

Not as documented. The gap between the two is where the exceptions live, and the exceptions are what determine whether automation is viable.

Observed, not assumed

03

Integration path confirmed before build

Which systems accept a write, which need a browser-level approach, and what that will cost to maintain. Legacy systems integration is cited by 28 percent of middle-market respondents as a top inhibitor.

Named, with the maintenance cost stated

04

Build with a confidence threshold and a human queue

Extraction and matching with an explicit threshold. Below it, work routes to a person. Above it, it proceeds with an audit record. The threshold is agreed, not chosen by us.

Threshold agreed in writing

05

Measure against the count, then widen

Same four measures as the baseline. The exception rate is reported honestly, including where it is higher than expected, and scope widens only after the first procedure holds.

Widened only on evidence

05

Platform native

Delivered natively on whichever platform you are already on. All four, no reseller relationship on any.

Azure

Document Intelligence, Logic Apps and Functions, Power Automate where the estate is Microsoft-centric, Service Bus for queueing.

AWS

Textract, Step Functions and Lambda, EventBridge, SQS for exception queues, A2I for human review.

Google Cloud

Document AI, Workflows and Cloud Run, Pub/Sub, Cloud Tasks for retry and queueing.

DigitalOcean

Functions and App Platform, managed Postgres for state, managed Kafka or Redis for queueing, open-source OCR where a managed service is not warranted.

Where an API exists we use it, which is more robust and cheaper to maintain. We say which category each integration falls into before you commit.

06

The evidence, if you want it

You do not need these numbers to recognise the problem. They are here because somebody in your approval chain will ask.

1 Jan 2027

is the date by which impacted payers must primarily meet the API requirements of the CMS Interoperability and Prior Authorization Final Rule, having implemented certain other provisions by 1 January 2026.

Source: CMS, CMS-0057-F.

$4.4tn

in illicit financial activity was estimated for 2025, up 1.3 trillion dollars since 2023, with fraud scams and bank fraud causing 579.4 billion dollars in losses.

Source: Nasdaq Verafin, 2026 Global Financial Crime Report.

28%

cite legacy systems integration as a top inhibitor to AI deployment, level with the talent gap.

Source: RSM US Middle Market AI Survey, 2026.

There is no reliable market size for this work, and we are not going to quote one. Named research firms published 2026 base-year figures for intelligent document processing ranging from roughly 3.1 billion to 14.2 billion dollars, a four-and-a-half-fold spread for the same category in the same year. Dated regulatory deadlines are a far better guide to where this work is genuinely being bought.

07

The dates that force the work

Five dates, each verified against the body that issued it.

Artifact / Dated obligations that force document and process work Verified against the issuing body

1 January 2026

CMS-0057-F

Certain provisions of the CMS Interoperability and Prior Authorization Final Rule required of impacted payers.

Source: CMS

2 December 2026

EU AI Act

Transparency and machine-readable labelling obligations for AI-generated content apply to systems placed on the market before 2 August 2026.

Source: European Parliament, 11 June 2026

1 January 2027

CMS-0057-F

The main API requirements for impacted payers.

Source: CMS

2 December 2027

EU AI Act

High-risk obligations for stand-alone Annex III systems, moved out from 2 August 2026 under the Digital Omnibus on AI.

Source: Council of the EU

2 August 2028

EU AI Act

High-risk obligations for AI embedded in regulated products under Annex I.

Source: Council of the EU

These five dates are verified against the issuing body. India's Digital Personal Data Protection Rules were notified on 13 November 2025 with phased enforcement, and reports of a proposal to compress that window were not confirmed in the Gazette at the time of writing, so we have deliberately left those dates off this timeline rather than publish a date we cannot stand behind.
08

Questions we are actually asked

How accurate is document extraction?

It depends on the document, and any number quoted before seeing yours is marketing. What we commit to is a measured accuracy figure on your own documents, scored against a labelled set before it ships, plus a confidence threshold below which work goes to a person rather than through.

What happens to the exceptions?

They go to a named queue with a named owner. There will always be an exception queue, and a provider who tells you otherwise is selling something. The measure that matters is whether the queue shrinks over time, and that is reported monthly.

Our system has no API. Is this still possible?

Usually yes, through browser-level automation, but it is scoped as an ongoing maintenance commitment rather than a one-off build, because the vendor will change the screen eventually. We will price that honestly rather than treat it as free.

Is this the same as robotic process automation?

No. Where an API exists we use it, which is more robust and cheaper to maintain. Screen-level automation is a last resort for systems that give us no other option, and we say which category each integration falls into before you commit.

We are a payer facing the CMS deadlines. Where does this start?

With the audit record and the exception path, not with the extraction. The API requirements land primarily on 1 January 2027, with certain provisions from 1 January 2026, and work designed backwards from the audit obligation is the work that survives a review.

Can you automate a process we have not documented?

We can, but the first step is observing how it is actually performed rather than how it is written down. The gap between the two is where the exceptions live. If that observation shows the procedure itself is broken, we will tell you to fix the procedure first.

09

Next step

Bring the procedure and the count.

Thirty minutes, no obligation. Tell us what is being processed by hand and roughly how much of it there is. If nobody knows the volume, that is the first thing worth establishing.

Diagnostic duration
Three to five days. One day for a solo operator or small business.
Investment
Scoped on the introductory call.
If nobody knows the volume
That is the first thing the Diagnostic establishes: volume, cycle time, error rate and hours on the target procedure.