How to run an AI pilot that reaches production
Most pilots stall before they go live. Here is the four-stage pattern that takes one workflow from idea to running system.
An AI pilot reaches production when it is run as a small production system from the start: one workflow, a measured baseline, a single success metric, real traffic in a narrow scope, and a human reviewing every output until the error rate earns autonomy. Most pilots stall for organisational reasons, because nobody owns the scope, the metric or the route to go-live, and a promising demo sits in evaluation until the budget moves on. The four stages below are designed to remove each of those failure points.
How do you scope a pilot that can succeed?
Stage one is scope. Pick a single workflow, and make it one with volume, clear rules and a visible owner. Then measure the baseline before anything is built: two weeks of real numbers covering how many items arrive, how long each takes to handle, how fast the first response goes out, and what the current error or miss rate is. A pilot without a baseline cannot prove anything, because there is nothing to measure the result against.
Agree one success metric and its threshold in writing, such as first response inside two hours, or 80% of documents classified correctly before review. One metric keeps the pilot honest; a scorecard of ten lets everyone find a number that flatters the result. And name the owner: one person who can accept the scope, read the dashboard and call the go-live.
What should be built before go-live?
Stage two builds against a test set rather than live work. Assemble a few dozen real historical cases, including the awkward ones, and configure the prompts, the classification rules, the approval points and the exception routing against them. Behaviour is measured on the test set first, so by the time live traffic arrives the error rate is already known rather than hoped for.
Build the operating surfaces at the same time: a simple dashboard showing volume, decisions, exceptions and approvals, and the exception queue itself, with a named person who works it. These are small pieces of work, and they are the difference between a system that can be operated and a model call with no way to see what it is doing.
How should the pilot go live?
Stage three is go-live on a small, real subset from day one: one inbox, one document type, one client segment. A human reviews every output before it counts, every exception is logged, and the workflow runs on genuine work rather than a rehearsal. Because the scope is narrow, the blast radius of any error is small, and because the traffic is real, the evidence is real.
The first week ends with evidence rather than opinion: live items handled against the agreed metric, a log of what the system declined to touch, and a first read on where its confidence boundaries sit. Refusals at this stage are information rather than failure; a system that declines the cases it is unsure about is behaving exactly as designed.
How do you move from pilot to production?
Stage four is operate and expand. Review the real exception log weekly, tune the prompts and rules against the cases that actually failed rather than the ones anyone imagined, and widen the scope in steps: more volume, another document type, a second inbox. Approval gates are relaxed only as the error-rate record justifies it, one action type at a time.
Decide the operating ownership before expansion: who reads the dashboard, who works the exception queue, who signs off changes. A pilot that reaches this stage already has controls, logging, live traffic and an owner, which means there is no separate leap to production left to make. The next workflow starts the same four stages with the lessons of the first already applied.
The pattern works because every stage produces the evidence the next stage stands on: a baseline to beat, a tested error rate, a live result and an exception log. Run the stages in order, keep the scope narrow, and let the record rather than the calendar decide when autonomy increases.
How to automate your back office with AI
A practical, audit-first way to remove the repetitive admin that quietly costs you every week, and to measure what each fix is actually worth.
A board's guide to governing AI
What directors should ask before approving AI in the business, and the controls that make the answer yes.