How to run an AI pilot properly and know what it proved
- 5 days ago
- 3 min read
Updated: 2 days ago
Introduction
Pilots exist to reduce the cost of being wrong. Done properly, one produces a clear answer for a small investment. Done as they usually are — open-ended, unmeasured, and evaluated on how people feel about it — a pilot produces neither a decision nor a saving, and it consumes the enthusiasm that would have funded the next attempt.
The failures are consistent and avoidable. No baseline, so nothing can be compared. No stated measure, so the result is whatever the advocate says it is. No end date, so the pilot becomes the permanent arrangement by default. Each of these is fixed by a decision made before the pilot starts, and none of them can be fixed afterwards.
1. How to run an AI pilot properly begins with the stopping rule
Decide it first.
The measure, the threshold and the date on which you decide. Written down, agreed by everyone who will be in the room. This single step prevents most of the ways pilots fail.
2. Record the baseline before switching anything on
Otherwise there is nothing to compare against.
Two or three numbers for the current process, over the preceding weeks. Ten minutes of work, and it is impossible to reconstruct once the change has been made.
3. Keep the scope small and real
One process, actual work.
A pilot on a sandbox or on artificial cases proves nothing about your business. One team, one process, real volumes, for a fixed period. Small scope is what makes a genuine test affordable.
4. Run a comparison group where you can
The strongest design available.
Half the team, half the region, alternate weeks. Something changing at the same time will otherwise be credited or blamed, and a comparison group removes almost all of that ambiguity.
5. Give it long enough to pass the novelty
Six weeks minimum, usually.
Early results are distorted in both directions: enthusiasm inflates them and unfamiliarity depresses them. The steady state is what you are trying to observe, and it takes several weeks to appear.
6. Choose participants who will tell you the truth
Not only the enthusiasts.
A pilot staffed entirely by advocates produces a positive result and no information. Including a sceptic who will use it properly and report honestly is worth more than another supporter.
7. Record the problems as they occur
The most valuable output.
Every error, workaround, confusion and manual correction, logged at the time. This list is what tells you the real cost of the change, and it is invariably forgotten by the time of the review.
8. Decide on the date, and be willing to stop
The discipline that makes pilots worth running.
Extending an inconclusive pilot is the standard outcome and it is a decision not to decide. Stopping is a legitimate and cheap result, and a business that can stop can afford to try more things.
9. Write down what you learned either way
The compounding benefit.
A short record of what was tried, what happened and why it was stopped or continued. Organisations retry the same failed idea every two years because nobody wrote the first attempt down.
Be careful with pilots on live customer-facing processes. A test that degrades service is not free, and the customers involved did not agree to be part of it, so the exposure should be limited deliberately.
Conclusion
Fix the measure, the threshold and the decision date before switching anything on.
Record the baseline first because it cannot be recovered later, keep the scope to one process with real volumes, run a comparison group wherever it is possible, allow at least six weeks for novelty to wear off, include a sceptic among the participants, log every problem and workaround as it happens, make the decision on the agreed date and be genuinely willing to stop, and write down what you learned whichever way it went.
.png)



Comments