Run an AI pilot that produces a decision, not a demo
Most pilots end in a positive write-up and no change. Designing for a decision fixes that.
Before you start
- A specific process that is slow or expensive
- Someone able to say yes or no at the end
What you will be able to do
- Define a success threshold before results can influence it
- Measure against a real baseline rather than an estimated one
- End the pilot with a decision instead of a recommendation
The familiar pattern: a team runs a three-month AI pilot, reports promising results, and nothing changes. Nobody is at fault — the pilot was never designed to produce a verdict.
The fixes are all in the setup, and they take an afternoon.
Pick one task, narrowly defined
A pilot that spans a department measures nothing you can attribute.
Choose a single task with a clear start and end, done often enough to gather data in weeks, by people you can actually talk to. "Support triage for tier-one tickets" is a pilot. "AI in customer service" is a programme.
Breadth is what makes results unattributable — when six things changed at once, nobody can say which one helped.
Write down the number that means yes
Before you see any results, and agreed by the person who decides.
State it plainly: "we adopt this if median handling time drops by 20% with no increase in reopened tickets". Get the decision-maker to agree to that sentence in advance.
Done afterwards, the threshold gets set to just below whatever you measured. This is not dishonesty, it is how people read evidence — which is exactly why the number goes first.
- Choosing a metric that only improves. Time saved always improves; pair it with a quality measure that can get worse, or you have not tested anything.
Measure the baseline properly
Two weeks of the current process, measured the same way.
Measure the existing process before you change anything, with the same instrument you will use afterwards. Remembered baselines are consistently wrong and always flattering to the new thing.
This also catches the awkward case where the current process is already better than assumed, which is a genuinely useful result and one that never emerges from a pilot with no baseline.
Include the costs people forget
Review time, exception handling, and the fixed cost of running it.
Count the checking. If output needs review, that review is part of the cost and it is where most of the projected saving quietly goes.
Count the exceptions too — the cases that fall out of the automated path and now need someone to handle them specially, often at higher cost than before. And count integration and maintenance, which do not stop when the pilot does.
Tip Set the decision date in the calendar
A pilot without an end date becomes a permanent parallel process.
Book the meeting when you start the pilot, with the decision-maker in it. The agenda is one item: adopt, drop, or run one more defined round with a stated reason.
Without it the pilot never formally ends; it just runs alongside the old process forever, costing both.
One task, a number agreed in advance, a real baseline, and a named decision date. A pilot without all four produces a write-up.
Common questions
Was this guide useful?
92% of readers found this useful
Read next
Estimate what an AI feature will cost you per month
Per-token pricing looks trivially cheap and routinely surprises people at the invoice. The gap is almost always retries, context a…
Write an AI usage policy your team will actually follow
A restrictive AI policy does not stop people using AI. It stops them telling you, which is considerably worse.
Connect two apps with an AI step in the middle
The genuinely useful automations are not the clever ones. They are a trigger, one AI step that makes a small judgement, and a writ…
Write ad variants and test them properly
AI removes the cost of writing variants, which makes it very easy to run tests that cannot teach you anything. The discipline is i…