
Most legal AI pilots do not fail. They simply never end. Six months in, a handful of enthusiasts are using the tool daily, most of the licences are dormant, nobody has written down what success would look like, and the renewal date arrives before the decision does. The vendor calls this adoption. It is not.
A pilot is a decision-making instrument. If it is not built to produce a yes or a no by a fixed date, it is a free trial with extra meetings.
Decide what you are testing before you decide what you are buying
The most common structural error is picking a product and then looking for work to put through it. Reverse that. Start from a task that is genuinely painful, frequent enough to measure, and bounded enough to judge.
Good candidates share three properties. The task happens at least weekly. The output can be judged right or wrong by someone senior in under fifteen minutes. And somebody currently doing it would be glad to stop.
First-pass review of inbound commercial contracts against a playbook qualifies. Summarising a document set for a case team qualifies. “Improve efficiency across the litigation department” does not, because nothing about it can be measured inside a quarter.
Pick a baseline, and pick it before you start
You cannot report a saving against a number you never recorded. Before the tool is switched on, capture the current state for the specific task: how long it takes, who does it, how often it comes back for correction, and what it costs in either billed or unbilled time.
Two weeks of honest baseline data is worth more than six months of enthusiastic anecdote. This is the step teams skip, and skipping it is why only 18 per cent of organisations collect return on investment metrics on AI, according to Thomson Reuters research from 2026. They cannot, because the before was never written down.
Size the group deliberately
Between six and twelve people is the range that works. Fewer than six and one person’s enthusiasm or scepticism dominates the result. More than twelve and you cannot support them properly, so usage decays and you learn nothing except that people are busy.
Composition matters more than size. Include at least one confirmed sceptic. A pilot staffed entirely with volunteers who already wanted the tool produces a result you cannot generalise from, and everyone in the room knows it.
Run it for eight to twelve weeks, with a date in the diary
Shorter than eight weeks and you are measuring novelty. Longer than twelve and the pilot becomes the status quo, which is how tools get renewed by default.
Put the decision meeting in calendars on day one, with the decision-maker present. The single highest-leverage thing you can do at the start of a pilot is book its ending.
Measure four things, not fourteen
- Time on task against the baseline, measured on the same work type, not on whatever the tool happened to be good at.
- Quality, judged by a reviewer who does not know whether the output was AI-assisted. Blind review is the only version of this that means anything.
- Rework rate: how often output needed substantive correction rather than light editing. This is usually the number that decides the case.
- Actual usage, pulled from the vendor’s own logs. Ask for this access in the pilot agreement, before signature.
Resist the urge to add a satisfaction survey as a fifth measure. People report liking tools they barely use.
Write down the failure conditions in advance
This is the part that separates a pilot from a purchase with extra steps. Before starting, agree in writing what result would mean no. For example: fewer than 50 per cent of participants using it weekly by week six, or a rework rate above 30 per cent, or no measurable time saving on the target task.
Committing to the kill criteria before you are emotionally invested in the tool is the whole trick.
Handle the confidentiality question first, not last
A pilot is still live client work. The ethics position in the United States is set out in ABA Formal Opinion 512, which treats confidentiality under Model Rule 1.6 as engaged when client information goes into a generative tool, and which requires informed client consent in circumstances where the information would be disclosed.
Before any live matter data enters the system, settle four points: whether inputs are used for training, where data is stored and under whose jurisdiction, retention and deletion terms, and who at the vendor can access it. Get the answers in the contract. A sales engineer’s assurance on a call is not a contractual term.
If any of that cannot be resolved in time, run the pilot on synthetic or closed-matter documents. It weakens the test slightly. It is far better than the alternative.
Budget for the work around the tool
The licence is rarely the expensive part. Configuration, playbook building, template mapping, training, and the internal time to run the pilot itself routinely exceed it. Where the tool needs to reach your document management system, integration is its own project.
Costing the pilot at licence price alone produces a business case that collapses the moment you try to scale it.
What a finished pilot looks like
One page. The task tested, the baseline, the four measures with before and after numbers, the usage data, the failure conditions and whether they were hit, and a recommendation with a price.
If you cannot write that page, the pilot did not finish. Extend it once with a new end date, or stop. Do not let it drift into a renewal.
Related reading: the evolution of legal technology in law firms, and our guide to what legal AI actually costs.
Where the legal industry reads first.
Enjoyed this article? Get the biggest legal industry updates, deals, appointments, insights and expert interviews in your inbox, free.
No spam. Unsubscribe anytime.