The short version
- Abandonment is now the norm rather than the exception: 42 percent of companies scrapped most of their AI initiatives in 2025, up from 17 percent the year before.
- Pilots fail on four things, and capability is rarely one of them: no owner, no measure, too broad a scope, and no user who wanted it.
- A pilot that cannot name the habit it is changing will produce a demo and a decision to revisit next quarter.
- Scope to one team of six to fifteen people and one measurable habit. Everything else is a proof of concept pretending otherwise.
Most companies now have a history of AI pilots rather than a portfolio of AI systems. The interesting question is not whether the technology worked in the pilot, because it usually did. It is why the pilot could not produce a decision.
"We ran three pilots. None shipped."
Abandonment is the norm now
In 2025, 42 percent of companies scrapped most of their AI initiatives, up from 17 percent the year before. That is a striking move in a single year, and it is not a statement about model quality, which improved over the same period.
It is a statement about how the initiatives were set up. A rate that more than doubles in a year points at a repeated design error rather than at a technology limit, which is encouraging, because design errors are fixable and technology limits are not.
The four failures
| Failure | How it shows up |
|---|---|
| No owner | A committee sponsors it, so nobody's quarter depends on it |
| No measure | Success is "learnings", so the pilot cannot pass or fail |
| Scope too broad | Three departments, six use cases, four months, no conclusion |
| No user who wanted it | Bought for the buyer's problem, used by people with a different one |
The second row is the most common and the easiest to fix. A pilot whose success criterion is not written down before it starts will be assessed afterwards by whoever has the strongest opinion, and the safe verdict is always to revisit next quarter.
The failure that matters most
The fourth row is the one that kills pilots that otherwise look successful. Workplace software is generally bought to solve a manager's problem and used by people who have a different problem, and usage collapses the moment the pilot's novelty and social pressure fade.
The test is straightforward and uncomfortable: in week four, without anybody reminding them, are the intended users still using it? Usage in week one measures compliance. Usage in week four measures whether the thing is useful to the person doing it.
If the answer is no, the pilot has produced a real finding, which is more than most pilots manage. It is not a technology finding.
How to design a pilot that can conclude
Four decisions, all made before it starts.
- One named owner whose quarter depends on it. Not a steering group.
- One number, written down in advance, with a pass threshold. Ideally a behaviour rather than a satisfaction score.
- One team, six to fifteen people. Small enough to observe directly, large enough that the result is not one enthusiast.
- One habit being changed. If you cannot name the habit in a sentence, the scope is still too broad.
And decide in advance what a failure looks like, because a pilot that cannot fail cannot succeed either. "If fewer than half the team is still doing this unprompted in week four, we stop" is a real criterion.
The shape we use
For transparency, this is how StandIn pilots are scoped, and the reasoning applies whoever you are piloting with.
One team of six to fifteen, one habit, and one measured number. The habit is the ninety-second brief at the end of the day. The number is the share of people still writing it unprompted in week four, because that single measure tells you whether the thing is useful to the person doing it rather than to the person who bought it.
The value being tested is equally specific: when someone is off, their colleagues get answers from what they already wrote, in their words, with a source under every answer, instead of waiting. If the habit holds and the waiting falls, it worked. If the habit does not hold, no amount of capability rescues it, and we would rather learn that in week four than in month six. See StandIn for leadership teams.
Common Questions
Why do AI pilots fail to reach production?
Usually for four reasons unrelated to capability: no single owner, no pre-agreed measure, too broad a scope to produce a conclusion, and no end user who wanted the thing in the first place.
How long should an AI pilot run?
Four to six weeks with one team. Long enough for novelty to wear off, which is when you learn whether it is genuinely useful, and short enough that the outcome still matters to the owner.
What should we measure in a pilot?
A behaviour rather than a sentiment. The share of intended users still doing the new thing unprompted in week four is more informative than any satisfaction survey, because it cannot be answered politely.
Should a pilot include several teams?
No. One team of six to fifteen gives you observable behaviour and a clean result. Several teams multiply the variables and produce a report rather than a decision.
When you're off, your StandIn is on.
It answers your teammates' questions from work you've already done, in your words, with a source under every answer.