Nearly every organisation is doing something with AI. Very few can show what it bought them. The difference is not the technology.

There is a conversation happening in most boardrooms right now that goes something like this. We have Copilot licences. We ran a pilot in the summer. Everyone says it is going well. And then somebody asks what it saved, and the room goes quiet.

That silence is the most interesting thing in enterprise AI at the moment because it is nearly universal and has almost nothing to do with model quality.

The evidence is unusually consistent.

Research findings in this area often conflict with one another—this time they do not. MIT’s study of enterprise deployments found that 95% of AI pilots produced no measurable impact on profit and loss. Figures widely attributed to Gartner put the share of AI agent pilots that never reach production at 89%, with IDC at 88%. PwC surveyed 4,454 chief executives across 95 countries and found 56% saying AI had delivered no cost or revenue improvement at all in the previous twelve months.

Four studies, different methods and samples, landing in roughly the same place.
Four studies, different methods and samples, landing in roughly the same place.

A caveat is worth entering here, because these numbers get repeated a great deal and rarely questioned. The MIT figure comes from a study with a fairly small sample, and it attracted legitimate methodological criticism when it was published. Several of the agent pilot percentages circulate mainly through secondary reporting rather than a primary publication you can read. I would treat the precise values as indicative rather than exact.

What survives that scepticism is the direction, which is consistent across every source and matches what most people observe in their own organisations. And the more interesting number is the tail rather than the headline: the pilots that did reach production are reported to return well over 100%. So this is not a story about AI failing to work. It is a story about most organisations never getting far enough to find out whether it works.

Where the value actually leaks

In the programmes I have seen up close, the loss is not one big failure. It is four smaller ones stacked on top of each other, and each one is boring enough that nobody escalates it.

The attrition is cumulative. Fix one gate and you still lose at the next three.
The attrition is cumulative. Fix one gate,e and you still lose at the next three.

Nobody in the business owns it. The idea came from IT, an innovation team, or a vendor workshop. It has a sponsor but no owner: nobody whose actual job gets easier or harder depending on whether it works. Ideas in this state do not die, which is worse. They drift.

The data is not reachable. The proof of concept ran on a curated extract that somebody assembled by hand. Production needs the live source, with real permissions, and it turns out the permissions are a mess. This is the point where most projects quietly slip a quarter.

There is no route from prototype to run. Somebody built something good in Copilot Studio or Foundry. Now it needs a support model, a change process, a budget line and a name on a rota. No such route exists, so the thing stays a demo and gradually stops working.

Nobody baselined the benefit. This one is fatal and entirely self-inflicted. If you did not measure how long the task took before, you cannot prove anything afterwards. You end up arguing from anecdote against a CFO who deals in numbers, and you lose.

The pilot did not fail. It was never set up in a way that could succeed or fail, which is a different and more expensive problem.

What the survivors do differently

The organisations getting real returns are not using better models. Everyone has access to broadly the same frontier models now, and on the Microsoft stack the gap between what a large enterprise and a mid-market business can deploy has narrowed considerably. What separates them is closer to project hygiene than to technology.

They pick fewer things. A shortlist of three genuinely useful use cases beats a portfolio of fifteen interesting ones, because the constraint is rarely ideas. It is the attention of the people who understand the process well enough to change it.

They start from a task, not a tool. The question is not what can Copilot do. It is which job in this business takes too long, happens often, and produces something a person then checks anyway. That last part matters more than it sounds, because a task with a natural human review step is one where AI can occasionally be wrong without anyone getting hurt.

And they measure before they build. Not elaborately. Ten people, two weeks- how long does this take you today? That single step converts every later conversation from opinion into arithmetic.

For the practitioners

  • Microsoft’s 2026 Work Trend Index found that organisational factors, meaning culture, manager support and talent practices, drive roughly twice the AI impact of individual mindset and behaviour. If your programme is mostly training and enthusiasm, you are working on the smaller half of the problem.
  • Baseline with something you already own. Viva Insights and the Microsoft 365 admin usage reports give you a defensible starting point without commissioning a study.
  • Instrument the pilot from day one. Copilot Studio and Foundry both surface usage and cost telemetry, and Purview audit captures Copilot and agent interactions. Turning these on afterwards means you lose the comparison period permanently.
  • Write the success criterion before the build starts and get it agreed by whoever will eventually challenge it. A number nobody agreed to in advance is a number that gets argued away.

The uncomfortable implication

If the failure rate were about technology, the fix would be procurement. Consider buying better tools, hiring better engineers, and waiting for the next model. That is a comfortable diagnosis because it is somebody else’s job.

The evidence points somewhere less convenient. The failures cluster around ownership, data access, operating model and measurement, all of which are management problems that existed before AI arrived and will still be there afterwards. AI just made them expensive enough to notice.

That is not a counsel of despair. It is quite encouraging, because those problems are tractable in a way that inventing better models is not. But it does mean the first useful AI conversation in most organisations is not about AI.

The objection worth taking seriously

There is a reasonable counter-argument to everything above, and it deserves a hearing rather than a dismissal. It runs like this: demanding a measurable business case before every experiment is how organisations talk themselves out of new technology. Early computing, early internet and early cloud all looked like poor investments under strict return analysis. Some capability has to be built speculatively, and the learning has value even when the pilot does not.

I think that is largely right, and it changes the conclusion less than it appears to. The argument for exploratory work is an argument for a small, explicitly labelled portion of the portfolio, funded as learning and judged on what was learned. What it is not is a justification for fifteen unmeasured pilots, each of which was presented to a board as a business improvement.

The problem in most organisations is not that they experiment. It is that they run production-shaped projects with experiment-shaped rigour, and then are surprised when nobody can say whether it worked. Label the exploration as exploration and the discipline question resolves itself.

What a credible first year looks like

If you wanted a shape for the first twelve months that would survive contact with a sceptical finance director, it is less exciting than most roadmaps and considerably more likely to produce evidence.

PeriodWhat is happeningWhat you should have at the end
Months 1 to 3Data and permissions assessment. Baseline measurement on three candidate tasks. Governance configured while the estate is small.A defensible before picture and a shortlist with named owners
Months 4 to 6Build and pilot the first use case with the team that does the work. Instrument everything.One task genuinely done differently, with before and after numbers
Months 7 to 9Second and third use cases. First honest review of what the first one actually saved.Evidence that generalises, or a clear reason why it did not
Months 10 to 12Scale what worked. Retire what did not. Make the case for the next tranche.A funding argument built on your own data rather than vendor case studies
Deliberately unambitious. The point is that each quarter produces something the next quarter can be argued from.

The temptation is to compress this, and the pressure to do so is usually external. But the compression almost always comes out of the first quarter, which is the data and measurement work, which is precisely the part that determines whether anything later can be proven.

What to take away

  1. Assume your pilot will produce no measurable benefit unless you designed it to do so. That is the base rate.
  2. Find the owner before you find the use case. We need someone whose work changes, not someone who approves the budget.
  3. Baseline the task before you touch it. Two weeks of timings are enough, and it is the highest-return activity in the whole programme.
  4. Treat the route from prototype to production as part of the build, not as something that happens later.
  5. Judge the programme on tasks changed, not on licences deployed or pilots launched.

Share.
Leave A Reply