
Why Most AI Projects Never Reach Production
MIT researchers reviewed hundreds of enterprise AI initiatives and found the overwhelming majority produced nothing measurable. The reason is not the technology, and it is the same reason in almost every case.

A pilot runs, everybody agrees it worked, and eighteen months later nothing about the business has changed.
This is common enough to have been measured. It is worth understanding why, because the failure is structural rather than technical and it is avoidable.
What the research found
Researchers at MIT’s Project NANDA published The GenAI Divide in July 2025, drawing on a review of over 300 publicly disclosed AI initiatives, 52 structured interviews and 153 survey responses from senior leaders, gathered between January and June 2025.
Their headline finding: against an estimated $30 to $40 billion of enterprise investment, roughly 95% of generative AI projects produced no measurable return. Around 5% were delivering real value.
Two honest caveats belong with that number. The study looks at large organisations rather than companies between five and fifty million. And “no measurable return” frequently means nobody built a way to measure, not that nothing happened. Both of those make the finding less damning and more useful, because both are fixable before you start.
The pattern underneath it
What separates the 5% is not model choice, budget or technical sophistication. It is whether the thing became part of how work happens.
A pilot is something people opt into. A system is something that runs whether or not anybody remembers it exists. Almost every project that produces nothing stops at the first and never becomes the second, and the gap between them is not technical work. It is triggers, permissions, ownership and the decision to switch off the old path.
Which is unglamorous, which is why it does not get done.
The four places it dies
It never got a trigger. The capability exists and somebody has to remember to use it. Adoption follows enthusiasm, enthusiasm fades, and within a quarter it is two people. Anything depending on a human remembering is not automated, it is available.
It writes nowhere. Output arrives in a chat window or a document and a person retypes it into the system that runs the business. The judgement was automated and the labour was not, so the hours never come back.
The old path stayed open. Both routes exist, people use whichever is familiar under pressure, and the data ends up in two places. Nothing is ever decommissioned, so nothing is ever measured.
Nobody owns it. It was built by whoever was interested. They moved on, something broke quietly, and the business learned that these things do not last.
Why measurement is the one to fix first
If a project cannot be measured it cannot be defended, and undefended projects lose their budget the first time the year is difficult.
The fix costs an afternoon. Before anything is built, write down the current state of one number: elapsed time from trigger to done, hours consumed, error rate, percentage handled within a day. One number, recorded, dated.
Most of the organisations reporting no return never did this. They cannot show a gain, which is not the same as not having made one, but from a board’s point of view it is indistinguishable.
What AI is genuinely good at, and where to aim it
The results that hold up share a shape. They sit on a step that was previously blocked because it needed judgement, and judgement always meant a person.
Reading a document that arrives in any format and pulling out what matters. Deciding which of forty enquiries deserves attention first. Drafting a reply that accounts for what the client said three months ago. Watching a data set and noticing the one thing that changed. Each of those was impossible to automate with rules, which is why they stayed manual, and each is now solvable.
Aim at those and the return is legible. Aim at a process that was already efficient and you get a marginally faster version of something nobody was complaining about.
Why the pilot framing causes the failure
A pilot is designed to be reversible, and everything about how it gets run follows from that.
It is given to volunteers rather than to the team that owns the process. It runs beside the real workflow instead of inside it. It is deliberately not connected to the systems of record, because connecting things is the expensive part and nobody spends that money on a trial. And it is scheduled to end, which means the moment it succeeds, it stops.
Every one of those choices is sensible in isolation and together they guarantee the outcome. You have proved a capability works in conditions that do not resemble your business, and you now have to fund the real project from scratch, against a result that no longer exists.
The organisations in the 5% skipped that stage. They picked something small enough that building it properly was cheaper than trialling it, and they built it properly the first time.
The question to ask in the first meeting
When somebody proposes an AI project, ask what will be switched off when it works.
If the answer is nothing, you are adding a capability rather than changing a process, and the hours will not appear. If the answer is a specific task that a named person currently does on specific days, the project has a shape and a measurable result.
It is a blunt question and it sorts proposals faster than any business case.
What a small company should take from this
The study covers large organisations, and the smaller you are the better your odds, for a reason worth stating plainly.
A fifty million dollar company has one owner, one process and one decision. The distance from “this works” to “this is how we do it now” is a conversation rather than a change programme. Large organisations lose their projects in that distance. You are not required to.
What you have to do is refuse the pilot framing. Do not run a trial. Build the narrow version, put it in the path of real work, switch the old route off, and record the number before and after. That is the entire difference between the 5% and everybody else.
Sources
Aditya Challapally, Chris Pease, Ramesh Raskar and Pradyumna Chari, The GenAI Divide: State of AI in Business 2025, MIT Project NANDA, July 2025.
AI Optimize builds systems that run on triggers, write into the software you already use, and tell somebody when they fail. Nothing we deliver depends on a person remembering it exists. That work sits under Workflow Automation.
Related reading

Your First Ninety Days With AI
A plan for a business that has decided to do something and does not want to waste the first attempt. One process, one number, and a deliberate order.

A Tool Is Not a System, and the Difference Costs You
Most companies that buy AI end up with more software and the same headcount. The line between a tool and a system is what decides which one you get.
WHAT WE BUILD



