Why we went and checked
This chapter was supposed to open with a frightening statistic. That is how the genre works: you lead with the share of AI projects that fail, the reader feels the ground move, and then you appear as the person who knows the way through.
We went to find the study behind the number so we could cite it properly. What we found is that the three statistics everyone quotes — the MIT 95%, the RAND 80%, and the Gartner 40% — are, in order, a different claim, a citation of a magazine article, and a prediction about the future. Not one of them is a measured failure rate for AI projects.
We are showing you the working rather than quietly picking a different opening, because the working is worth more to you than the statistic was. If you take one thing from this whole guide, take the habit below: go and look at what the number actually says.
| As quoted | What the source actually says | Who was measured | What kind of thing it is |
|---|---|---|---|
| MIT: 95% of AI projects fail | “95% of organizations are getting zero return” | 52 interviews, 153 conference survey responses, 300 public initiatives. Enterprises. | A profit-and-loss measure, not a failure rate |
| RAND: 80% of AI projects fail | “By some estimates, more than 80 percent of AI projects fail”, footnoted to a magazine article | Nobody. RAND interviewed 65 engineers and data scientists about why projects fail, not how many. | A citation of a magazine, not a measurement |
| Gartner: 40% will be cancelled | “Over 40% of agentic AI projects will be canceled by the end of 2027” | Nobody yet. It has not happened. | A forecast, not evidence |
| Small businesses like yours | No figure exists. | Nobody has measured it. | The gap this chapter reports rather than fills |
The rest of this chapter is that table with its working shown. If the table is enough, skip to what the good evidence actually shows, which is the part that would change what you do.
The MIT 95%
You have seen this one. It is usually rendered as “MIT found 95% of AI projects fail”.
The report is The GenAI Divide: State of AI in Business 2025, from MIT’s Project NANDA, published in July 2025. Its actual sentence is that, despite thirty to forty billion dollars of enterprise investment, 95% of organizations are getting zero return.
That is a claim about profit-and-loss impact inside roughly six months at large companies. It is not a project failure rate, and it is not about small businesses at all. The evidence base is 300 or so publicly disclosed initiatives, 52 organisation interviews, and 153 survey responses collected at four industry conferences. The report describes itself on its own cover as “Preliminary Findings”. It has not been peer reviewed.
The part that almost nobody quotes is the funnel in the same report: 60% of organisations evaluated a custom AI tool, 20% reached a pilot, and 5%reached production. Read as a conversion rate from pilot to production, that is roughly one in four pilots making it, with most organisations never having piloted anything. That is a materially different picture from “95% fail”, and it is in the same document.
A number that has been repeated ten thousand times has been checked by approximately none of them. That is not cynicism, it is how citation works: the second person cites the first, and the thousandth cites the nine hundred and ninety-ninth.
The RAND 80%
The second most-quoted version is “RAND found that 80% of AI projects fail, twice the rate of ordinary IT projects”.
RAND’s report RR-A2680-1 does contain that sentence, near the front. The wording is “By some estimates, more than 80 percent of AI projects fail”, and the sentence after it says that is twice the failure rate of information technology projects which do not involve AI. The words “by some estimates” carry a footnote, and the footnote points at a magazine article rather than at a study.
We should say how far we got. The report itself returned an access error to us, so the wording above is what independent secondary accounts of it agree on rather than something we read off the page. That is a weaker position than we would like on a chapter about checking sources, and stating it is better than implying we opened a document we did not.
So RAND is repeating a figure, not producing one. What RAND actually did was interview 65 experienced data scientists and engineers, drawn from industry and academia, about why projects fail. That work is genuinely useful and it is the part worth reading. It is simply not a measurement of how many fail.
The Gartner 40%
The third is “Gartner says 40% of agentic AI projects will be cancelled”. This one is quoted accurately and then read wrongly.
Gartner’s press release of 25 June 2025 says: “Over 40% of agentic AI projects will be canceled by the end of 2027, due to escalating costs, unclear business value or inadequate risk controls.” That is a forecast. It is a firm of analysts stating an expectation about something that has not happened yet. It may well turn out to be right. It is not evidence about anything that has already occurred, and using it as though it were is a category error.
The same is true of the 2024 Gartner line about 30% of generative AI projects being abandoned after proof of concept by the end of 2025. Prediction, not measurement.
The number nobody has
Here is the finding that actually changed how we talk about this.
The most defensible abandonment figure we could find is S&P Global Market Intelligence: 42%of companies abandoning most of their AI initiatives, up from 17% the year before, with the average organisation scrapping 46% of proofs of concept. More than 1,000 respondents across North America and Europe, reported by CIO Dive on 14 March 2025. We should say plainly that S&P’s own page returned an access error to us, so we read that figure through CIO Dive rather than at source.
Now notice what all of these have in common. MIT: enterprises. RAND: corporate IT. Gartner: enterprise agentic projects. S&P: companies with executives who answer analyst surveys. BCG’s September 2025 study of 1,250 firms, which found only 5% achieving value at scale, surveyed CxOs across 68 countries.
There is no measured abandonment rate for AI projects in small, owner-led businesses. We looked for it specifically. It does not appear to exist. Which means every time someone tells you that most AI projects like yours fail, they are transferring a number from organisations with procurement departments and change-management budgets onto a business where the owner does the ordering.
We are not going to fill that gap with a guess. A missing number reported as missing is worth more than an invented one, and if we invented this one you would have no reason to believe the calculator in the previous chapter either.
What the good evidence actually shows
Underneath the noise there is a body of properly designed work, and it tells a coherent story that is neither hype nor doom. It is more useful than any of the three headline numbers.
It works well in narrow, high-volume, low-skill-ceiling tasks
Customer support agents with an AI assistant resolved 14% more issues per hour on average, and 34% more among the least experienced agents, with almost no effect on the most experienced ones (Brynjolfsson, Li and Raymond, Quarterly Journal of Economics, staggered rollout across 5,179 agents). Mid-level professional writing tasks got 40% faster with quality up 18% in a randomised trial of 453 professionals (Noy and Zhang, Science, 2023).
It does approximately nothing in aggregate, so far
Linking adoption surveys to Danish administrative records for roughly 25,000 workers across 7,000 workplaces, Humlum and Vestergaard found precise null effects on earnings and hourstwo years after ChatGPT — effects larger than 2% ruled out (NBER working paper 33777). Not a small gain. A carefully measured nothing.
It can make a weaker operator worse
This is the finding most relevant to a small business and the one you will never see in a pitch. In a randomised field experiment with 640 Kenyan entrepreneurs given an AI business mentor over five months, the average effect could not be distinguished from zero. The strongest performers at baseline improved by about 15%. The weakest performers got about 8% worse (Otis, Clarke, Delecourt, Holtz and Koning, Harvard Business School working paper 24-042).
And people are badly wrong about which of these they are in
METR ran a randomised controlled trial with 16 experienced developers across 246 real issues on codebases they knew well. With AI allowed they were 19% slower. Beforehand they predicted they would be 24% faster. Afterwards, having been slower, they believed they had been 20% faster.
The sample is small and the authors say so. The transferable finding is not the 19%. It is the 39-point gap between what people felt and what happened, in a group who were, by any measure, expert.
Narrow task, high volume, someone inexperienced doing it: good odds. Broad judgement, low volume, someone experienced doing it: the evidence says be careful. And your own sense of which one you are in is not reliable, which is the entire argument for measuring before and after.
Why any of this matters to you
You are not going to read the underlying papers, and you should not have to. What you should take is a working method, and it is the same one you would use on a supplier quoting you a delivery time.
- Ask what the number measured.“Zero return within six months” and “the project failed” are not the same claim, and the gap between them is where most of the fear lives.
- Ask who was in the sample. A figure from enterprises with change-management budgets says very little about a business with eleven staff.
- Ask whether it is a forecast. Analyst predictions get quoted in the past tense within about a fortnight of publication.
- Notice who is holding the number.Including us. We sell AI consulting, and the correct thing for you to do with a scary statistic from us is exactly what we just did to everyone else’s.
The next chapter is the one those four questions lead to: the cases where the honest answer is that AI does not help, which on the evidence above is a larger set than anybody selling it would like.
Where these numbers came from
Every figure on this page, with what it is and where it is from. If a number is illustrative rather than measured, it says so here and it says so in the text.
- 95%The famous MIT figure. Its actual sentence is that 95% of organisations are getting zero return, which is a measure of profit-and-loss impact, not a project failure rate.MIT Project NANDA, The GenAI Divide: State of AI in Business 2025, July 2025. Self-described as Preliminary Findings. Built on 300+ publicly disclosed initiatives, 52 organisation interviews and 153 survey responses collected at four industry conferences, research period January to June 2025. Not peer reviewed.From a named study
- 60 / 20 / 5The funnel inside that same MIT report: 60% of organisations evaluated a custom AI tool, 20% reached pilot, 5% reached production. Read as a conversion rate from pilot, that is roughly a quarter succeeding, not 95% failing.MIT Project NANDA, The GenAI Divide, July 2025, same report.From a named study
- 80%The RAND figure. RAND did not measure it. Their sentence is "By some estimates, more than 80 percent of AI projects fail," footnoted to a magazine article rather than to a study.RAND Corporation report RR-A2680-1, Ryseff, De Bruhl and Newberry, 2024. RAND own research in that report is interviews with 65 data scientists and engineers with at least five years of experience, in industry and academia, about why projects fail.From a named study
- 40%The Gartner figure is a forecast about 2027, not a measurement of anything that has happened.Gartner press release, 25 June 2025: "Over 40% of agentic AI projects will be canceled by the end of 2027, due to escalating costs, unclear business value or inadequate risk controls."From a named study
- 42%Share of companies abandoning most of their AI initiatives, up from 17% the year before. The most defensible abandonment figure we found, and it is an enterprise sample reported secondhand.S&P Global Market Intelligence survey of more than 1,000 respondents in North America and Europe, reported by CIO Dive on 14 March 2025. S&P Global own page returned an access error to us, so we read this through CIO Dive.From a named study
- not measuredThe share of AI projects abandoned by small and owner-led businesses specifically. Every abandonment figure we could find is drawn from enterprise or executive samples. We did not find this number and we are not estimating it.Our own search, August 2026, across the studies listed on this page. Stated as a gap rather than filled.We counted it
- ruled out above 2%Effects of AI chatbots on earnings and hours, two years after ChatGPT. Not a small positive effect. A precisely estimated nothing.Humlum and Vestergaard, NBER working paper 33777, linking adoption surveys to Danish administrative records across roughly 25,000 workers in about 7,000 workplaces, 11 occupations, 2023 to 2024. Issued May 2025, revised March 2026.From a named study
- minus 8%Effect of an AI business mentor on the weakest half of small entrepreneurs. The average effect could not be distinguished from zero; the strongest performers gained 15% and the weakest lost.Otis, Clarke, Delecourt, Holtz and Koning, The Uneven Impact of Generative AI on Entrepreneurial Performance, Harvard Business School working paper 24-042, randomised field experiment with 640 Kenyan entrepreneurs over five months, December 2023.From a named study
- 19% slowerExperienced developers took 19% longer to complete real tasks with AI allowed, while predicting beforehand they would be 24% faster and believing afterwards they had been 20% faster.METR, Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity, randomised controlled trial, 16 developers across 246 real issues on mature repositories, July 2025.From a named study
- 5% of 1,250 firmsThe share achieving AI value at scale, with 60% not achieving material value at all. Another enterprise and executive sample, which is the point being made about all of them.Boston Consulting Group, The Widening AI Value Gap: Build for the Future 2025, survey of 1,250 senior executives across more than 25 sectors in 68 countries, September 2025.From a named study
- 40% fasterTime taken on mid-level professional writing tasks, with quality rated 18% higher. A randomised trial, though on short artificial tasks rather than real job output.Noy and Zhang, Experimental evidence on the productivity effects of generative artificial intelligence, Science, July 2023, randomised controlled trial with 453 college-educated professionals.From a named study
- plus 14%Customer support issues resolved per hour with an AI assistant, rising to 34% for the least experienced agents and close to nothing for the most experienced.Brynjolfsson, Li and Raymond, Generative AI at Work, Quarterly Journal of Economics 140(2), 2025, pages 889 to 942. Staggered rollout across 5,179 support agents; first circulated as NBER working paper 31161 in April 2023.From a named study