
Most AI deployments fail. What the numbers say — and why
If you deployed AI in your company and have the feeling it “somehow didn’t pay off,” you did not buy a bad product and you did nothing stupid. You are in the majority — an overwhelming one. Before we say how we ourselves see it, let’s look at the data. No hype, sources given, because that is the only way to talk about this seriously.
Let us make one thing clear up front: we do not write this as experts who know what you should do. No honest party has a single recipe for AI — we don’t have one either. We can only show what we see from our side and how we interpret it. The decision, as always, stays with you.
The scale: this is not the exception, it’s the rule
Hard numbers from the past year, each from a serious institution:
- MIT (the report “The GenAI Divide: State of AI in Business 2025”) — 95% of corporate generative-AI pilots produce no measurable return on the bottom line. Only about 5% genuinely accelerate the business. The study was based on 150 interviews with leaders, 350 surveys and an analysis of 300 public deployments.
- RAND Corporation (2024) — more than 80% of AI projects fail to deliver their intended value. That is twice as much as comparable IT projects without AI. (RAND is one of the oldest and most respected American research institutes — independent analysis for government, defense and healthcare; not a company selling AI.)
- S&P Global (2025) — 42% of companies abandoned most of their AI initiatives, after only 17% did so a year earlier. Nearly half of pilots never reach production.
- McKinsey (late 2025) — over 80% of organizations see no real impact of AI on their operating result, despite having deployed it.
Different methods, different research firms, different countries — and the same result: an AI deployment that genuinely pays off is a rarity today, not the norm. If yours did not work out, you are in the company of most of the market.
Why they fail — and this is the most important part
Here begins the part most conversations about AI skip. Since so many deployments fail, the natural conclusion is: “the models are still too weak, let’s wait for better ones.” That is not true — and all the serious research says the same.
RAND, after talking to experienced engineers, points to the causes: no clear definition of success, weak data foundations, integration gaps, and chasing the latest technology instead of the business outcome — and underneath it all, fading commitment from leadership. MIT calls it a “learning gap”: the problem is not the model’s quality, but how the company deploys it and whether the system even learns its processes.
In other words: AI deployments fail for organizational reasons, not technical ones. The model usually works. What fails is everything around it — no structure, no oversight, no clear goal, no data discipline. That is good news, though it sounds blunt: organizational things can be fixed; you don’t have to wait for a technological miracle.
There is one more figure from the MIT study that says plainly where the difference lies: deployments based on partnership with a specialized vendor succeed about 67% of the time, while those built internally “from scratch” succeed three times less often. It is not that an in-house team is worse. It is that deploying AI requires structure and experience that cannot be conjured up with a first pilot.
There is also a market mechanism here that is rarely mentioned. Around every hot technology comes the temptation to sell it — and to buy it — like a ready-made product off the shelf, because that is the quickest and easiest. But in most applications AI is not an off-the-shelf service for everyone. The same tool that works beautifully in one company will not work in another unless it is tailored to its data, its processes and its reality. MIT put it plainly: generic tools (like popular chatbots) shine in the hands of an individual thanks to their flexibility, but they get stuck in a company because they do not learn its processes. Whoever counts on a quick profit from a trendy technology — whether by buying it without the work of tailoring, or selling it as a finished product — hits the same wall. This is not a flaw of AI. It is a misunderstanding of its nature.
Where it fails most — and one loud example
The share of failed deployments varies by sector. A 2025 RAND analysis shows that the most trouble is in financial services, followed closely by healthcare — in the latter almost 79% of projects never move beyond the pilot phase (a figure corroborated independently by industry analyses). Interestingly, in medicine the cause is again not “technical”: only about 22% of clinical-AI projects involve doctors from the start, and validation “on the side” — running the model in parallel with the real process before it touches a patient — turns out to be the most effective way of catching errors. Again: structure and oversight, not the model itself.
The loudest documented example of this mechanism is Zillow — an American real-estate giant. Its algorithm valued homes and made purchase offers on that basis. The model was good — but it was trained on a calm market and could not keep up with a sharp change in prices; for a while it kept overvaluing homes while the market was already cooling. The result: over 500 million dollars in losses, the shutdown of the entire business line in 2021, and the layoff of a quarter of the company (figures from the company’s reports and a Stanford GSB analysis). This is not a story about a stupid algorithm. It is a story about what happens when nobody watches the limits within which a model is reliable — and that is a problem of oversight, not of technology.
What it looks like on the Polish market
We also gathered what Polish companies actually stumble on — from public statements over the past year. This is not a list of accusations against anyone; these are recurring patterns that, under the pressure of the AI fashion and management expectations, anyone can fall into. The picture is consistent with the global research, with local specifics:
- “AI doesn’t pay off” — the costs of tokens, infrastructure and energy outweigh the benefits.
- Legacy systems and data silos block scaling; a deployment works in one place and goes no further.
- Hallucinations and errors in production — a model can generate falsehoods and needs human oversight; the trouble arises when it reaches production without that oversight.
- The “we have to have AI” pressure — pilots launched for the sake of having them, with no clear goal or measure of value; sometimes too many tools at once, which blocks decisions instead of helping.
- Data protection — pasting client data into chatbots before any processing policy exists.
- A drop in use after the first enthusiasm — a strong start, then silence (in one account the number of users fell from three thousand to a hundred).
One statement sums it up best: “The biggest problem is law, ethics, choosing a vendor; AI has to be under a specialist’s supervision.” This is not the voice of an AI opponent. It is the voice of someone who deployed it and saw what was missing. And that is the right tone for this whole conversation: not “companies are naive,” but “a layer is missing that nobody talked about out loud.”
Honestly: we write this as a player, not a judge
We do not pretend to be a neutral auditor. We are a company that deploys AI — we have a stake in it and we say so plainly. But that is exactly why these numbers matter to us rather than being inconvenient: they reinforce our belief that it is not worth promising “magical autonomy,” but worth starting from understanding why deployments fail, and building so that they don’t fail for those very reasons. We do not say this to elevate ourselves above others — we say how we ourselves approach it.
Since the causes are organizational, we too look for the answer on the side of organization, not the technology itself. We do not claim this is the only right way — it is how we approach the matter, and our other texts are about it:
- No oversight and hallucinations in production → a human in the loop, verification of the result by an independent party, a gate on irreversible actions.
- No clear goal and scope creep → a closed set of tasks, proof of completion instead of “seems to work.”
- Who is responsible when AI gets it wrong → responsibility built into the mechanism, not into a policy (operator consent + a ledger of decisions).
- “Build it yourself or take it ready-made” → the data is clear: a deployment with structure and experience beats improvisation.
Summary
Most AI deployments fail — and not because the technology is too weak, but because it lacks structure, oversight and discipline around it. The numbers from MIT, RAND, S&P and McKinsey say the same, and the Polish market confirms it. If your deployment did not work out, it is not a sign to give up on AI — it is a sign that a layer was missing that few people talk about, because it is harder to sell than another model.
That layer is where we start. See how we put it together — from simple tools to systems for coordinating AI work under human supervision. Because AI that pays off looks less flashy than the promises sound: less magic, more discipline.
MafiaAI — a team of people and AI agents building tools, websites and solutions. We deploy AI so that it pays off — with oversight, proof and honesty about the limits. More: t8.pl