Short answer. MIT’s 2025 study of 300 enterprise deployments found 95% of generative AI pilots delivered zero measurable profit-and-loss impact. The transformations fail on execution, not technology: siloed data, no owner accountable for production, and pilots scoped to impress a demo instead of run in a workflow. The ones that ship treat AI as an operational system with a named owner, governed data, and a production target from day one. At Ableneo, roughly 4 of 5 projects reach production because that discipline is built in before the first model is trained.
The gap between AI ambition and AI in production is now measured, not guessed. MIT’s State of AI in Business 2025 report, based on 52 executive interviews and 300 public deployments, put the pilot failure rate at 95%. S&P Global’s 2025 enterprise survey found 42% of companies had scrapped most of their AI initiatives, up from 17% a year earlier. RAND research places general AI project failure near 80%.
These are not model-accuracy problems. The models work in the notebook. They fail on the way to the workflow. Three patterns repeat across the wreckage:
Ableneo shipped 34 production AI projects in 2025 across 20 clients in four countries. The delivery rate, roughly 4 of 5 projects reaching production, comes from refusing to start a build without a named owner and a governed data source. That single rule removes the two most common causes of failure before code exists.
In banking and insurance the cost of a stalled transformation is now regulatory as well as commercial. From 2 August 2026 the EU AI Act’s high-risk obligations apply to AI used in credit scoring and creditworthiness assessment, listed in Annex III. Providers and deployers must run a risk management system, govern the training data, keep technical documentation, guarantee human oversight, and register the system in the EU database. Penalties reach 15 million euro or 3% of global annual turnover.
A pilot that never reached production leaves a bank exposed twice: it burned the budget and it built no governed, documented, auditable system to show a supervisor. Under DORA, the same institution must also prove operational resilience for the ICT that runs those models. A transformation that ships with governance built in produces the audit evidence as a byproduct. A transformation that fails produces a slide deck and a compliance gap.
MIT found 95% of enterprise generative AI pilots deliver zero measurable P&L impact, and 42% of companies have scrapped most AI initiatives.
The causes are organizational, not technical, and they cluster. The OECD’s 2025 work on AI adoption names skills and data as the two hardest barriers: half of firms report employees lack the skills to use AI, and data sits siloed across heterogeneous IT systems that were never built for integration. Layer on three execution failures and the 95% number stops being surprising: budgets aimed at visible functions like marketing rather than the operations and finance workflows where ROI is higher, external tools abandoned in favour of internal builds that succeed half as often, and a pilot scope written to demo rather than to run. None of these is a machine learning problem. Every one is a decision made before the model was built.
The ones that reach production share a shape. First, a single accountable owner for the running system, not a committee. Second, a narrow first use case inside a real workflow where the output changes a decision, not a dashboard nobody reads. Third, governed data with a defined source, lineage, and refresh, decided before the build. Fourth, a production definition of done: monitoring, retraining triggers, and human oversight specified up front, not bolted on. MIT’s data shows tools built with an external partner reach value roughly twice as often as pure internal builds, because the partner brings the production discipline the internal team is learning for the first time. Shipping is a design choice made at the start, not a rescue attempted at the end.
Three questions predict the outcome early. Who is on call when this model fails in production, and is that person named today? What is the exact data source, and is it governed or is it a one-time extract? Which decision in which workflow changes when the model is right? A pilot that cannot answer all three by day 90 is on the 95% path, regardless of how good the accuracy metric looks. A pilot that answers all three has already done the hard part, because the technical build is rarely where transformations die.
MIT found external-partner deployments reach value about twice as often as internal builds. The reason is not talent, it is repetition. A team building its first production AI system meets every failure mode for the first time: data governance, drift monitoring, oversight design, incident response. A partner that has shipped dozens of systems has met them before and designs around them. For a regulated institution the calculation is sharper, because the partner also brings the audit and documentation patterns the EU AI Act now demands. Build when the capability is core and you have production experience in house. Partner when the goal is to reach production fast with governance evidence intact.
Ableneo architects AI as an operational system, governed, observable, and accountable, from the first workshop. That is why roughly 4 of 5 of our projects reach production while the market average scraps nearly half before launch. Across 34 production projects in 2025, 94% using large language models, the pattern holds: name the owner, govern the data, define production before you build. The same discipline that ships a project also produces the evidence a supervisor asks for. For the operational side of the same problem, see our detailed guide on how to get a stalled AI pilot into production.
Key takeaways
Planning AI in a regulated business? Ableneo takes systems from classification to governed production.