AI Strategy / Founder Note

Why 95% of AI pilots fail (and what the 5% do differently)

The best personal trainer in the world is now free, instant, and endlessly patient. Gyms are still empty in March. That gap is currently costing enterprises tens of billions of dollars in AI spend.

A personal note first

For the last twelve months I have been on my own fitness journey. Nothing dramatic, no transformation photos. Just the ordinary work of showing up several times a week, tracking what I eat, and watching numbers that move slowly.

What surprised me was not that it was hard. It was that the hard part was not the part I expected. I had access to more good information than any generation before me, and none of it was the constraint. What actually changed things was picking one specific outcome, cutting down the number of things I was trying to do at once, and building a rhythm I did not have to renegotiate with myself every morning.

Somewhere around month four it occurred to me that I was describing, almost word for word, the problem I watch enterprises struggle with every week in their AI programs. This piece is what came out of that.

The best personal trainer in the world is now free

Open ChatGPT, Claude, or Grok and ask for a twelve week program to lose twenty pounds and add strength. You will get something extraordinary in about nine seconds.

It will periodize your training across mesocycles. It will set your protein at 1.6 to 2.2 grams per kilogram of bodyweight and explain the research behind the range. It will program progressive overload with a deload every fourth week, warn you that the scale will move sideways for the first ten days as glycogen and water shift, tell you which knee-friendly substitutions to use for barbell back squats, and generate a grocery list, a sleep protocol, and a plan for what to do at a wedding buffet.

Ask a follow up question and it adjusts everything. Ask it to explain the physiology like you are twelve. Ask it at 2am.

This is genuinely better than what most people could buy from a certified trainer at two hundred dollars an hour in 2015. It is free, instant, patient, and it never gets bored of you.

And yet: gyms are not full of lean, strong people. Most of the people who downloaded that perfect plan will not be following it in March.

Information was never the constraint

This is the part worth sitting with, because it is the part everyone skips.

The fitness industry did not have an information problem in 2015 either. Every principle in that AI-generated program was already published, already free, already sitting in a thousand videos and a hundred books. What large language models did was compress and personalize the advice. They made it faster to get, easier to understand, and specific to you.

What they did not do, and cannot do, is close the gap between a recommendation and a result.

That gap has a name in behavioral science. It is the intention to action gap, and it is not made of ignorance. It is made of Tuesday. It is made of a work crisis, a sick kid, rain, a late meeting, a bad night of sleep, and the entirely reasonable decision to skip today and start properly on Monday.

No amount of additional prescription closes it. A better plan does not make you go. A more detailed plan sometimes makes it worse, because researching the plan feels like progress, and it costs nothing.

What actually separates the people who get results

Look at anyone who genuinely transformed their fitness and you will find the same four things, none of which are informational.

INTENT

A reason, not a wish

Something specific enough to survive a bad week. “Get in shape” fails. “Carry my kid up two flights of stairs without stopping, by my fortieth birthday” survives.

FOCUS

Three things, not thirty

They did not simultaneously start keto, CrossFit, cold plunges, a run streak, and intermittent fasting. They picked a few levers and left the rest alone until the first ones were automatic.

CADENCE

Showing up on the bad days

Not perfectly. Just more often than not, for long enough that the compounding worked.

DOING

Walking the walk

Every serious lifter knows a guy who can explain the entire literature on hypertrophy and has not changed shape in four years. Talking about the program is not the program.

Why you hire a coach at the plateau

Here is the interesting turn. The people most likely to hire a human coach are not beginners. They are the people who already know a lot and have stopped progressing.

Why would you pay someone to tell you things you already know?

Because that is not what you are paying for. A good coach at that stage supplies four things a chat interface structurally cannot:

  • Accountability. Someone is expecting you at 6am and will notice if you are not there.
  • Cadence. A fixed rhythm of check-ins, measurements, and adjustments that you do not have to generate willpower to sustain.
  • Judgment about tradeoffs. Not “here are twelve options,” but “for you, this quarter, do these two and ignore the rest.”
  • Enforced showing up. The highest-leverage intervention in all of fitness is attendance, and attendance responds to obligation.

This is not a soft claim. Dr. Gail Matthews at Dominican University of California ran a study with 267 participants across five conditions, from simply thinking about a goal up to writing the goal, committing to specific actions, sharing them with a supportive person, and sending that person a weekly progress report. The group that added the weekly progress report achieved substantially more than every other group, including the group that had written goals and an accountability partner but no weekly cadence. The active ingredient was not the plan. It was the reporting rhythm. (Matthews, Dominican University)

Weekly check-ins. Written commitments. A named owner. If that sounds less like a gym and more like a program management office, you are already ahead of the rest of this article.

Now replace “fitness” with “enterprise AI”

Every dynamic above is currently playing out inside large companies at a cost of tens of billions of dollars.

MIT’s NANDA initiative published The GenAI Divide: State of AI in Business 2025, based on a review of roughly 300 publicly disclosed AI initiatives, 52 organizational interviews, and surveys of senior leaders. The headline finding: despite thirty to forty billion dollars of enterprise investment, about 95 percent of generative AI pilots produced no measurable impact on the P&L. Only around 5 percent crossed into production with real value. (The Hill, Forbes)

Read the diagnosis rather than the headline and it is entirely familiar. The researchers did not attribute the failure to model quality or regulation. They attributed it to what they called a learning gap: brittle workflows, tools that do not adapt to how the organization actually works, and misalignment with day to day operations. Budgets skewed heavily toward the visible functions, sales and marketing, while the clearer returns sat in operations and back office work. And externally built, vendor-partnered deployments succeeded at roughly twice the rate of internal builds.

In fitness terms: everyone downloaded the plan, nobody ran the program, and the people who hired a coach did about twice as well.

One honest caveat

Credibility matters more than a scary number. The MIT study has been criticized for a narrow definition of success, essentially requiring deployment past pilot with measurable KPIs and ROI visible six months out. (Marketing AI Institute)

The precise percentage is debatable. The mechanism is not. Almost nobody who has run an enterprise AI program in the last three years disputes that the models are the easy part.

The three things that convert AI capability into AI ROI

1. Intent: name the number before you name the tool

Most AI initiatives begin with a technology decision and reverse-engineer a business case. That is backwards, and it is why so many end as impressive demos with no owner.

Intent means a single sentence a CFO would recognize: reduce executive escalations by 30 percent this fiscal year, cut average time to resolution on P1 cases by a third, deflect 10 percent of inbound volume through better knowledge retrieval. A number, a timeframe, and a named executive whose bonus is affected.

“Explore GenAI use cases” is the corporate equivalent of “get in shape.”

2. Focus: one workflow, one metric, one team

The fastest way to guarantee no measurable ROI is to attempt eleven use cases simultaneously so that no single one gets enough attention to work.

8×8 did the opposite. They started with a thirty day pilot applying sentiment analysis to a defined group of high-touch, high-value customers rather than the entire book of business. (8×8 case study) One signal, one segment, one question: can we see trouble before the customer escalates? That constraint is what made the result legible.

Narrow scope is not a lack of ambition. It is how you generate the first proof point that funds the next three.

3. Program management: the cadence is the product

This is the part nobody puts on a slide, and it is the part that decides everything.

A real AI program has a named owner who is not the CIO, a weekly operating review, a defined baseline captured before go-live, a scorecard with leading and lagging indicators, a set of internal champions inside the frontline teams, and an explicit decision cadence for what gets expanded, changed, or killed.

It looks unglamorous. It looks like the coach who texts you on Sunday night asking for last week’s numbers. That is exactly the point.

What this looks like when it actually works

The following are not vendor selection stories. Every one of these companies bought roughly the same category of technology that thousands of other companies also bought. The difference is in how they ran it.

Salesforce Sequenced rollout

Prove one signal, then expand

Salesforce began with a focused problem: long-running cases where nobody could tell in advance which ones would turn ugly. They anchored on customer sentiment as the leading indicator, and the turning point was an observed pattern where a sharp rise in negative sentiment preceded a drop in weekly CSAT by about two weeks.

56%fewer escalations
1 hrreturned per manager per day
60swarm leads covered

The program detail that matters: they did not start company-wide. They proved the pattern, rolled out to the entire customer-facing organization in Spring 2023, and only then expanded into quality assurance and automated coaching. Sequence, prove, expand. (Salesforce case study)

Informatica Adoption by design

Champions on the frontline, not a training program

Informatica went live in 45 days. The speed is the headline. The adoption design is the lesson.

45 daysto live
18%fewer escalations
8%lower case volume

Informatica identified internal champions who introduced use cases to their peers and fielded questions from the rest of the team, which is why the platform landed with, in their description, no learning curve. Their support managers were carrying 125 to 150 open cases at any moment, so anything requiring a formal training program would have died on contact.

They also had clear intent. They already understood the technical drivers of escalation. What they lacked was the sentiment side, and combining the two gave them a defensible picture of why cases escalated rather than just that they did. (Informatica case study)

NICE Workflow embedding

Put the AI where the work already happens

NICE started narrow, deploying Resolve SX with Precision RAG in the customer portal, against a knowledge estate fragmented across Salesforce, portals, and documentation inherited through acquisitions.

98%search accuracy
4%lower monthly case volume
20-30%of cases are internal users

Then came the program work. They expanded to internal knowledge sources. They embedded the same generative search inside the Salesforce internal portal so agents got identical answers in the tool they already lived in. They integrated with Microsoft Teams so engineers could get a referenced answer without switching systems. And they used the Assist module to auto-draft knowledge articles from resolved cases, with humans reviewing and publishing.

The knowledge base finally stopped decaying, because knowledge creation was no longer competing with case closure for an engineer’s time. (NICE case study, Resolve SX)

Basware Signal into workflow

Pair the insight with the thing that acts on it

Look at what Basware combined. Sentiment analysis and escalation prediction told them which cases were in trouble. Intelligent Case Assignment then routed those cases to the right skills, which mattered because self-service was absorbing the routine issues and leaving agents with unfamiliar, harder problems.

80%lower escalation rate
30%faster resolution
100%sentiment visibility

Signal without a workflow is a dashboard. Signal wired into routing is an outcome. Basware also used the program to build a standing bridge between support, customer success, and product, so that what the signals surfaced actually reached the people who could fix the underlying cause. (Basware case study)

And the shorter versions

Nutanix reduced escalations and backlog by 40 percent while protecting an NPS that had averaged above 90 for seven consecutive years. Databricks cut SLA misses by 40 percent and lifted CSAT. Demandbase built a daily practice around reviewing sentiment alerts and acting on negative trends, which is possibly the most underrated sentence in our entire customer library, because a daily practice is precisely what a coach installs. (All customer stories)

The pattern behind the numbers

Put them side by side and the common thread has nothing to do with model architecture.

CompanyResultThe program decision behind it
Salesforce56% fewer escalationsProved one signal, then sequenced the expansion into QA and coaching
InformaticaLive in 45 daysInternal champions drove adoption from inside the frontline teams
NICE98% search accuracyEmbedded AI into Salesforce and Teams where work already happened
Basware80% lower escalationsWired the signal directly into skill-based routing
8×8Reactive to predictiveScoped a 30 day pilot to one high-value customer segment

Every one of them narrowed the scope, named the outcome, assigned owners, and built a rhythm. None of them succeeded because they had access to a model their competitors did not have.

A 90 day program blueprint

If you want the coach’s version, here it is. This is roughly how our most successful deployments actually run.

DAY 0-15

Baseline and intent

Capture the current numbers before anything changes, because a program without a baseline can never prove ROI. Name one primary metric and at most two secondary ones. Assign a business owner in the support organization, not in IT. Identify three to five frontline champions.

DAY 15-45

Connect and observe

Get the platform live on your data. SupportLogic is typically live in 45 days, extracting signals across every interaction rather than the 2 percent that manual review reaches. Do not change any workflow yet. Watch what the signals say and validate them against cases your leaders already know were painful. This builds the trust you will need later.

DAY 45-75

One workflow, instrumented

Pick a single workflow and wire the signal into it. Escalation prediction into a daily manager review. Sentiment alerts into the account team’s Slack channel. Prioritization into the queue. One workflow, measured weekly against the baseline.

DAY 75-90

Review, prune, expand

Hold a formal review against the scorecard. Kill what did not move the number. Expand what did. Then choose the second workflow, and only the second.

ONGOING

The weekly rhythm

Fifteen minutes, same time, same people, same scorecard. This is the entire secret and it costs nothing.

What to measure

Split your scorecard the way a coach splits body composition from the training log.

Leading / weekly

  • Signal coverage across interactions
  • Alert acknowledgement rate
  • Flagged cases actioned within SLA
  • Active user rate among frontline managers
  • Sentiment trend by account tier

Lagging / monthly

  • Escalation rate
  • Time to resolution
  • Case backlog age
  • CSAT and NPS
  • Renewal and expansion on flagged accounts

If your leading indicators are flat, your program has an adoption problem, not a technology problem. That is the AI equivalent of stepping on the scale while never stepping into the gym.

The uncomfortable conclusion

We build AI infrastructure for CX. Autonomous agents, a Snowflake-backed data cloud, an MCP server, a CRM-less architecture. We think it is very good, and the results above suggest our customers agree.

And we will tell you plainly: buying it will not, by itself, produce ROI.

The models are commoditizing. Frontier capability is available to your competitors at the same price it is available to you, and the gap between the best model and the fourth-best model narrows every quarter. Whatever advantage exists in 2026 does not live in the model. It lives in what an organization is willing to do consistently.

The equipment is not the fitness. The program is the fitness. The companies pulling real returns out of AI are not the ones with better technology. They are the ones who named a number, narrowed the scope, assigned an owner, and showed up every week.

Nobody ever got in shape by reading about squats.

Ready to build the program?

The technology takes 45 days. The operating rhythm is what we help you keep.

Frequently asked questions

Why do most enterprise AI projects fail to deliver ROI?

Research from MIT’s NANDA initiative found that roughly 95 percent of generative AI pilots produced no measurable P&L impact, and attributed the failure not to model quality but to brittle workflows, weak integration with daily operations, and organizational adoption gaps. In practice, projects fail because no one owns a specific number, the scope is too broad to measure, and there is no operating cadence to sustain usage after launch.

What is the difference between AI capability and AI ROI?

Capability is what the technology can do. ROI is what your organization actually does with it, repeatedly, at scale, measured against a baseline. The gap between the two is closed by intent, focus, and program management, not by a better model.

How long should an enterprise AI deployment take before showing results?

Faster than most companies assume. Informatica was live on SupportLogic in 45 days and saw an 18 percent escalation reduction and an 8 percent case volume reduction in that window. Mid-market organizations tend to move from pilot to implementation in around 90 days, while large enterprises often take nine months or more, which is itself a program management problem rather than a technical one.

Who should own an AI program in a support organization?

A business owner inside the support organization whose performance metrics are directly affected, supported by frontline champions. IT should own integration and governance. If the only owner is technical, adoption almost always stalls after launch.

What should I measure to prove AI ROI in customer support?

Track leading indicators weekly (signal coverage, alert action rate, active users among managers) and lagging indicators monthly or quarterly (escalation rate, time to resolution, backlog age, CSAT, retention). Capture all of them before go-live, because without a baseline you cannot attribute any improvement.

Is it better to build AI internally or buy from a specialized vendor?

The MIT research found externally partnered deployments succeeded at roughly twice the rate of internal builds, largely because vendors bring workflow fit and adoption experience alongside the technology. Informatica estimated that building equivalent functionality internally would have taken eight or nine months instead of the two it took to see results.

Don’t miss out

Want the latest B2B Support, AI and ML blogs delivered straight to your inbox?

Subscription Form