Why 61% of AI Service Projects Miss Year One
The slide always lands the same way. Cost-per-contact drops off a cliff, a hockey-stick curve bends toward the ceiling, and somewhere in the room a VP of support feels the knot in their stomach loosen for the first time in two quarters. The deck is not lying. The numbers under it are real. Then twelve months pass, the renewal meeting gets scheduled, and the same VP is squinting at a dashboard that can't quite say what the slide promised. The model worked. The math didn't show up.
That gap has a number now. McKinsey put it at 61% of AI customer-service projects missing their year-one targets (McKinsey, 2025). Not failing outright. Missing. Which is worse in a way, because a clean failure gets killed and a near-miss gets defended, re-scoped, defended again, and quietly starved.
So the question is the interesting one. The economics are not in dispute. A human-handled ticket runs eight to twelve dollars; an AI-handled one runs fifty cents to a dollar and change, a twelve-to-twenty-four-times spread (Gartner / Forrester, 2025). Support costs fall around 30% on average across deployments (IBM, 2025), and the best programs are pulling 53% out (McKinsey, 2025). The ceiling is real even for the projects that underperform. So why do most of them underperform?
The model was never the problem
Walk into the post-mortem expecting to hear about hallucinations or a weak vendor, and you'll leave confused. The model usually did its job. AI handles roughly 80% of routine inquiries when it's pointed at the right ones (IBM, 2025), and the best-in-class programs deflect 62% of total volume (Forrester Wave, 2025). The capability is there, on the shelf, priced and proven.
What's missing happens before a single line of automation ships. Two unglamorous things, both on the front end, both easy to skip when everyone's excited about the demo.
The first is the baseline. The second is the scope.
You can't bank a saving you never measured
Here's the trap that swallows the most projects. A team rolls out AI deflection, it works, tickets drop, and then renewal day arrives and finance asks the only question that matters: what did we save? The honest answer is a shrug. Nobody wrote down the cost-per-contact before the build, so there's no before to measure the after against. The savings are real and completely unprovable, which in a budget review is the same as not existing.
Cost-per-contact isn't a vanity metric. It's the unit the whole case is denominated in. Volume times cost-per-ticket times deflection rate is the dollar figure, and if you never pinned the first two numbers down, the third one is just a percentage floating in space. A program that captured its baseline can walk in and say tickets went from $9.40 to $1.10. A program that didn't can only say things feel faster. One of those gets renewed.
The discipline is boring and that's exactly why it gets cut. Spend the first two weeks instrumenting the current state, by channel, by ticket type, by hour. Phone is the expensive channel, so it's the one worth measuring hardest. The teams that hit their numbers treat the baseline as the deliverable that protects every other deliverable. The teams that miss treat it as paperwork and skip to the fun part.
Scope to the slice that deflects clean
The second failure is greedier. A team looks at the full queue, sees a mountain of tickets, and points the AI at all of it. Then the messy 20% goes to work on the average.
Most support volume is routine and genuinely deflectable. Password resets, balance checks, where's-my-order, the same five tenant questions all day. That slice deflects cleanly, cheaply, and at high confidence. The other slice is the angry edge case, the multi-part complaint, the thing that needs a human and a little grace. Aim automation at that, and it fails in public, drags the satisfaction score down, and hands the skeptics in the room their told-you-so.
The programs in the top quartile do something that looks almost timid. They scope narrow. They take the cleanest, highest-volume, lowest-risk inquiries first, prove the deflection, bank the measured saving, then widen. Deflection rates climb from a low baseline into the high double digits precisely because the early wins were chosen, not inherited. Restraint reads as weakness on a roadmap and shows up as strength on a renewal.
The two weeks that decide the year
Strip it down and the pattern is almost rude in its simplicity. Measure first so the saving is provable. Scope to the deflectable slice so the metrics don't get poisoned by the cases that were never going to automate. Do those two things and you're in the 39% that hit. Skip them for the demo and you're in the majority, defending a near-miss that started losing the day the baseline went unwritten.
The model will keep getting better. It was never the bottleneck. The bottleneck is two weeks of unglamorous work that nobody applauds and everybody needs. The slide promises the ceiling. The front end decides whether you ever reach it.
Sources: McKinsey 2025 (61% of AI customer-service projects miss year-one targets; top-quartile programs cut support costs 53%); IBM 2025 (support costs fall ~30% on average; AI handles ~80% of routine inquiries); Gartner / Forrester 2025 ($8-12 human-handled ticket vs $0.50-1.05 AI-handled, a 12-24x gap); Forrester Wave 2025 (best-in-class deflection reaches 62%).
