AI customer service ROI: what Gartner's 432 use cases actually show
Gartner analysed 432 AI use cases in customer service and found that only a quarter produced positive ROI. The more useful number is the one underneath it: 42% could not be scored at all.
TL;DR: Gartner's 432-use-case analysis found 25% of AI customer service use cases produce positive ROI and 25% produce negative returns — but the largest bucket, 42%, is "unclear," which is a measurement failure, not a technology failure.
Methodology & sources
- Gartner's four-way ROI split and budget figures, as reported by CX Dive (17 August 2026) and CX Dive (11 September 2026). Gartner has not published the underlying use-case list.
- Gartner's 2029 rehiring prediction from its own press release (9 September 2026).
- Aissist data: first-party production statistics from 288,866 conversations across 17 organisations (Aug 23–Sep 8, 2026).
- Vendor pricing: published rate cards compiled on Aissist's cost benchmark of AI service, plus Intercom's live pricing page.
- All figures verified September 2026. Disclosure: this is Aissist's blog. Aissist.io sells an AI operational layer for customer service and sales, and the argument below is one we have a commercial interest in.

This article is about what the ROI data says. For how to build the measurement itself, see how to measure the true ROI of agentic AI. For the organisational causes behind the failures, see why AI projects fail to deliver ROI.
How many AI customer service use cases actually produce ROI?
One in four. Gartner's analysis of 432 AI use cases in customer service found 25% delivered positive ROI, 25% delivered negative returns, 42% produced unclear value, and 11% broke even, as reported by CX Dive in August 2026. That is the whole population, not a selected sample of failures.
| ROI outcome | Share of use cases (Gartner, Aug 2026) | Implied count of 432 |
|---|---|---|
| Positive ROI | 25% | ~108 |
| Negative ROI | 25% | ~108 |
| Unclear ROI | 42% | ~181 |
| Break-even | 11% | ~48 |
As reported, those four figures sum to 103%, so they are rounded and the implied counts are approximate — which is itself a small lesson in the precision available on this subject.
The exhaustive split matters more than the adverse half of it. "A quarter of AI deployments lose money" is a statistic anyone can wield. "A quarter lose money, a quarter pay back, and 42% cannot be scored either way" is a description of a market, and it points at a different problem. Three-quarters of these use cases are not failures. They are unproven.
The spending sits on the other side of that uncertainty. Service and support functions spent a median of almost $6 million on AI in 2025, committing 13% of functional budget to it, and more than 75% of leaders plan to increase AI investment in 2026, Gartner told CX Dive in August 2026. Budgets are being raised against a scorecard that is 42% blank.
"Too many AI rollouts begin with pressure to demonstrate a credible AI strategy to the board, rather than with a clearly defined business problem." — Julie Geller, Principal Research Director, Info-Tech Research Group
What separates the use cases that lose money from the ones that pay back?
The definition of success chosen before anything was built. A deployment optimised to close conversations without a human books its savings once and its costs somewhere else; a deployment optimised to solve the issue end to end books both in the same column. The technology in the two cases is frequently identical, which is why "does AI work in customer service" is the wrong question.
The industry has two different metrics wearing one word. Deflection counts conversations that never reached a human — including the abandoned and the unresolved. Resolution counts issues actually solved. Aissist's breakdown of deflection vs resolution rate walks through how eight vendors define and bill on each; the short version is that they are not interchangeable and the gap between them is where negative ROI lives.
| Deflection-shaped deployment | Resolution-shaped deployment | |
|---|---|---|
| Optimised for | Closing the conversation without a human | Solving the issue end to end |
| Counted as success | Containment; no handover | Problem fixed, customer confirms |
| First metric to move | Deflection rate up | Resolution rate up |
| Where cost reappears | Second contact, longer human handle time, CSAT drop | Escalation rate — visible and budgeted |
| Visible in the AI line item? | No. It lands in headcount and CSAT | Yes |
The asymmetry in the last row is the whole mechanism. A deflection metric has never once reported bad news about itself. The second contact arrives days later, is logged as a new ticket, is handled by a human, and is attributed to nothing.
"Containment is also too often mistaken for success. Delaying contact with a human agent is not the same as resolving the customer's problem. The real test is much simpler: did the customer get what they needed, with less effort?" — Julie Geller, Principal Research Director, Info-Tech Research Group
Klarna is the named, dated version of this. In February 2024, Klarna and OpenAI published that Klarna's AI assistant had handled 2.3 million conversations in its first month, two-thirds of Klarna's customer service chats, doing "the work of 700 full-time agents," with resolution time down from 11 minutes to under 2 and an estimated $40 million profit improvement for 2024. In May 2025, more than a year later, Klarna began recruiting human customer service representatives again, with CEO Sebastian Siemiatkowski saying the company had underestimated the tradeoffs.
"I just think it's so critical that you are clear to your customer that there will always a human if you want." — Sebastian Siemiatkowski, CEO, Klarna, quoted by CX Dive (May 2025)
Note what did not happen: Klarna did not switch the AI off. Its assistant still handled two-thirds of inquiries at the time of that reversal. The volume claim held. The staffing conclusion drawn from it did not.
What does a negative-ROI AI deployment actually cost?
A negative-ROI AI deployment is a customer service AI use case whose total cost — licence fees, integration and maintenance labour, and the downstream cost of contacts it closed without solving — exceeds the labour and handle time it removed, measured over the same period. The third term is the one that is almost never measured, and it is usually the one that flips the sign.
Start with the portfolio. Gartner reports service teams running an average of nearly five AI use cases and spending a median of almost $6 million a year on them — call it $1.2 million per use case. Apply the four-way split to five use cases and the median organisation has roughly one that pays back, one that loses money, and two it cannot score. In money: about $1.5 million a year on negative returns, about $2.5 million on no verdict. Gartner reports the split by use-case count rather than by spend, so the even spread is our assumption, not theirs.
Now the unit economics, which is where a deflection-shaped deployment quietly overcharges. Aissist's own production data across 288,866 conversations (first-party, measured) shows 63.1% resolved on the inclusive definition — a complete answer was given, the customer never confirmed — but only 38.2% resolved strictly end to end, where every escalation and transfer counts as unresolved. Bill against the looser number, get the tighter one, and the effective cost per genuinely resolved issue is 1.65× the rate card.
The table below applies that gap to $0.99 per resolution, the most common published rate in the market — Intercom's pricing page lists $0.99 USD per Fin outcome on every plan and standalone with no seats required (read September 2026), though the page does not define what makes an outcome billable. The resolution rates are Aissist's own measured production figures, not Fin's; this is an illustration of the arithmetic, not a performance claim about any vendor.
| Conversations / month | Billed at 63.1% | Annual bill @ $0.99 | Genuinely resolved at 38.2% | Effective cost per genuine resolution | Annual spend on conversations billed but not solved |
|---|---|---|---|---|---|
| 10,000 | 6,310 | $74,963 | 3,820 | $1.64 | $29,581 |
| 50,000 | 31,550 | $374,814 | 19,100 | $1.64 | $147,906 |
| 200,000 | 126,200 | $1,499,256 | 76,400 | $1.64 | $591,624 |
The same arithmetic applies to Aissist, and we would rather write it down than have a buyer discover it. Aissist bills $0.09 USD per AI interaction (vendor-published rate card, read September 2026). At the four-interactions-per-session default on Aissist's cost benchmark page, that is $0.36 per session — $0.57 per inclusively-resolved conversation, and $0.94 per strictly-resolved one. The same 1.65× markup, on our own numbers.
None of the figures above include the second contact. At 10,000 conversations a month, 2,490 conversations a month are billed as handled and are not finished; where those return, they consume an agent, and that cost lands in headcount rather than in the AI line. Payback requires that removed handle time exceed the rate card plus the returning contacts — which is why measuring the true ROI of agentic AI has to start on the resolution side of the ledger.
"There's an outstanding question about whether AI prices are going to rise, and in the meantime, you still have to do all of this work from the people that you have replaced. I would say that paints a picture where you are going to have to rehire people to do the same work that you attempted to eliminate." — Emily Potosky, Senior Director Analyst, Gartner, quoted by CX Dive (September 2026)

Why can't 42% of deployments say whether they worked?
Because nobody instrumented the outcome. The 42% "unclear" bucket is the largest in Gartner's split, and an organisation lands there by measuring what the AI did — conversations handled, deflection rate, response time — without measuring what happened to the customer afterwards. Nobody files a win under "unclear."
This is the first-hand part. When Aissist measured 288,866 production conversations across 17 customer organisations between 23 August and 8 September 2026, genuine resolution across ten of those organisations ranged from 13.9% to 77.9% under identical measurement criteria. Same platform, same definition, a 64-point spread. That range is not a technology variable; it is a function of how much of each customer's system the AI was actually connected to and what escalation policy sat behind it. An organisation that reports one blended number for a portfolio like that has not measured anything — which is precisely how a deployment ends up unscored rather than scored badly.
The measurement gap runs market-wide. Aissist's AI customer service benchmark puts the verified cross-programme median for genuine end-to-end resolution at about 41%, with the top quartile near 59%, and notes that vendor marketing claims typically run 20 to 40 percentage points above independently verified figures. Aissist's own headline figures — 83% average resolution and 4.8 CSAT on the resolution–CSAT tradeoff page — are vendor-claimed and sit well above that verified median. Both numbers are ours. Only one of them is measured the strict way.
"AI does not follow one cost curve, and it does not produce one uniform type of value. CFOs need to stop looking for a single ROI formula and instead build a balanced portfolio." — Twisha Sharma, Senior Principal, Research, Gartner Finance Practice, Gartner press release (March 2026)
The uncomfortable implication for buyers is that the 42% is mostly recoverable and the 25% mostly is not. An unmeasured deployment can be instrumented in a quarter. A deployment that has been running on deflection for two years has already spent the money.
What has to be true for an AI service deployment to pay back?
An AI service deployment pays back when three conditions hold: the vendor's definition of a billable resolution is written into the contract, repeat contacts are measured within seven days, and escalation is counted rather than suppressed. All three are checkable before signing. All three are procurement questions, not engineering ones.
- If your vendor bills per "resolution," get the definition in the contract, because the billing unit and the outcome are not the same object. Ask specifically whether a conversation counts as resolved when the customer does not confirm, and whether a handover to a human refunds the charge.
- If you cannot report repeat-contact rate within 7 days, do not report ROI, because the cost of a deflected contact appears on that timeline. This single metric moves most of the 42% out of "unclear."
- If your deflection rate is rising while CSAT is flat or falling, treat it as a cost increase, not a saving — Aissist's resolution–CSAT tradeoff analysis puts the workable range at roughly 60–80% genuine resolution, above which satisfaction bought through suppressed handoff starts to collapse.
- If the business case rests on headcount reduction, price the rehire. Gartner predicts that "by 2029, 30% of employees laid off due to replacement by AI will need to be rehired, often at a significantly higher cost" (Gartner press release, September 2026), and Emily Potosky of Gartner has said contact centre staff were "hit first and maybe hardest by the AI hype."
The common failure is starting from the board rather than from a queue, a pattern this market has now repeated often enough to have a literature — Aissist's own account of why AI projects fail to deliver ROI covers the organisational half of it.
"What we're seeing with the deployment, the backstory here is like everyone is trying to get AI. This is a top-down initiative: We need AI and customer support and customer experience." — Antoine Nasr, Head of AI, Forethought AI Agents by Zendesk, quoted by CX Dive (August 2026)
Negative ROI is a measurement outcome before it is a technology outcome
Gartner's 432 use cases do not show that AI fails in customer service. They show that 25% of deployments were built to close conversations and were charged for it, that 42% were never instrumented well enough to know either way, and that a quarter worked. The variable separating the buckets is the definition of "resolved" chosen before anything was built, and that definition is set in procurement, not in the model.
For a buyer, the practical consequence is narrow. Ask the vendor what triggers a charge. Measure repeat contacts at seven days. Refuse to accept a deflection rate as a savings figure. Do those three things and the deployment can still fail — but it will fail visibly, in a quarter, for a few thousand dollars, rather than invisibly for two years at $1.5 million a year. On Gartner's numbers, 42% of this market currently cannot tell the difference.
Measuring resolution rather than deflection is the whole argument. Pulse™ reads every ticket and scores what actually happened to the customer, which is what moves a deployment out of the "unclear" bucket. Book a consultation →
Frequently asked questions
What percentage of AI customer service projects have negative ROI?
Twenty-five percent. Gartner's analysis of 432 AI customer service use cases found 25% delivered negative returns, alongside 25% positive, 42% unclear and 11% break-even, as reported by CX Dive in August 2026. The four figures are rounded and sum to 103%.
Is negative ROI in AI customer service caused by the technology?
Rarely. Negative ROI in AI customer service is usually caused by billing and measuring on deflection — conversations closed without a human — rather than on resolution, so the cost of returning contacts lands in headcount instead of the AI budget. The same model produces either outcome.
How do I calculate the real cost per AI resolution?
Divide what you are billed by the number of issues genuinely solved end to end, not by the number of conversations the AI closed. On Aissist's production data, a 63.1% inclusive rate against a 38.2% strict rate makes the effective cost 1.65× the rate card.
Did Klarna's AI customer service actually fail?
No, but its staffing conclusion did. Klarna's AI assistant handled two-thirds of customer service chats and still did at the point of reversal, but in May 2025 Klarna began rehiring human representatives more than a year after claiming the assistant did the work of 700 agents.
How much are companies spending on AI in customer service?
Service and support functions spent a median of almost $6 million on AI in 2025 and committed 13% of functional budget to it, according to Gartner as reported by CX Dive. More than 75% of service leaders plan to increase that investment in 2026.
What is the difference between deflection rate and resolution rate?
Deflection counts conversations that never reached a human, including abandoned and unresolved ones. Resolution counts issues actually solved end to end. Deflection is always the higher number, which is why vendors prefer it and why it is the wrong basis for an ROI case.
Will companies rehire the support staff they replaced with AI?
Some will. Gartner predicts that by 2029, 30% of employees laid off due to replacement by AI will need to be rehired, often at significantly higher cost, and predicted separately in February 2026 that half of companies that cut customer service headcount for AI will rehire by 2027.
Sources
- CX Dive, Only one-quarter of AI customer service use cases produce ROI — Gartner analysis of 432 AI use cases; published 17 August 2026. Read 17 September 2026.
- CX Dive, Are AI-driven customer service cuts here to stay? Gartner predicts rehiring — median AI spend and Emily Potosky interview; published 11 September 2026. Read 17 September 2026.
- Gartner, Gartner Identifies 4 Shifts Shaping the Future of Work, 9 September 2026. Read 17 September 2026.
- Gartner, Gartner Says CFOs Need to Rethink the ROI of AI Investments, 24 March 2026. Read 17 September 2026.
- OpenAI, Klarna's AI assistant does the work of 700 full-time agents, February 2024. Read 17 September 2026.
- CX Dive, Klarna changes its AI tune and again recruits humans for customer service, 9 May 2025. Read 17 September 2026.
- Intercom, Pricing — $0.99 USD per Fin outcome. Read 17 September 2026.
- Aissist.io, AI Customer Service Statistics 2026, Deflection vs Resolution Rate, Cost Benchmark of AI Service, AI Customer Service Benchmark, The Resolution–CSAT Tradeoff. Read 17 September 2026.
Changelog
- 17 September 2026 — Published. All figures verified against primary sources on this date.



