Skip to content

Deflection Is Not Resolution: The Real Economics of AI Customer Service

Published: at 03:05 AMSuggest Changes

A service director opened her quarterly review with a number she had spent a year building: 64% containment. Almost two-thirds of customer contacts, she told the room, were being handled without an agent touching them. The CFO let the slide sit, then asked a narrower question. Of the customers the assistant had “contained”, how many got back in touch within three days about the same problem? Nobody could answer. By the next quarter, total contact volume was flat, frontline headcount had not fallen, and the containment figure had drifted three points in the wrong direction.

Containment and deflection are attractive because they measure something real and easy: the contacts that stopped short of a human. What they leave out is whether the customer’s problem went away. Deflection is a routing outcome. Resolution is a business outcome, and the two part company the moment an assistant is rewarded for ending conversations rather than closing issues. For a large share of service organisations, the metric on the board measures the wrong event — and the gap between the two widens as the technology improves, because a smoother conversation is easier to mistake for a solved problem.

What deflection actually counts

Deflection counts the contacts that did not reach an agent. That is a reasonable engineering signal about channel capacity and a poor economic signal about service quality. Gartner’s survey of 5,728 customers found that only 14% of service issues were fully resolved in self-service, even though 73% of customers used self-service at some point in their journey. For issues customers themselves called “very simple”, the resolution rate reached only 36%. [Gartner, August 2024] An interaction that ends is not the same thing as an issue that closes.

That gap is the whole argument, and it has not closed with better models. Gartner’s 2026 research found customers were roughly three times more likely to take a service issue to a third-party GenAI tool than to a company-provided chatbot, and that a separate survey of 1,303 service leaders put the median AI share of the 2025 budget at 12%, the highest of any function Gartner assessed. Only 24% of those leaders reported positive financial returns across their AI use cases. [Gartner, July 2026] Spending at the top of the league table with returns near the bottom is what happens when the success metric and the value metric describe different things.

Containment became the default because it is the one number an AI vendor and a service director can both point at without agreeing on anything else. It is measurable from logs, it moves quickly, and it survives a demo. Resolution requires evidence from the customer, arrives later, and occasionally contradicts the dashboard. The metric that is easiest to collect has quietly become the metric that is hardest to defend.

The repeat contact is where the money is lost

Every unresolved first contact returns as a second one. SQM Group’s January 2026 analysis of repeat contacts catalogues the costs that never surface on a containment dashboard: the same issue handled twice, longer waits for new callers because agents are re-handling old ones, and contact volume that looks like demand but is the same demand arriving again. SQM’s benchmarking puts industry first-call resolution near 70%, and its research finds customer satisfaction falls by about 15% each time a customer has to make contact again about an issue that should already be settled. [SQM Group, January 2026]

The unit economics turn on this. Cost per contact is a channel statistic. Cost per resolved issue is the number that reaches the P&L. When 30% of issues need two contacts to close, every “resolution” costs 1.3 contacts, and the effective cost per resolution runs roughly 30% above the headline cost-per-contact figure. A containment programme can push the first number down and the second one up in the same quarter, because the contacts it removes from the agent queue reappear later, more expensive and less patient.

The measurement window decides whether you see this at all. A repeat rate defined over 24 hours misses the customer told to wait 48 hours. Three days — a 72-hour window — is long enough to catch the customer who tried self-service, failed, and came back, and short enough to attribute the return to the original interaction. Report repeats by issue type rather than in aggregate. A blended repeat rate hides the two or three contact reasons that generate most of the avoidable volume.

Escalation, churn and the cost of a wrong answer

Unresolved contacts return in two forms. Some come back as a repeat; others escalate to a more expensive tier. When a frontline assistant cannot resolve an issue, the contact moves to a senior agent, a specialist or a supervisor, each costing more per hour than the frontline. The escalation bill is larger than the pay difference: it includes queue time, the rebuild of context the first contact already gathered, and the concession often needed to repair the relationship. An escalation that ends in a refund the first contact could have issued is pure loss.

The most expensive event in the ledger is a wrong answer the customer accepts. An assistant that misstates a refund policy, quotes an ineligible price or promises a delivery the business cannot make does not create a repeat contact — it creates an action taken on false information. The recovery cost lands outside the service budget, in concessions, rework and, at the extreme, churn. Klarna’s own 2024 announcement claimed its AI assistant handled two-thirds of customer service chats and the work of 700 agents, worth an estimated $40 million in profit improvement; by May 2025 the chief executive told Bloomberg that cost had become “a too predominant evaluation factor” and that the result was lower quality, prompting a return to human support. [Bloomberg, May 2025] A contained conversation that ends in a confident error is the most expensive outcome a service operation can produce.

Trust follows the same path. Gartner’s 2026 survey of 3,566 customers found that 87% of customers expect companies using GenAI for service to provide access to a human agent, and that customers forced through several unsuccessful AI interactions before reaching a person become less willing to use the tool again. [Gartner, August 2026] Churn rarely announces itself as a service failure. It appears months later in attrition data owned by another function, long after the containment dashboard has been reset for the next quarter.

A measurement model a service leader can adopt

Replace the containment target with a resolution ledger. Five measures, read together, describe whether the service is genuinely working:

Add one trust signal: the repeat and churn rate of customers who used the AI channel compared with those who did not, matched on issue type. The comparison is crude and it is decisive. If AI-channel customers contact the company again more often than the rest, the assistant is generating work rather than absorbing it.

Two ratios carry the argument to an executive audience. Resolution rate against containment shows whether the assistant is solving issues or ending conversations. Cost per resolved issue against cost per contact shows whether the savings are real or borrowed from a future quarter. Neither ratio needs new technology to compute. Both need the service organisation to stop reporting contacts and start reporting outcomes.

The question to put to the service leader

Ask for two numbers side by side, for the same period and the same issue types: containment and the 72-hour repeat contact rate. If containment is climbing and repeats are climbing with it, the AI is re-routing the workload, not reducing it. If repeats are falling while cost per resolved issue falls as well, the deployment is doing real work and has earned more authority.

Then apply the test that settles the question. Can your service leader produce, within a day, the 72-hour repeat contact rate for AI-handled issues, the cost per resolved issue across every channel, and the concession spend attributable to incorrect AI answers? If any of the three is missing, the containment figure on the dashboard is not evidence of performance. It is a routing statistic wearing the costume of a business result.

Sources


Previous Post
Prompt, Retrieve or Fine-Tune: How to Choose Where Model Behaviour Lives
Next Post
Internal Audit Has to Assure Systems That Change Themselves