All Posts

CX Tech

Contact Center AI: What the Data Actually Shows

Bhavika J

Techshorts Editorial Team

What "contact center AI" covers

The category spans three different jobs that get marketed as one thing. Voice AI answers or routes phone calls without a human on the line. Support automation, usually a chatbot or an agent-assist layer, handles or accelerates written tickets and chats. CSAT tooling scores those interactions afterward, often using the same large language models that ran the conversation to grade it.

Vendors sell these as a bundle: deploy the assistant, cut headcount, watch satisfaction hold steady or improve. The last twelve months produced enough real-world attempts to check that pitch against outcomes, rather than against a demo.

The measured record so far

Klarna is the most cited case because it ran the experiment in public. In February 2024 the buy-now-pay-later company launched an OpenAI-powered assistant it said handled the equivalent workload of about 700 full-time agents, and it froze customer-service hiring for more than a year. By May 2025, CEO Sebastian Siemiatkowski told Bloomberg the company was reopening those roles. He said the AI-first approach produced "lower quality" support and that Klarna was "hiring humans again" to guarantee a person was reachable when customers wanted one (Bloomberg, as reported by Forbes, 2025). Klarna's own description of the failure was not technical. It was that customers noticed the drop in quality and the company decided that mattered more than the cost saved.

Gartner's research points the same direction at a market level, not just one company. A March 2025 poll of 163 customer service and support leaders found 95% intend to keep human agents in the loop rather than move to an agent-less model (Gartner, June 10, 2025). Gartner used that data to predict that half of the organizations planning a significant AI-driven workforce reduction will abandon those plans by 2027. Three months later, a separate Gartner poll of more than 3,400 organizations backed a related prediction: over 40% of agentic AI projects will be canceled before the end of 2027, citing unclear return on investment and immature autonomous decision-making as the leading causes (Gartner, June 25, 2025). In September, Gartner analyst Kathy Ross went further, predicting that no Fortune 500 company will have eliminated human customer service entirely by 2028, calling a fully agentless model "both unlikely and undesirable" (Gartner, September 10, 2025).

The most detailed look at why comes from outside the CX category entirely. MIT's NANDA initiative at the Media Lab reviewed more than 300 disclosed generative AI deployments across industries, plus structured interviews and a leader survey, covering January to June 2025. It found that despite $30 to $40 billion in enterprise spending, 95% of generative AI pilots showed no measurable effect on profit and loss (MIT NANDA, "The GenAI Divide: State of AI in Business 2025," July 2025, covered by Fortune, August 18, 2025). The report's most useful finding for a buyer is not the failure rate. It is the gap between build approaches: pilots built on a purchased, specialized vendor tool succeeded roughly 67% of the time, while pilots built in-house succeeded at about a third of that rate. The tool mattered less than whether the team building on it could adapt the workflow as real conversations exposed gaps the demo never showed.

Why the gap between pitch and result exists

A vendor demo is scripted against clean inputs. A live phone queue or chat window is not. The failure mode Gartner and MIT both describe is the same one Klarna hit: a system that resolves the easy, high-volume, low-stakes contacts well and silently degrades on the contacts that carry account risk, emotional weight, or an edge case the training data never covered. Containment rate, the share of contacts an AI system resolves without escalating to a person, rises in exactly those conditions. CSAT does not rise with it. A system can look highly effective on the containment dashboard while quietly pushing dissatisfied customers to give up rather than escalate, which is a different outcome than resolution.

What to look at before buying

Ask for containment rate and CSAT reported together, by contact type, not blended into a single company-wide number. Ask what happens to the contacts the system does not resolve: how fast they reach a human, and whether the handoff carries the conversation history or starts the customer over. Ask whether the vendor's own success numbers come from a pilot on a subset of contacts or a full production queue, since the MIT NANDA gap between pilot and full deployment is where most of the reported failures occurred.

What commonly goes wrong

Teams measure the automation's success by how much volume it took off human agents, not by what happened to the customers routed through it. Klarna's reversal and Gartner's abandonment forecasts both trace back to that same substitution: cost per contact went down while a slower-moving signal, actual customer sentiment, went unmeasured until it showed up in brand damage or churn.