All Posts

CX Tech

Why Contact Centre AI Metrics Look Best in the First 30 Days

Bhavika J

Editorial Team

The Number That Gets Published

When a company puts AI into its support queue, the numbers that come out first are usually measured in the first weeks: chat volume handled, average resolution time, satisfaction scores captured right after the first wave of conversations. These are real numbers. They are also the easiest numbers a deployment will ever produce, because the queries that arrive first are disproportionately the simple ones. Every vendor knows this. Not every buyer does.

Klarna's First 30 Days

Klarna gave the industry an unusually complete data set to check this against, because it published its own numbers and then, later, let a reporter watch what happened to them.

In a press release dated February 27, 2024, Klarna said its OpenAI-built assistant had handled 2.3 million customer service conversations in its first month, doing the equivalent work of 700 full-time agents, and covering two-thirds of all chat volume across 23 markets (Klarna, Feb. 27, 2024). Klarna reported average resolution time falling from 11 minutes to under two, a 25% drop in repeat inquiries, and customer satisfaction "on par" with human agents, though it published no numeric score alongside that claim. The company estimated the assistant would add $40 million to 2024 profit (Klarna, Feb. 27, 2024).

What Eighteen Months Showed

By May 2025, the picture had changed enough that CEO Sebastian Siemiatkowski discussed it publicly with Bloomberg. "As cost unfortunately seems to have been a too predominant evaluation factor when organizing this, what you end up having is lower quality," he said, adding that Klarna would guarantee customers a route to a human agent going forward (as reported in Fortune, May 9, 2025, and Customer Experience Dive). Both outlets reported that Klarna had begun rehiring for customer service roles it had cut, after complaints that the assistant handled routine account and payment questions well but produced generic or unresolved answers on disputes, fraud claims and hardship cases.

Forbes contributor Bernard Marr, writing in July 2026, put the shift in the context of Klarna's broader customer service staffing, which he reported the company had reduced from roughly 5,000 to 3,500 roles around the AI rollout (Forbes, July 16, 2026). By mid-2026, Klarna was running a hybrid model: AI still handles the bulk of routine chat volume, but customers with complex or sensitive issues are routed to a person, and the company has described this as a permanent design choice rather than a rollback.

Why the Gap Opens

The mechanism is not mysterious. A support queue's easiest tickets arrive first because they are the ones customers can describe clearly and AI can pattern-match against. The harder tickets, the ones involving a dispute, a fraud claim, an emotionally charged complaint, take longer to accumulate in the data and longer for a company to notice it is losing them. A 30-day number measures the queue at its most favorable mix. A number taken 18 months in measures the queue as it actually is.

This is not unique to Klarna or to customer service. MIT's NANDA initiative, studying enterprise generative AI deployments more broadly, found that roughly 5% of pilots were producing measurable revenue impact by mid-2025, with the large majority stalling after an initial rollout (as reported in Fortune, Aug. 18, 2025). That figure spans far more than contact centres, but it points at the same pattern: early numbers and durable numbers are frequently different populations of data.

What to Ask a Vendor For

A buyer evaluating support automation should ask for the metric measured at a specific, stated interval, not "results so far." Ask what share of deployment volume is complex versus routine, since containment and resolution rates mean little without that split. Ask whether a satisfaction score covers all resolved interactions or only the ones where the customer responded to a survey, since response bias skews toward the calmer end of the ticket queue. And ask what happened to the metric between month one and month twelve, since that is the number a launch announcement will never volunteer.

Klarna's public reversal is useful precisely because it is rare. Most companies do not let anyone watch their contact centre metrics degrade in public. The absence of a similar story from a given vendor is not evidence the pattern did not happen there too.