All Posts

CX Tech

What Four Field Experiments Measured When Contact Centres Added AI

Bhavika J

Editorial Team

Most performance figures for contact centre AI come from vendors describing their own customers. A smaller body of evidence comes from field experiments, where researchers compare agents with and without AI inside a live operation. Four such studies, run at three companies, give buyers a measured baseline to hold vendor claims against.

What agent assist is

Agent assist, often sold as a copilot, is software that reads a live customer conversation and suggests replies, diagnoses or knowledge articles to a human agent. The agent stays responsible for the conversation and can accept, edit or ignore each suggestion. The category exists because contact centres carry wide performance gaps between new and experienced staff, and much of what top agents know never reaches a knowledge base.

An autonomous AI agent is different. It handles the conversation itself and hands off to a human when it fails or the customer asks. Vendors often market both under the same label, but the evidence on them is not the same.

What the assist studies measured

The largest study followed 5,179 support agents at a Fortune 500 software company as a GPT-based chat assistant was rolled out in stages. Access to the tool raised issues resolved per hour by 14% on average. Novice and low-skilled agents improved by 34%, while experienced, highly skilled agents saw minimal effect (Brynjolfsson, Li and Raymond, 2025). Customers were also less likely to ask for a supervisor, and attrition fell, driven by newer agents staying.

A randomized experiment with 5,940 chat agents at Taobao, Alibaba's main marketplace, found a similar pattern with an important limit. The assistant drafted diagnoses and solutions at the start of each chat. It improved service speed and customer ratings, but had no significant effect on customer retrials, the researchers' objective measure of whether the problem was solved (Ni et al., 2024). Low performers gained most. Top performers saw little speed gain and a decline in both rated and objective quality, which the authors link to increased multitasking.

A third randomized experiment, at a meal delivery company, found AI-assisted agents responded faster and lifted customer sentiment, again with the largest benefit for less experienced agents (Zhang and Narayandas, 2025). The effect depended on the conversation. It helped most on subscription cancellations and least on repeat complaints, where the cause sat outside anything the agent could fix. Customers who had already hit a chatbot comprehension failure reacted worse to AI-assisted agents, because unusually fast replies made them think they were still talking to a bot.

What the autonomous agent study measured

The evidence on autonomous AI is thinner. A second Alibaba experiment, run on Taobao in August 2024 with 647 workers and 680,676 chats, had treated workers supervise an agentic AI system that resolved eligible chats while they handled the rest. The deployment cut average chat duration and had limited effect on retrial rates, but substantially lowered customer ratings on AI-eligible chats (Wang et al., 2026; Tuck School of Business, 2026).

Human intervention preserved service quality when the AI hit a technical problem beyond its capability. It was far less effective once the customer had become frustrated or sceptical, and those emotionally escalated chats ended with lower ratings and more follow-up contacts.

Customer surveys point the same way. In a Gartner survey of 3,566 B2B and B2C customers conducted in February and March 2026, 87% said companies using GenAI for service must offer a way to reach a human agent (Gartner, 2026). In the same survey, only 27% said they would try a chatbot again after a negative experience (Gartner, 2026).

What to measure in a pilot

The studies suggest four checks before accepting a vendor's result.

  • Measure resolution, not only speed. The Taobao assist study recorded faster service and better ratings with no change in retrials. A pilot that tracks handle time alone can log a gain the customer never feels.
  • Segment by agent tenure. Average gains in these studies came mostly from less experienced agents. A team of long-tenured agents should expect less, and should watch top performers for quality loss.
  • Segment by contact type. Cancellations and repeat complaints responded differently in the meal delivery study. A blended average can hide a queue where AI makes outcomes worse.
  • Track the handoff for autonomous agents. Measure ratings and repeat contacts on chats the AI passed to a human, and how long the customer spent with the AI before the handoff.

Where possible, compare agents with and without the tool over the same period, as these studies did, rather than comparing this quarter with last.

What commonly goes wrong

The most common error is reading a speed gain as a quality gain. Speed improved in every study here; objective resolution moved much less. The second is generalising from someone else's average. These results come from chat operations at a software firm, a Chinese marketplace and a meal delivery company. None measured voice, and none tells a buyer what to expect in a regulated or highly technical queue.

The third is spending ahead of measurement. Service leaders have increased AI spending by 38% while overall service budgets rose 2%, according to a Gartner survey of 199 leaders conducted in April and May 2026 (Gartner, 2026). The field evidence supports AI assistance for newer agents on routine chat. It is far less settled on whether autonomous agents can hold customer ratings once a conversation goes wrong.

Sources

  • Brynjolfsson, E., Li, D. and Raymond, L. "Generative AI at Work." The Quarterly Journal of Economics, 140(2). 2025. https://academic.oup.com/qje/article/140/2/889/7990658
  • Ni, X., Wang, Y., Feng, T., Lu, L. X., Wang, Y. and Zhou, C. "Generative AI in Action: Field Experimental Evidence from Alibaba's Customer Service Operations." SSRN working paper. 2024. https://papers.ssrn.com/sol3/papers.cfm?abstract_id=5012601
  • Zhang, S. and Narayandas, N. D. "Engaging Customers with AI in Online Chats: Evidence from a Randomized Field Experiment." Management Science. 2025. https://pubsonline.informs.org/doi/10.1287/mnsc.2022.03920
  • Wang, Y., Zhu, C., Feng, T., Lu, L. X. and Jia, B. "Agentic AI and Human-in-the-Loop Interventions: Field Experimental Evidence from Alibaba's Customer Service Operations." arXiv working paper. 2026. https://arxiv.org/abs/2605.14830
  • Tuck School of Business. "Even With Humans in the Loop, Agentic AI Systems Struggle." 2026. https://tuck.dartmouth.edu/news/articles/even-with-humans-in-the-loop-agentic-ai-systems-struggle
  • Gartner. "Gartner Survey Finds 87% of Customers Say Companies Using GenAI for Customer Service Must Provide Access to a Human Agent." 2026. https://www.gartner.com/en/newsroom/press-releases/2026-08-04-gartner-survey-finds-87-percent-of-customers-say-companies-using-genai-for-customer-service-must-provide-access-to-a-human-agent0
  • Gartner. "Gartner Survey Finds Only 27% of Customers Would Try a Chatbot Again After a Negative Experience." 2026. https://www.gartner.com/en/newsroom/press-releases/2026-09-02-gartner-finds-only-27-percent-of-customers-would-try-a-chatbot-again-after-a-negative-experience
  • Gartner. "Gartner Survey Finds AI Spending by Customer Service Leaders Has Surged by 38%, Despite Overall Service and Support Function Budgets Rising by Just 2%." 2026. https://www.gartner.com/en/newsroom/press-releases/2026-08-26-gartner-survey-finds-ai-spending-by-customer-service-leaders-has-surged-by-38-percent-despite-overall-service-and-support-function-budgets-rising-by-just-2-percent