AI Helped the Workers Who Needed It Most—and Distracted the Best
A randomized Alibaba field experiment found faster service and better customer ratings overall, with the largest gains among lower performers and a warning for top performers.

Scope note: This essay covers a four-week AI-assistant experiment with newer customer-service agents at Alibaba. It does not establish long-term effects on learning, employment, or experienced workers outside this setting.
The same AI tool can be a ladder for one worker and loose gravel for another.
Alibaba tested an AI assistant with 5,940 customer-service agents on Taobao. The company randomly gave 2,895 agents access to suggested issue diagnoses and responses while 3,045 agents continued with the standard tools. The four-week experiment covered roughly 2.56 million chats and 390,000 customer ratings.
On average, the assistant improved speed and customer ratings. The gains were not evenly distributed. Lower-performing agents improved most. Top performers gained little speed and sometimes lost service quality as their attention moved across too many chats.
The average result was positive
Simply receiving access reduced issue-identification time by 8.2 percent and total chat time by 1.1 percent. Customer dissatisfaction fell by 3.4 percent relative to the earlier average, and ratings rose by 1.2 percent. The rate at which customers returned within three days did not change significantly.
The difference between ratings and return behavior matters. Customers felt better about the service, but the experiment did not show a broad improvement in that harder follow-up measure.
The assistant did more than paste a canned reply. It diagnosed the likely problem and proposed a solution. Agents could use, edit, or ignore the text. When they used the assistant, they responded faster, sent more informative messages, and required less explanatory effort from customers.
This was assistance inside a live operation, not a writing exercise conducted in a quiet lab. The system had to survive actual queues, actual customers, and workers handling several conversations at once.
Lower performers gained the most
The researchers divided agents into five groups based on their customer ratings before the test. The lowest-performing groups saw the largest improvements in speed and customer satisfaction.
The assistant gave those workers a usable starting point for diagnosis and response. It reduced the cost of recalling the right procedure while a customer waited. In effect, it made some of the operation’s stored expertise available at the moment of need.
This narrowed the gap between lower and higher performers. That is a more important result than a small rise in the overall average. An organization may get greater value by strengthening the workers who lack an established method than by giving the same intervention to everyone and hoping the mean moves.
It also suggests a better rollout question. Do not ask only, “Did the tool improve performance?” Ask, “Whose performance changed, and why?”
The best agents paid an attention tax
Top performers were more skeptical of the suggestions and used them less. The researchers found no evidence that their quality fell because they blindly copied bad answers. Their language remained informative and clear.
The problem appeared in the workflow. Top performers spent more time shifted away from the active customer chat after the assistant arrived. They seemed to use the freed capacity to monitor more conversations and verify more material. Customers waited longer, disengaged, and returned more often. The top group saw a 4.1 percentage-point increase in three-day return contacts.
Efficiency created room. The operation filled that room with more simultaneous work. Quality then paid the bill.
This is not a reason to withhold AI from experienced staff. It is a reason to design a different intervention for them. A newer worker may need a strong proposed answer. An expert may need better review tools, harder cases, mentoring work, or protection from a queue that expands every time a second is saved.
Short-term evidence has a boundary
The experiment involved agents with less than one year of tenure and ran for four weeks. It cannot tell us whether the initial gains persist, whether lower performers learn from the system, or whether constant assistance weakens their independent judgment. It also took place in one large Chinese e-commerce operation with a specific chat workflow.
Still, the causal evidence is unusually strong for workplace AI. The test shows that a useful assistant can improve a live service operation. It also shows why a single deployment policy is lazy.
The average can rise while one group advances and another slips. Deploy to the worker, the task, and the surrounding queue—not to an imaginary employee who exists only in a dashboard.
