InsightsWhat others do2 min read
What support teams actually shipped
Klarna, Intercom, Autodesk, Starbucks. Read for the mechanism, not the headline: an operator wired into backend APIs, with evals. What a COO should measure.
Cliff des Ligneris
The support deployments that worked are not chatbots. They are operators wired into the systems that hold the answer.
This is the first post in a series where I read public enterprise AI cases for the mechanism rather than the headline. The figures below are the companies’ own public claims. I have not verified them, and wildcard* had nothing to do with any of these projects. Treat the numbers as marketing and the architecture as the lesson.
Four cases.
Klarna, in its own corporate announcement, says its OpenAI-based assistant handled 2.3 million chats in its first month, two thirds of the volume, and cut resolution time from 11 minutes to under two. The sentence that matters is further down: the agent was integrated into the order, billing and refund backend APIs. It did not explain the refund policy. It issued the refund.
Intercom’s Fin AI Agent report describes a similar shape with a different emphasis. Fin is grounded in the customer’s help centre and, by Intercom’s account, resolves millions of tickets. The part I would copy is what sits around the model: a strict evaluation framework for accuracy and brand tone. They did not ship on vibes.
Autodesk’s case study with IBM is older and smaller in ambition. A virtual agent on Watson Assistant handles roughly 100,000 conversations a month across 60 defined intents. Sixty. Not “anything a customer might ask”. A bounded set of things, each with a known resolution path.
Starbucks points the same idea inwards. Green Dot Assist, by Starbucks’ own description, is an iPad assistant for baristas: recipes, store routines, inventory guidance, piloted in 35 coffeehouses. The user is the employee, not the customer. The knowledge is operational, not conversational.
The pattern across all four: a narrow scope, real system access, and a test harness before scale. None of the four put a general chat window in front of customers and hoped.
For a COO at a company with 30 to 300 people, this reduces to four things to measure before and after.
- Share of conversations resolved with no human touch, by intent. Not overall. Per intent, because the average hides the failures.
- Time to resolution, median and 90th percentile.
- Reopen rate within seven days. A fast wrong answer comes back.
- Escalation accuracy: of the conversations handed to a person, how many needed one.
If a vendor cannot report the first and the fourth from day one, the system is a chatbot with a support badge. Klarna’s refund API connection is the whole difference, and it is an engineering task, not a prompt.
Where does this break? Tell me about a support deployment that worked without system access.