What Klarna Teaches, and What It Doesn't
Vishal Sachar
Co-Founder & CEO of CLRT
In February 2024, Klarna announced that its AI assistant had handled 2.3 million conversations in its first month, two thirds of the company's service chats, doing what it described as the equivalent work of 700 full time agents. Fourteen months later its chief executive told Bloomberg that "there will be always a human if you want". The market promptly split into two camps: one declared the automation of support a failure, the other dismissed the walk back as public relations. Both camps are looking at the most useful case study in applied AI and taking home the wrong lesson.
Start with what actually happened, because the details carry the argument. By Klarna's own account, the February 2024 numbers were striking. The assistant handled 2.3 million conversations in its first month. Resolution time fell from eleven minutes to under two. Repeat inquiries dropped by a quarter, and customer satisfaction held level with human agents. On the back of it the company projected a $40 million profit improvement for 2024. These are vendor numbers, published by a company with every incentive to make them look good, and they should be read that way. But nothing that followed suggests the volume figures were wrong. The assistant really did absorb an enormous share of the contact load, and it kept absorbing it. Whatever Klarna discovered in the year that followed, it was not that the machine could not handle the chats.
Then came May 2025, and the Bloomberg interview that launched a thousand told-you-sos. Sebastian Siemiatkowski said the quiet part precisely: "cost unfortunately seems to have been a too predominant evaluation factor", and "what you end up having is lower quality". Klarna began a rehiring pilot and pledged that a customer who wants a human will always get one. Read carefully, though, and this is not a reversal. The assistant stayed. The volume it handles stayed. What changed was the position at the edges: the company stopped treating the automated share as the whole of support and started, in Siemiatkowski's words, "really investing in the quality of the human support" as "the way of the future". Klarna did not walk back the automation. It walked back the claim that automation was the entire product.
Here is the reading both camps miss. Support was never one workflow. It is two products sold under one name. In most contacts the customer is buying an answer: where is my order, when does the refund land, what does this charge mean. Those are high volume, checkable against a system of record, and cheap to get right, which is why the assistant genuinely absorbed them. In a small minority the customer is buying judgment: the disputed charge, the exception the policy never anticipated, the loyal customer who is angry for the first time. That minority is a rounding error in the volume figures and most of the brand. Cost per contact prices the first product precisely and prices the second at zero, which is exactly what an evaluation dominated by cost quietly did. The walk back is simply the moment the unpriced half presented its bill.
The transferable lesson is not the position Klarna landed on. Always a human is the right answer for a consumer credit brand with millions of small contacts; it may be the wrong answer for your business, or the right one for entirely different reasons. The transferable lesson is the method Klarna arrived at backwards: split the workflow by what the customer is actually buying, and only then decide what to automate. Every function has this seam. Sales has outreach volume and the relationship moment. Finance has reconciliation and the judgment call that decides what a number means. Legal has boilerplate and the clause that carries all the risk. The seam never announces itself, because on a volume dashboard the two halves look identical. You either find it deliberately, before pointing the automation, or you find it the way Klarna did: in production, in public, with the brand as the measuring instrument.
The volume was automatable. The judgment was the brand.
A deeper dive
The engineering underneath the split is where the real difficulty lives, because the two halves cannot be separated once and left alone. They have to be routed between, contact by contact, and misrouting is the failure mode that costs. A volume contact sent to a human is waste, irritating but survivable. A judgment contact absorbed by the volume machine is the expensive direction: the customer with a genuine dispute meets a fluent system that resolves the ticket and loses the relationship, and the dashboard records the exchange as a success. The routing cannot lean on the model's own confidence, which is poorly calibrated exactly where it matters, and it cannot lean on keywords, because the angry loyal customer often writes politely. The triggers that actually work are specific to the business, learned from its real contacts, and they need maintaining, because the seam moves as products, policies, and customers change. That work is unglamorous, ongoing, and decisive, and it is precisely what a per contact price does not include.
Notice also what the metrics were doing during the year the quality eroded. Every number Klarna published in February 2024 was, as far as anyone can tell, true. Conversation volume, resolution time, repeat contacts, satisfaction on par: each measured the automatable half, and each moved in the right direction. The quantity that broke was not on the dashboard, because trust erodes on a lag and reveals itself late, in retention, in escalations that never happen because the customer left instead, in the tone of the contacts that remain. A business planning the same move should assume its own dashboard has the same blind spot and do the work Klarna's evaluation skipped: name the unmeasured variable the judgment half protects, decide how it will be watched, and treat any business case built purely on cost per contact as an argument about half the workflow. That is judgment plus instrumentation. It does not come in the box with the agent, and it is the difference between Klarna's 2024 and Klarna's 2025.
Work with CLRT
Finding this seam is diagnostic work, and it is what CLRT does before anyone builds anything. We map where volume ends and judgment begins in your specific workflows, in numbers, then engineer the routing, escalation, and verification that keep the machine on its side of the line. Klarna ran the experiment with its own brand as the test case. You do not have to. Start with the CLRT Ascent diagnostic at ascent.clrtstudio.com and see where your line actually sits, before the market draws it for you.

Vishal Sachar
Vishal Sachar is the Co-Founder and CEO of CLRT, where he helps UAE businesses make sense of applied agentic AI and put it to work. He writes on agentic systems, AI governance, and the economics of automation. Reach him at vishal@clrtstudio.com or on LinkedIn.


