A customer service team working at a support desk, representing the human role restored after an AI rollout was scaled back
← Back to Blog
Agentic AI

Klarna Rehired Its Humans. That Is the Agentic AI Test

Head portrait of Alex Goryachev
Alex Goryachev·August 3, 2026· min read

Key Takeways

  • Klarna's AI chatbot once handled 75% of customer chats and let the company cut headcount 22%, then the CEO publicly reversed course and started rehiring humans because quality suffered.
  • MIT's NANDA research found 95% of enterprise generative AI pilots show no real financial payoff, and Gartner projects more than 40% of agentic AI projects will be canceled by the end of 2027.
  • The tools that actually deliver value stay narrow on purpose: IT ticket triage, procurement work, and coding agents that a person reviews before anything ships.
  • Developers already model the right pattern: Stack Overflow found 84% use AI tools regularly while trust in AI-generated code accuracy fell to about a third, because they keep a human in the loop at the one point a mistake actually costs something.

At its peak, Klarna's AI chatbot handled 75% of all customer service chats. That's roughly 2.3 million conversations a month. The company said it did the work of 700 human agents. Klarna cut headcount by 22% during that stretch, down to 3,500 employees. It called this a preview of the future.

Then the company's own CEO walked it back, in public, on the record. Sebastian Siemiatkowski told Bloomberg the quality wasn't there. His words: "Really, investing in the quality of human support is the way of the future for us." Klarna is rehiring now. It's recruiting students, people in rural areas, and dedicated remote workers to answer the phone again. The cheaper option worked, technically. It just didn't work well enough for a brand that depends on customers trusting it with their money.

I keep coming back to something else Siemiatkowski said, because I think it's the sharper insight in the whole story. He wants customers to always know a human is available if they want one. That's a business decision about which parts of the customer relationship can tolerate an agent's mistakes, and which parts can't. It has nothing to do with nostalgia for the old way of doing things. Klarna handles money. A customer who feels brushed off by a bot during a billing dispute doesn't file a polite complaint. They close the account.

Klarna isn't an outlier here. MIT's NANDA research found that 95% of enterprise generative AI pilots show no real financial payoff at all. Most of that spending shows up on a balance sheet as a cost, and nothing else. The tools that actually deliver stay narrow on purpose. IT ticket triage. Procurement work. Coding agents that a person still reviews before anything ships. None of those touch a customer directly at the moment something can go wrong in front of them.

Gartner's broader forecast for agentic AI tracks the same pattern from a different angle. The firm expects more than 40% of agentic AI projects to be canceled by the end of 2027. The reasons it names are cost, unclear ROI, and weak risk controls, not weak models. Projects fail because nobody built an exit ramp before scope grew past what the team could manage. That's the exact trap Klarna walked into, and then walked back out of, at real cost, in full public view.

Developers, of all people, have already made peace with this tradeoff. I think their instinct is the right one to copy. Stack Overflow's most recent survey found 84% of developers now use AI tools regularly. In that same survey, trust in the accuracy of AI-generated code fell to about a third, the lowest level on record. Read those two numbers side by side and they don't contradict each other. They describe the human-plus-agents stack working as it should. Constant use, paired with a standing habit of checking the work before it ships.

That's the real difference between Klarna's first attempt and what developers do every day without thinking twice about it. Developers never handed the agent the whole job and walked away. They kept a human in the loop at the one point where a mistake actually costs something. They let the agent do the rest. Klarna's original rollout removed the human from that same point: the live conversation with a paying customer. The company only found out how much that cost once the quality numbers came in.

I think about this the way I think about every technology that got adopted faster than anyone could test where it actually belonged. The printing press didn't replace every scribe overnight, and it didn't need to. It replaced the copying, not the judgment about what deserved to be copied in the first place. Agentic AI is running into that same boundary now, one industry at a time. Klarna just hit it first, loudly, in front of everyone.

Before your company deploys an agent anywhere a customer can feel the difference, ask the question Klarna learned the hard way. Can we switch this off by noon if the quality isn't there, and will anyone notice fast enough to make that call?

Sources: Entrepreneur/Bloomberg, June 2025 (Klarna reversal, Siemiatkowski quotes) · Gartner, June 2025 (40% project cancellation forecast) · MIT NANDA, August 2025 (95% no-P&L-impact finding) · Stack Overflow Developer Survey, December 2025 (84% AI tool usage, trust decline).

Which agentic AI use cases are safe to deploy right now?

The ones that stay narrow and reversible. IT ticket triage, procurement automation, and coding agents paired with mandatory human review are the categories MIT's research found actually delivering results, largely because they are bought from vendors rather than built from scratch and because a person can catch a bad decision before it reaches a customer.

Why did Klarna reverse its AI customer service rollout?

Klarna replaced roughly 700 customer service agents with AI in 2024 and called it a success. By May 2025, CEO Sebastian Siemiatkowski told Bloomberg the company had gone too far, that service quality had suffered, and Klarna began rehiring humans so customers could talk to a person again.

Does low trust in AI-generated code mean developers should stop using it?

No. Stack Overflow's 2025 survey found 84% of developers use AI tools regularly even as trust in the accuracy of AI-generated code fell to about a third, an all-time low. The two numbers together describe the human-plus-agents stack: constant use, paired with a standing habit of checking the work.

Alex advises enterprise leaders on which agentic AI use cases are ready now and which still need a human in the loop.

book a conversation →

← Back to Blog
Head portrait of Alex Goryachev
Alex Goryachev

WSJ-bestselling author · Former Managing Director of Innovation, Cisco · Advisor, CSU AI Working Group · LinkedIn Top AI Voice

Work with Alex

Bring this thinking to your organization

Alex works with executive teams at global enterprises on AI strategy, governance frameworks, and organizational readiness. Available for keynotes, C-suite workshops, and advisory engagements.