Are You AI Enabled?
Part-Time Chief of Staff ← Back to overview

Field Notes · AI

Are You AI Enabled?

7 min read · case notes

Lower cost, better service, more output. Four companies ran the experiment for real, and one of them got the order of operations wrong in public. Here's what actually happened in each case.

"AI is going to save us money" is usually said in a meeting where nobody has decided what "us" gets to keep and what the customer loses in exchange. That's the wrong question. The companies that actually pulled this off didn't treat cost, speed, and customer experience as three things to trade against each other. They treated them as three readouts on the same instrument, and refused to call the project a win until all three moved together. One of the four learned that the hard way, which is exactly why its story belongs first.

The one that worked, until the third number caught up with it

In early 2024, Klarna rolled out an OpenAI-powered assistant to handle customer service, and for a company its size the numbers were startling. Within a month it was handling two-thirds of all support chats. Klarna said the assistant was doing the equivalent work of 700, later reported as 853, full-time agents. Average resolution time fell from eleven minutes to under two. Repeat contacts on the same issue dropped by a quarter. The company projected a $40 million improvement to 2024 profit from the change alone.

2/3of support chats handled by AI within the first month
11 min → 2 minaverage resolution time, before and after
$40Mprojected profit improvement for 2024

By 2025, Klarna was quietly rehiring humans. CEO Sebastian Siemiatkowski said plainly that cost had been weighted too heavily in how the rollout was judged, and that the quality customers actually experienced had slipped as a result. "As cost unfortunately seems to have been a too predominant evaluation factor when organizing this, what you end up having is lower quality," he said, adding that "really investing in the quality of the human support is the way of the future for us."

AI gives us speed. Talent gives us empathy. Together, we can deliver service that's fast when it should be, and empathic and personal when it needs to be.Klarna, 2025, on its revised AI-plus-human model

None of the original numbers were fake. The mistake was optimizing two columns, cost and speed, while under-watching the third. The fix wasn't abandoning the assistant. It was putting a human back in the loop for the cases that needed one, and measuring customer experience with the same discipline as the cost savings from day one.

The one where customers preferred the machine

Octopus Energy's AI platform, Kraken, ran a three-month trial of a generative assistant called Arlo on real customer emails in the UK. Arlo handled around 8,000 emails a week, about 4% of the company's total customer email volume, drafting full responses rather than canned replies. When customers rated the responses they received, Arlo's came back at a 76% satisfaction score. Human agents handling comparable queries scored 72%.

That result is easy to misread as "the robot is better than people," which isn't quite the point. The actual mechanism is that Arlo took the routine, well-understood queries off the desks of human agents, who were then free to spend their attention on the harder, less scripted cases where a person genuinely adds more value. Volume went to the machine. Judgment stayed with the humans. Both numbers, cost and satisfaction, moved in the right direction at once because the split was deliberate.

The one that turned a cost center into a growth engine

IKEA's starting point was the least glamorous version of this problem: a large, expensive customer service operation, the kind every retailer treats as pure overhead. Rather than cutting the roughly 8,500 people staffing it, IKEA retrained them, from handling routine calls into interior design advisors, while its chatbot Billie took over more and more of the repetitive volume. Billie's own resolution rate grew from 47% of inquiries in its first two years to 74% today.

47% → 74%of customer queries Billie resolves without a human
60% → 89%IKEA's own customer satisfaction score, before and after
€1.08B → €1.25Bannual sales from the remote centers those retrained staff now run

The retrained staff didn't disappear into a smaller org chart. They moved into remote sales centers that now generate over a billion euros a year, growing 15 to 20% annually for three years running. A department that used to be a line item on the cost side of the ledger became one that shows up on the revenue side, and customer satisfaction rose right along with it. That's the trifecta done properly: automation absorbed the volume, people moved up the value chain instead of out the door, and the number that used to only go one direction, service cost, started paying for itself.

The one nobody notices, because it's just quietly working

Moderna built 750 custom internal AI tools across the company within two months of rolling out an enterprise AI platform, and its employees now average around 120 AI-assisted conversations a week each. CEO Stephane Bancel has said the company believes it can keep operating at a fraction of the headcount a traditional pharmaceutical company would need for the same output, using AI to absorb the research, drafting, and coordination work that used to require a much larger staff.

Morgan Stanley took the same logic to its financial advisors: a GPT-4-based assistant trained on the firm's own research library, built specifically to handle the searching and drafting that ate into an advisor's day, so that the advisor's actual time went back to clients instead of to lookup work. Neither company is claiming a headline productivity multiplier. They're just quietly running fewer people through more output, which is a less exciting story and a much more common one than the viral chatbot headlines suggest.

What the pattern actually is

In every case that held up, AI took over the repetitive, high-volume, low-judgment part of the job, the routine email, the drafted response, the research lookup, and left the judgment calls with a person. In the one case that didn't hold up at first, the same technology was pointed at the same kind of work, but the project was scored on cost and speed alone until the customer experience number forced a correction. The technology wasn't the variable. The discipline of tracking all three ledgers at once was.

That's precisely the discipline a special-projects mandate needs before it touches AI at all: find the specific repetitive volume eating a team's time, price out what automating it is actually worth in the first year, and don't call it a win until cost, output, and the human experience on both sides of the transaction have all moved the right way together. Most companies can name the AI tool they should be piloting. Fewer of them have anyone whose whole job is watching all three numbers at once before they scale it.

Referenced: Klarna, "Klarna AI assistant handles two-thirds of customer service chats in its first month" (2024) and subsequent 2025 statements from CEO Sebastian Siemiatkowski on reinvesting in human support. · Octopus Energy / Kraken, "Octopus Energy's AI trial wins customer approval" (Arlo customer service trial). · Fortune, "Inside Ikea's big bet on humans in the age of AI" (2026), on the Billie chatbot and retrained remote-sales workforce. · OpenAI, "Moderna" case study, on ChatGPT Enterprise adoption and custom GPT usage. · Morgan Stanley, "Morgan Stanley Research Announces AskResearchGPT" and reporting on its GPT-4-based advisor assistant.

If you can't yet say which three numbers your own repetitive work would move, and by how much, that's exactly what the audit is for.

Take the audit → Back to the overview