Perspectives

The Metric Measured the Tool. The People With the Problem Were Never Asked.

Estimated reading time: 12 minutes

Perspectives is a new weekly series, pulled straight from the book I’m writing, AI Debt. Here’s what it is, and why I’m doing it in public.

Every AI decision turns into one of two things: a debt or a dividend. Which is this one?

In June 2025, Commonwealth Bank of Australia switched on a voice bot to take incoming calls in its Customer Service Direct unit. On July 29, the bank told 45 people in that unit their jobs would end, and started working out their exit dates. The number the bank gave the staff and their union: the bot had cut calls by 2,000 a week.

The 45 staff kept answering phones while the bank worked out who would go, and when. The calls did not slow down. They rose. Team leaders were told to help pick up the slack and answer calls, too. The bank asked the same 45 people it was letting go to work overtime, to cover the calls the bot was supposed to have removed.

Their union asked for the call-volume data, was refused, and took the dispute to the Fair Work Commission. On August 20, ahead of the next hearing, the bank reversed itself: "CBA's initial assessment that the 45 roles in our Customer Service Direct business were not required did not adequately consider all relevant business considerations and this error meant the roles were not redundant. We have apologised to the employees concerned and acknowledge we should have been more thorough in our assessment of the roles required."

The Real Outcome, Said Out Loud on the Same Day

The bot did what it was built to do and the bank counted calls the bot took. The true ROI KPI was something else: how many customers got their problem solved, and how fast, once the bot took the simple calls and the people took the hard ones. The bank had said so itself on the day it announced the cuts: "By automating simple queries, our teams can focus on more complex customer queries that need empathy and experience." That was the outcome. It was never written down as the measure.

If you approved a rollout this year, a number like the bank's is in this month's status report. It arrived with the tool, it looks fine, and you signed off on it the way the rest of us did: the tool was new, the quarter was short, and nobody handed you a better number. That is not a failure of leadership. That is the job now.

The Number Described the Bot's Output

Adoption debt is the cost of technology that moved faster than the people using it. In a lot of buildings it starts in a similar manner. The tool works in the demo and the pilot. Then the problem owners receive it, their work unchanged, and it gets measured by a number that describes the tool and says nothing about the problem to be solved.

Assumption debt is what you owe for not asking. Adoption debt is what you owe for not asking the people who have the problem.

The problem rarely sits in one department. Finance sees the cost. Operations sees the queue. Customer service sees the caller, and hears the question the bot could not handle. Each of them holds a piece of it. A rollout that did not collaborate across the business, united in the defined problem and the requirements to solve it, is a solution looking for a problem everywhere.

The 7 Steps

One large group in concept, many people in practice: finance, operations, customer service, HR, whoever's work the tool touches and whoever owns the outcome. They belong in the whole conversation, from the first question to the last review. Two questions organize it, in order. Fitness for purpose: does this do what we need? Fitness for use: does it do it reliably enough, under real conditions, to be usable? The people who have the problem are the only ones who can answer the first, and the only ones who can decide the second.

7 Steps the Problem Owners Take

1. They name the problem, in their own words. The first meeting is a premortem. Gary Klein described it in Harvard Business Review in September 2007: the people who will use the tool are told the rollout has failed, and each writes down, alone, why. Daniel Kahneman's reason it works: the premortem "overcomes the groupthink that affects many teams once a decision appears to have been made."

2. They frame it. Which part of the work is overly complex or manual, what solved looks like, and which higher-value work is stuck behind it.

3. They write the requirements. Requirements written two floors up are the gap the tool falls into.

4. They try the prototype, as sponsored users, against the requirements they wrote. User acceptance testing enables them to decide what is working and what is not, in rounds, until the output meets the outcome they named. That is fitness for use, checked under real conditions.

5. They write the KPIs that prove the outcome. A vendor's metric measures the tool. Theirs measures the company's problem.

6. They name the higher-value work the time will go to, their managers agree it matters, and it gets an ROI KPI of its own.

7. They stay in the conversation. Fit gets re-checked as the work changes, and training keeps going after launch day. The Conference Board found in July 2026 that 28.3% of workers say their organization provides no AI training at all.

Your Part: Two Refusals

Your part in the 7 steps is not to write the number. It is two refusals. Refuse a rollout that skipped the room, and refuse a number nobody in that room wrote. Everything else on the list belongs to the people with the problem, and it only works if you are the one who holds the door open.

The Corporate Reality

  • McKinsey asked C-suite leaders in January 2025 whether they would involve nontechnical employees in the early development of AI tools. Fewer than half, 48%, said yes.
  • PwC and the Manufacturing Institute asked manufacturers in March 2026 why their AI initiatives had failed. 45% said, in part, because frontline leaders were not sufficiently included in design or rollout.
  • The Institute for the Future of Work, studying 8 organizations with Chartered Institute of Personnel and Development(CIPD) in April 2026, put it in one sentence: "Where employees were excluded from decision-making, organisations reported low uptake and poor return on investment. Where organisations prioritised transparency, dialogue and co-creation, adoption was stronger and more sustainable."

A Better Outcome Is Available

Stanford's Digital Economy Lab reviewed 51 enterprise deployments the same month and found the same thing from the other side: "When users genuinely want the solution, adoption friction disappears."

There is a human reason inside that research. In 2012, Norton, Mochon and Ariely showed that people put a higher value on things they helped build, and called it the IKEA effect. One condition: the build has to get finished. A half-built tool the people helped design earns no loyalty at all.

The Hour the Tool Gave Back

At the bank, the hour was the whole promise. Simple calls to the bot, so the people could take the ones that "need empathy and experience." The bank named that work itself, from the top. The number it chose, calls the bot took, measured the machine and said nothing about those calls. The 45 people who would have done that work were never asked what it was worth, or which calls they would take first.

They are not alone. BCG asked close to 12,000 people in June 2026, and 66% of frontline employees "still receive limited or no guidance on what to do with the time they save." The hour has a purpose when the person doing the work names the higher-value activity it goes to, and their manager agrees it matters. Where that conversation never happens, the hour goes back into the old work, and the return goes with it.

AI Mandates in 2025, Human Boomerang in 2026

Leadership's answer to a people gap in 2025 was to push more AI usage. Within a year the companies that pushed hardest were taking the pressure back off, and the people they pushed had a name for what they were doing in between.

April 7, 2025. Shopify's chief executive, Tobi Lutke, posted a memo: "Before asking for more headcount and resources, teams must demonstrate why they cannot get what they want done using AI." AI use went into performance and peer reviews.

April 28, 2025. Duolingo's chief executive, Luis von Ahn, told the company it would be "AI-first," and that "AI use will be part of what we evaluate in performance reviews."

August 2025. Coinbase's chief executive, Brian Armstrong, described what happened to engineers who had not onboarded to the AI tools by his deadline: "Some of them had a good reason. Some of them didn't, and they got fired." His own word for the approach was "heavy-handed."

The Walk-Backs

Spring 2026. Amazon set a target of more than 80% of its developers using AI weekly and put usage on internal leaderboards. Staff had a name for what happened next: "tokenmaxxing," running the tools to move the number. On May 29, 2026, Amazon confirmed the leaderboard was gone. Senior vice president Dave Treadwell: "Please don't use AI just for the sake of using AI. Use AI to help you solve customer problems, to help you solve business problems, to innovate."

April 2026. Von Ahn took AI use back out of Duolingo's reviews. His reason is this whole edition in one sentence: "It felt like rather than being held accountable for the actual outcome, we're trying to just push something that in some cases did not fit."

The Law Is Catching Up

The rules are arriving behind the walk-backs. Since January 1, 2026, the state of Illinois has required employers to give notice whenever AI is used to influence an employment decision, disciplinary actions included. In California, SB 947, the No Robo Bosses Act of 2026, passed both houses on August 31 and awaits the governor's signature by September 30. It bars relying solely on an automated system to fire or discipline anyone.

Usage in a performance review counts the tool. It cannot count the quality of the work. The people being counted know the difference, and now the law is catching up to it.

The Dividend: The Measure Was Written Before the AI Was Built

In August 2025, Smarter Balanced, the consortium whose assessments millions of students take, published 8 student-centric design principles for any AI used in its assessments, developed with IBM Consulting. I facilitated the outreach. Who took part, in the release's own words: "multidisciplinary groups including educational assessment and measurement experts; those serving in state, district, and local educational contexts; higher education experts; and policymakers." Feedback came in at 3 national conferences before anything was agreed upon.

Set that against the 7 steps. The people the outcome was created for named what the AI owed them before any AI was built. Those principles became the requirements. They became the test: the consortium's stated use for them is to "evaluate whether AI-based solutions are appropriate for teachers and students." And they became the foundation, guidance for every assessment design that followed.

The AI Success Metric: SLA 0 to 4.

What makes it a KPI: each principle carries a service-level agreement, written for that organization, and scored on one scale.

  • SLA 0: the AI does not meet the principle in any way.
  • SLA 1: it minimally meets the principle.
  • SLA 2: it builds on SLA 1 and adds more capability toward the principle.
  • SLA 3: it builds on SLA 1 and SLA 2 and meets the principle more fully.
  • SLA 4: it includes all the capabilities of SLA 1, 2 and 3, and maximizes them to fully meet the intended outcome.

A vendor cannot produce that score. Only the people who wrote the principle can.

Swap "student-centric principle" for the problem your own people named, and the scale still works.

What exists here is the part most rollouts skip: a written measure of the outcome, and a way to show, level by level, how far the AI meets it. The people who would live with the outcome wrote the exam before any developer took it.

What 45 People Could Have Told the Bank

The call logs were the bank's own. The 45 people answering the phones could have said, in the first minute of a premortem, which way the number was going. But those people were not asked when the number was chosen. Everything after that was the interest.

If the people with the problem had written your rollout's number, what would it have counted?

You are running more tools with fewer people than you were two years ago, and each one arrives with a number you did not write. The people with the problem can write a better one, and that can start this week.

Next in the series: accountability debt, an AI decision no one owns.

Where Does Adoption Debt Sit in Your Organization?

Start with the AI Debt Diagnostic. 9 questions, about 4 minutes, free. It scores your organization across the nine debts and shows where to start.

Leave a Comment

Your email address will not be published. Required fields are marked *

This site uses Akismet to reduce spam. Learn how your comment data is processed.

Before you go

Find out which AI debt is costing you the most.

The AI Debt Diagnostic is free. Nine questions, about four minutes. It scores you across all nine debts and names the one charging you the most right now. Results land on screen and in your inbox.

Take the AI Debt Diagnostic

New here? Debt or Dividend publishes weekly on LinkedIn, subscribe. Or browse every perspective.

Ready to talk now? Book a strategy conversation.

Scroll to Top