Perspectives

You Bought a Number. It Was True Somewhere Else.

Estimated reading time: 12 minutes

Perspectives is a new weekly series, pulled straight from the book I’m writing, AI Debt. Here’s what it is, and why I’m doing it in public.

Every AI decision turns into one of two things: a debt or a dividend. Which is this one?

A patient's chart lights up on a night shift. Possible sepsis, which is what happens when the body's response to an infection turns on its own organs. It kills fast, so that alert is not a suggestion. The nurse draws the labs and pages the physician. The physician confirms, and the antibiotic clock starts. Over a year, nurses move thousands of times.

How often does that alert have to be right to be worth having? The hospitals that bought it could not tell you. They didn't set what good looks like.

Due Diligence Versus Outcome

These are two stages of the same purchase.

The first stage is pre-sales work that proved the vendor was a real company, that it would still be there in three years, and that its contract would hold up. It ran across procurement, legal, and security. It helped your organization reduce the risk of doing business with a new entity.

It did not prove the product would solve your defined problem.

The second stage belongs to whoever owns the problem, and it is the part that takes time: a proof of concept run against the actual problem, a bake-off when more than one vendor claims to solve it, and a number of your own that defines what solving it would look like. Urgency debt is usually why you skip that second stage.

Miss it, and operations takes delivery and inherits the vendor's KPIs, built for a market, not your problem.

That is where assumption debt starts. The vendor's number was true somewhere else, and you didn't define what good looks like for your problem.

Epic Sold a Score. Hospitals Bought a Promise.

Let's revisit the night shift. That alert came from a prediction tool built into the hospital's medical records system. Epic's system runs a large share of American hospitals. The promise was that the tool would spot patients sliding toward sepsis before clinical staff saw it coming.

Epic scored the tool on how well it distinguished between the two groups: patients who would develop the illness and patients who would not. That score meant the tool should identify the right patient about 8 times out of 10. Compare that to a coin toss, which would get the right answer 50% of the time.

By the end of 2021, about 170 Epic customers were running it. Those customers are health systems across dozens of hospitals.

The Tool Missed 67% of the Patients It Was Bought to Catch

A research team at Michigan Medicine tested the tool on their own patients, 27,697 of them across 38,455 hospital stays, and published the results in JAMA Internal Medicine in June 2021. On the same test Epic used, the tool identified the right patient about 6 times out of 10, not the 8 out of 10 Epic reported.

2,552 of those patients developed sepsis. The tool missed 1,709 of them, or 67%, and those were the patients it was bought to find. Every hour that passes before the right antibiotic reaches a patient in septic shock, more of them die.

It also sent alerts on 18% of all hospital stays, 6,971 alerts in total, and most were for patients who were fine. Nurses spent most of that year answering warnings about people who were not in danger.

The conclusion: the tool could not reliably identify who was getting sick, and its scores did not match what happened to the patients.

Early Detection: Doctor – 1, Tool – 0

Three years later, a second study explained why. A University of Michigan team published a follow-up in NEJM AI in February 2024. Part of what the model detected was that a clinician had ordered a blood culture. It was scoring patients a doctor had grown suspicious about. What Epic sold as early detection was mostly late confirmation.

That is what makes this failure hard to see from inside a hospital. Compare the alerts against what the clinicians concluded, and you find agreement, and agreement looks like proof. Michigan Medicine only found the gap by testing the tool against what happened to the patients afterward.

Zero Minutes of Warning

This is not a 2021 problem that aged out. Across 2023, two county emergency departments in Harris County, Texas ran the model over 145,885 patient encounters. The results were published in JAMIA Open in November 2024. The tool caught 14.7% of cases.

Half the alerts arrived after the patient was sick. The median warning time was zero minutes. An alert that arrives at the same moment as the diagnosis warns no one. The authors compared it against an alert that fires at random and reported only marginally better performance.

The Data Was There. The Measure Was Not.

Every hospital running that tool had the patients, the records, and the outcomes sitting in the same system that produced the alerts. No one in those buildings had set what outcome was acceptable and tested it against what the tool could do.

The data was never the constraint. What was missing was a number a team had agreed on in advance, and the standing decision that checking it against that number was part of the job.

The Number Belonged to the Vendor

Hundreds of hospitals. 3 years. One arithmetic problem.

An organization with an unchecked assumption has not assigned anyone to check it. The analyst who can measure the tool cannot switch it off. The executive who can switch it off is not shown the measurement. And the number, the benchmark, and the definition of working all come from the vendor.

The vendor is not immune. Vendors get sued, they lose renewals, they answer for it in the press. The difference is timing and exposure. Your cost starts the day the tool goes live, and it lands on your customers and your staff. The vendor's cost arrives years later, argued by lawyers, and the contract usually caps it.

That contract is the one procurement, legal and security validated in stage one. It made the vendor safe to buy from. It also set the ceiling on what you get back when their number turns out to have been true somewhere else.

Your customers pay first. Your business pays most.

Governance debt is what you owe for not watching. Assumption debt is what you owe for not asking.

Behind the Vendor's Claim Sit 3 Assumptions of Your Own

The vendor's number is the one you can point at. The expensive ones are in your own description of the problem, and here are the three that cost the most.

Your Data Changed, and the Model Did Not Notice

On Unity's Q1 2022 earnings call, chief executive John Riccitiello named two failures in the company's ad-targeting business. The second one: “we lost the value of a portion of our data, training data due in part to us ingesting bad data from a large customer.”

Unity put the cost at $110 million in 2022.

The model kept running the whole time. What moved was the data underneath it, and it moved because of what one customer sent. In Riccitiello's words, the repair ran “data rebuilding, model training and improvement, and then revenue recovery.”

Unity's data changed, and no measure was watching for it. Before you buy, name the business unit that owns the data this tool depends on, what they check it against, and how often. If that answer is a shrug, the purchase is resting on shaky ground.

Klarna Measured Speed. Customers Wanted a Person.

In February 2024, Klarna announced that its AI assistant had handled 2.3 million conversations, two-thirds of its customer service chats, “the equivalent work of 700 full-time agents.” Resolution time fell from 11 minutes to under 2, satisfaction matched human agents, and Klarna projected $40 million in profit improvement for the year.

15 months later, in May 2025, chief executive Sebastian Siemiatkowski said the company had leaned too far into cost-saving and started hiring people back. His words: “investing in the quality of human support is the way of the future for us.”

The numbers were real. They measured speed and volume. What Klarna took them to prove was that customer service was working. The customer who wanted a person was not counted, so nothing in the reporting could show them leaving.

A Pilot With No Definition of Success Cannot Report Failure

McDonald's spent 3 years testing AI voice ordering in more than 100 restaurants. In June 2024, it ended the test, said the work gave it confidence that voice ordering would be part of its restaurants' future, and did not say whether this one worked.

A project ran for 3 years, across more than 100 sites, and ended with no public answer to one question: did it work? There was no answer because there had been no target. “We should be doing something with AI” is not a problem statement, and a pilot that starts there has nothing it can be right or wrong about.

An Output Is Not an Outcome

Governance debt usually surfaces when someone outside forces a look. Assumption debt does not, and it is not quiet either. It is expensive in the places nobody is reviewing.

The sepsis tool produced an output every hour of every day it ran. The outcome it was bought for was patients reached before they got dangerously ill. Nobody wrote that down as the measure, so the output stood in for the outcome, and the reporting looked healthy while the tool missed two-thirds of the people it existed to find.

That is why the review does not catch it. A review of the output finds a system doing exactly what it was built to do. Only a review against a defined outcome finds the gap, and the outcome is the part nobody defined.

Nothing alarms inside the system, because the system is not broken. The tool scored every patient, every hour, and printed a number. Unity's model kept placing ads right up to the guidance cut. The alarm goes off somewhere else: in a revenue line, in a patient's chart, in a bias nobody was watching for.

So who decided? Not the vendor. Finance, operations, and the function that will live with the tool meet in one room, agree on which problem is worth solving, and define what a good, realistic outcome would look like. Naming the outcome up front does not make that meeting easy. It makes it scheduled work instead of organizational risk.

The Dividend: Duke Named the Outcome First

Duke started somewhere Epic's customers did not. Before the model touched a patient, they named the outcome they were after: earlier treatment for patients, and they built the measurement that could prove or disprove it. They tested their tool, Sepsis Watch, against 4 years of their own hospital's records. It went live in November 2018 across 3 Duke hospitals, attached to a named clinical workflow rather than a bare alert.

Trust But Verify

Then came the harder test. A different health system, Summa Health in Ohio, took Duke's model into 4 emergency departments and 205,005 patient encounters without changing a thing. It held. It flagged patients roughly 3 to 5 hours before they became critically ill. Those results were published in 2025, where the next buyer can read them.

Asking Became a Team's Job

Then Duke did the part that pays. It turned the one check into a standing requirement. Duke Health's Algorithm-Based Clinical Decision Support Oversight Committee, chaired by its chief health information officer, puts pre- and post-deployment checkpoints on any algorithm used in patient care. In its own words: “our governance applies equally to internally developed solutions and those procured from external sources.”

Epic's customers bought a tool. Duke built the standing ability to judge one, then pointed it at everything that came after, including the tools it did not build. The difference is not that Duke found better people. Duke created a team whose job is to weigh risk against outcome, and put that question on the agenda before anyone needed the courage to raise it.

What Outcome Are We Buying, and How Will We Know?

That is the question to put to the room before you buy, with finance, operations, and whoever will sign the contract. What outcome are we expecting, how would this tool get us there, and what would we have to see to conclude that it did not?

If that room cannot answer it, you are not buying a system. You are adopting an assumption. And an assumption fails plenty loudly. It just does not fail on a review of the output, which is the review most organizations run.

Technology will answer any question you ask it. Ask it what you skipped, and it will tell you that too. The asking is the part that must be assigned to somebody. Finance, operations, and the executive who signs decide together that the question belongs on the agenda, and that decision is the difference between creating a debt and earning a dividend.

Somewhere tonight a chart lights up on a night shift. Whether that alert is worth the next four minutes of a nurse's night was settled in a room she never entered, by people who either wrote the number down or did not.

Next in the series: adoption debt, technology that outran the people meant to use it.

The One Assumption You Own

Every assumption in this piece was handed to you by someone else. One is yours, and it comes before all of them: is this problem worth solving at all?

That is what the AI Use-Case Gate does. It runs before a vendor is in the room. Name the business problem in one sentence without using the word AI, price what it costs you today, and score whether it is a use case at all. It ends where this article ends, with a decision written down and a date to revisit it.

Download digitalBard-AI-Use-Case-Gate.pdf

Leave a Comment

Your email address will not be published. Required fields are marked *

This site uses Akismet to reduce spam. Learn how your comment data is processed.

Before you go

Find out which AI debt is costing you the most.

The AI Debt Diagnostic is free. Nine questions, about four minutes. It scores you across all nine debts and names the one charging you the most right now. Results land on screen and in your inbox.

Take the AI Debt Diagnostic

New here? Debt or Dividend publishes weekly on LinkedIn, subscribe. Or browse every perspective.

Ready to talk now? Book a strategy conversation.

Scroll to Top