Audit Your AI Stack in 10 Minutes: The Metric Nobody Is Tracking
Every healthcare AI program eventually produces the same slide. Logins are up. Alerts fired are up. Model executions per week are climbing. Leadership nods, the renewal gets approved, and the program is declared a success.
None of those numbers answer the question that matters: did the AI change a human decision last month, and if so, which one?
Healthcare organizations have become skilled at measuring AI adoption. They remain poor at measuring AI impact. The result is a growing portfolio of tools that are technically live and functionally invisible, running quietly inside clinical and operational workflows without ever touching the decisions those workflows exist to support.
This is not primarily a story about AI failing. Most of the models behind these dashboards perform close to their design specification. It is a story about organizations that never defined what a changed decision would look like before they turned the tool on, and are only now discovering that adoption and impact are different variables entirely.
The Agreement Trap
If an AI tool agrees with a clinician, many organizations see that as proof it works. In reality, it may suggest the opposite. An AI system that rarely challenges clinical decisions is not influencing them, it is simply confirming them. And confirmation creates far less value than most AI business cases expect.
Value concentrates in the minority of cases where the system and the clinician disagree. Cognitive divergence, the moment an algorithm’s output conflicts with what a person was about to do, is the only place a decision can actually change. Everywhere else, the AI is simply a faster route to a destination the clinician had already chosen.
Medication-related clinical decision support shows this dynamic at scale. A 2018 study of alerts at a 793-bed teaching hospital, published in the Journal of the American Medical Informatics Association, found that nearly three in four alerts were overridden, and that roughly 40 percent of those overrides did not hold up as clinically appropriate on review. A later systematic review spanning 23 separate studies put override rates for computerized order-entry alerts anywhere between 46 and 96 percent. Somewhere inside that range sits a healthy friction zone where AI is genuinely earning its place in the workflow. Near the top of it, the tool has become background noise that clinicians route around by habit.
Activity Is Not the Same as Change
Most dashboards measure activity because it is easy to track. Metrics like logins, alert volumes, report views, and AI model runs are automatically captured by the systems that generate them. But these metrics cannot tell you whether a clinician, case manager, or claims reviewer actually changed a decision or took a different action because of the AI.
This challenge is not unique to healthcare, although the cost of getting it wrong is much higher. According to McKinsey’s latest global AI survey, 88% of organizations use AI in at least one business function, but only 39% can measure a financial return from those investments. Even among those, most say AI contributes less than 5% of their overall business impact. Another 2025 McKinsey study found that the biggest driver of AI success is redesigning workflows, not simply deploying AI. Yet fewer than one in four organizations using generative AI have changed even part of a workflow to take advantage of it. Healthcare follows the same trend. An industry analysis based on McKinsey’s 2025 technology findings reported that only 1% of health systems consider their AI adoption fully mature.
The organizations sitting in that gap are not lacking AI. They are lacking a way to see past adoption into the decisions adoption was supposed to change.
Where Informing Ends and Deciding Begins
Not every interaction between AI and a clinician is meant to change a decision, and assuming it is often leads to misleading evaluations. Sometimes AI simply provides information that helps clinicians confirm what they were already planning to do. Other times, it changes the course of action entirely.
The difference is easiest to see in operational workflows. For example, an eligibility check that changes an approval decision, a claims model that overturns a denial, or an AI alert that triggers an escalation that otherwise would not have happened are all examples of AI influencing a decision.
Clinical use cases are often less obvious. If AI summarizes a patient’s chart and confirms a diagnosis the physician had already reached, it is informing the decision, not changing it. But if AI identifies an unexpected finding on a scan that causes the physician to cancel a planned discharge and request a specialist consultation, it has changed the decision.
The key question is: Did the AI change what happened next? If the workflow or outcome stayed the same, the AI provided useful information but did not influence the decision, regardless of how accurate or confident its recommendation appeared.
Why the EHR Cannot See What Changed
Electronic health records (EHRs) were built to record what happened, mainly for billing, compliance, and continuity of care. They were not designed to record why a clinician made a particular decision. That creates a major challenge for measuring AI. If you cannot see the reasoning behind a decision, you also cannot determine whether the AI actually changed it.
This is known as the EHR telemetry gap. It explains why many healthcare organizations cannot answer a simple but important question: If this AI recommendation had been wrong or had never appeared, would the outcome have been any different?
To answer that question, organizations need to connect an AI recommendation to a specific action, such as a changed medication order, a cancelled discharge, or an escalation in care. Without that connection, they are forced to judge AI by activity metrics like alert volume or usage instead of real clinical impact.
A well-known sepsis prediction tool integrated into a major EHR platform highlights this problem. An independent evaluation published in JAMA Internal Medicine in 2021 found that the tool identified sepsis correctly only 63% of the time, compared with the developer’s reported 76% performance. It also missed about two-thirds of patients who later developed sepsis, even though it generated alerts for 18% of all hospitalized patients.
More alerts did not lead to earlier detection. Instead, they contributed to alert fatigue. In fact, a companion commentary describing a similar deployment reported that staff at one hospital covered a monitoring camera to reduce constant interruptions. The AI was active, but it rarely changed clinical decisions in the way it was intended to.
The Ten-Minute Audit
An organization does not need a research team to find out where it stands. It needs three checks, run honestly, on any AI tool that has been live long enough to generate a real usage history.
1. The override and dismissal ratio
Pull the usage logs for a given AI tool and calculate how often its output is overridden or dismissed outright. A rate near 100 percent means clinicians have learned to ignore it. A rate near zero means it functions as a rubber stamp. The signal worth acting on sits in between, where disagreement is frequent enough to prove the tool is being genuinely evaluated rather than reflexively accepted or reflexively dismissed.
2. The order-modification delta
Ask for a report of clinical or operational actions, medication orders, imaging orders, discharge plans, claims determinations, that were changed or canceled within a defined window of an AI output. If that report cannot be produced at all, the organization has zero decision telemetry for that tool, no matter how many times it fired.
3. The transparency and ownership check
Confirm whether clinicians can see the basis for a given recommendation inside their existing workflow, and confirm who owns the resulting impact number. The ONC’s HTI-1 rule, in effect for predictive decision support tools since January 2025, now requires certified health IT to expose a defined set of source attributes covering a tool’s development, validation, and measured performance. That infrastructure already exists inside most certified systems. Few organizations have connected it to an actual audit.
What the Audit Usually Finds
Run this audit honestly and the number that comes back is rarely a modest correction to the story leadership has been telling. Published override literature suggests why. Depending on the alert category and care setting, independent studies have found override rates anywhere from the high 40s to well above 90 percent, with entire categories of high-confidence alerts, duplicate-drug checks among them, sometimes overridden almost universally. Layer in the fact that a meaningful share of the remaining accepted recommendations are cases where the clinician was already headed in that direction, and the true rate of AI-driven behavioral correction across a typical stack tends to land far below what a usage dashboard implies.
This discovery does not mean the AI is broken. It means the organization has been measuring the wrong layer of the system.
Why Most AI Doesn’t Fail
When an AI project fails to deliver results, the first reaction is often to blame the model. In most cases, that is the wrong conclusion. The model is usually not the problem.
Stroke triage is a good example of why. AI tools that detect large vessel occlusions (LVOs) can identify a suspected stroke directly from CT scans and immediately alert the neurointervention team. This removes delays caused by waiting for a radiologist’s review and follow-up phone calls.
Multiple studies have shown that this approach significantly speeds up treatment. One study reported that the median time from patient arrival to specialist notification dropped from 40 minutes to 25 minutes. A larger meta-analysis also found meaningful reductions in both door-to-groin-puncture time and door-in-door-out time, helping patients receive treatment sooner.
These improvements were not simply because the AI model was more accurate than other clinical AI systems. They happened because the AI delivered the right information to the right person at the right time, when action could still make a difference.
Compare this with the sepsis prediction tool discussed earlier. Both systems use AI to predict clinical events and support decision-making. Yet the sepsis tool became an example of alert fatigue, while the stroke triage system is widely recognized for improving patient care and reducing treatment delays.
The key difference was not the algorithm. It was how the AI was integrated into the clinical workflow and whether its recommendations led to timely action.
Decision Telemetry Is the Next Maturity Metric
Usage telemetry tells you how often an AI tool was used. Decision telemetry tells you whether the tool actually changed a decision compared with what would have happened without it.
Most healthcare organizations have invested heavily in measuring usage because those metrics, such as logins, alerts, and model executions, are already available in vendor dashboards. Measuring decision impact is much harder because organizations must intentionally build the systems needed to connect AI recommendations to real clinical or operational actions.
This is beginning to change. Recent research in radiology no longer measures AI success only by diagnostic accuracy. It also evaluates whether an AI explanation actually changed a radiologist’s or physician’s decision, and by how much. One multi-site study involving more than 200 radiologists and physicians took exactly this approach.
That represents a shift from measuring AI activity to measuring AI impact. Healthcare organizations that build decision telemetry today will be able to answer a much more important question for every AI tool they deploy: Is this investment improving outcomes, or is it simply generating more activity?
AI Adoption Metrics vs. AI Decision Impact Metrics
| AI Adoption Metrics | AI Decision Impact Metrics |
|---|---|
| Logins and active users | Cases where the AI’s output changed the plan of action |
| Alerts fired per period | Orders modified or canceled within a defined window of an alert |
| Model executions per day | Override rate, tracked against a defined healthy range |
| Dashboard and report views | Documented before-and-after outcome for a specific decision |
| Vendor-reported accuracy | Independently measured behavioral shift rate |
The Question Worth Asking at the Next Steering Committee
Healthcare organizations do not need another AI dashboard. They need evidence that AI changed a decision, produced through a defined, owned process that runs on a fixed cadence rather than only when a renewal or a board question forces the issue.
Organizations that cannot measure decision impact cannot accurately measure AI value, regardless of how sophisticated the underlying model is or how many licenses have been deployed. That is an uncomfortable finding for a program that has spent two years reporting adoption numbers. It is also the more useful finding, because it points toward a fix, rescoping a tool, changing where in the workflow it fires, assigning clear ownership of the impact number, rather than a rip-and-replace decision that blames an algorithm for what is actually a measurement problem.
Pegasus One works with healthcare organizations on exactly this layer of the problem: connecting AI outputs to the operational workflows they are meant to influence, building the decision telemetry that EHRs were never designed to capture, and helping leadership teams define what a changed decision looks like before the next tool goes live, not two years after. The audit above takes ten minutes. Building the infrastructure to run it every quarter, on every tool in the stack, is the actual work.
Talk to the Pegasus One team today about moving beyond the ten-minute audit to a decision telemetry layer that runs every quarter, on every tool in the stack.