Your AI absorbed the easy cases. Look at what it handed back.
By Craig Joseph, MD ·
Every clinical AI deployment I have reviewed came with a slide about what the tool takes off the clinician’s plate. I have never seen the companion slide: what it puts back on. There always is one. The tool absorbs the routine work, and what remains for the human is a smaller number of more difficult calls, made later, with less of the surrounding reasoning intact. Health systems are instrumenting the first slide obsessively and the second not at all.
What does a clinical AI actually take away?
Less than the vendor deck implies, and something different from what it claims. Consider three tools most large systems now run.
Ambient documentation did not remove the work of deciding what mattered in a visit. It changed the task from composing that judgment to reviewing someone (or something) else’s draft of it, which is a different cognitive job and, as I will argue, a more complicated one to do well.
Predictive alerting did not remove the decision to treat. It converted it into a decision about the alert: is this firing, for this patient, at this moment, the one that deserves a response, given the forty that did not?
Imaging AI did not remove the radiologist’s read. It moved the radiologist from producing a read to ruling on one already sitting on the screen with a confidence score attached.
In each case, a decision was relocated, not retired. The behavioral scientist Michael Hallsworth, writing about video review in football, describes technology as making errors “fractal”: they recur at ever-smaller scales instead of disappearing. The offside line is his example. A camera system now proposes the exact frame in which the ball was played, an official confirms it, and the arguments have simply migrated to whether a glancing touch counts as playing the ball and where a skeletal model thinks a shoulder ends. Precision moved but did not resolve the judgment.
Why is the leftover decision harder than the original?
Three reasons, and they compound.
It is rarer. The model absorbed the common cases, so the human now sees the residual ones at a fraction of the old frequency. Skill that is exercised less often degrades; more on the evidence for that below.
It is more difficult by construction. The cases the model handles confidently are, by definition, the ones with a clean signal. What reaches the clinician is the ambiguous remainder, selected for being complicated.
And it arrives stripped of context. When a clinician worked the whole problem, the reasoning that produced the decision was available at the moment of deciding. When the tool has done the front half, the clinician is asked to adjudicate a conclusion without having built the path to it. We retired the reasoning along with the task, then asked the person to supply it on demand.
How well do experts actually handle the decision they were left?
Worse than any governance committee assumes, and the best evidence comes from radiology, the specialty with the longest exposure to algorithmic support.
In a 2023 experiment published in Radiology, Thomas Dratsch and colleagues at the University of Cologne had 27 radiologists assign BI-RADS categories to 50 mammograms aided by a system presented to them as AI. Twelve suggestions had been planted as errors. When the suggestion was correct, inexperienced readers got 79.7% of cases right; when it was wrong, they got 19.8%. Moderately experienced readers went from 81.3% to 24.8%. The very experienced group, the people a health system would point to as its safeguard, went from 82.3% to 45.5%.
Read that last pair again. Two decades of subspecialty expertise bought about twenty-five percentage points of resistance to a machine that happened to be wrong, and still left the expert wrong more often than right. That is the measured performance of the human in the loop on precisely the decision the loop left them.
Nor is this a matter of naivety. A survey of the radiologists who read screening mammograms for Norway’s national program found 47% considered automation bias a high risk. In the same survey, 68% anticipated that AI would raise their own detection rate. They see the trap clearly. They also expect to step around it.
Does the skill survive when the tool is switched off?
There is now a direct measurement, and it says no, or at least not intact.
In 2025, Krzysztof Budzyń and colleagues published in The Lancet Gastroenterology & Hepatology an observational study from four Polish endoscopy centers that had introduced AI polyp detection. They looked only at the colonoscopies performed without the AI, three months before and three months after the tool arrived. The adenoma detection rate in those unassisted procedures fell from 28.4% to 22.4%, an absolute drop of six points. In the adjusted analysis, prior exposure to the AI was independently associated with lower detection (odds ratio 0.69).
The study is retrospective and the authors are appropriately cautious about causation. But the direction is exactly what the relocation argument predicts. Adenoma detection is a skill that improves with reps. The tool took the reps. When the tool was absent, the residual skill was measurably smaller than it had been a few months earlier. Nobody running that endoscopy suite had a dashboard for it.
Does a good tool make this problem go away?
No. A good tool moves the judgment to a different person, usually one who is not a clinician and whose decision is made once, in advance, out of sight.
The Swedish MASAI trial is the strongest randomized evidence we have that AI-supported mammography screening works: 6.4 cancers detected per 1,000 women against 5.0 with standard double reading, no significant increase in false positives, and a 44% reduction in screen-reading workload. It is a real result, and I do not want to diminish it.
Look at the mechanism, though. The AI assigns each mammogram a score, and the score decides whether one radiologist reads it or two. The cutoffs that route a woman’s mammogram to one reader or two were chosen before the trial enrolled its first participant. A 2025 systematic review in BMJ Open covering 31 studies puts the general condition plainly: triage configurations reduce reading volume by 40% to 90% while maintaining detection, provided the thresholds are “conservatively calibrated.” The benefit of the whole program hangs on a calibration decision that no radiologist reading a case will ever see, and that decision was a human judgment. The trial did not eliminate the frame-picker. It hired one.
What should a health system measure instead of override rate?
Almost every organization I speak with can tell me its override rate. Almost none can tell me whether the overrides were right. Override rate is attendance, not performance. It records a disagreement with the tool and says nothing about whether the clinician had ninety seconds or nine, whether they had information the model lacked, or what happened to the patient.
A decision dashboard, as opposed to a model dashboard, would contain at least four things. A sampled audit of overrides with chart review and outcomes, because an override is a hypothesis about the model being wrong and hypotheses can be checked. A parallel sample of acceptances, because Dratsch’s data says the dangerous behavior is agreeing, not disagreeing. A measure of time-to-decision at the moment the tool fired, because a judgment made in nine seconds is a different judgment from one made in ninety. And a periodic unassisted-performance check, on the Budzyń model, so you learn that a skill is eroding before the day the tool is down and you need it.
This is much harder to build than an accuracy dashboard for the model. It is also the only one that reports on the judgment you kept. I have argued before that clinical judgment is trainable and that health systems have largely declined to train it. The relocation problem is the other half of that argument: we are now actively un-training it and not looking.
Which decisions should a machine never be allowed to touch?
That question should be answered in writing, before deployment, and almost no governance committee has done it. Football’s video review protocol did. For its first eight years it allowed video review of four kinds of decision: goals, penalties, direct red cards, and mistaken identity. Everything else stayed with the referee by design, and the list predates the first review. This summer the list grew, which is its own small demonstration of the fractal point: even the boundary document is subject to relocation pressure.
Your organization has an approval pipeline for AI. It almost certainly has no document that says which clinical decisions a model may participate in and which belong to a human because you decided they do. Without that document, the answer gets written one procurement at a time, by whichever vendor shows up next with a good slide about what their tool takes off the plate.
Ask them for the other slide. Then ask who chose the frame.