Your best early warning system is a nurse who charts too much
By Craig Joseph, MD ·
A deterioration model that contains no vital signs, no laboratory values, and no words from any note reduced inpatient mortality by roughly a third in a randomized trial across four hospitals. It did that by reading one thing: how often, and in what pattern, nurses were charting. This is great except for one thing: your ambient AI roadmap is designed to change exactly that pattern, and nobody has yet checked what happens to the model when it does.
What was the model actually reading?
The CONCERN early warning system, built by Columbia nursing informaticists with colleagues at Mass General Brigham, scores the frequency and timing of nursing surveillance documentation: the optional set of vitals, the added comment, the reassessment nobody required. Its designers started from an old observation on the wards: when a nurse gets uneasy about a patient, she checks more, and she writes more of it down, usually well before the numbers cross any threshold. The extra documentation is a message. The trouble, as the CONCERN investigators put it, is that the message has no reliable audience, since the rest of the team does not see the pattern, only the entries.
The trial randomized 74 units at two health systems, 60,893 encounters in all. Where the care team could see the score, the adjusted hazard of dying in hospital fell by 35.6% and length of stay by 11.2%, and unanticipated transfers to intensive care rose by about a quarter, which is what you would expect if the system was surfacing patients earlier rather than later. The investigators excluded hospice patients, anyone with a DNR or DNI order, and specialty units such as oncology, so the effect was measured on the general floors where we expect people to recover.
The model added no diagnostic insight of its own. It made one nurse’s worry visible to the physician, the charge nurse, and the next shift. That is the whole mechanism which means the model is only as good as the behavior feeding it.
Is this a nursing quirk, or does every clinician do it?
Every clinician does it, and the largest study of the physician version reached a conclusion that should unsettle anyone who thinks of a timestamp as metadata. Denis Agniel, Isaac Kohane, and Griffin Weber at Harvard looked at 669,452 patients seen at two Boston hospitals over one year and asked, for 272 common lab tests, which predicted three-year survival better: the value that came back, or the circumstances of the order, meaning the hour, the weekday, and how long it had been since the previous draw.
The circumstances won. Of the 174 tests in which one dimension or the other carried information, the ordering variables beat the value in 118. For tests ordered in pairs, the interval since the prior draw was the single strongest predictor more often than the result itself. A normal white cell count drawn at 4 a.m. carried a worse prognosis than an abnormal one drawn at 4 p.m., and the authors’ explanation is one every hospitalist will recognize: “doctors generally only see sick patients in the middle of the night.” The order is an assessment. The physician has already decided the patient is sick, and the timestamp is where that decision leaked into the record.
Nursing had its own version of this finding years earlier, when a Columbia group mined 15 months of documentation and found that patients who died had accumulated more optional vital signs and more free-text comments in their last two days than patients who lived. The data-quality people had a name for the phenomenon before the clinicians noticed it. Goldstein and colleagues at Duke called it informed presence: being in the record at all, and how often, is not random, and George Hripcsak and David Albers had earlier described EHR data as an indirect measure of the patient filtered through the healthcare process. This treasure trove of data is a resource that may be about to go missing because ambient AI and similar tools might blunt the signal.
How much of a nurse’s worry can you actually measure?
A Dutch group followed 3,522 surgical patients on three wards, asking nurses to record whether they were worried and which of nine indicators explained why. A conventional early warning score built from vital signs alone identified unplanned ICU admission or unexpected death with an AUROC of 0.86, meaning it ranked a patient headed for trouble above a stable one 86% of the time. Adding the nine worry indicators lifted that to 0.91. Two indicators dominated: a change in breathing, with an odds ratio of 15.2, and the nurse’s own gestalt, recorded as a change in behavior, “doesn’t look good,” or something in the eyes, at 14.6. Free-text unease matched the respiratory exam.
The same study contains a finding that nobody expected. Of the patients who deteriorated, 29% had gaps in their recorded vitals. Of the patients who stayed well, 76% had gaps. Respiratory rate, the vital sign most often skipped by nurses, was missing for 22.5% of the patients who went downhill and 70.3% of those who did not. The authors filed this under limitations. Read it the other way and it is the CONCERN signal showing up on three ordinary surgical floors, years before anyone trained a model on it. Nurses fill in the whole set for the patients they are uneasy about.
What happens to a model that was trained on behavior we are about to remove?
It degrades, quietly, and the literature already has a name for it. Before the ambient part, though, the corollary: if clinician behavior is this informative, the commercial models your health system licensed are almost certainly consuming it. A team led by Brett Beaulieu-Jones and Kohane trained deep-learning models on 42.9 million admissions using nothing but what was ordered, performed, and billed on day one, with no vitals, labs, or notes, and predicted in-hospital mortality with an AUC of 0.89. Full-record models published by others reach about 0.95. The gap is small enough that the authors warn a model at that level may be reading clinicians rather than patients, encoding what the team already suspected rather than adding to it.
That is not an indictment. I argued last month that AI relocates judgment rather than removing it; here is a case where the relocation is the product. A model that broadcasts one nurse’s concern to the whole team is doing something useful with judgment that already exists. But a model built on behavior is hostage to that behavior. Samuel Finlayson and colleagues wrote to the New England Journal of Medicine in 2021 to name the failure mode for clinicians: dataset shift, the gap between the data a model learned from and the data it now sees. Their headline example was the University of Michigan pulling Epic’s sepsis model in April 2020, when the pandemic broke the link between fever and bacterial sepsis. They sort the causes into three groups, and one group is “changes in behavior,” with new reimbursement incentives as the example. Nothing in their logic confines it to incentives.
Now hold the roadmap up against that column: ambient documentation for nurses, continuous vital-sign capture from wearables, AI-drafted assessments. Each one, by design, replaces a choice a clinician used to make with a process that runs on its own. The reassessment at 3 a.m. becomes a scheduled reading. The blank respiratory rate that meant “this patient is fine” is filled in for everyone by a sensor. The optional comment is subsumed into a generated narrative. CONCERN’s inputs do not become wrong; they become uniform, which for a model that reads variation is the same thing. As far as I can find, no one has published on whether ambient capture flattens the surveillance signal.
What should a health system do before it goes ambient?
First, take inventory of the breadcrumbs before any tool that automates documentation goes live. For each one, write down the clinician behavior it eliminates, what that behavior was communicating, and which models, alerts, and dashboards downstream consume it. Put the question to your vendors directly: which of your model’s features derive from what clinicians do, as opposed to the patient’s physiology? A vendor who cannot answer has told you something about how well they understand their own model. This is the same discipline I asked for when the topic was override rates: instrument the human decision, not just the machine’s.
Second, when a signal is going to disappear, replace it deliberately. If ambient capture makes worry invisible in the pattern of charting, give nurses a one-tap field for it that lives inside the new tool rather than beside it, and the Dutch indicators are a defensible starting list, since “doesn’t look good” has stronger evidence behind it than most of the alerts we ship. Then decide, in writing, where that signal goes. Routed to the care team, it is a rescue tool, and CONCERN shows what it can do. Routed to a productivity dashboard, it is surveillance of the workforce, built from the same keystrokes. Clinicians can tell the difference within a week, and once they can, the signal is gone for a different reason.
What does a quieter chart cost?
Possibly nothing, if you know what it was saying. But a metric that goes obsolete is not the same as a signal that goes silent, and the people best placed to tell the two apart are the nurses whose habits generated the signal in the first place. Ask them what they do differently when a patient scares them. Then check whether anything in the new workflow can still hear it.