Why the $20 test lost to the echocardiogram
By Craig Joseph, MD ·
Every order set I have helped build rests on an assumption nobody writes down: that clinicians do the cheap, fast, frictionless thing more often than the expensive, slow, gated thing, and that our job is to move the right option closer to the top of the list. Remove clicks. Pre-check the box. Put the test one keystroke away.
A new study in Pediatrics runs that assumption against 87 million patients, and the assumption loses badly enough to be worth a governance meeting.
What did the study actually measure?
Whether children newly diagnosed with high blood pressure got the workup the guideline asks for. James Nugent at Yale and David Kaelber at Case Western Reserve and MetroHealth queried the TriNetX network — 51 health care organizations, 87 million patients — for patients aged 1 to 17 carrying a first coded diagnosis of primary hypertension, then counted what was ordered in the six months that followed.
Not much of anything. Urinalysis, 19.9%. Electrolytes and creatinine, 17.7%. Lipid profile, 8.1%. Renal ultrasound, which the guideline asks for in children under six, reached 12.2% of that age group. Ambulatory blood pressure monitoring, in the children old enough to tolerate the cuff, reached 7.7%.
The most frequently completed test, at 21.1%, was the echocardiogram.
Why is the echocardiogram the wrong winner?
Because of sequence, not because it is a bad test. The 2017 AAP guideline reserves echocardiography for the point at which you are considering starting a medication, which typically arrives six to twelve months into lifestyle management. The urinalysis and the chemistry panel are meant to happen at the front of the process, on everybody.
So the ordering pattern is inverted relative to the recommendation, and inverted in the direction that costs the most. A urinalysis requires a specimen cup and about twenty dollars. An echocardiogram requires a referral, frequently a prior authorization, an appointment, a technologist, and a pediatric cardiologist to read it. The test that is fast, cheap, and indicated for every patient finished behind the test that is slow, expensive, and indicated later for some of them.
Every implementation framework I have used predicts the opposite result.
We build order sets on the theory that friction suppresses ordering. If friction were the binding constraint here, the specimen cup would have won by a mile.
Are your order sets still running the 2004 guideline?
This is the possibility I would check first, because it is the cheapest one to rule out and the easiest to fix. The 2017 guideline replaced the 2004 Fourth Report, and one of the specific things it changed was echocardiography. The Fourth Report recommended an echocardiogram and a renal ultrasound for every hypertensive child. The 2017 update pulled the echo back to the medication decision point.
Guidelines change. Order sets, preference lists, referral templates, and the mental model of the physician who trained in 2009 do not change on the same schedule. The study found that echocardiography remained the most common test and that renal ultrasound rates did not move at all across the decade — which is exactly the fingerprint you would expect from a retired recommendation still running in production.
There is a decent chance that somewhere in your build there is a pediatric hypertension order set assembled before 2017 and never revisited, quietly making the old guideline the path of least resistance. That is not a clinician problem. That is a maintenance problem, and it belongs to informatics.
What if the real variable is what happens when the result comes back?
Sort the tests not by cost or effort but by how much unfinished work each result creates, and a second pattern appears. The echocardiogram produces a disposition. Normal means the structural question is closed and the family gets good news. Abnormal means cardiology takes the next several decisions. Either way, the physician opening that result knows what happens next and it frequently is not their problem anymore.
Now take the same child’s urinalysis showing trace protein, or a creatinine sitting at the top of the reference range. There is no disposition waiting. There is a repeat test, a judgment call about whether the number means anything, an unclear threshold for calling nephrology, and a conversation with a parent that has to happen before anyone knows what they are talking about. The result does not route anywhere. It stays in the physician’s head.
Ambulatory monitoring is the purest case. A positive study hands the child back to you, confirms you are now managing a chronic condition, and comes with no handoff attached. It answers a diagnostic question and opens an operational one.
The tests that got completed were the ones whose results resolve. The tests that got skipped were the ones whose results open a loop somebody has to keep carrying between patients. No workflow map that counts clicks will ever surface that distinction, and it may explain more of the variance than the clicks do.
I want to be careful here, because the mechanism is not conscious. No physician has ever declined to order a urinalysis on the grounds that an abnormal result would generate work. The deferrals are individually reasonable. We can check the pressure again at the next visit. We can order labs then. There is no need to frighten this family today. Aggregate a few hundred thousand defensible small decisions and they surface in a national database as 19.9%.
Are we measuring adherence to a diagnosis nobody made?
Largely, yes, and this is the part that should worry a quality committee more than the testing rates do. Only 0.5% of the children with an ambulatory visit in the cohort ever received a hypertension diagnosis code. Expected prevalence runs somewhere between 2% and 5%.
I would not lean too hard on that arithmetic — an incident coding rate across a seven-year window and a point prevalence estimate are not the same measurement, and treating them as interchangeable is how honest findings get oversold. But the direction is corroborated by older and more direct work. A 2007 JAMA cohort study that went back to the actual blood pressure readings rather than the diagnosis codes found that of the children who met criteria for hypertension, 26% had it documented anywhere in the chart. Three quarters of them were carrying an undiagnosed chronic condition. Kaelber was an author on that paper too, which means he has now spent nearly twenty years demonstrating the same failure from two different angles.
Elevated blood pressure in a child is the least-resolved result in the entire sequence. It is not a diagnosis and it points at no single action. It arrives mid-visit, in a schedule with no slack, and the correct response to it is to open yet another loop: confirm the measurement, pull the prior readings, decide when to recheck, explain a concern without frightening anyone, and make sure this child does not vanish into the roomy gap between “follow up later” and someone actually being responsible.
We are measuring the workup of a diagnosis that mostly never gets made, and when it does get made, it mostly does not get confirmed by the test the guideline requires. The denominator has already collapsed before the quality measure starts counting.
What would a system that finishes the work look like?
Two changes, and neither is an alert.
Centralize the workup, because the volume is too low to distribute. The study found roughly 50,000 incident cases across 51 organizations over seven years. That is about 140 per organization per year, which means a given pediatrician in a large system diagnoses this perhaps once annually. Nobody builds muscle memory at that frequency, and expecting them to is the same category error I have written about when a health system tries to nudge its way out of a judgment problem. Healthcare already accepts this logic for anticoagulation clinics, oncology navigation, and heart failure bridge programs. Pediatric hypertension has the same profile: rare enough that distributed expertise is a fantasy, protocolized enough that most of the work does not need a physician. A pooled team can own device logistics, prior authorizations, scheduling, family education, and follow-up without taking the primary care relationship away from anyone.
Give the software a job, not a message. Traditional decision support interrupts a person and asks that person to do something. A more useful system would find the elevated readings, check what has already been done, stage the orders, start the ambulatory monitoring logistics, and — this is the part that matters — carry the disposition logic for the results. Trace protein arrives with a defined next step. A borderline creatinine arrives with a protocol instead of a question mark. An abnormal ambulatory study starts a pathway instead of generating one more in-basket message to be considered between patients. Real judgment stays with the physician. Reconstructing a routine process from memory once a year does not qualify as real judgment.
Most of the current excitement about AI in the EHR is about drafting text. Drafting text is useful. It is also the least interesting thing you could ask software to do, compared with making it accountable for moving a clinical process to a finished state.
What should the dashboard show?
Not alert firings, and not order counts. Report the funnel. Of the children with an elevated reading, how many were recognized? Of those, how many were diagnosed? How many had the right tests ordered, and then completed? How many results produced a documented action? How many loops are still open right now, and who is holding them?
Most systems already have this data. The hard part is the agreement that a clinical process gets measured through completion rather than at the moment the EHR captures a click.
A caveat the authors are honest about and I will repeat: they could not see the specialty of the ordering clinician, so some of those echocardiograms were surely ordered by cardiologists after referral, and echocardiography is not a foolish test in this group — the literature they cite finds left ventricular hypertrophy in 33% to 41% of young people at diagnosis. Renal ultrasound has the highest yield of anything for secondary causes. The guideline’s objection is about sequence, not value. None of that explains an ambulatory monitoring rate under 8%, and none of it rescues the denominator.
The build I keep thinking about is one I did years ago: a pediatric blood pressure alert that calculated the percentile silently, fired when it should, got acknowledged, and produced orders. For a long time I blamed alert fatigue for what happened next, which was nothing. I think now the alert was doing precisely what we designed it to do. It found something that might be wrong, told the busiest person in the room, and left. We never wrote down whose job the next step was, so it became the job of whoever we had just interrupted.
Before your next decision support build ships, somebody should have to write down what happens after every plausible result and who owns it. “The ordering clinician decides” is a permissible answer. It should never be the answer nobody chose.