top of page

Leave Some Stones Unturned

Writer: Graham Walker, MD
Graham Walker, MD
5 days ago
7 min read

The Future of Clinical Decision Support is a new series from MDCalc exploring how AI, evidence, and physician judgment are reshaping medicine. These essays are intended to spark discussion about where technology is taking clinical decision support – and where physicians feel it should go.


Eight weeks ago, my left ankle looked like this. ⬅️


Every AI tool I showed it to told me to get an x-ray. I didn't.


I missed a step hiking outside Lisbon and knew immediately: uh oh, this is a bad one. It swelled fast in my shoe, I couldn't bear weight, and I bunny-hopped into a Portuguese pharmacy to buy an ankle brace. It took a wheelchair in Lisbon and a cane across three airports to get me home.


I'm also an emergency physician, so I did what doctors do when we become patients: I examined myself repeatedly and decided it was probably a sprain. No x-ray. No doctor. Motrin, ice, and waiting.


Yes, I know. Inability to bear weight is an Ottawa ankle rule criterion, and by the book I was a candidate for imaging. But the rule is built to miss almost nothing, which means most people who trip it don't have a fracture. I made a judgment call.


Eight weeks later, here I am, standing on it — and attempting pistol squats at the gym — with no pain whatsoever.


Generative AI never runs out of Things


I keep thinking about that ankle because finding reasons to do A Thing is about to get a lot easier. The Thing could be an x-ray, a lab, another diagnosis to entertain, another prescription. Generative AI is spectacularly good at coming up with Things.


Imagine I'd gone to an emergency department that first day and an AI reviewed my chart: significant swelling, ecchymosis, lots of pain with weight bearing. Given the degree of swelling and difficulty bearing weight, consider radiographs to exclude occult fracture.


"Consider" an x-ray is a very easy recommendation to defend. So is "order an x-ray." The marginal cost to the AI of suggesting it is basically zero. The marginal cost of actually doing everything AI can reasonably suggest is not.


The math of 15,000 extra x-rays


Imagine 100,000 patients show up with ankle injuries. Before AI, 50,000 get x-rays. Then we add an assistant that's extremely good at noticing every reason an x-ray might help. It doesn't order anything. It doesn't force the clinician to do anything. It just helpfully points things out, and imaging goes from 50% to 65%.


Maybe it catches some fractures we used to miss. Fantastic. Put those cases in the PowerPoint. Say it finds 50. We've also just done 15,000 additional x-rays, or 300 films for every fracture found. Those cost patients and the system money, some generate findings that require more imaging or a referral, and all of it has to be weighed against whatever those extra catches were worth.


And the patient may even appreciate the x-ray. "Better safe than sorry," they often say to me. Or say we do find a tiny, clinically-meaningless avulsion fracture (a tiny speck of bone that got ripped off that we don't do anything about). Now they're a believer: I had a fracture and we almost missed it!


So, look, there are incentives everywhere to do more. To just get one more x-ray. I get it.


But the question isn't whether the AI found something. It's:

  1. Did the thing it found actually matter? and

  2. What did we have to do to everyone else in order to find it?


The last cases are the expensive ones


"Don't miss anything" gets harder as medicine gets better. Take appendicitis. Emergency medicine is already really, really good at it. Not perfect, but between history, exam, labs, decision tools, imaging, observation, and a century of physicians growing progressively more terrified of missing it, we've gotten pretty damn good. (In one recent study of more than 120,000 pediatric ED visits for abdominal pain in Michigan, there were just 141 delayed diagnoses of appendicitis, about 1 in every 850 visits.)


Now suppose we want to get even better. Where are the remaining cases? Not among the patients with classic right lower quadrant tenderness, leukocytosis, vomiting, and a CT showing an inflamed appendix. We found those. The rest are buried in the vastly larger population of patients who don't look much like they have appendicitis: vague pain, atypical exam, the kid we figured had a stomach bug.


To find more of them, you have to test more of those people. Way more. None of that is an argument for missing appendicitis. It's just unfavorable math: the lower the prevalence, the more tests it takes to find each additional case.


"Why not?" times a million


This is where "just to be safe" stops being so obviously safe. For one patient, an extra test often seems trivial. Why not get the x-ray? Why not add the troponin? Why not get the CT? Why not ask the AI to make sure we're not missing something? Each decision can be perfectly defensible. But multiply "why not?" across millions of encounters and I promise you, we have a very different problem. Why not has no natural stopping point.


This is why I've gotten uncomfortable with healthcare AI evaluations that reward detection and ignore what happens downstream. Finding more disease isn't the same as improving health, and generating more possibilities isn't better reasoning. Preventing one miss is a bad trade if it costs a mountain of low-value care elsewhere.


The denominator matters.


Restraint leaves no receipts


And restraint is almost invisible. There's no insurance claim showing I appropriately skipped an x-ray. No CPT code for "emergency physician correctly decided to leave his ankle alone." Nobody puts me on a dashboard under X-RAYS PREVENTED BY NOT DOING ANYTHING. My ankle just got better.


Healthcare is full of these non-events: the CT not ordered, the antibiotic not prescribed, the incidentaloma never discovered, the unnecessary referral, the low-risk diagnosis not chased to absolute certainty. Good clinicians make these decisions constantly. They're just much harder to count than the things we do. Action leaves receipts. Restraint doesn't.


But wait... now add money


We have lots of ways to reward doing something and remarkably few ways to reward correctly doing nothing. Now add money. A hospital gets paid when the x-ray happens, not when it doesn't. The radiology group gets paid to read it. A fee-for-service system earns revenue by providing care, not avoiding it. The doctor can bill at a higher level for doing tests. And an AI company trying to prove that its product improves care will have a far easier time showing the fractures it found than explaining how often it correctly recommended nothing.


None of this requires anyone to behave badly. That's what makes it interesting.

  • The patient wants reassurance the ankle isn't broken. Reasonable. ✅

  • The physician doesn't want to miss a fracture. Also reasonable. ✅

  • The AI vendor wants to show it caught something. Fair. ✅


Nobody is waking up plotting to make American healthcare more expensive. Hospital executives, insurers, physicians, and AI companies mostly want it to cost less, sincerely, while working inside systems that pay them when it costs more. If you earn a percentage, a bigger number is a bigger paycheck.


Then we drop generative AI into the middle of that, with a nearly unlimited ability to suggest more Things.


The oldest diagnostic tool we have


There's an alternative, and medicine has used it forever: time. It's what I did with my ankle. I waited it out.


We do the same with abdominal pain all the time. If I think your probability of appendicitis is low enough, sometimes the right answer is: I don't think you have appendicitis right now. Here's what I'm worried about. Here's what to watch for. If you're worse tonight or tomorrow, come back, I'll examine you again, and maybe imaging makes sense then.


That's not failing to make a diagnosis. That's clinical judgment under uncertainty. Biology keeps happening after the patient leaves the room. Symptoms evolve, diseases declare themselves, injuries heal, pretest probabilities shift. Sometimes six or twelve hours of biology tells you more than another $5,000 of diagnostics right now.


The problem is that we've made the CT incredibly easy and the second evaluation incredibly annoying. The patient most definitely gets a second bill in America. They may wait another six hours in the waiting room, take another unpaid day off work, and see a different clinician who starts from scratch. If we want clinicians and patients to tolerate appropriate uncertainty, we have to make reassessment cheap and easy: same clinician where possible, no second bill, a clear way back in. Make waiting a legitimate diagnostic strategy instead of an incomplete encounter.


We should make it easier to buy time than to buy another test.


Teaching clinical AI when not to speak


I could wait because I knew exactly what would make me change my mind. Most patients don't. That safety-netting is a job AI could actually be good at.


But AI could just as easily push the other way, if we tell it to. The easiest clinical AI to build is the one that always has something else to say: another possibility, another warning, another test, another rare diagnosis not yet excluded. It will happily leave no stone unturned, if you let it. Every recommendation can be individually reasonable while the aggregate is terrible.


In my last piece, I argued that clinical decision support should be quieter: monitoring far more of medicine while interrupting clinicians far less. Restraint goes one step further. The best clinical AI shouldn't just know when not to speak. It should know when not to recommend another Thing.


That's a lot harder than generating a comprehensive differential. It means understanding pretest probability, thresholds, downstream consequences, and the harms of the interventions themselves. And it means accepting that the goal of medicine was never to drive uncertainty to zero. It's to make good decisions, despite the uncertainty.


Eight weeks ago, an AI could have handed me a perfectly defensible list of reasons to image my ankle. Instead, I gave it time.


Voltaire said the art of medicine consists of amusing the patient while nature cures the disease.


(Nature, notably, won't send you a bill.)






About MDCalc

Since 2005, MDCalc has built clinical decision support around transparent evidence, physician judgment, and trust. We believe those same principles should guide the next generation of clinical AI.

Have a perspective to share? If you'd like to contribute an essay or start a conversation, we'd love to hear from you. Contact us at team@mdcalc.com.
 
 

Don't Miss an Update!

Thanks for signing up!

Unsubscribe at any time.

  • LinkedIn
  • facebook
  • Bluesky
bottom of page