Medicine Was Never Clean
- Graham Walker, MD
- 22 hours ago
- 3 min read
And that's exactly why clinicians are the right people to judge clinical AI
The Future of Clinical Decision Support is a new series from MDCalc exploring how AI, evidence, and physician judgment are reshaping medicine. These essays are intended to spark discussion about where technology is taking clinical decision support – and where physicians feel it should go.
I've been practicing medicine long enough to remember when we thought evidence-based medicine would end clinical uncertainty. It didn't. We got better tools for quantifying and communicating that uncertainty — which is a different, more honest thing.
The same conversation is happening now with AI, and I think we're making the same mistake: expecting the technology to resolve what the technology fundamentally cannot resolve.
Medicine has always run on noisy, imperfect inputs.
The troponin that's borderline.
The monitor that says "V Tach" but the patient has an essential tremor and is sitting there reading a book.
The chest X-ray a radiologist calls "possible infiltrate."
The patient who doesn't fit any risk stratification tool because she's 81, has nine comorbidities, and the original study enrolled 43-year-old men.
Clinicians are not passive recipients of data, information, or evidence. We are professional filters of it. That's not a workaround; it's the job.
The problem with the newest generation of AI clinical tools is not that they're imperfect. Everything in medicine is imperfect. The problem is that some of them are imperfect in a way that defeats the filter.
When a lab value is borderline, you know it's borderline. When an imaging report hedges, you can read the hedge. When a risk score is validated only in a narrow population, the methods section tells you so — if you look. The signal of uncertainty is legible, and legible uncertainty is something physicians know how to handle.
Generative AI fails differently. It produces fluent, confident prose that sounds like a senior attending — even when it's drawing on case-series-grade data, or extrapolating beyond its training, or simply wrong in a way that's invisible unless you happen to already know the answer. The output looks like signal. And a filter calibrated for legible noise doesn't catch a confident hallucination the way it catches a hedged imaging report.
This isn't an argument against AI in clinical decision support. But it is an acknowledgement that it doesn't even work at all in medicine without physicians who know how to exercise their own clinical judgment.
A clinical AI tool that presents case-series evidence and RCT evidence in the same confident tone is not helping physicians practice evidence-based medicine, it's attempting to replace evidence hierarchy.
Clinicians, I don't think, are not demanding perfect AI. We have settled for less from our EKG machines and our lab value standard deviations and our motion artifact during CT scans. We're demanding honest AI. We've worked with imperfect tools our entire careers; we're actually quite good at it. What we cannot work with is a tool that hides its imperfections behind fluency, because that's the one kind of noise our filters weren't built for.
The standard that should follow every AI-generated clinical output is the same one we've always applied to evidence: What was this based on? How strong is that evidence? Where does it not apply?
Trust in medicine has never come from certainty. It comes from showing your work.

