top of page

The Immortality of Bad Ideas

Writer: Justin Healy
Justin Healy
Sep 25
5 min read

The Future of Clinical Decision Support is a new series from MDCalc exploring how AI, evidence, and physician judgment are reshaping medicine. These essays are intended to spark discussion about where technology is taking clinical decision support – and where physicians feel it should go.


Imagine we’re in 1972 and LLMs have (somehow) just been invented.


In 1972, doctors are feeling pretty good. In the past 20 years, a nascent pharmaceutical industry has produced new antibiotics, new psychiatric drugs, and new antihypertensives. Propranolol—with its physiology-based pharmacology—suggests a revolution in drug development. The contraceptive pill has ushered in a revolution in society more generally. Patients with end-stage kidney disease can be offered dialysis, and surgeons are now able to operate on a living human heart.


Imagine that into this confident, rapidly progressing field of medicine comes this near-magical technology. GPT-’72 can access and synthesise the world’s medical knowledge and provide context-specific responses to some of the most difficult medical cases.

But the doctors of 1972 aren’t stupid. They know that this new technology will need to be evaluated. So they design question banks, unstructured clinical scenarios, and rigorous benchmarks across a range of indicators.


And the models perform astonishingly well. When asked about children with fever, they are quick to recommend aspirin. When quizzed about post-MI care, they suggest prophylactic lidocaine. And when queried about homosexuality, they direct you to the DSM. This excellent performance across rigorous medical benchmarks has doctors around the world anxiously contemplating their professional futures.


Seems ridiculous, right?


One response to this thought experiment is that 2026 is not like 1972. After decades of evidence-based medicine, we have the data to make better decisions. We have learnt not to blindly trust convention and tradition, and we have legions of researchers producing data to objectively support our decision-making. However, that data can be patchy, inconclusive, or may simply not apply to the patient sitting in front of you. Good luck finding the guideline that tells you exactly how to manage the anxious adolescent sitting in primary care. Or the paper that guides management of a borderline troponin with non-specific symptoms.


So, in the absence of a clear evidence base for our decisions, we rely on much the same tools as they might have used in 1972: expert consensus.


Many AI papers use expert consensus to evaluate model performance. They get a group of experts and ask them to rate how the model does, typically across several domains, such as accuracy and safety. But medicine is hard. And even subject-matter experts can, and do, disagree.


In a study conducted in Kenya, an LLM was deployed inside an electronic medical record to support clinical decision-making. External clinicians evaluated the quality of documentation with and without AI support. You might hope that two experts reading the same note against the same marking scheme would broadly agree. But the team found substantial disagreement between raters—up to one third in some domains.[1]


An even more extreme example comes from a recent Stanford study that asked psychiatrists to evaluate LLM outputs in mental health settings. Despite recruiting three psychiatrists and carefully briefing them on the marking criteria, agreement between the three was consistently poor.[2] The psychiatrists had fundamentally different ideas of what a “good response” looked like.


In cases where expert agreement is low, this may suggest that the answer is fundamentally ambiguous or that it depends on the values different clinicians bring to the problem. One approach some studies take is to treat this disagreement as noise to be averaged out: split the difference between perspectives and evaluate the model on that basis. But, as the authors of the psychiatry paper point out, this noise contains real information. It may reflect genuine disagreement among subject-matter experts.


A model’s performance against a contested ground truth may be less interesting than the fact that the ground truth is contested in the first place.


Not all AI research reveals fundamental epistemic differences between experts. The NoHarm benchmark, developed by researchers from Stanford and Harvard, tests LLMs across a range of clinical scenarios and looks for examples of harmful acts or omissions. When developing the benchmark, the researchers reported very high levels of agreement between clinicians when asked to rate the harmfulness of various management plans.[3]


But inter-rater agreement answers only one question: whether experts agree with one another. It does not tell us whether they are right.


Our imaginary team of clinical AI researchers in 1972 may have found high levels of agreement on the role of paediatric aspirin, prophylactic lidocaine, or medicalised sexuality. Models evaluated on those frameworks may have scored extremely well—far better, perhaps, than pesky humans who had doubts about some of the medical shibboleths of the day. The long list of medical reversals tells us that eminent consensus is not always synonymous with truth.


This is not an argument against expert consensus. In a world full of unknowns, the considered view of specialists is usually the best we have. But we should be clear-eyed about its limitations. Low levels of agreement between experts may not be noise to be averaged away, but rather a signal to be interrogated. High levels of agreement may be superficially reassuring, but may simply reflect the assumptions of the time. In either case, evaluating medical AI against expert consensus risks obscuring genuine uncertainty and reinforcing the assumptions embedded in that consensus.


There are a few things that can reduce this risk. Researchers using expert consensus should be explicit about rates of inter-rater agreement and about how rating scales were developed. If inter-rater agreement is relatively low, this should not be hidden in an appendix, but discussed as an important finding in and of itself. Lastly, benchmarks should be designed to be flexible enough to accommodate a range of expert opinion, so they can be updated as evidence and ideas evolve.


There’s an idea, often attributed to Max Planck, that “science advances one funeral at a time”: old ideas sometimes lose their authority only when their defenders do. Whatever we think of that rather bleak account of scientific progress, we should make sure medical AI doesn’t give our bad ideas immortality.


Dr Justin Healy

August 2026


[1] Reference: https://cdn.openai.com/pdf/a794887b-5a77-4207-bb62-e52c900463f1/penda_paper.pdf. Substantial disagreement was defined as raters disagreeing by more than 1 point on a Likert scale. Furthermore, Fleiss’ kappa for the study indicated only fair agreement across all domains.

[2] Jafari K, Rust PUN, Eddy D, Fraser R, Vasan N, Djordjevic D, Dadlani A, Lamparth M, Kim E, Kochenderfer M. Expert evaluation and the limits of human feedback in mental health AI safety testing. In: Proceedings of the 2026 ACM Conference on Fairness, Accountability, and Transparency (FAccT ’26). New York: Association for Computing Machinery; 2026. doi:10.1145/3805689.3812332.

[3] Wu D, Nateghi Haredasht F, Maharaj SK, et al. First, do NOHARM: a medical safety benchmark and randomized study of physician and AI teaming on clinical consultations. arXiv. Published December 1, 2025. Revised July 13, 2026. doi:10.48550/arXiv.2512.01241.



About MDCalc

Since 2005, MDCalc has built clinical decision support around transparent evidence, physician judgment, and trust. We believe those same principles should guide the next generation of clinical AI.

Have a perspective to share? If you'd like to contribute an essay or start a conversation, we'd love to hear from you. Contact us at team@mdcalc.com.
 
 

Don't Miss an Update!

Thanks for signing up!

Unsubscribe at any time.

  • LinkedIn
  • facebook
  • Bluesky
bottom of page