Simar Bajaj and Joseph Sakran 

What happens when medical students rely on AI – and never develop their own judgment?

AI’s danger isn’t just in experts losing the ability to reason. It’s that trainees may never learn how to do so in the first place
  
  

dcotor using tablet computer
‘Patients will need doctors who can stand apart from the machine long enough to know when it is wrong, incomplete, or right for the wrong reason.’ Photograph: gorodenkoff/Getty Images/iStockphoto

In healthcare, there’s growing concern over doctors becoming less clinically adept as they increasingly rely on AI tools. But what about the trainees – medical students, residents and fellows – who are using these tools before they’ve built their own clinical judgment? The idea of deskilling implies that someone possessed an ability and then lost it. Here, the danger is not just deskilling but never-skilling. Although a doctor who has forgotten how to reason is recoverable, one who never learned how may not be.

OpenEvidence, essentially an AI chatbot for clinicians, has given this concern its most concrete form. About two-thirds of US doctors actively use OpenEvidence, asking about puzzling symptoms, drug interactions, and clinical guidelines, getting responses within seconds, anchored in the latest research. Trainees, unsurprisingly, have also begun to use this AI tool in many of the same ways – but at a far more formative stage.

For example, trainees once asked to build a list of potential diagnoses might struggle and offer an incomplete set, learning what they missed, sometimes painfully. Now, trainees can simply ask OpenEvidence and get a nearly perfect answer, complete with possibilities they might have never considered and none of the embarrassment of having overlooked them. Repeating this answer on the wards may make the trainee look prepared and even impress the supervising doctor.

However, this performance can also conceal the very deficit that training is meant to reveal: that the struggle is the point. Medical training, more than most professions, is an apprenticeship. A student becomes a resident, a resident becomes a fellow, and a fellow becomes an attending – every step shaped by failure, uncertainty and increasing responsibility. With years of repetition and watchful supervision, the habits of clinical reasoning slowly become part of the physician’s inner architecture.

Technology has long shifted how people learn medicine, from advanced imaging to electronic medical records. But AI is different, not just expanding what doctors can see but inserting itself into the cognitive machinery that training is meant to build. As these tools become more capable and the physician’s role increasingly involves supervising them, experienced clinicians may have enough intuition and independent judgment to critically evaluate the machine’s answers. But for trainees whose understanding of medicine is being formed alongside AI, the relationship is more fraught. Can they really question the reasoning that shaped their own? What happens when the generation trained by AI becomes the generation responsible for catching its mistakes? With unchecked use among trainees, we risk creating supervisors of reasoning before we create reasoners.

The stakes of that question are growing: a recent study in Nature Medicine found that tools pulling from the latest medical literature, like OpenEvidence does, can be less reliable than they appear and, in some cases, less accurate than general-purpose AI chatbots. The problem of misplaced trust is already embedded in the AI that trainees are using today.

To be clear, many trainees sense the trap, telling us they know that tools such as OpenEvidence can become a crutch. But these trainees also feel stuck in an arms race: if everyone else is using AI to sound more prepared, opting out feels like unilateral disarmament. The solution, then, cannot rest on individual restraint. It has to be structural.

That is why medical schools and residency programs need to shape not just whether trainees use AI, but when. No one can police every search on every phone, nor should they. But supervising doctors can build a simple expectation – reason first, consult AI second – and assess accordingly. Trainees should have to make their unaided first pass visible, committing to a leading diagnosis, naming the dangerous possibilities to rule out, and explaining what to do next. In practice, that might mean a resident who admits a patient overnight first writes a brief “pre-AI assessment” after the history and physical exam. On rounds, when a new lab result or symptom changes the case, the attending might need to pause the team – before anyone can consult AI – to ask how this changes the diagnosis or treatment plan.

As AI becomes more deeply integrated into medicine, this will feel cumbersome and inefficient. But such friction is purposeful: the learning scientists Elizabeth and Robert Bjork describe how “desirable difficulties” slow performance in the moment but improve retention and transfer of skills over time. In fact, used after an independent attempt, AI could actually serve as a powerful tutor, showing trainees what they missed and what they overemphasized.

Sequencing, however, may not be enough on its own. Aviation thus offers a useful precedent: pilots in training are not taught to avoid autopilot but to preserve their manual competence. The Federal Aviation Administration even advises pilots to maintain manual flying skills by periodically disengaging automation and hand-flying. Medicine needs similar discipline, with trainees required to periodically work through no-AI cases and assessed on their unaided reasoning to reveal potential drift.

Finally, trainees should be taught to interrogate AI itself. Programs could run the medical equivalent of flight simulator drills, built from real clinical cases: for example, a polished AI-generated assessment with a subtle flaw. Afterward, attendings could debrief not only whether the trainee reached the right answer but also when they trusted the tool, when they questioned it, and when they found the flaw. Just as important, attendings should mix in AI outputs that are perfectly accurate, so students learn not reflexive skepticism but disciplined judgment.

None of this is an argument for making medical training harder for its own sake or romanticizing humiliation as pedagogy. In every generation of medicine, there is a temptation to confuse difficulty with virtue, but the struggle to independently reason through a patient’s case is not hazing but a core competency.

AI is here to stay, and patients stand to benefit from its speed and reach. But patients will also need doctors who can stand apart from the machine long enough to know when it is wrong, incomplete, or right for the wrong reason – doctors whose reasoning is not subordinated to it. Although AI can reason over the facts it is given, a trainee who has seen pneumonia that looks like pneumonia, then pneumonia that looks like heart failure, then heart failure that looks like pneumonia, develops a richer bedside judgment: what to notice, what to question, and when a familiar pattern should be distrusted. That is what medical training is trying to produce. AI should help augment this, not replace it.

  • Simar Bajaj is a medical student and Knight-Hennessy Scholar at Stanford University School of Medicine, as well as an award-winning journalist

  • Dr Joseph V Sakran is a trauma surgeon and public health expert who serves as executive vice chair of surgery at Johns Hopkins Medicine

 

Leave a Comment

Required fields are marked *

*

*