Clinical reasoning and AI

Clinical reasoning depends on what a clinician asks, which findings they consider relevant, and how they revise a working explanation as evidence arrives. A correct answer on a medical exam captures only part of that process.

Evidence relevance and physician-AI collaboration

MedPAIR, by Yuexing Hao and colleagues, compares physician trainees' sentence-level relevance judgments with those of language models in medical question answering. The available arXiv preprint finds substantial disagreement. Filtering information according to trainee relevance labels improves accuracy in many settings, with exceptions for some model and dataset combinations. A correct final answer therefore does not establish that a model selected the same evidence as a physician trainee.

Relevance labels are contextual judgments. Agreement with a trainee is not proof of safe clinical reasoning, and disagreement is not automatically a model error. MedPAIR studies written questions, not live consultations or learning outcomes. Its relevance findings motivate a question for educational design: can we inspect the evidence behind each decision while an encounter unfolds?

Physician-AI collaboration also needs its own evaluation. In Goh and colleagues' randomized trial of LLM assistance for diagnostic reasoning, access to an LLM did not significantly improve physicians' diagnostic reasoning scores over conventional resources. Providing a capable model does not by itself establish an effective collaboration workflow.

VeReaFine: checking relevance and retrieving missing evidence

VeReaFine: Iterative Verification Reasoning Refinement RAG for Hallucination-Resistant on Open-Ended Clinical QA, Phasook and colleagues, August 2025, describes a retrieval, verification and generation pipeline for clinical questions. A verifier checks retrieved biomedical passages for relevance and identifies missing evidence. If the context is insufficient, it generates a focused follow-up query. The system can repeat this process for up to three iterations before generating and checking its answer.

VeReaFine searches an indexed evidence corpus rather than following live web links as MedBrowseComp does. Its ClinIQLink evaluation includes open-ended clinical reasoning tasks and questions about erroneous reasoning steps. Results concern that benchmark and its models; they do not establish expert-level reliability in patient care or educational effectiveness.

For Med Abhyaas, the useful design question is whether a learner can identify a missing fact, seek relevant evidence, and check which claims that evidence supports before updating a differential or plan. MedPAIR's relevance judgments and VeReaFine's verification loop offer ways to study this process. Any automated feedback still needs clinician review and comparison with independent learner performance. Read the full paper.

MedBrowseComp: following evidence across sources

MedBrowseComp: Benchmarking Medical Deep Research and Computer Use, Chen and colleagues, 20 May 2025, tests medical information retrieval across linked sources. Its deep-research questions require one to five hops through oncology resources, PubMed, ClinicalTrials.gov, FDA records and, for some tasks, company or financial data.

One question asks an agent to identify an effective regimen ingredient from trial NCT00974311, then find its FDA exclusivity date. Reported errors include using secondary announcements instead of the FDA Orange Book and confusing initial approval with exclusivity expiration. The benchmark therefore tests whether an agent can locate the right source and extract the requested fact, beyond producing a plausible answer.

MedPAIR asks which evidence a model treats as relevant within a supplied case. MedBrowseComp asks whether an agent can find needed evidence across sources. Together they motivate practice in identifying missing information, choosing a source, checking what it supports and revising a decision. For Med Abhyaas, recording these searches alongside the encounter is an educational design proposal that needs evaluation.

MedBrowseComp chiefly evaluates oncology fact retrieval. Guideline concordance is proposed future work. Its benchmark scores do not demonstrate faithful internal reasoning, improved learning or patient benefit. Read the paper and examples or explore the dataset.


Use of AI for Medical Education

Virtual patients offer a setting for low-stakes practice before high-stakes clinical encounters. A learner can elicit a history, explain a concern, make an incomplete decision, receive feedback and try again. AI can make patient dialogue more responsive, but the case facts, clinical expectations and feedback still need clinician review. A fluent simulated patient can give inconsistent answers or disclose information that a real patient would not volunteer.

Connect relevant evidence to the reasoning trajectory

MedPAIR suggests looking beyond answer accuracy to the information that supports a decision. Extending this idea to virtual patient practice is a proposed educational approach, not a method validated by MedPAIR. Each case can record the information available at a particular turn, what the learner requested, and which findings they cited when revising their plan.

History elicitation
Assess whether follow-up questions clarify the presenting concern and seek discriminating findings. Distinguish evidence the learner actively elicited from facts the simulator volunteered. Review missed safety-relevant questions as well as unnecessary questioning.
Differential updates
Ask the learner to state a working differential before and after new evidence. Review which findings changed the ranking, whether contradictory evidence received attention, and whether the learner closed the differential prematurely.
Test selection
Require a reason for each proposed test and an explanation of how its result could change the next decision. Assess unnecessary testing, omitted investigations and the consequences of waiting in the context of the authored case.
Uncertainty
Compare expressed confidence with the available evidence. Ask what remains unknown and which finding would change the learner's mind. Confident wording alone should not earn a higher score.
Escalation
Review recognition of urgency, limits of competence and the point at which the learner seeks senior help. Case-specific escalation criteria should come from clinician-authored expectations, with room for more than one defensible approach.
Patient communication
Assess explanations, consent, understanding, response to concerns and safety-net advice. Review what the learner actually said and how the patient responded. Language fluency, accent or stylistic similarity to an AI answer should not substitute for clinical or communication quality.

A trajectory assessment should allow acceptable alternatives and preserve uncertainty about its own ratings. Sentence relevance, a written rationale and conversational behavior are observable evidence; none gives direct access to a learner's internal reasoning. Automated scores need comparison with clinician ratings before use in consequential assessment.

Feedback that a learner can act on

After an encounter, feedback should point to a specific question, decision or explanation, identify the supporting or missing evidence, and propose a next practice task. For example, an authored persistent-cough case might show where the learner offered reassurance before clarifying duration and associated symptoms. The learner can revisit that turn, explain the revised differential and then attempt a different case without hints.

Keep supported practice and assessment distinct. Hints can help during practice; independent assessment should record unaided decisions before displaying AI suggestions. Faculty should review disputed feedback and audit the simulator for factual drift, answer leakage and inconsistent responses across languages or learner groups.

What virtual patient studies show

These studies evaluate particular training systems and comparators. They support testing virtual patient practice, with mixed results across settings. They do not validate every AI simulator, establish improved patient outcomes, or demonstrate that AI should replace human standardized patients and supervised clinical teaching.

Practising physician-AI collaboration

A useful training encounter can ask the learner to commit to a provisional judgment, examine an AI suggestion, and explain why they accept, modify or reject it. Feedback can compare the evidence used by each and identify unsupported recommendations, missed contradictions and uncritical acceptance. The physician educator remains responsible for reviewing the case and assessment expectations.

This workflow should be evaluated against both unaided practice and conventional teaching. Measure performance on new cases without AI, retention after a delay, disagreement with faulty suggestions, and whether feedback changes behavior. A better score during an assisted encounter is not sufficient evidence of independent competence.

Med Abhyaas as a research direction

Med Abhyaas explores virtual patient encounters, clinical reasoning and patient communication in Indic languages. Its scripted previews illustrate how turn-level evidence and feedback could support repeated practice. The education approach here connects those encounters to MedPAIR's question about evidence relevance.

Med Abhyaas is a research project. The current examples and ratings are illustrative and unvalidated. Clinical review, native-speaker review, assessment reliability and controlled learner studies remain necessary before claiming educational effectiveness. Neither MedPAIR nor the virtual patient trials above establish clinical benefits for Med Abhyaas.

Explore the Med Abhyaas practice concept or see the clinical AI open problems.


Medical authority, independent judgment and learner agency

AMIE's supervised primary-care feasibility study

The supplied Lancet paper, Conversational diagnostic artificial intelligence in ambulatory primary care: a prospective feasibility study, reports AMIE conversations before urgent primary-care visits. Of 114 enrolled patients, 98 completed both the AI interaction and appointment. Physician safety supervisors monitored every conversation. No prespecified conversation safety stops were required, but supervisors recorded one hallucination and supplied clinical information in five interactions.

This was a single-centre, single-arm study in English-speaking adults. It provides initial evidence about feasibility, conversation quality and acceptability under supervision. Zero safety stops does not establish comprehensive safety or justify removing supervision. The study does not demonstrate improved patient outcomes, workforce substitution or medical-student learning. Its educational relevance is the need to study evidence elicitation and clinician review in actual workflows as well as written benchmarks.

Krumholz on expertise and the patient relationship

In Beyond Information Asymmetry in Medicine: The Changing Nature of Medical Authority and Expertise, Harlan M. Krumholz argues that patients increasingly arrive with AI-generated interpretations and plans, alongside clinicians who use related tools. Medical authority therefore needs to rest on accountable commitment to the patient's welfare, values and choices. Deep knowledge remains necessary to recognize omissions, communicate uncertainty and disagree with a fluent but unsuitable recommendation.

For medical education, this means practising how to hear a patient's existing understanding, explain the evidence behind a recommendation and acknowledge when the patient or AI has identified something the clinician missed. This is an editorial argument about professional responsibility, not an experimental demonstration of AI's benefits.

Read the complete JACC essay PDF or Krumholz's accompanying LinkedIn post.

Bedard and discussion about learning before assistance

Rachael Bedard's New York Times opinion piece on AI and medical students and the supplied author discussion on X are further reading. The original article and thread were not accessible during this update. KFF Health News' roundup identifies Bedard and frames the opinion as a contrast between AI's value for experienced clinicians and concerns about medical students.

Informal discussion supplied with this page raises a related concern: clinicians who learned before AI may use it differently from learners still acquiring judgment. Participants also argue that curricula can adapt and that tool design should deliberately preserve human agency, critical thinking and meaning-making. These are discussion perspectives, not measured educational outcomes.

For Med Abhyaas, that concern motivates a practice design: ask for the learner's own question, differential and justification before revealing suggestions; provide specific feedback afterward; and evaluate a new case without assistance. Whether this design reduces dependence or improves learning needs testing. Neither the opinion nor the informal discussion establishes that AI causes cognitive decline.

Khullar on tasks, skills and the clinical workforce

Dhruv Khullar's Artificial Intelligence and the Future of the Clinical Workforce asks why automating clinical tasks may not translate directly into fewer clinical jobs. The publisher's abstract and Weill Cornell's account of the perspective support these distinctions:

  • Jevons paradox suggests that greater efficiency can increase use. Lower-cost care may expand demand.
  • The lump-of-labor fallacy assumes a fixed amount of work. Technology can change what care is possible and create new work.
  • A job combines tasks and skills. O-ring theory emphasizes how failure in one step of a connected process can undermine the whole result; successful task automation leaves questions about coordination and responsibility.

This is an economic perspective, not a causal study or a guaranteed forecast of clinician employment. The full publisher text was not accessible during this update; the summary uses its abstract and the author-institution account. In education, a reasonable implication is to train learners for evidence gathering, supervision, communication and accountable decisions as task boundaries change.

Khullar's LinkedIn post about the NEJM perspective repeats the efficiency, changing demand, and tasks-versus-jobs arguments, while allowing that roles could change and some jobs could disappear. This summary uses indexed post text; no author X thread about this paper was verified.

Wachter and Khullar in conversation

What If Your AI Is the Better Doctor?, from Why Should I Trust You?, brings Robert Wachter and Dhruv Khullar together to discuss autonomous and physician-assisted care, trust, access and the work physicians do beyond a diagnosis. The episode is a discussion of possible futures, not a trial of education or patient outcomes.

Watch on YouTube, listen to the Apple Podcasts episode or supplied Spotify episode, or browse the Apple Podcasts show. See also Khullar's supplied LinkedIn discussion of the interview. YouTube metadata and automatic captions confirmed the interview; Spotify and LinkedIn content were not independently accessible.

Wachter on the facts supplied to AI and the work beyond diagnosis

In On Straw Men and Doormen, Robert Wachter distinguishes reasoning over a curated set of facts from gathering and organizing those facts during care. He calls the clinician's concise account of salient findings a problem representation. His doorman analogy argues that automating a conspicuous task can leave other valuable parts of a job, including communication, coordination and managing practical constraints.

The connection to MedPAIR is a question about evidence selection. A virtual patient encounter can assess how a learner builds a problem representation before receiving an AI recommendation, then revises it when new information arrives. This extends a sentence-relevance benchmark into an educational design proposal; it does not establish that physician-written summaries are always better. Wachter's essay is commentary and includes disclosures of AI-company relationships.


MedPAIR reading group recording

Watch Yuexing Hao present MedPAIR at the Snorkel AI reading group. The recording discusses annotation, human-model relevance agreement and evaluation methodology. The talk and the linked arXiv preprint describe different study versions; their counts and results should not be combined.

Editorial chapters below were derived from the recording's automatic English captions, retrieved with yt-dlp. They are approximate topic boundaries, not uploader-provided chapters. Captions contain transcription errors in names and technical terms.

  1. 00:11 路 Introduction and open benchmark support
  2. 03:39 路 Correct medical answers and evidence use
  3. 08:25 路 Sentence relevance labels and study design
  4. 15:45 路 Spurious rate and relevant context
  5. 21:00 路 Physician workload and human oversight
  6. 25:14 路 Relevance disagreement and filtering results
  7. 29:15 路 Sentence selection and next steps
  8. 33:30 路 Evaluation with simulated users and Matrix
  9. 41:43 路 Q&A on domain evaluation and clinical feedback
  10. 58:51 路 Implications for clinicians and trainees

Papers and further reading

Newest publications first, using the paper's original posting or online publication date. The recording chapters above remain in playback order.

Updated 9 October 2026. This page describes research and educational design, not guidance for patient care.