Clinical reasoning and AI
Clinical reasoning depends on what a clinician asks, which findings they consider relevant, and how they revise a working explanation as evidence arrives. A correct answer on a medical exam captures only part of that process.
Evidence relevance and physician-AI collaboration
MedPAIR, by Yuexing Hao and colleagues, compares physician trainees' sentence-level relevance judgments with those of language models in medical question answering. The available arXiv preprint finds substantial disagreement. Filtering information according to trainee relevance labels improves accuracy in many settings, with exceptions for some model and dataset combinations. A correct final answer therefore does not establish that a model selected the same evidence as a physician trainee.
Relevance labels are contextual judgments. Agreement with a trainee is not proof of safe clinical reasoning, and disagreement is not automatically a model error. MedPAIR studies written questions, not live consultations or learning outcomes. Its relevance findings motivate a question for educational design: can we inspect the evidence behind each decision while an encounter unfolds?
Physician-AI collaboration also needs its own evaluation. In Goh and colleagues' randomized trial of LLM assistance for diagnostic reasoning, access to an LLM did not significantly improve physicians' diagnostic reasoning scores over conventional resources. Providing a capable model does not by itself establish an effective collaboration workflow.
VeReaFine: checking relevance and retrieving missing evidence
VeReaFine: Iterative Verification Reasoning Refinement RAG for Hallucination-Resistant on Open-Ended Clinical QA, Phasook and colleagues, August 2025, describes a retrieval, verification and generation pipeline for clinical questions. A verifier checks retrieved biomedical passages for relevance and identifies missing evidence. If the context is insufficient, it generates a focused follow-up query. The system can repeat this process for up to three iterations before generating and checking its answer.
VeReaFine searches an indexed evidence corpus rather than following live web links as MedBrowseComp does. Its ClinIQLink evaluation includes open-ended clinical reasoning tasks and questions about erroneous reasoning steps. Results concern that benchmark and its models; they do not establish expert-level reliability in patient care or educational effectiveness.
For Med Abhyaas, the useful design question is whether a learner can identify a missing fact, seek relevant evidence, and check which claims that evidence supports before updating a differential or plan. MedPAIR's relevance judgments and VeReaFine's verification loop offer ways to study this process. Any automated feedback still needs clinician review and comparison with independent learner performance. Read the full paper.
MedBrowseComp: following evidence across sources
MedBrowseComp: Benchmarking Medical Deep Research and Computer Use, Chen and colleagues, 20 May 2025, tests medical information retrieval across linked sources. Its deep-research questions require one to five hops through oncology resources, PubMed, ClinicalTrials.gov, FDA records and, for some tasks, company or financial data.
One question asks an agent to identify an effective regimen ingredient from trial NCT00974311, then find its FDA exclusivity date. Reported errors include using secondary announcements instead of the FDA Orange Book and confusing initial approval with exclusivity expiration. The benchmark therefore tests whether an agent can locate the right source and extract the requested fact, beyond producing a plausible answer.
MedPAIR asks which evidence a model treats as relevant within a supplied case. MedBrowseComp asks whether an agent can find needed evidence across sources. Together they motivate practice in identifying missing information, choosing a source, checking what it supports and revising a decision. For Med Abhyaas, recording these searches alongside the encounter is an educational design proposal that needs evaluation.
MedBrowseComp chiefly evaluates oncology fact retrieval. Guideline concordance is proposed future work. Its benchmark scores do not demonstrate faithful internal reasoning, improved learning or patient benefit. Read the paper and examples or explore the dataset.
Use of AI for Medical Education
Virtual patients offer a setting for low-stakes practice before high-stakes clinical encounters. A learner can elicit a history, explain a concern, make an incomplete decision, receive feedback and try again. AI can make patient dialogue more responsive, but the case facts, clinical expectations and feedback still need clinician review. A fluent simulated patient can give inconsistent answers or disclose information that a real patient would not volunteer.
Connect relevant evidence to the reasoning trajectory
MedPAIR suggests looking beyond answer accuracy to the information that supports a decision. Extending this idea to virtual patient practice is a proposed educational approach, not a method validated by MedPAIR. Each case can record the information available at a particular turn, what the learner requested, and which findings they cited when revising their plan.
- History elicitation
- Assess whether follow-up questions clarify the presenting concern and seek discriminating findings. Distinguish evidence the learner actively elicited from facts the simulator volunteered. Review missed safety-relevant questions as well as unnecessary questioning.
- Differential updates
- Ask the learner to state a working differential before and after new evidence. Review which findings changed the ranking, whether contradictory evidence received attention, and whether the learner closed the differential prematurely.
- Test selection
- Require a reason for each proposed test and an explanation of how its result could change the next decision. Assess unnecessary testing, omitted investigations and the consequences of waiting in the context of the authored case.
- Uncertainty
- Compare expressed confidence with the available evidence. Ask what remains unknown and which finding would change the learner's mind. Confident wording alone should not earn a higher score.
- Escalation
- Review recognition of urgency, limits of competence and the point at which the learner seeks senior help. Case-specific escalation criteria should come from clinician-authored expectations, with room for more than one defensible approach.
- Patient communication
- Assess explanations, consent, understanding, response to concerns and safety-net advice. Review what the learner actually said and how the patient responded. Language fluency, accent or stylistic similarity to an AI answer should not substitute for clinical or communication quality.
A trajectory assessment should allow acceptable alternatives and preserve uncertainty about its own ratings. Sentence relevance, a written rationale and conversational behavior are observable evidence; none gives direct access to a learner's internal reasoning. Automated scores need comparison with clinician ratings before use in consequential assessment.
Feedback that a learner can act on
After an encounter, feedback should point to a specific question, decision or explanation, identify the supporting or missing evidence, and propose a next practice task. For example, an authored persistent-cough case might show where the learner offered reassurance before clarifying duration and associated symptoms. The learner can revisit that turn, explain the revised differential and then attempt a different case without hints.
Keep supported practice and assessment distinct. Hints can help during practice; independent assessment should record unaided decisions before displaying AI suggestions. Faculty should review disputed feedback and audit the simulator for factual drift, answer leakage and inconsistent responses across languages or learner groups.
What virtual patient studies show
- Artificial intelligence-powered virtual standardized patients in teaching history-taking skills to medical students reports a randomized trial involving 67 third-year Vietnamese medical students. Both the AI patient and conventional standardized-patient groups improved; between-group differences in knowledge gains, OSCE scores and satisfaction were not statistically significant. This result does not establish equivalence or noninferiority.
- SOPHIE: AI Standardized Patient Improves Human Conversations in Advanced Cancer Care reports a randomized study of 51 healthcare students and practitioners. AI patient practice with personalized feedback produced greater short-term gains in rated empathy, explicit communication and empowerment than a reading module, assessed in conversations with human standardized patients. This is a preprint and does not establish durable transfer to patient care.
These studies evaluate particular training systems and comparators. They support testing virtual patient practice, with mixed results across settings. They do not validate every AI simulator, establish improved patient outcomes, or demonstrate that AI should replace human standardized patients and supervised clinical teaching.
Practising physician-AI collaboration
A useful training encounter can ask the learner to commit to a provisional judgment, examine an AI suggestion, and explain why they accept, modify or reject it. Feedback can compare the evidence used by each and identify unsupported recommendations, missed contradictions and uncritical acceptance. The physician educator remains responsible for reviewing the case and assessment expectations.
This workflow should be evaluated against both unaided practice and conventional teaching. Measure performance on new cases without AI, retention after a delay, disagreement with faulty suggestions, and whether feedback changes behavior. A better score during an assisted encounter is not sufficient evidence of independent competence.
Med Abhyaas as a research direction
Med Abhyaas explores virtual patient encounters, clinical reasoning and patient communication in Indic languages. Its scripted previews illustrate how turn-level evidence and feedback could support repeated practice. The education approach here connects those encounters to MedPAIR's question about evidence relevance.
Med Abhyaas is a research project. The current examples and ratings are illustrative and unvalidated. Clinical review, native-speaker review, assessment reliability and controlled learner studies remain necessary before claiming educational effectiveness. Neither MedPAIR nor the virtual patient trials above establish clinical benefits for Med Abhyaas.
Explore the Med Abhyaas practice concept or see the clinical AI open problems.
MedPAIR reading group recording
Watch Yuexing Hao present MedPAIR at the Snorkel AI reading group. The recording discusses annotation, human-model relevance agreement and evaluation methodology. The talk and the linked arXiv preprint describe different study versions; their counts and results should not be combined.
Editorial chapters below were derived from the recording's automatic English captions, retrieved with yt-dlp. They are approximate topic boundaries, not uploader-provided chapters. Captions contain transcription errors in names and technical terms.
- 00:11 路 Introduction and open benchmark support
- 03:39 路 Correct medical answers and evidence use
- 08:25 路 Sentence relevance labels and study design
- 15:45 路 Spurious rate and relevant context
- 21:00 路 Physician workload and human oversight
- 25:14 路 Relevance disagreement and filtering results
- 29:15 路 Sentence selection and next steps
- 33:30 路 Evaluation with simulated users and Matrix
- 41:43 路 Q&A on domain evaluation and clinical feedback
- 58:51 路 Implications for clinicians and trainees
Papers and further reading
Newest publications first, using the paper's original posting or online publication date. The recording chapters above remain in playback order.
- 8 October 2026. Conversational diagnostic artificial intelligence in ambulatory primary care: a prospective feasibility study. Lancet, prospective feasibility study.
- 29 September 2026. Beyond Information Asymmetry in Medicine: The Changing Nature of Medical Authority and Expertise. Krumholz, JACC editorial.
- 25 September 2026. Rachael Bedard on AI and medical students. New York Times opinion; original text inaccessible during this update.
- 12 September 2026. Artificial Intelligence and the Future of the Clinical Workforce. Khullar, NEJM perspective, online publication.
- 10 September 2026. What If Your AI Is the Better Doctor?. Wachter and Khullar, podcast discussion.
- 27 August 2026. On Straw Men and Doormen. Wachter, commentary.
- 30 April 2026. Artificial intelligence-powered virtual standardized patients in teaching history-taking skills to medical students: a randomized controlled trial. BMC Medical Education.
- August 2025. VeReaFine: Iterative Verification Reasoning Refinement RAG for Hallucination-Resistant on Open-Ended Clinical QA. Phasook et al., BioNLP workshop paper.
- 29 May 2025. MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering. Hao et al., arXiv preprint; paper date, distinct from the later recording.
- 20 May 2025. MedBrowseComp: Benchmarking Medical Deep Research and Computer Use. Chen et al., arXiv preprint.
- 5 May 2025. AI Standardized Patient Improves Human Conversations in Advanced Cancer Care. SOPHIE, arXiv preprint.
- 28 October 2024. Large Language Model Influence on Diagnostic Reasoning: A Randomized Clinical Trial. Goh et al., JAMA Network Open.
Updated 9 October 2026. This page describes research and educational design, not guidance for patient care.