Large language models in clinical decision-making

7 trials
  • 2026O'Sullivan LLM (cardiology)Physician decision-makingIn 9 general cardiologists assessing complex suspected genetic cardiomyopathy, an LLM assistant alongside the raw investigations improved blinded subspecialist-rated quality of triage, diagnosis and management compared with cardiologists working unassisted.
  • 2026Tao LLM (care transitions)Patient-facing useIn 2,069 patients attending specialist clinics, an LLM chatbot taking the history and drafting a referral before the visit reduced specialist consultation duration compared with no chatbot before the specialist visit.
  • 2026Agweyu LLM (primary care, Kenya)Standalone consultationIn 9,691 patients at primary care facilities in Kenya, an LLM decision support system in the medical record did not reduce treatment failure within 14 days compared with the electronic medical record alone.
  • 2025Goh GPT-4 (management reasoning)Physician decision-makingIn 92 physicians answering clinical vignettes, GPT-4 plus conventional resources improved management reasoning scores compared with conventional resources alone.
  • 2025AMIEStandalone consultationIn 159 text consultations with validated patient-actors, an LLM-based conversational diagnostic AI (AMIE) improved diagnostic accuracy and rated consultation quality compared with primary care physicians.
  • 2025Bolton AI antimicrobial prescribingPhysician decision-makingIn 42 antibiotic-prescribing clinicians judging case vignettes, AI decision support with explanations for intravenous-to-oral switching did not differ from standard-of-care information alone for most switching decisions and completion times.
  • 2024Goh LLM (diagnostic reasoning)Physician decision-makingIn 50 physicians working up clinical vignettes, access to a large language model alongside conventional resources did not improve diagnostic reasoning scores compared with conventional resources alone.