Goh LLM (diagnostic reasoning)
Large Language Model Influence on Diagnostic Reasoning: A Randomized Clinical Trial
Patient / Population Intervention / Exposure Comparison Outcome
In 50 physicians working up clinical vignettes, access to a large language model alongside conventional resources did not improve diagnostic reasoning scores compared with conventional resources alone.
N
50patients
Design
Single-blind RCT, physicians randomised, multicentre
Endpoint
Diagnostic reasoning score per case on a standardised rubric (differential accuracy, supporting/opposing factors, next diagnostic steps), graded by blinded expert consensus
Relevance
2Important — one of several pillars.
ResultMedian diagnostic reasoning score 76% vs 74% per case (adjusted difference 2 percentage points, 95% CI −4 to 8; P=0.60); time per case 519 vs 565 s (P=0.20). The LLM alone scored 16 points higher than the conventional-resources group (95% CI 2–30; P=0.03).
Goh E, et al. JAMA Netw Open. 2024;7(10):e2440969. 10.1001/jamanetworkopen.2024.40969
Discussion & critique
The first randomised test of whether handing physicians an LLM improves their reasoning — it did not, yet the LLM alone outperformed both physician groups, exposing a human–AI collaboration gap rather than a model deficit. Small (50 physicians), vignette-based, one month of recruitment and no patient outcomes. Its companion trial on management reasoning (Goh GPT-4 (management reasoning)) found a benefit.