Goh LLM (diagnostic reasoning)

Large Language Model Influence on Diagnostic Reasoning: A Randomized Clinical Trial

Paciente / Población Intervención / Exposición Comparación Desenlace

In 50 physicians (attendings and residents in family, internal or emergency medicine) randomised to work up clinical vignettes, access to a large language model alongside conventional diagnostic resources did not improve diagnostic reasoning scores compared with conventional resources alone.

N
50
Diseño
Single-blind RCT, physicians randomised, multicentre
Desenlace
Diagnostic reasoning score per case on a standardised rubric (differential accuracy, supporting/opposing factors, next diagnostic steps), graded by blinded expert consensus
Relevancia
2
ResultadoMedian diagnostic reasoning score 76% vs 74% per case (adjusted difference 2 percentage points, 95% CI −4 to 8; P=0.60); time per case 519 vs 565 s (P=0.20). The LLM alone scored 16 points higher than the conventional-resources group (95% CI 2–30; P=0.03).
Goh E, et al. JAMA Netw Open. 2024;7(10):e2440969. 10.1001/jamanetworkopen.2024.40969
Discusión y crítica

The first randomised test of whether handing physicians an LLM improves their reasoning — it did not, yet the LLM alone outperformed both physician groups, exposing a human–AI collaboration gap rather than a model deficit. Small (50 physicians), vignette-based, one month of recruitment and no patient outcomes. Its companion trial on management reasoning (Goh GPT-4 (management reasoning)) found a benefit.