Goh LLM (diagnostic reasoning)
Large Language Model Influence on Diagnostic Reasoning: A Randomized Clinical Trial
In 50 physicians (attendings and residents in family, internal or emergency medicine) randomised to work up clinical vignettes, access to a large language model alongside conventional diagnostic resources did not improve diagnostic reasoning scores compared with conventional resources alone.
The first randomised test of whether handing physicians an LLM improves their reasoning — it did not, yet the LLM alone outperformed both physician groups, exposing a human–AI collaboration gap rather than a model deficit. Small (50 physicians), vignette-based, one month of recruitment and no patient outcomes. Its companion trial on management reasoning (Goh GPT-4 (management reasoning)) found a benefit.