Agweyu LLM (primary care, Kenya)

Generative AI-Enabled Clinical Decision Support System in Primary Care: A Pragmatic, Cluster-Randomized Trial

Patient / Population Intervention / Exposure Comparison Outcome

In 9,691 patients at primary care facilities in Kenya, an LLM decision support system in the medical record did not reduce treatment failure within 14 days compared with the electronic medical record alone.

N
9,691patients
Design
Pragmatic cluster RCT, clinical officers randomised, 16 primary care facilities, Kenya
Endpoint
Expert-adjudicated composite of treatment-failure events within 14 days of enrolment
Relevance
1Practice-defining — the trial the guideline rests on.
ResultTreatment failure 102/4,693 (2.2%) vs 94/4,654 (2.0%) (aOR 0.77, 95% CI 0.55–1.08; P=0.13). No serious adverse events judged related to the intervention and no safety signal on independent review.
Agweyu A, et al. Nat Med. 2026;32(8):3032-3039. 10.1038/s41591-026-04503-6
Discussion & critique

The first large pragmatic randomised trial of an LLM in real-world, low-resource primary care — and a null result on the patient-level primary outcome, with any benefit 'probably modest' in the authors' words. Reassuring safety, but a 2% event rate leaves little power, follow-up was 14 days, and the randomised units were clinicians rather than patients. A caution against extrapolating vignette-based LLM gains (Goh GPT-4 (management reasoning)) to patient outcomes.