NetGoodIndexSubmit a correction

Benefit Ledger · verified · Medicine & Health

AMIE outperforms primary-care physicians on a text OSCE study

Google’s AMIE diagnostic-dialogue system beat 20 primary-care physicians on most specialist and patient-actor axes in a randomized text-chat OSCE, posted in 2024 and published in Nature on 9 April 2025.

9 Apr 2025Tier 2 NotableMethodology 0.1

Current score

+0.29

3 base · Notable (tier 2 of 5, 3 pts)
× 0.8500 attribution · Primary causal contribution
× 0.7000 evidence · Peer review or independent validation
× 0.4000 realization · Demonstrated
× 0.4000 durability
Event-level product before credit split: 0.29

Notable simulated diagnostic-dialogue result (tier 2). High attribution. Nature RCT-style OSCE. Realization is lab/OSCE only. Low durability until clinical deployment.

What happened

AMIE used self-play and automated feedback. The study included 159 scenarios from Canada, the UK, and India. Authors warned that physicians were constrained to unfamiliar text chat and that the system is not deployed in care. The result is a controlled communication/diagnosis study, not patient outcomes.

Model attribution

+0.29

Gemini/PaLM medical

LLM system optimized for diagnostic dialogue and evaluated against PCPs.

AMIE is the evaluated system.

Attribution 0.8500 · Credit share 100% · Google DeepMind

Claims

  • Specialist raters preferred AMIE to PCPs on 30 of 32 axes in the Nature OSCE study.

    outcome · supported

  • AMIE is in routine clinical use.

    outcome · disputed

Sources

primary sources

Secondary domains: Computer Science

Revision history

  • 13 Sep 2026 · 0.00 0.29

    Initial adjudicated seed score under methodology 0.1.