Back to timeline

Benchmark · May 16, 2023

Physicians prefer the model's answers

On 16 May 2023 Google published Med-PaLM 2: 86.5% on MedQA, more than 19 points above Med-PaLM, and on 1,066 consumer questions physicians preferred its answers to physicians' answers on eight of nine axes of clinical utility.

Why it matters

In five months the score on medical exam questions rose from what the authors call just passing to a level where physician raters chose the model's answer over their colleagues' more often. The evaluation was by Google's authors; the peer-reviewed version appeared twenty months later.

The model is built on PaLM 2 with medical fine-tuning and a new prompting method, ensemble refinement. The evaluation added 240 adversarial questions probing the model's limits; on them Med-PaLM 2 beat Med-PaLM on every axis. The first version gives Med-PaLM's MedQA score as 67.2% and says it was the first to pass the threshold. The peer-reviewed paper in Nature Medicine repeats 86.5% and eight axes of nine. The comparison is with physicians' written answers, not with the treatment of patients.

Event record

Event date
May 16, 2023
Timeline date
Event date
Verification
Sources gathered automatically · September 25, 2026
Lines
ID
evt-0750

The first version of the preprint; the peer-reviewed paper appeared in Nature Medicine on 8 January 2025.

Sources

Related events