Back to timeline

Benchmark · December 26, 2022

A language model graded like a physician

On 26 December 2022 Google Research and DeepMind published Med-PaLM and the MultiMedQA suite: Flan-PaLM scored 67.6% on US Medical Licensing Examination-style questions, more than 17 points above the previous best, and Med-PaLM, tuned with prompts, answered in line with scientific consensus almost as often as physicians.

Why it matters

Clinical knowledge of models had been measured automatically on limited benchmarks; here clinicians also rated the answers for factuality, possible harm and bias, and the model came close to physicians. The authors themselves write that it remains inferior to clinicians.

MultiMedQA joins six existing sets, MedQA, MedMCQA, PubMedQA, MMLU clinical topics and others, with the new HealthSearchQA of 3,375 commonly searched health questions. On 140 questions clinicians judged consistent with consensus 61.9% of Flan-PaLM's answers, 92.9% of clinicians' answers, and Med-PaLM's answers 92.6% by the first version's introduction and 92.9% by its own results section. Potentially harmful: 29.7% of Flan-PaLM's answers, 5.8% of Med-PaLM's and 6.5% of clinicians'. The claim that Med-PaLM was the first to pass the threshold appears only in the next paper; it is not in this one.

Event record

Event date
December 26, 2022
Timeline date
Event date
Verification
Sources gathered automatically · September 25, 2026
Lines
ID
evt-0749

The first version of the preprint; the peer-reviewed paper appeared in Nature on 12 July 2023.

Sources

Related events

Records that link to this one