A language model graded like a physician
On 26 December 2022 Google Research and DeepMind published Med-PaLM and the MultiMedQA suite: Flan-PaLM scored 67.6% on US Medical Licensing Examination-style questions, more than 17 points above the previous best, and Med-PaLM, tuned with prompts, answered in line with scientific consensus almost as often as physicians.
Why it matters
Clinical knowledge of models had been measured automatically on limited benchmarks; here clinicians also rated the answers for factuality, possible harm and bias, and the model came close to physicians. The authors themselves write that it remains inferior to clinicians.
MultiMedQA joins six existing sets, MedQA, MedMCQA, PubMedQA, MMLU clinical topics and others, with the new HealthSearchQA of 3,375 commonly searched health questions. On 140 questions clinicians judged consistent with consensus 61.9% of Flan-PaLM's answers, 92.9% of clinicians' answers, and Med-PaLM's answers 92.6% by the first version's introduction and 92.9% by its own results section. Potentially harmful: 29.7% of Flan-PaLM's answers, 5.8% of Med-PaLM's and 6.5% of clinicians'. The claim that Med-PaLM was the first to pass the threshold appears only in the next paper; it is not in this one.