Physicians prefer the model's answers
On 16 May 2023 Google published Med-PaLM 2: 86.5% on MedQA, more than 19 points above Med-PaLM, and on 1,066 consumer questions physicians preferred its answers to physicians' answers on eight of nine axes of clinical utility.
Why it matters
In five months the score on medical exam questions rose from what the authors call just passing to a level where physician raters chose the model's answer over their colleagues' more often. The evaluation was by Google's authors; the peer-reviewed version appeared twenty months later.
The model is built on PaLM 2 with medical fine-tuning and a new prompting method, ensemble refinement. The evaluation added 240 adversarial questions probing the model's limits; on them Med-PaLM 2 beat Med-PaLM on every axis. The first version gives Med-PaLM's MedQA score as 67.2% and says it was the first to pass the threshold. The peer-reviewed paper in Nature Medicine repeats 86.5% and eight axes of nine. The comparison is with physicians' written answers, not with the treatment of patients.