Chatbot Arena: a score that cannot be studied in advance
On 3 May 2023 LMSYS opened a site where a visitor writes their own prompt, receives two answers from unnamed models and votes for the better one. The first leaderboard rested on 4.7 thousand votes gathered in a week and put Vicuna-13B on top with 1169 Elo points.
Why it matters
Until then every evaluation of a conversational system rested on a fixed set, and a fixed set can be studied - deliberately, or because it ended up in the training data. A prompt that did not exist until now cannot be. The ordering of models came for the first time from people actually talking to them, and it updated daily rather than once per publication.
The mechanics: anonymous randomised pairs, the prompt written by the visitor, model names revealed after the vote. Ratings are computed with the chess Elo system, because it provides the three properties a pairwise comparison needs: it scales to many models without data for every pair, it places a new model with relatively few trials, and it gives a unique order. The opening leaderboard, collection window 24 April to 1 May 2023, nine open models on 4.7 thousand votes: Vicuna-13B 1169, Koala-13B 1082, OpenAssistant Pythia-12B 1065, Alpaca-13B 1008, ChatGLM-6B 985, FastChat-T5-3B 951, Dolly-v2-12B 944, LLaMA-13B 932, StableLM-tuned-alpha-7B 858. The post names plainly what this is built against: HELM and lm-evaluation-harness give measurements on fixed sets but do not answer which model is better in a live conversation. The authors also note that the Anthropic LLM paper had already adopted the Elo system for model evaluation. What this record does not claim: the post contains no Bradley-Terry model - that arrived in the arena's later work. The authors' university affiliations are not named in the post either, so they are not named here. And the method carries its own cost, named later: it measures what visitors like rather than what is right.