A first exam for open models in Ukrainian
In May 2024, at the UNLP workshop in Turin, Oleksiy Syvokon, Mariana Romanyshyn and Roman Kyslyi reported the first shared task on fine-tuning large language models for Ukrainian: open weights only, fitting in 16 GB of GPU memory. On 751 questions from the national school-leaving exam of 2020-2023 the best team reached 0.49 accuracy; GPT-4, outside the competition, 0.61.
Why it matters
An open, reproducible measure appeared of how much a small open model knows of Ukrainian language, literature and history and how well it writes Ukrainian, with a hidden test, a public Codabench leaderboard and human evaluation. It also recorded the gap to a closed model.
The training set is 3,063 exam questions of 2006-2019 in Ukrainian history and Ukrainian language and literature, without questions that reference images; 100 open tasks were written by two native speakers, and the answers were rated by 63 volunteers, over 300 judgements per model, with TrueSkill. Twenty-two teams registered and three submitted. On the open questions GPT-4 scores 28.48 and the best open system 25.05; the paper's conclusion says 26.77 instead, and the record takes the table's figure. The winners fine-tuned Mistral 7B, others Gemma and Vicuna-13B. The author order here follows the PDF; the anthology page gives another. The research report behind the candidate linked this work to speech recognition of 1968; there is no connection between them.