Masakhane: translation made by the speakers
On 5 October 2020 the Masakhane community posted a paper on participatory research in machine translation for African languages: over 400 participants from at least 20 countries had within a year gathered new translation data and benchmarks for over 30 languages, a third of them evaluated by people. The authors come from several dozen institutions, from Pretoria and Kano to Stellenbosch.
Why it matters
The paper starts from the view that a low-resourced language lacks not only data but the people and conditions to gather and evaluate it, and shows that participants without formal training in the field can make a contribution that gets published. Models, data and code were released openly.
Most models were trained on the JW300 corpus, which is of missionary origin; the paper names this limit itself and tests translation outside religious topics separately. The barrier to entry was lowered by a tutorial for training a JoeyNMT transformer on Google Colab; language pairs were chosen by 32 participants for their own needs; over ten participants went on to publish their own work. What the record does not claim: exact numbers of languages and participants - the paper gives only "over".