A language model on 0.94 gigabytes of text
On 11 November 2021 Kelechi Ogueji, Yuxin Zhu and Jimmy Lin presented AfriBERTa, a multilingual model trained from scratch on eleven African languages alone, on a corpus of 0.94 gigabytes, or 108.8 million tokens. XLM-R by comparison was trained on about 2395 gigabytes and 164 billion tokens, mBERT on roughly 100 gigabytes and 12.8 billion.
Why it matters
Until then a multilingual model for a language with a small digital trace was built as an annex to a large one: the language was laid beside English and transfer was hoped for. Here it was shown for the first time that for four of these eleven languages a first model could be made at all without English and without terabytes, on a corpus that fits on one disk.
The corpus was drawn mainly from BBC World Service news, scraped up to 17 January 2021, and topped up from Common Crawl; it holds 5,448,911 sentences. AfriBERTa large has 126 million parameters against 270 million for XLM-R base and 172 million for mBERT. Evaluation ran on named entity recognition with MasakhaNER and on news classification, spanning ten languages, four of which were absent from the pretraining corpus. What the record does not claim. The model does not beat mBERT and XLM-R outright: the paper says AfriBERTa large outperforms mBERT and is very competitive with XLM-R, and that it beats both on several languages the three share, namely Hausa, Amharic and Swahili. On Nigerian Pidgin both larger models beat it, which the authors attribute to that language's closeness to English. The report this candidate came from presented the result as a win over both. The report cited arXiv:2103.11811. That identifier belongs to MasakhaNER, the benchmark AfriBERTa was measured on rather than AfriBERTa itself; the date of 22 March 2021 that the report gave as the release date is the submission date of that other paper.