Three per cent of a training budget, spent on one language
On 16 April 2023 the Maritaca AI group presented Sabia: LLaMA 7B, LLaMA 65B and GPT-J further pretrained in Portuguese on 3 per cent or less of their original training budget. On Poeta, a suite of 14 Portuguese datasets, Sabia-65B marginally outperformed GPT-3.5-turbo, the base model of the ChatGPT of the day.
Why it matters
Until then the choice for a language outside the top ten was to wait until it was added to the next multilingual model, or to build one from scratch. This work showed a third route and its price, a few per cent of somebody else's training budget, and took apart where the gain comes from: more from facts about the country than from grammar.
The data comes from the Portuguese subset of ClueWeb22, filtered with the MassiveText quality filters adapted to Portuguese. The GPT-J tokeniser yields 7.8 billion tokens on this corpus and the LLaMA tokeniser 7.3 billion; training ran over 10.4 billion tokens, or 1.52 epochs. The 7-billion model reached 124,000 tokens per second at 45.2 per cent model FLOPs utilisation, the 65-billion one 14,000 at 47.4 per cent. The authors separate two reasons for the gain: datasets conceived in Portuguese gained more than translated ones, so most of the improvement falls on knowledge of a domain and a culture rather than on linguistic nuance. What the record does not claim. There is no 14-billion-token corpus named Claudio in the paper, as the report had it: the string Claudio occurs zero times and 14 billion occurs zero times. Nor is there a claim that Sabia-65B beat LLaMA-65B on all 14 tasks; the paper reports a considerable gain overall and says that for the 65-billion model the gain came mostly from the native rather than the translated subset. The record does not claim that weights were opened on this day.