Arabic as a third of the training set
On 30 August 2023 Inception, MBZUAI and Cerebras Systems released Jais and Jais-chat, models of 13 billion parameters under the Apache 2.0 licence. Training ran over 395 billion tokens with Arabic at 33 per cent of the mixture: 72 billion collected tokens passed 1.6 times to reach 116 billion, plus 232 billion English tokens and the remainder code.
Why it matters
Until then Arabic in a multilingual model was one row among fifty, a share of a percentage point. Here it was made one of two languages, and what was missing turned out to be data rather than attention: the 72-billion-token corpus the authors call the largest Arabic collection of its time still had to be repeated 1.6 times.
The architecture is a GPT-3 decoder with SwiGLU activation and ALiBi positional bias. Training ran on 16 CS-2 systems within Condor Galaxy 1, the Cerebras supercomputer built in partnership with G42. The authors deliberately added no other languages so as not to make Arabic a minority, while still adding English, because Arabic alone is not enough for a model able to show emergent behaviour. On their own evaluation the authors claim better knowledge and reasoning in Arabic than any existing open Arabic or multilingual model, and competitiveness in English with English-centric open models of the same size despite far less English data. What the record does not claim. The four-exaflop figure for Condor Galaxy 1 that the report gave appears in the paper only inside the title of a footnoted link to a Cerebras blog, making it the manufacturer's claim about its own machine rather than a measurement; it does not enter. The comparisons against other models were made by the authors themselves.