Llama 3
Meta released models of 8 and 70 billion parameters, trained on fifteen trillion tokens, seven times the data Llama 2 received.
Why it matters
Open models stopped being a generation behind: the smaller of the two held the level the largest closed models had reached a year earlier.
The amount of data, not of parameters, became the main lever: the 8-billion model saw more text than its 70-billion predecessor, and that is where the gain came from. Meta also widened the tokeniser vocabulary and used grouped-query attention to make inference cheaper. The licence was not quite open, carrying a restriction for very large services, but in practice the weights became the base for most derived models of the year.