Back to timeline

Availability · December 11, 2023

Mixtral 8x7B

Mistral released a model in which only part of the parameters work on each token, and it beat the 70-billion Llama 2 at far lower compute cost.

Why it matters

The mixture-of-experts scheme reached open release, showing that quality does not require every parameter to work every time.

The model holds 46.7 billion parameters but engages about 13 billion per token: a router picks two experts out of eight. That gives the quality of a larger model at the cost of a smaller one. The scheme was assumed to be in use inside large closed models; the open release made it checkable and reproducible.

Event record

Event date
December 11, 2023
Timeline date
Event date
Verification
Sources gathered automatically · September 22, 2026
Lines
ID
evt-0291

The model was announced on 11 December 2023.

Sources

Related events

Records that link to this one