Availability · December 11, 2023
Mixtral 8x7B
Mistral released a model in which only part of the parameters work on each token, and it beat the 70-billion Llama 2 at far lower compute cost.
Why it matters
The mixture-of-experts scheme reached open release, showing that quality does not require every parameter to work every time.
The model holds 46.7 billion parameters but engages about 13 billion per token: a router picks two experts out of eight. That gives the quality of a larger model at the cost of a smaller one. The scheme was assumed to be in use inside large closed models; the open release made it checkable and reproducible.
Event record
- Event date
- December 11, 2023
- Timeline date
- Event date
- Verification
- Sources gathered automatically · September 22, 2026
- Lines
- ID
- evt-0291
The model was announced on 11 December 2023.
Related events
- Extends Mistral 7B
The same company moved to a mixture of experts two and a half months later.
Mixtral of experts - Related Llama 2
Open weights without restrictions, outperforming Llama 2.
Records that link to this one
- Related Llama 3
Two open models of the same season, taking different routes to quality.
- Related DeepSeek-V3
The same mixture-of-experts scheme, carried to 671 billion parameters.
- Related Llama 4
The mixture-of-experts scheme first opened in 2023 became the basis of the whole family here.
- Related Fine-grained and shared experts
A month after Mixtral's open release, this work took apart why a mixture of experts works and changed how one is built.
- Related The sparsely-gated mixture-of-experts layer
Mixtral 8x7B works on the same sparse-gate principle: at every layer, for every token, a router picks two of eight parameter groups. Its announcement does not cite this paper.
Mixtral of experts