The sparsely-gated mixture-of-experts layer
On 23 January 2017 Noam Shazeer and co-authors at Google Brain, among them Jeff Dean and Geoffrey Hinton, posted a layer of up to thousands of expert networks of which a trainable gate selects only a few for each example (noisy top-k gating). Placed between LSTM layers, it gave models of up to 137 billion parameters and, in the authors' words, greater than 1000x improvements in model capacity with only minor losses in computational efficiency. On WMT'14: 40.56 BLEU English-French and 26.03 English-German.
Why it matters
Model capacity was separated from the computation spent per example: parameters could grow by orders of magnitude while each token still passes through a handful of experts. Conditional computation, proposed before in theory, worked at scale on GPU clusters. DeepSeekMoE in 2024 names this paper as the one that introduced the mixture of experts into language model training.
Language modelling on 100 billion words: test perplexity kept improving up to 65,536 experts (68 billion parameters), 39 per cent lower than a computationally matched baseline, and degraded at 131,072 experts, possibly from too much sparsity. The translation results are 1.34 and 1.12 BLEU above the strong baselines of Wu et al. 2016 (GNMT). One author, Krzysztof Maziarz, was at the Jagiellonian University in Krakow. The layer sits between LSTMs, not in a Transformer. Mixtral 8x7B in December 2023 works on the same principle, but its announcement does not cite this paper, and the record does not claim a direct line.