Fine-grained and shared experts
On 11 January 2024 DeepSeek published an architecture in which the FFN block is split into many small experts and some of them are always on. DeepSeekMoE 16B beat LLaMA2 7B on most benchmarks while spending 39.6% of the computation.
Why it matters
Until then a mixture of experts was a way to make a large model cheaper. This work named why a cheaper model can also be better: when experts are large and each must know a little of everything, knowledge is duplicated across them, and fine splitting with a dedicated shared core removes that duplication.
Two devices. The first is fine-grained segmentation: the intermediate FFN dimension is split into m parts, and instead of one or two large experts mK small ones activate, which widens the space of specialisations combinatorially. The second is dedicated shared experts, which run on every token regardless of the router and hold the common knowledge that would otherwise be repeated inside each expert. Figures from the first version. DeepSeekMoE 16B has 245% of the parameters of LLaMA2 7B but needs 39.6% of the computation, and outperforms it on the majority of internal benchmarks; both were trained on 2 trillion tokens. Against DeepSeek's own dense 7B, comparable quality at 40.5% of the computation. At small scale, DeepSeekMoE 2B matches GShard 2.9B, which has 1.5 times the expert parameters and computation. The paper names its own ceiling: DeepSeekMoE 2B only "nearly approaches" the dense model with the same total parameter count, that is with 16 times the FFN parameters, and the authors treat that as the upper bound for mixture-of-experts models.