Partly re-initialising an expert's weights
On 26 February 2025 Institute of Science Tokyo and Sakana AI showed how to start training a mixture of experts from a finished dense model by re-initialising only part of each expert's weights. A model with 5.9 billion active parameters matched a 13-billion dense one at about a quarter of the training compute.
Why it matters
Until then a mixture of experts was either trained from nothing or assembled from copies of a finished dense block. The first is expensive, the second yields experts that are all alike and never specialise. Partial re-initialisation found the middle: the start stays cheap and the experts diverge.
The device is that when the dense block is copied into experts, part of each copy's weights is dropped and learned again. That breaks the symmetry between copies without losing everything the model already knows. Figures from the first version: a mixture with 5.9 billion active parameters reaches the quality of a 13-billion dense model from the same family while spending about a quarter of the training FLOPs. Code and weights are open. What the record does not claim. The authors limit the result to long training outright: the advantage shows up "in the long term, specifically when training on hundreds of billions of tokens or more". On short runs it is not claimed.