Model collapse on a model's own output
On 24 July 2024 Nature published a peer-reviewed paper describing model collapse: when generation after generation is trained on what the previous one produced, the tails of the original distribution disappear irreversibly. Shown for language models, variational autoencoders and Gaussian mixture models.
Why it matters
Until then synthetic data was discussed as a way round the exhaustion of text. The paper moved the question from practice to theory and named the price: data about genuine human interaction becomes more valuable the more machine text there is online.
The authors separate early from late collapse. In early collapse the model loses information about the tails of the distribution; in late collapse the distribution narrows to small variance around a mode and bears little resemblance to the original. The cause is not one but three: finite sampling error, limited model expressivity, and error in the learning procedure itself. The experiment: Meta's OPT-125m fine-tuned on wikitext2, with each later generation trained on the text the previous one produced, generation by five-way beam search, each experiment run five times. To keep conditions as favourable as possible, each generation started from the best model by the original validation set — so in life the effect may be stronger. What the record does not claim. Collapse is not inevitable under all conditions: when a random 10% of the original data was mixed in at every generation, the degradation was only minor. The phenomenon is shown at the scale of OPT-125m, not on frontier models; the paper generalises to those by theory rather than measurement.