Scaling laws for sparse autoencoders
On 6 June 2024 OpenAI showed that decomposing activations into features itself obeys scaling laws. A 16-million-latent autoencoder was trained on GPT-4 activations for 40 billion tokens, and 7% of its latents stayed dead.
Why it matters
Two weeks earlier it had become clear that features can be pulled out of a working model. This work turned that into engineering: instead of tuning a sparsity penalty, the number of active features is set directly, and how much data and compute it will take can be worked out in advance.
Replacing the L1 penalty with a TopK function that keeps exactly the k largest activations removes two troubles at once: sparsity is now set by a number rather than by tuning a coefficient, and the shrinkage of magnitudes that L1 introduces by construction disappears. Dead latents are prevented by initialising the encoder as the transpose of the decoder and by an auxiliary loss over the largest dead latents. The scale of the problem they started from: without these measures their own ablations reached up to 90% dead latents, and Anthropic's 34-million-latent autoencoder had only 12 million alive. With them, the largest 16-million-latent autoencoder has 7% dead. The convergence law: the number of tokens to convergence grows roughly as Θ(n^0.6) for GPT-2 small and Θ(n^0.65) for GPT-4, that is sublinearly in the number of latents. The authors say themselves that this cannot hold: under a sublinear budget each latent eventually receives too little gradient signal, so the law must break somewhere.