Features pulled out of a working model
On 21 May 2024 Anthropic decomposed the activations of Claude 3 Sonnet with sparse autoencoders and obtained millions of features, each answering to one comprehensible concept. Forcing a feature's activation changed the model's behaviour with no change to the weights.
Why it matters
Until then dictionary learning had been shown on a one-layer toy network, and the main doubt was whether it scaled at all. Here it was applied to a model that is on sale — and the features turned out to be not only readable but causal: they can be pulled on.
The autoencoders were trained on residual stream activations halfway through the network — a choice explained by the residual stream being smaller than the MLP layer and by the middle likely holding more abstract features. Three sizes: 1,048,576, 4,194,304 and 33,554,432 features. The tie to real intervention: the Golden Gate Bridge feature (34M/31164353), clamped to ten times its maximum activation, makes the model talk about the bridge in any context. A separate result is that features fire on corresponding images too, although dictionary learning was done on text alone. The record does not claim that the safety features found say anything about how dangerous the model is. The authors warn against exactly that reading: "there's a difference between knowing about lies, being capable of lying, and actually lying in the real world." Features were found for unsafe code, bias, sycophancy, deception and power-seeking, and criminal content — and the paper stresses that what is interesting is not that they exist but that they can be found and acted on.