methods / Dictionary Learning / Transcoders
Transcoders
A sparse dictionary trained to approximate an entire MLP sublayer's input-to-output function (rather than reconstructing a single site's own activations, as an SAE does), whose ReLU-gated decoder-vector sum turns weights-based circuit attribution through the MLP into an exact linear factorization of input-dependent (feature activation) and input-invariant (decoder-encoder dot product) terms.