Transcoders linearize MLP circuit attribution and match SAE quality
measured in 1 paperDunefsky et al. introduce transcoders: SAE-shaped networks trained to reconstruct an MLP sublayer's output from its input rather than the MLP's own activations [dunefsky-etal-2024-transcoders] Circuit attribution then factorizes exactly into an input-dependent activation term times a fixed decoder-encoder dot product, needing no gradient/Taylor linearization of the MLP [dunefsky-etal-2024-transcoders] On GPT2-small, Pythia-410M, and Pythia-1.4B, transcoders sit on reconstruction/sparsity Pareto frontiers equal to or better than SAEs, with the gap widening at larger scale [dunefsky-etal-2024-transcoders] Applying the exact-factorization method to GPT2-small's greater-than circuit surfaces a finer feature-level account and isolates one anomalous flat direct-logit-attribution feature [dunefsky-etal-2024-transcoders]