MATH · IN · MODELS

Transcoders

Techniqueadvanced

A sparse dictionary trained to approximate an entire MLP sublayer's input-to-output function (rather than reconstructing a single site's own activations, as an SAE does), whose ReLU-gated decoder-vector sum turns weights-based circuit attribution through the MLP into an exact linear factorization of input-dependent (feature activation) and input-invariant (decoder-encoder dot product) terms.

Used in (3 observations)

structure: Linear Direction · models: Qwen3-0.6B, Qwen3-1.7B, Qwen3-4B, Qwen3-8B, Qwen3-14B · paper: Latent Planning Emerges with Scale
structure: Linear Direction · models: FLUX.1 [schnell] · paper: DifFRACT: Diffusion Feature Reconstruction and Attribution for Circuit Tracing
structure: Linear Direction · models: GPT-2 Small, Pythia-410M, Pythia-1.4B · paper: Transcoders Find Interpretable LLM Feature Circuits