MATH · IN · MODELS
methods / Dictionary Learning / Gated Sparse Autoencoders

Gated Sparse Autoencoders

Techniqueadvanced

Splits an SAE's encoder into two weight-tied sub-networks — a gate deciding which features are active and a magnitude estimator sizing them — applying the L1 sparsity penalty only to the gate's pre-activations so magnitude estimation is no longer biased toward zero, provably equivalent to a single-layer JumpReLU encoder once the two sub-networks' weights are tied.

Used in (2 observations)

structure: Linear Direction · models: GELU-1L, Pythia-2.8B, Gemma-7B · paper: Improving Dictionary Learning with Gated Sparse Autoencoders
structure: Linear Direction · models: DrugAssist, GeLLM3O-LLaMA3, GeLLM3O-Mistral, MolGen · paper: SLIM: Sparse Latent Steering for Interpretable and Property-Directed LLM-Based Molecular Editing