Gemma Scope releases 2,000+ JumpReLU SAEs; feature splitting persists with width
measured in 1 paperLieberum et al. train and release over 2,000 JumpReLU SAEs across every layer and three sites of Gemma 2 (2B/9B/27B plus instruct 9B), totaling 30M+ features [lieberum-etal-2024-gemma-scope] Sweeping dictionary width from 2^14 to 2^20, feature splitting persists across the whole range rather than saturating [lieberum-etal-2024-gemma-scope] A cluster of ultra-high-frequency latents appears in their JumpReLU SAEs and in TopK SAEs but is reported absent from Gated SAEs [lieberum-etal-2024-gemma-scope] Reconstruction fidelity is comparable across sites, but downstream loss damage is consistently higher for residual-stream SAEs, whose small errors compound across later layers [lieberum-etal-2024-gemma-scope]
Structure
Context
feature-splitting persisting (not saturating) across a wide SAE-width scaling range, ultra-high-frequency latent cluster shared by JumpReLU/TopK, absent from Gated SAEs, attributed importance becoming more diffuse as L0 grows, most pronounced in the residual stream, site-dependent dissociation between reconstruction fidelity (FVU) and downstream loss damage
Confirmed in models
Papers
Gemma Scope: Open Sparse Autoencoders Everywhere All At Once on Gemma 2 — Lieberum, Tom, Rajamanoharan, Senthooran, Conmy, Arthur, Smith, Lewis, Sonnerat, Nicolas, Varma, Vikrant, Kramár, János, Dragan, Anca, Shah, Rohin, Nanda, Neel