MATH · IN · MODELS
methods / Dictionary Learning / Sparse crosscoders

Sparse crosscoders

Techniqueadvanced

A sparse dictionary-learning method — a generalization of SAEs — trained jointly across layers or models to recover shared, causally-relevant feature directions that persist across sites rather than one dictionary per site.

Used in (5 observations)

structure: Linear Direction · models: Claude 3 Sonnet · paper: Sparse Crosscoders for Cross-Layer Features and Model Diffing
structure: Linear Direction · models: Llama-3.1-8B-Instruct, Gemma-2-9B-it, Qwen2.5-7B · paper: Delta-Crosscoder: Robust Crosscoder Model Diffing in Narrow Fine-Tuning Regimes
structure: Linear Direction · models: Gemma-2-2B, Gemma-2-2B-it · paper: Overcoming Sparsity Artifacts in Crosscoders to Interpret Chat-Tuning
structure: 1D continuum manifold · models: Gemma-2-2B, EmbeddingGemma, word2vec (trained on Wikipedia), Claude 3.5 Haiku · paper: Symmetry in Language Statistics Shapes the Geometry of Model Representations, When Models Manipulate Manifolds: The Geometry of a Counting Task
structure: Linear Direction · models: DeepSeek-R1-Distill-Llama-8B, Llama-3.1-8B · paper: Internal states before "wait" modulate reasoning patterns