Crosscoders recover shared cross-layer and cross-model features
measured in 1 paperLindsey et al. introduce sparse crosscoders: one dictionary reads and writes activations at multiple layers or models at once, with the L1 penalty weighted by summed per-layer decoder norms [lindsey-etal-2024-crosscoders] Against matched per-layer SAEs a crosscoder achieves lower eval loss per feature but needs ~2x more training FLOPs, evidencing consolidated cross-layer structure [lindsey-etal-2024-crosscoders] Decoder directions drift across layers even where decoder norm persists, ruling out passive residual-stream relaying [lindsey-etal-2024-crosscoders] A cross-model crosscoder on Claude 3 Sonnet base vs finetuned splits features into shared, base-only, and finetuned-only (~4,000-5,000 model-specific per side), surfacing a refusal and a code-review feature [lindsey-etal-2024-crosscoders] Shared features' decoder directions are highly aligned between the two models, with a minority low or negatively aligned as candidate repurposed concepts [lindsey-etal-2024-crosscoders]