MATH · IN · MODELS

Affine maps transfer SAEs, probes, and steering across model sizes

measured in 1 paper

Chen, Merullo, Stolfo & Pavlick fit affine maps between residual streams of differently-sized pretrained models and transfer whole SAEs, probes, and steering vectors [chen-etal-2025-transferring-linear-features-across-language-models-with-model-stitching] For the GPT-2 pair (small to medium) stitching preserves cross-entropy within roughly 9-11% overhead [chen-etal-2025-transferring-linear-features-across-language-models-with-model-stitching] Overhead is larger for other families (Pythia deduped 17-27%, Gemma-2 8.3-39%), so the 9-11% figure is GPT-2-specific [chen-etal-2025-transferring-linear-features-across-language-models-with-model-stitching] Using a transferred SAE as initialization for a larger target model cuts SAE training cost by roughly 50% [chen-etal-2025-transferring-linear-features-across-language-models-with-model-stitching]

Context

a single cross-model affine map transferring an entire family of already-extracted objects (SAE dictionaries, probes, steering vectors) rather than one direction at a time, using a transferred object as a cheaper initialization for a larger model, rather than only as a static evaluation of transfer quality

Papers

Transferring Linear Features Across Language Models With Model Stitching — Chen, Zhenyu, Merullo, Jack, Stolfo, Alessandro, Pavlick, Ellie2025 · arXiv:2506.06609