MATH · IN · MODELS

Orthogonal Procrustes alignment maps a Gemma Scope SAE 'Concept Atlas' onto three independently-trained Llama-3.1-8B models' activation spaces, with quantified retrieval/translation quality (AUROC 0.82-0.86, AP 0.39-0.49 vs. a 0.046 random baseline), and the transferred concept directions can be added back into the subject model's residual stream to steer generation

measured in 1 paper

Puri, Berend, Lapuschkin & Samek (2025/2026) build a "Concept Atlas" from a Gemma Scope sparse autoencoder trained on Gemma-2-2B's residual stream (layer 20), then fit an orthogonal Procrustes map from three distinct "subject" models' activation spaces (Llama-3.1-8B base, Llama-3.1-8B-Instruct, DeepSeek-R1-Distill-Llama-3.1-8B-Instruct) into this atlas. Translation quality is quantified via AUROC (0.82-0.86) and average precision (0.39-0.49, versus a random baseline of 0.046) across five subject-model layers, and via near-perfect mean-reciprocal-rank retrieval on 454 independently-validated concept features. Atlas concept directions mapped back into a subject model and added (norm-preserving) to its residual stream at multiple layers simultaneously demonstrably steer generation toward the target concept, though this steering effect is reported qualitatively rather than with an inline quantified success rate.

Context

quantified cross-model transfer of SAE-derived concept directions via an orthogonal Procrustes map, extending the map's existing Procrustes/relative-representation cross-model-alignment evidence to sparse-autoencoder feature space specifically

Papers

Atlas-Alignment: Making Interpretability Transferable Across Language Models — Puri, Bruno, Berend, Jim, Lapuschkin, Sebastian, Samek, Wojciech2025 · arXiv:2510.27413