Orthogonal Procrustes alignment maps a Gemma Scope SAE 'Concept Atlas' onto three independently-trained Llama-3.1-8B models' activation spaces, with quantified retrieval/translation quality (AUROC 0.82-0.86, AP 0.39-0.49 vs. a 0.046 random baseline), and the transferred concept directions can be added back into the subject model's residual stream to steer generation
measured in 1 paperPuri, Berend, Lapuschkin & Samek (2025/2026) build a "Concept Atlas" from a Gemma Scope sparse autoencoder trained on Gemma-2-2B's residual stream (layer 20), then fit an orthogonal Procrustes map from three distinct "subject" models' activation spaces (Llama-3.1-8B base, Llama-3.1-8B-Instruct, DeepSeek-R1-Distill-Llama-3.1-8B-Instruct) into this atlas. Translation quality is quantified via AUROC (0.82-0.86) and average precision (0.39-0.49, versus a random baseline of 0.046) across five subject-model layers, and via near-perfect mean-reciprocal-rank retrieval on 454 independently-validated concept features. Atlas concept directions mapped back into a subject model and added (norm-preserving) to its residual stream at multiple layers simultaneously demonstrably steer generation toward the target concept, though this steering effect is reported qualitatively rather than with an inline quantified success rate.