MATH · IN · MODELS
methods / Representation Alignment / Orthogonal Procrustes alignment

Orthogonal Procrustes alignment

Techniqueintermediate

Fits the orthogonal matrix $W^\star = \arg\min_{W\in O(d)} \|WX-Y\|_F = UV^\top$ (via the SVD of $YX^\top$) that best maps one representation onto another, then evaluates the fit functionally — e.g. via nearest-neighbor retrieval or downstream transfer — rather than only reporting a single aggregate similarity number.

Used in (11 observations)

structure: Linear Direction · models: Llama-3.1-8B, Qwen2.5-14B, Qwen-2.5-7B, Qwen2.5-0.5B, Llama-3.2-1B · paper: Intrinsic Guardrails: How Semantic Geometry of Personality Interacts with Emergent Misalignment in LLMs
structure: Platonic Representation Hypothesis, Linear Subspace · models: Conneau et al. Monolingual MLM (English), Conneau et al. Monolingual MLM (French), Conneau et al. Monolingual MLM (German), Conneau et al. Monolingual MLM (Russian), Conneau et al. Monolingual MLM (Chinese) · paper: Emerging Cross-lingual Structure in Pretrained Language Models
structure: Platonic Representation Hypothesis, Linear Subspace · models: GTR-base, E5-base-v2, Stella-base-en-v2, Granite-Embedding-278M-Multilingual · paper: mini-vec2vec: Scaling Universal Geometry Alignment with Linear Transformations
structure: Platonic Representation Hypothesis, Linear Subspace · models: ViT-S/16 (I-JEPA objective, independently trained per view: smallNORB/nuScenes/ImageNet-1k) · paper: Social-JEPA: Emergent Geometric Isomorphism in Independently Trained World Models
structure: Linear Subspace · models: CLIP ViT-B/32, CLIP ViT-L/14, SigLIP, FLAVA · paper: Canonicalizing Multimodal Contrastive Representation Learning
structure: Platonic Representation Hypothesis, Linear Subspace · models: BERT-base-cased, mBERT (BERT-base, Multilingual Cased) · paper: Cross-Lingual BERT Transformation for Zero-Shot Dependency Parsing
structure: Linear Subspace · models: Llama-3-8B-Instruct, Llama-3.1-8B-Instruct, Aya Expanse 8B, Gemma 2 9B IT, Qwen2.5-7B-Instruct · paper: Understanding Subword Compositionality of Large Language Models
structure: Linear Subspace · models: GPT-2-small, GPT-2-Medium, GPT-2-Large · paper: Multilinguality as Sense Adaptation