Othello-move representations from seven architecturally distinct real models (GPT-2, BART, T5, Flan-T5-XL, LLaMA-2-7B, Mistral-7B, Qwen2.5-7B) Procrustes-align into a shared geometric space at cosine similarity up to 97.2% (unsupervised), with PCA trajectories converging across models and nearest-neighbor tile embeddings recovering the board's actual physical adjacency structure
measured in 1 paperYuan & Søgaard (2025) apply supervised and unsupervised (adversarial + iterative refinement) Procrustes alignment, adapted from cross- lingual word-embedding literature, to final-hidden-layer Othello-move representations from seven differently-architected real models (GPT-2-small, BART-base, T5-base, Flan-T5-XL/3B, LoRA-fine-tuned LLaMA-2-7B, LoRA-fine-tuned Mistral-7B, Qwen2.5-7B). Cross-model cosine similarity after alignment reaches up to 93.1% (supervised, GPT-2 <-> BART) and 97.2% (unsupervised, BART <-> Mistral); PCA visualization of the per-game move trajectory shows convergent geometric structure across models; layer-wise similarity heatmaps peak along the diagonal at corresponding depths; and a "latent move projection" analysis finds the nearest-neighbor tile embedding to any given tile is consistently its actual spatial neighbor on the physical board -- a genuine spatial/geometric isomorphism claim beyond simple linear decodability. Clears scope on criterion (a): a measured, quantified cross-model shared-manifold structure in real pretrained-and-fine-tuned models, going beyond the original linear- probing Othello-GPT results already in this map. No causal intervention is performed (purely representational/alignment analysis). Cross-model convergence of this kind is platonic- representation-adjacent; see [[platonic-representation]].