The honest-vs-deceptive activation cloud collapses to near-rank-1 in some real architectures but not others
measured in 1 paperZolfaghari LoRA-fine-tunes five real model families (Pythia-1.4B, Gemma-2-2B/9B, Qwen2.5-7B, Llama-3.1-8B) into honest and deceptive variants and applies a mechanistic geometry suite (effective rank, participation ratio, Fisher Discriminant Ratio, centroid distance, adjacent-layer direction cosine) to the labeled honest-vs-deceptive activation cloud [zolfaghari-2026-multi-model-study-linear-representations-synthetic-deception] Pythia, Llama, and Qwen collapse the deception-direction cloud to near-rank-1 effective rank (1.06-1.07), while Gemma-2 retains a much higher effective rank (60-234) for the same labeled contrast, an architecture-dependent bifurcation [zolfaghari-2026-multi-model-study-linear-representations-synthetic-deception] Linear-probe decodability of the same honest-vs-deceptive contrast is uniformly near-ceiling (AUC >= 0.99) across all five architectures regardless of the rank-collapse split, showing decodability and rank structure are separate properties [zolfaghari-2026-multi-model-study-linear-representations-synthetic-deception]