MATH · IN · MODELS

Topic subspaces sharpen with depth; a centroid direction induces CoT

measured in 1 paper

Saglam et al. study 11 autoregressive models on six arXiv topics, measuring per-layer hard-margin SVM separability and PCA intrinsic dimensionality [saglam-etal-2025-linear-subspaces] Under 10% of principal components explain nearly all variance, so topics occupy a near-affine low-dimensional subspace; separability rises toward final layers, reaching 100% SVM accuracy on all topic pairs in the largest models [saglam-etal-2025-linear-subspaces] A chain-of-thought framing distinction becomes linearly separable even more sharply than topic identity [saglam-etal-2025-linear-subspaces] Adding the CoT-vs-non-CoT centroid-difference direction to the hidden state reliably induces CoT-style responses (flagged preliminary), and a lightweight MLP guardrail built on the finding halves harmful responses [saglam-etal-2025-linear-subspaces]

Context

hard-margin SVM separability per layer across topic pairs, low-dimensional affine subspace via 90%-variance PCA threshold, chain-of-thought framing separability exceeding topic separability, centroid-difference steering direction causally inducing CoT behavior, lightweight latent-space MLP guardrail built on the separability finding

Papers

Large Language Models Encode Semantics and Alignment in Linearly Separable Representations — Saglam, Baturay, Kassianik, Paul, Nelson, Blaine, Weerawardhena, Sajana, Singer, Yaron, Karbasi, Amin2025 · arXiv:2507.09709