Topic subspaces sharpen with depth; a centroid direction induces CoT
measured in 1 paperSaglam et al. study 11 autoregressive models on six arXiv topics, measuring per-layer hard-margin SVM separability and PCA intrinsic dimensionality [saglam-etal-2025-linear-subspaces] Under 10% of principal components explain nearly all variance, so topics occupy a near-affine low-dimensional subspace; separability rises toward final layers, reaching 100% SVM accuracy on all topic pairs in the largest models [saglam-etal-2025-linear-subspaces] A chain-of-thought framing distinction becomes linearly separable even more sharply than topic identity [saglam-etal-2025-linear-subspaces] Adding the CoT-vs-non-CoT centroid-difference direction to the hidden state reliably induces CoT-style responses (flagged preliminary), and a lightweight MLP guardrail built on the finding halves harmful responses [saglam-etal-2025-linear-subspaces]