MATH · IN · MODELS

LLM semantic-axis directions reproduce human semantic-differential structure

measured in 1 paper

Kozlowski & Boutyline construct 32 semantic-axis directions (e.g. beautiful-ugly, soft-hard) as mean contrastive-pair differences from residual streams of Llama-3.2-3B, Llama-3.1-70B, Qwen3-1.7B and Qwen3-32B [kozlowski-boutyline-2026-semantic-feature-geometry] Word projections onto these axes correlate with a 1,750-respondent human semantic-differential survey (Pearson r from >0.8 down to >0.3 across the 32 scales) [kozlowski-boutyline-2026-semantic-feature-geometry] The 32 axes are meaningfully non-orthogonal: their pairwise cosine similarities reproduce the pairwise correlations among the corresponding human survey scales [kozlowski-boutyline-2026-semantic-feature-geometry] PCA puts >45% (3B) / >33% (70B) of variance in the top 3 components, far above the 3.1% expected under orthogonality, matching the Evaluation-Potency-Activity triad [kozlowski-boutyline-2026-semantic-feature-geometry] Canonical correlation analysis aligns this LLM-derived 3D subspace with the equivalent subspace from the human survey [kozlowski-boutyline-2026-semantic-feature-geometry] Additive steering along one axis produces spillover on off-target axes proportional to their cosine similarity, replicated (weaker) in the 70B and both Qwen3 sizes [kozlowski-boutyline-2026-semantic-feature-geometry]

Context

contrastive-pair semantic-axis construction (10 antonym pairs/axis), cosine similarity between axes reproduces human survey correlation structure, PCA on raw semantic-axis vectors (top-3 components >45%/>33% variance vs. 3.1% orthogonal baseline), canonical correlation analysis (CCA) alignment with human survey PCA subspace, norm-relative additive steering with cross-axis spillover proportional to cosine similarity

Papers

Semantic Structure of Feature Space in Large Language Models — Kozlowski, Austin C., Boutyline, Andrei2026 · arXiv:2604.27169