Diff-in-means moral-foundation directions, discovered across 14 real LLMs, are causally steerable and rewired (not newly formed) by post-training
measured in 1 paperYu, Yi, Karimi-Malekabadi, Abdurahman, Ye, Narayanan, Zhao & Dehghani extract difference-in-means moral-foundation directions across 14 real pretrained checkpoints (Llama-3.1-8B/70B, Qwen2.5-7B/14B/32B, Qwen3-30B-A3B, Mistral-7B-v0.3, base and instruct), finding significant linear separability in all 35 (model, foundation) pairs (Wasserstein distance 0.16-0.71, AUC greater than 0.55) [yu-yi-etal-2026-tracing-moral-foundations-in-large-language-models] Direction-reversal rate drops sharply with post-training (e.g. Llama-3.1-8B 33 percent to 4 percent), indicating the directions emerge during pretraining and are selectively rewired rather than formed de novo by instruction tuning [yu-yi-etal-2026-tracing-moral-foundations-in-large-language-models] Causal steering with the extracted directions passes a dose-response criterion in 68 of 70 pairs, and SAE decoder directions cosine-aligned with the dense vectors give finer-grained micro-steering that exceeds macro-vector steering in 17 of 20 cells while better preserving general capability [yu-yi-etal-2026-tracing-moral-foundations-in-large-language-models]