MATH · IN · MODELS

Agentic-trait probe directions share R^2 yet are orthogonal, revealing one agency axis

measured in 1 paper

Yap trains 9 TopK SAEs across depths and sublayers of Qwen3.5-35B-A3B and fits ridge probes per agentic trait in SAE latent space, projected through the decoder into steering vectors [yap-2026-behavioral-steering-moe-agency-axis] Risk-calibration and tool-use-eagerness probes share nearly identical R^2 (0.795 vs 0.792) yet are nearly orthogonal (cosine -0.017), dissociating predictive geometry from causal specificity [yap-2026-behavioral-steering-moe-agency-axis] Fewer than 1% of SAE features explain 50% of each steering vector's norm [yap-2026-behavioral-steering-moe-agency-axis] Autonomy steering at prefill achieves Cohen's d=1.01 (shifting ask_user behavior to near-zero) while decode-only steering is null, and no steering vector achieves specificity >1.0, so every vector modulates one shared agency axis [yap-2026-behavioral-steering-moe-agency-axis]

Context

near-orthogonal probe directions with shared causal axis, prefill-vs-decode steering dissociation, feature-concentration in steering vectors

Papers

Behavioral Steering in a 35B MoE Language Model via SAE-Decoded Probe Vectors: One Agency Axis, Not Five Traits — Yap, Jia Qing2026 · arXiv:2603.16335