Sparse logit-lens directions in VLA FFN value vectors steer robot actions
measured in 1 paperHaon et al. project FFN value vectors in OpenVLA-7B and Pi0 onto the vocabulary basis (logit-lens), revealing sparse directions for concepts like "speed", "direction", "up", and "slow" [haon-etal-2025-mechanistic-interpretability-steering-vla] Fewer than 25% of OpenVLA's 352,255 FFN value vectors are rewired for action prediction, the rest inherited semantic directions from the vision-language backbone [haon-etal-2025-mechanistic-interpretability-steering-vla] Upweighting a semantic-direction cluster steers both models zero-shot: "fast" interventions give larger end-effector displacement than "slow" [haon-etal-2025-mechanistic-interpretability-steering-vla] Full-depth "up" cluster injections give the largest mean Y-displacement, validated in LIBERO-Long (OpenVLA) and on a physical UR5 arm (Pi0-FAST) against a random-cluster control [haon-etal-2025-mechanistic-interpretability-steering-vla]