MATH · IN · MODELS

Ablating general SAE feature directions in a VLA collapses grasp success

measured in 1 paper

Swann et al. train sparse autoencoders on the residual-stream activations of pi0.5's PaliGemma backbone and of OpenVLA, identifying feature directions for motion primitives and semantic concepts [swann-etal-2026-sparse-autoencoders-reveal-interpretable-and-steerable-features-in-vla-models] On real-world DROID hardware, ablating the most-general cross-episode-transferable feature directions collapses grasp success, while ablating episode-specific memorized directions barely changes performance [swann-etal-2026-sparse-autoencoders-reveal-interpretable-and-steerable-features-in-vla-models] Additive steering toward specific object-feature directions shifts grasp outcomes substantially more than an FFN-neuron control, and steering a "close gripper" feature raises the gripper-closure rate [swann-etal-2026-sparse-autoencoders-reveal-interpretable-and-steerable-features-in-vla-models] The steering and ablation results are reported qualitatively (without per-trial success fractions), validated on LIBERO simulation and DROID real-world hardware [swann-etal-2026-sparse-autoencoders-reveal-interpretable-and-steerable-features-in-vla-models]

Context

causal dissociation between generalizable and episode-specific SAE feature directions via selective ablation, in a real-world robot policy, single-feature-direction additive steering producing large, targeted behavioral shifts in a non-language modality (motor action)

Papers

Sparse Autoencoders Reveal Interpretable and Steerable Features in VLA Models — Swann, Aiden, McGranahan, Lachlain, Buurmeijer, Hugo, Kennedy, Monroe, III, Schwager, Mac2026 · arXiv:2603.19183