MATH · IN · MODELS

SAE latents in a frozen multirobot policy are steerable via closed-loop affine edits

measured in 1 paper

Das et al. train a sparse autoencoder over a frozen multi-quadrotor navigation policy's activations to identify behavior-relevant latent features [das-etal-2026-steering-multirobot-behavior-via-closed-loop-affine-activation-editing] A lightweight RL steering policy applies state-dependent affine edits (scale and shift, not just addition) to selected SAE latents at every inference step [das-etal-2026-steering-multirobot-behavior-via-closed-loop-affine-activation-editing] Applied purely at the activation level without changing the frozen policy's weights, the edits steer velocity profiles and coordinate formation-preserving behavior [das-etal-2026-steering-multirobot-behavior-via-closed-loop-affine-activation-editing] They also induce a novel emergent behavior (reduced camera-surveillance exposure) absent from the unedited policy, establishing the latent directions as causally sufficient [das-etal-2026-steering-multirobot-behavior-via-closed-loop-affine-activation-editing]

Context

multirobot control, sparse autoencoders, closed-loop steering, affine activation editing, emergent behavior

Papers

Steering Multirobot Behavior via Closed-Loop Affine Activation Editing — Das, Satyajeet, Chiu, Darren, Hegde, Shashank, Sukhatme, Gaurav S.2026 · arXiv:2606.11489