SAE latents in a frozen multirobot policy are steerable via closed-loop affine edits
measured in 1 paperDas et al. train a sparse autoencoder over a frozen multi-quadrotor navigation policy's activations to identify behavior-relevant latent features [das-etal-2026-steering-multirobot-behavior-via-closed-loop-affine-activation-editing] A lightweight RL steering policy applies state-dependent affine edits (scale and shift, not just addition) to selected SAE latents at every inference step [das-etal-2026-steering-multirobot-behavior-via-closed-loop-affine-activation-editing] Applied purely at the activation level without changing the frozen policy's weights, the edits steer velocity profiles and coordinate formation-preserving behavior [das-etal-2026-steering-multirobot-behavior-via-closed-loop-affine-activation-editing] They also induce a novel emergent behavior (reduced camera-surveillance exposure) absent from the unedited policy, establishing the latent directions as causally sufficient [das-etal-2026-steering-multirobot-behavior-via-closed-loop-affine-activation-editing]