methods / Causal Validation / Causal interventions (steering) / Closed-Loop Affine Activation Editing (CLAE)
Closed-Loop Affine Activation Editing (CLAE)
Trains a sparse autoencoder over a frozen policy's activations to identify behavior-relevant latent features via post-hoc probing, then trains a lightweight RL steering policy that applies state-dependent affine edits (scale and shift, not just addition) to selected SAE latents at every inference step, closing the loop between the edit and the policy's own changing state.