Claude-3-Sonnet SAE features causally steer safety behaviors and split with scale
measured in 1 paperTempleton et al. train three SAEs (~1M/4M/34M features) on the middle-layer residual stream of Claude 3 Sonnet, the first SAE interpretability at frontier scale [templeton-etal-2024-scaling-monosemanticity] Clamping named features causally steers behavior: the Golden Gate Bridge feature at 10x makes the model self-identify as the bridge, and an unsafe-code feature at 5x produces a buffer overflow [templeton-etal-2024-scaling-monosemanticity] Further clamped features drive sycophantic praise, deception/secrecy, honesty, and racist invective, spanning safety-relevant domains [templeton-etal-2024-scaling-monosemanticity] Feature splitting is quantified: a "San Francisco" feature in the 1M SAE splits into 2 (4M) then 11 (34M) finer features, each other's cosine nearest neighbors [templeton-etal-2024-scaling-monosemanticity]