MATH · IN · MODELS

Claude-3-Sonnet SAE features causally steer safety behaviors and split with scale

measured in 1 paper

Templeton et al. train three SAEs (~1M/4M/34M features) on the middle-layer residual stream of Claude 3 Sonnet, the first SAE interpretability at frontier scale [templeton-etal-2024-scaling-monosemanticity] Clamping named features causally steers behavior: the Golden Gate Bridge feature at 10x makes the model self-identify as the bridge, and an unsafe-code feature at 5x produces a buffer overflow [templeton-etal-2024-scaling-monosemanticity] Further clamped features drive sycophantic praise, deception/secrecy, honesty, and racist invective, spanning safety-relevant domains [templeton-etal-2024-scaling-monosemanticity] Feature splitting is quantified: a "San Francisco" feature in the 1M SAE splits into 2 (4M) then 11 (34M) finer features, each other's cosine nearest neighbors [templeton-etal-2024-scaling-monosemanticity]

Context

feature clamping (steering), feature splitting across SAE dictionary scale, cosine-similarity feature neighborhoods, safety-relevant causal features (deception, sycophancy, code vulnerabilities, bias)

Confirmed in models

Papers

Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet — Templeton, Adly, Conerly, Tom, Marcus, Jonathan, Lindsey, Jack, Bricken, Trenton, Chen, Brian, Pearce, Adam, Citro, Craig, Ameisen, Emmanuel, Jones, Andy, Cunningham, Hoagy, Turner, Nicholas L., McDougall, Callum, MacDiarmid, Monte, Freeman, C. Daniel, Sumers, Theodore R., Rees, Edward, Batson, Joshua, Jermyn, Adam, Carter, Shan, Olah, Chris, Henighan, Tom2024