Latent Scaling corrects crosscoder sparsity artifacts and isolates chat latents
measured in 1 paperMinder et al. show the standard crosscoder's L1-on-decoder-norm loss can misattribute a latent's direction as chat-unique when it genuinely exists in the base model too [minder-etal-2025-crosscoders-chat-tuning-latent-scaling] Latent Scaling, a post-hoc diagnostic, re-measures each latent's true presence in each model rather than trusting the trained decoder norm [minder-etal-2025-crosscoders-chat-tuning-latent-scaling] Replacing the L1 crosscoder with a BatchTopK one substantially mitigates the artifact on Gemma-2-2B base/chat [minder-etal-2025-crosscoders-chat-tuning-latent-scaling] The corrected crosscoder isolates genuinely chat-specific, causally-steerable latents (e.g. "false information," multiple distinct refusal-trigger latents) rather than one monolithic refusal feature [minder-etal-2025-crosscoders-chat-tuning-latent-scaling]