Llama Scope's 256-SAE suite reveals feature geometry at 8B scale
measured in 1 paperHe et al. train 256 TopK sparse autoencoders (32K and 128K widths) across every layer and sublayer of Llama-3.1-8B-Base [he-etal-2024-llama-scope] A feature-geometry analysis finds nearest-neighbor structure and semantically coherent neighborhoods (e.g. a "Threats-to-Humanity" cluster) among learned SAE latents [he-etal-2024-llama-scope] Feature splitting -- wider SAEs learning genuinely finer rather than duplicated features -- is confirmed at open-weight 8B production scale [he-etal-2024-llama-scope] The result extends the smaller-scale JumpReLU/Gemma-Scope width-ladder and the original Claude-3-Sonnet demonstration to a fully open model and SAE suite [he-etal-2024-llama-scope]
Structure
Context
sparse-autoencoders, feature-splitting
Confirmed in models
Papers
Llama Scope: Extracting Millions of Features from Llama-3.1-8B with Sparse Autoencoders — He, Zhengfu, Shu, Wentao, Ge, Xuyang, Chen, Lingjie, Wang, Junxuan, Zhou, Yunhua, Liu, Frances, Guo, Qipeng, Huang, Xuanjing, Wu, Zuxuan, Jiang, Yu-Gang, Qiu, Xipeng