MATH · IN · MODELS

Llama Scope's 256-SAE suite reveals feature geometry at 8B scale

measured in 1 paper

He et al. train 256 TopK sparse autoencoders (32K and 128K widths) across every layer and sublayer of Llama-3.1-8B-Base [he-etal-2024-llama-scope] A feature-geometry analysis finds nearest-neighbor structure and semantically coherent neighborhoods (e.g. a "Threats-to-Humanity" cluster) among learned SAE latents [he-etal-2024-llama-scope] Feature splitting -- wider SAEs learning genuinely finer rather than duplicated features -- is confirmed at open-weight 8B production scale [he-etal-2024-llama-scope] The result extends the smaller-scale JumpReLU/Gemma-Scope width-ladder and the original Claude-3-Sonnet demonstration to a fully open model and SAE suite [he-etal-2024-llama-scope]

Context

sparse-autoencoders, feature-splitting

Confirmed in models

Papers

Llama Scope: Extracting Millions of Features from Llama-3.1-8B with Sparse Autoencoders — He, Zhengfu, Shu, Wentao, Ge, Xuyang, Chen, Lingjie, Wang, Junxuan, Zhou, Yunhua, Liu, Frances, Guo, Qipeng, Huang, Xuanjing, Wu, Zuxuan, Jiang, Yu-Gang, Qiu, Xipeng2024 · arXiv:2410.20526