Replaces a standard SAE's ReLU-plus-L1-penalty with a hard TopK activation that keeps exactly the k largest pre-activations and zeroes the rest, fixing sparsity directly (removing the L1-shrinkage bias) and enabling a dead-latent-revival auxiliary loss (AuxK) that keeps dead-latent fraction low even at 16-million-latent scale, yielding clean joint scaling laws in autoencoder size and sparsity.
Used in (11 observations)
structure: Linear Direction · models: Whisper small, HuBERT-base · paper: AudioSAE: Towards Understanding of Audio-Processing Models with Sparse AutoEncoders
structure: Linear Direction, Anisotropy · models: CLIP ViT-B/32, MetaCLIP ViT-B/32, OpenCLIP ViT-B/32, SigLIP 2 · paper: Same Concept, Different Directions: Cross-Modal Feature Heterogeneity in Sparse Autoencoders
structure: Linear Direction · models: Phi-3-mini, Llama-3.1-8B-Instruct · paper: Steering Language Model Refusal with Sparse Autoencoder Features
structure: Linear Direction · models: text-embedding-3-small · paper: Disentangling Dense Embeddings with Sparse Autoencoders
structure: Linear Direction · models: Whisper base · paper: Mechanistic Interpretability of ASR models using Sparse Autoencoders
structure: Linear Subspace · models: E5-large-v2 · paper: Aligning Sentence Embeddings to Human Concepts via Sparse Autoencoders
structure: Linear Direction · models: GPT-2 Small, GPT-4 · paper: Scaling and Evaluating Sparse Autoencoders
structure: Linear Direction · models: ESM-2 (650M) · paper: Sparse Autoencoders for Low-N Protein Function Prediction and Design
structure: Linear Centroids Hypothesis · models: ResNet-50 (supervised, ImageNet), DINOv2 ViT-L/14, DINOv3 ViT-B/16, GPT-2-Large, Llama-3.1-8B · paper: The Linear Centroids Hypothesis: Features as Directions Learned by Local Experts
structure: Linear Separability · models: DAC (Descript Audio Codec), SpeechTokenizer · paper: Towards Interpretable Framework for Neural Audio Codecs via Sparse Autoencoders: A Case Study on Accent Information
structure: Linear Direction · models: Qwen3.5-35B-A3B (MoE, ~3B active/token) · paper: Behavioral Steering in a 35B MoE Language Model via SAE-Decoded Probe Vectors: One Agency Axis, Not Five Traits