Relaxes TopK SAE's fixed per-sample active-latent count to a per-batch budget — keeping the top n×k activations across a batch of n samples rather than exactly k per sample — letting easy samples use fewer latents and hard samples use more, then restoring single-sample inference via a threshold estimated from the average minimum-positive-activation across batches.
Used in (6 observations)
structure: Linear Direction · models: Whisper small, Whisper large-v3 · paper: Whisper Hallucination Detection and Mitigation via Hidden Representation Steering and Sparse AutoEncoders
structure: Linear Direction · models: GPT-2 Small, Gemma-2-2B · paper: BatchTopK Sparse Autoencoders
structure: Linear Direction · models: scGPT (whole-human pretrained checkpoint), scFoundation (pretrained) · paper: Sparse Autoencoders Reveal Interpretable Features in Single-Cell Foundation Models
structure: Linear Direction · models: CosyVoice3 (Qwen2.5-0.5B backbone) · paper: Interpreting and Steering a Text-to-Speech Language Model with Sparse Autoencoders
structure: Linear Direction, Anisotropy · models: CLIP, SigLIP, SigLIP 2, AIMv2 · paper: Interpreting the Linear Structure of Vision-Language Model Embedding Spaces
structure: Linear Direction · models: Stable Diffusion v1.5 · paper: Residualized Temporal Sparse Autoencoders for Interpreting Diffusion Models