MATH · IN · MODELS
methods / Direction Extraction / Fixed-dictionary sparse recovery

Fixed-dictionary sparse recovery

Techniqueintermediate

Decomposes an embedding into a sparse, non-negative linear combination of a fixed, human-interpretable concept dictionary (e.g. common words/phrases passed through the model's own text encoder) via per-sample convex optimization (LASSO/ADMM) - no encoder/decoder network is trained, unlike a sparse autoencoder.

Used in (2 observations)

structure: Linear Direction · models: FLUX.1 · paper: Interpreting Large Text-to-Image Diffusion Models with Dictionary Learning
structure: Linear Direction, Anisotropy · models: OpenCLIP ViT-B/32, CLIP ResNet-50 · paper: Interpreting CLIP with Sparse Linear Concept Embeddings (SpLiCE)