MATH · IN · MODELS

SLIM SAE features linearly correlate with molecular properties and steer editing

measured in 1 paper

Zhang et al. train a Gated SAE on a single layer of four frozen LLM-based molecular editors (DrugAssist, GeLLM3O-LLaMA3, GeLLM3O-Mistral, MolGen) with learnable per-property importance gates [zhang-etal-2026-slim-sparse-latent-steering-molecular] Individual top SAE features correlate strongly with molecular properties (Spearman rho up to +0.93 for molecular weight, -0.85 for QED), with six of eight properties reaching |rho|>=0.52 from a single feature [zhang-etal-2026-slim-sparse-latent-steering-molecular] Properties like HBD and DRD2 show weaker single-feature correlation, explicitly attributed to distributed rather than monosemantic encoding [zhang-etal-2026-slim-sparse-latent-steering-molecular] A gradient-based direction projected through the SAE's top features and added to the frozen model's residual stream causally steers property-directed molecular editing [zhang-etal-2026-slim-sparse-latent-steering-molecular]

Context

Gated SAE with learnable per-property importance gates on frozen molecular-editor hidden states, single-feature Spearman correlation with chemical properties as a monosemanticity measurement, gradient-derived direction projected through SAE top-k features then decoded and added at inference, quantified property-directed editing accuracy gains up to +42.4 points on MolEditRL, vanilla vs. task-oriented SAE direction cosine-similarity ablation

Papers

SLIM: Sparse Latent Steering for Interpretable and Property-Directed LLM-Based Molecular Editing — Zhang, Mingxu, Li, Yuhan, Li, Lujundong, Shen, Dazhong, Xiong, Hui, Sun, Ying2026 · arXiv:2605.10831