MATH · IN · MODELS

TopK sparse-autoencoder features on a real fine-tuned ESM2 align with real protein structural sites, and steering them causally improves in-silico protein design

measured in 1 paper

Tsui, Talreja & Aghazadeh (2025) train a TopK sparse autoencoder (d=4096, k=128) on layer-24 embeddings of real ESM2-650M, LoRA-fine-tuned per assay on MSA sequences, and show the resulting sparse latents align with real biological structure (active-site residues, C-terminus, allosteric/binding/epistatic sites mapped onto AlphaFold3 structures); the top 5% of SAE probe weights (by magnitude) explain 37-38% of fitness-prediction variance versus 25-28% for raw ESM-layer/logit weights [tsui-talreja-aghazadeh-2025-sparse-autoencoders-for-low-n-protein-function-prediction-and-design] From as few as N=24 labeled sequences, SAE-based linear probes outperform raw-ESM2 baselines in 58-69% of extrapolation tasks across six DMS assays from ProteinGym [tsui-talreja-aghazadeh-2025-sparse-autoencoders-for-low-n-protein-function-prediction-and-design] Causally, amplifying predictive SAE latents (identified via probe-weight magnitude) and decoding back through the SAE and remaining ESM2 layers ("feature steering") generates new protein variants that outperform ESM2-based design in 88% of metric-by-assay combinations, producing the single best-fitness variant in 5/6 DMS assays; validation is in-silico (an MLP fitness proxy, plus one assay against ground-truth combinatorial fitness), not wet-lab [tsui-talreja-aghazadeh-2025-sparse-autoencoders-for-low-n-protein-function-prediction-and-design]

Context

sparse-feature steering, low-N fitness extrapolation

Confirmed in models

Papers

Sparse Autoencoders for Low-N Protein Function Prediction and Design — Tsui, Darin, Talreja, Kunal, Aghazadeh, Amirali2025 · arXiv:2508.18567