MATH · IN · MODELS

ESM attention maps are linearly readable as contact maps in one pass

measured in 1 paper

Thorstenson shows raw attention matrices from ESM-2 (35M-3B) and ESMC-600M, combined via a parameter-free mean read-out, recover residue-residue contact maps [thorstenson-2026-protein-contacts-are-already-in-the-attention] Accuracy is competitive with the categorical-Jacobian method, which requires an expensive combinatorial sweep of masked-residue mutations per protein [thorstenson-2026-protein-contacts-are-already-in-the-attention] Because the read-out is parameter-free and needs only a single forward pass, contact information is already linearly present in the raw attention patterns [thorstenson-2026-protein-contacts-are-already-in-the-attention]

Context

protein-structure, attention-geometry

Papers

Protein Contacts Are Already in the Attention: A Single-Forward-Pass Alternative to the Categorical Jacobian — Thorstenson, Rome2026 · arXiv:2606.21876