MATH · IN · MODELS

GemmaScope SAE directions predict and steer code correctness

measured in 1 paper

Tahimic & Cheng use GemmaScope SAEs on Gemma-2-2b to find single decoder directions predicting code correctness (F1=0.821) and directions steering toward correct or incorrect code [tahimic-cheng-2025-mechanistic-interpretability-of-code-correctness-in-llms-via-sparse-autoencoders] Steering the "correct" direction gives a 4.04% correction rate on incorrect code (p<0.001) with a 14.66% corruption side-effect on correct code [tahimic-cheng-2025-mechanistic-interpretability-of-code-correctness-in-llms-via-sparse-autoencoders] Weight orthogonalization of the direction corrupts 83.6% of correct solutions versus 19.0% for a matched control (4.4x, p<0.001) [tahimic-cheng-2025-mechanistic-interpretability-of-code-correctness-in-llms-via-sparse-autoencoders] Both directions, trained only on base Gemma-2-2b, retain effectiveness after instruction-tuning (F1=0.772), evidencing a pre-existing model feature [tahimic-cheng-2025-mechanistic-interpretability-of-code-correctness-in-llms-via-sparse-autoencoders]

Context

code-correctness, sparse-autoencoders

Papers

Mechanistic Interpretability of Code Correctness in LLMs via Sparse Autoencoders — Tahimic, Charles, Cheng, Meng-Chieh2025 · arXiv:2510.02917