GemmaScope SAE directions predict and steer code correctness
measured in 1 paperTahimic & Cheng use GemmaScope SAEs on Gemma-2-2b to find single decoder directions predicting code correctness (F1=0.821) and directions steering toward correct or incorrect code [tahimic-cheng-2025-mechanistic-interpretability-of-code-correctness-in-llms-via-sparse-autoencoders] Steering the "correct" direction gives a 4.04% correction rate on incorrect code (p<0.001) with a 14.66% corruption side-effect on correct code [tahimic-cheng-2025-mechanistic-interpretability-of-code-correctness-in-llms-via-sparse-autoencoders] Weight orthogonalization of the direction corrupts 83.6% of correct solutions versus 19.0% for a matched control (4.4x, p<0.001) [tahimic-cheng-2025-mechanistic-interpretability-of-code-correctness-in-llms-via-sparse-autoencoders] Both directions, trained only on base Gemma-2-2b, retain effectiveness after instruction-tuning (F1=0.772), evidencing a pre-existing model feature [tahimic-cheng-2025-mechanistic-interpretability-of-code-correctness-in-llms-via-sparse-autoencoders]