MATH · IN · MODELS

A probe on SAE-PCA activations detects reward hacking token-by-token

measured in 1 paper

Wilhelm, Wittkopp & Kao train a per-layer sparse autoencoder on the last four layers of Qwen2.5-Instruct-7B, LLaMA-3.1-8B, and Falcon3-7B, then fit a logistic-regression probe on PCA-reduced SAE features for a per-token hack probability [wilhelm-wittkopp-kao-2026-monitoring-emergent-reward-hacking-during-generation-via-internal-activations] Against a GPT-4o judge (metric F1), the probe reaches F1=1.000 only on the benign control-adapter data, while reward-hacking detection F1 is 0.760-0.961 (e.g. Llama 0.961, Qwen 0.784) [wilhelm-wittkopp-kao-2026-monitoring-emergent-reward-hacking-during-generation-via-internal-activations] Reward-hacking versus benign activations are linearly separable in the SAE+PCA-reduced space; no causal steering is performed [wilhelm-wittkopp-kao-2026-monitoring-emergent-reward-hacking-during-generation-via-internal-activations]

Context

a linearly-decodable activation-space signal for reward-hacking behavior, extracted via SAE decomposition + PCA reduction and measurable during generation rather than only post-hoc, cross-model-family replication of a behaviorally-relevant linear separability result

Papers

Monitoring Emergent Reward Hacking During Generation via Internal Activations — Wilhelm, Jonas, Wittkopp, Alexander, Kao, Angela2026 · arXiv:2603.04069