A probe on SAE-PCA activations detects reward hacking token-by-token
measured in 1 paperWilhelm, Wittkopp & Kao train a per-layer sparse autoencoder on the last four layers of Qwen2.5-Instruct-7B, LLaMA-3.1-8B, and Falcon3-7B, then fit a logistic-regression probe on PCA-reduced SAE features for a per-token hack probability [wilhelm-wittkopp-kao-2026-monitoring-emergent-reward-hacking-during-generation-via-internal-activations] Against a GPT-4o judge (metric F1), the probe reaches F1=1.000 only on the benign control-adapter data, while reward-hacking detection F1 is 0.760-0.961 (e.g. Llama 0.961, Qwen 0.784) [wilhelm-wittkopp-kao-2026-monitoring-emergent-reward-hacking-during-generation-via-internal-activations] Reward-hacking versus benign activations are linearly separable in the SAE+PCA-reduced space; no causal steering is performed [wilhelm-wittkopp-kao-2026-monitoring-emergent-reward-hacking-during-generation-via-internal-activations]