MATH · IN · MODELS

An eval/deploy probe generalizes but its steering direction is inert

measured in 1 paper

Nguyen et al. train linear probes on Llama-3.3-70B-Instruct activations that separate real-world evaluation prompts from deployment prompts (AUROC ~0.82) [nguyen-etal-2025-probing-and-steering-evaluation-awareness-of-language-models] The same probes classify current artificial safety evaluations as test-like (Insider Trading 100%, Sandbagging 94.2%, Roleplaying 80.6%), evidencing an internal eval/deploy axis [nguyen-etal-2025-probing-and-steering-evaluation-awareness-of-language-models] Steering along the probe direction had negligible behavioral effect (<5% sandbagging recovery), so the direction is not behaviorally load-bearing [nguyen-etal-2025-probing-and-steering-evaluation-awareness-of-language-models] Only a prompt-suffix intervention produced meaningful recovery (83%), while SAE-feature steering reached at most about 25% [nguyen-etal-2025-probing-and-steering-evaluation-awareness-of-language-models]

Context

evaluation-awareness, sandbagging

Papers

Probing and Steering Evaluation Awareness of Language Models — Nguyen, Jord, Hoang, Khiem, Attubato, Carlo Leonardo, Hofstätter, Felix2025 · arXiv:2507.01786