An eval/deploy probe generalizes but its steering direction is inert
measured in 1 paperNguyen et al. train linear probes on Llama-3.3-70B-Instruct activations that separate real-world evaluation prompts from deployment prompts (AUROC ~0.82) [nguyen-etal-2025-probing-and-steering-evaluation-awareness-of-language-models] The same probes classify current artificial safety evaluations as test-like (Insider Trading 100%, Sandbagging 94.2%, Roleplaying 80.6%), evidencing an internal eval/deploy axis [nguyen-etal-2025-probing-and-steering-evaluation-awareness-of-language-models] Steering along the probe direction had negligible behavioral effect (<5% sandbagging recovery), so the direction is not behaviorally load-bearing [nguyen-etal-2025-probing-and-steering-evaluation-awareness-of-language-models] Only a prompt-suffix intervention produced meaningful recovery (83%), while SAE-feature steering reached at most about 25% [nguyen-etal-2025-probing-and-steering-evaluation-awareness-of-language-models]