Evaluation-awareness probe depth shifts late-to-early with model scale
measured in 1 paperManek projects a diff-of-means evaluation-awareness direction (from 203 contrastive prompt pairs) onto residual-stream activations at every layer across 11 open-weight models (Qwen2.5 0.5B-32B, Gemma2 2B/9B/27B, Llama-3.2 1B/3B) [manek-2026-representational-depth-of-evaluation-awareness-shifts-with-scale] The relative depth at which per-layer decoding AUROC peaks shifts from late layers in small models to the earliest layers in large ones (Qwen2.5 1.5B/3B peak at 0.96-0.97 vs 14B/32B at 0.021-0.031) [manek-2026-representational-depth-of-evaluation-awareness-shifts-with-scale] Gemma2 shows the same late-to-early shift (0.885 to 0.304) while Llama-3.2 stays mid-layer across both tested sizes [manek-2026-representational-depth-of-evaluation-awareness-shifts-with-scale] Peak AUROC ranges 0.586-0.873 and is non-monotonic with scale within a family [manek-2026-representational-depth-of-evaluation-awareness-shifts-with-scale]