MATH · IN · MODELS

Evaluation-awareness probe depth shifts late-to-early with model scale

measured in 1 paper

Manek projects a diff-of-means evaluation-awareness direction (from 203 contrastive prompt pairs) onto residual-stream activations at every layer across 11 open-weight models (Qwen2.5 0.5B-32B, Gemma2 2B/9B/27B, Llama-3.2 1B/3B) [manek-2026-representational-depth-of-evaluation-awareness-shifts-with-scale] The relative depth at which per-layer decoding AUROC peaks shifts from late layers in small models to the earliest layers in large ones (Qwen2.5 1.5B/3B peak at 0.96-0.97 vs 14B/32B at 0.021-0.031) [manek-2026-representational-depth-of-evaluation-awareness-shifts-with-scale] Gemma2 shows the same late-to-early shift (0.885 to 0.304) while Llama-3.2 stays mid-layer across both tested sizes [manek-2026-representational-depth-of-evaluation-awareness-shifts-with-scale] Peak AUROC ranges 0.586-0.873 and is non-monotonic with scale within a family [manek-2026-representational-depth-of-evaluation-awareness-shifts-with-scale]

Context

evaluation-awareness, scaling-geometry

Papers

Representational Depth of Evaluation Awareness Shifts With Scale in Open-Weight LMs — Manek, Archit2026 · arXiv:2606.29196