Leela Chess Zero's correct moves appear early, then get overridden
measured in 1 paperSandmann, Lapuschkin & Samek extend the logit lens (functionally equivalent to zero-ablating later sublayers) to Leela Chess Zero's Post-LN policy network [sandmann-etal-2025-the-algorithm-is-not-the-behavior-learned-priors-override-look-ahead-in-a-chess-playing-neural-network] Playing strength and puzzle-solving ability rise monotonically with depth, but policy distributions follow non-smooth, non-monotonic trajectories [sandmann-etal-2025-the-algorithm-is-not-the-behavior-learned-priors-override-look-ahead-in-a-chess-playing-neural-network] Correct puzzle solutions are discovered in intermediate layers but subsequently discarded, with move rankings poorly correlated to the final output until late in the network [sandmann-etal-2025-the-algorithm-is-not-the-behavior-learned-priors-override-look-ahead-in-a-chess-playing-neural-network] This contrasts with the smooth distributional convergence typical of language models, evidencing iterative inference where a late stage overrides an already-computed answer [sandmann-etal-2025-the-algorithm-is-not-the-behavior-learned-priors-override-look-ahead-in-a-chess-playing-neural-network]