MATH · IN · MODELS

Probes decode a grid cognitive map that reasoning then reorganizes

measured in 1 paper

Arghal et al. use linear and MLP probes to decode a grid-position and goal-location cognitive map from GPT-OSS-20B's layer-15 pre-reasoning activations in a 2D grid-world navigation task [arghal-etal-2026-a-behavioural-and-representational-evaluation-of-goal-directedness-in-language-model-agents] The agent's chosen action agrees with the decoded map at 82.5% average accuracy across grid sizes [arghal-etal-2026-a-behavioural-and-representational-evaluation-of-goal-directedness-in-language-model-agents] Localization accuracy degrades with grid size, and post-reasoning activations reorganize the cognitive-map signal toward immediate action selection rather than a stable spatial code [arghal-etal-2026-a-behavioural-and-representational-evaluation-of-goal-directedness-in-language-model-agents]

Context

decoding a task-relevant spatial "cognitive map" from a real LLM agent's activations and validating it behaviorally via policy-agreement rather than probe accuracy alone, reasoning-induced reorganization of a decoded representation, tracked by comparing pre- and post-reasoning probe accuracy

Confirmed in models

Papers

A Behavioural and Representational Evaluation of Goal-Directedness in Language Model Agents — Arghal, Raghu, Chen, Fade, Dalton, Niall, Kortukov, Evgenii, McNamara, Calum, Nalmpantis, Angelos, Nirvaan, Moksh, Sarti, Gabriele, Giulianelli, Mario2026 · arXiv:2602.08964