MATH · IN · MODELS

Linear probes decode future Blocksworld planning steps from one pass

measured in 1 paper

Men et al. train linear probes (versus a nonlinear control) on hidden states of Llama-2-7b-chat and Vicuna-7B fine-tuned on Blocksworld planning [men-etal-2024-unlocking-the-future-look-ahead-planning-mechanistic-interpretability-in-llms] A linear probe predicts the 6th-step-ahead decision from the 1st-step representation at 0.51 accuracy, decaying smoothly with prediction distance [men-etal-2024-unlocking-the-future-look-ahead-planning-mechanistic-interpretability-in-llms] Linear and nonlinear probes track the same decay, evidence the look-ahead information is linearly encoded; MHSA key-masking confirms which attention paths carry it [men-etal-2024-unlocking-the-future-look-ahead-planning-mechanistic-interpretability-in-llms]

Context

planning, look-ahead-representations

Papers

Unlocking the Future: Look-Ahead Planning Mechanistic Interpretability in LLMs — Men, Tianyi, Cao, Pengfei, Jin, Zhuoran, Chen, Yubo, Liu, Kang, Zhao, Jun2024 · arXiv:2406.16033