MATH · IN · MODELS

Othello-GPT's board state is linear under the model's own relative frame

measured in 1 paper

Nanda et al. show Othello-GPT's board state, decodable only nonlinearly under absolute BLACK/WHITE labels, is near-perfectly linearly decodable under a relative MINE/YOURS/EMPTY frame (99.5% by the final layer vs 74-75% absolute) [nanda-etal-2023] The world model was linear all along under the basis the model itself uses, and the original nonlinear result was a labeling artifact [nanda-etal-2023] Adding the linear probe directions steers the model as effectively as Li et al.'s more complex gradient-based editing (0.02 vs 0.11 errors for erasing a tile) [nanda-etal-2023] Further linear sub-findings include an EMPTY direction, a causally-relevant FLIPPED direction, and monotonic layer-by-layer refinement of the board representation [nanda-etal-2023]

Context

world models, relative vs. absolute framing, probing methodology, linear representation hypothesis, toy/synthetic sequence model, multiple circuits, iterative refinement

Confirmed in models

Papers

Emergent Linear Representations in World Models of Self-Supervised Sequence Models — Nanda, Neel, Lee, Andrew, Wattenberg, Martin2023 · arXiv:2309.00941