MATH · IN · MODELS

Othello-GPT's board direction is decodable early but causally used only deep

measured in 1 paper

Hazineh et al. confirm Othello-GPT's linear player-relative board encoding via probing and introduce a learned linear inverse-map edit toward a target board state [hazineh-etal-2023-linear-latent-world-models-othello-gpt] Sweeping depth (1-8 layers), the linear board direction is decodable even in 1-layer models but is causally used for next-move logits only in deeper models [hazineh-etal-2023-linear-latent-world-models-othello-gpt] The causal effect concentrates in middle layers, a depth-dependence earlier Othello-GPT papers did not characterize [hazineh-etal-2023-linear-latent-world-models-othello-gpt] Representation presence and causal use are thus dissociable, with the transition from presence to use tracking network depth [hazineh-etal-2023-linear-latent-world-models-othello-gpt]

Context

world models, causal validation, model-depth scaling, toy/synthetic sequence model, linear representation hypothesis

Confirmed in models

Papers

Linear Latent World Models in Simple Transformers: A Case Study on Othello-GPT — Hazineh, Dean S., Zhang, Zechen, Chiu, Jeffrey2023 · arXiv:2310.07582