Stack depth is linearly decoded in tiny counter-language transformers
measured in 1 paperTiwari et al. train a tiny 1-layer transformer per formal language on Dyck-1 and Shuffle-2/4/6 counter languages, reframing next-token prediction as multi-label valid-token prediction [tiwari-etal-2025] Linear probes on the final layer decode each stack's current depth with high accuracy, while a Hewitt-Liang control-task check stays near chance [tiwari-etal-2025] Accuracy is systematically higher for Shuffle languages than Dyck-1 and higher for larger k, attributed to each individual stack updating less often as k grows [tiwari-etal-2025] The paper stops at the linear-direction claim and explicitly disclaims any causal characterization of the stacks [tiwari-etal-2025]
Structure
Context
formal languages, counter languages, stack representations, toy/synthetic sequence model, control tasks, circuit discovery
Confirmed in models
Method
Papers
Emergent Stack Representations in Modeling Counter Languages Using Transformers — Tiwari, Utkarsh, Gupta, Aviral, Hahn, Michael