Injects a fixed a priori basis (e.g. a Fourier or binary-digit matrix) into a probe's architecture as a candidate hidden representation scheme, learning only a low-dimensional projection into and out of it — testing whether a specific structured hypothesis fits activations far better than an unconstrained linear probe, rather than fitting the freest possible model.
Used in (3 observations)
structure: Lissajous Curves · models: OLMo 2 1B, OLMo 2 7B, OLMo 2 13B, Llama-3.2-1B, Llama-3.2-3B, Llama-3-8B, Llama-3.1-8B, Phi-4 (15B) · paper: Unravelling the Mechanisms of Manipulating Numbers in Language Models
structure: Circle · models: Llama-3-8B, Mistral-7B · paper: Language Models Encode Numbers Using Digit Representations in Base 10
structure: Lissajous Curves, Linear Direction · models: OLMo 2 1B, OLMo 2 7B, OLMo 2 13B, OLMo 2 32B, Llama-3.2-1B, Llama-3.2-3B, Llama-3-8B, Llama-3-70B, Phi-4 (15B) · paper: Pre-trained Language Models Learn Remarkably Accurate Representations of Numbers