MATH · IN · MODELS

A temporal-preference subgraph collapses continuous horizon geometry into a binary preference

measured in 1 paper

Rios-Sialer et al. causally localize a temporal-preference subgraph in Qwen3-4B-Instruct-2507 (layers 17-35, layer-24 attention) via four independent localization pipelines [riossialer-etal-2026-temporal-preference-concepts-and-their-functions-in-a-large-language-model] PCA within the subgraph shows time horizons form ordinal clusters whose separability is unstable until the user-to-assistant turn boundary, where attention collapses the continuous horizon into a binary preference [riossialer-etal-2026-temporal-preference-concepts-and-their-functions-in-a-large-language-model] The model's discount rates (k<0.005) are 3-8x below human controls (k~0.013) [riossialer-etal-2026-temporal-preference-concepts-and-their-functions-in-a-large-language-model] Contrastive Activation Addition with a probe-derived vector at layers 19-22 shifts temporal preference bidirectionally, revealing a probing-steering layer dissociation [riossialer-etal-2026-temporal-preference-concepts-and-their-functions-in-a-large-language-model]

Context

temporal-preference, manifold-collapse

Papers

Temporal Preference Concepts and Their Functions in a Large Language Model — Rios-Sialer, Ian, Darveshi, Shantanu, Jiang, Shuai, Paudel, Avigya, Pronina, Anastasiia, Bandyopadhyay, Ipshita, Shenk, Justin2026 · arXiv:2606.05194