MATH · IN · MODELS

Real GPT-2, TinyLlama, and Qwen2.5 representation trajectories show higher curvature for reasoning than lexical tasks, and a universal three-phase layerwise structure

measured in 1 paper

Five geometric metrics (trajectory length, curvature, semantic convergence index, layerwise cosine similarity, representational stability) are computed on the population trajectory traced by hidden representations across layers of three real trained transformers -- GPT-2, TinyLlama, and Qwen2.5 -- over five semantically controlled prompt families [pandey-etal-2026-trajectory-geometry-transformer-representations] Reasoning and analogy tasks produce trajectories of significantly greater curvature than lexical-variation tasks (0.71-0.83 rad vs. 0.27-0.31 rad) across all three architectures [pandey-etal-2026-trajectory-geometry-transformer-representations] Semantically related prompts show statistically significant trajectory convergence peaking in middle-to-late layers (convergence index 0.41-0.58, p<0.001, Mann-Whitney U), and ambiguous tokens show measurable trajectory bifurcation of up to 5.6x final-layer separation, absent in unambiguous controls [pandey-etal-2026-trajectory-geometry-transformer-representations] Layerwise cosine similarity reveals a universal three-phase structure (encoding, elaboration, output preparation) whose boundaries are consistent across all three architectures [pandey-etal-2026-trajectory-geometry-transformer-representations]

Context

trajectory curvature, semantic convergence, representational bifurcation, three-phase layerwise structure

Method

Papers

Trajectory Geometry of Transformer Representations Across Layers — Pandey, Vishal, Singh, Gopal, Mahdid, Yacine2026 · arXiv:2606.09287