MATH · IN · MODELS

A weak-perturbation response direction stays near-orthogonal across every layer

measured in 1 paper

Luick applies a weak scaling perturbation to a single token's residual activation and tracks the downstream response across all layers of Gemma-2-2B, Llama-3.2-3B-Instruct, and GPT-2-XL [luick-2024-universal-response-emergence-of-induction] The response-direction-to-state cosine similarity stays below 0.1 in magnitude across the entire residual stream [luick-2024-universal-response-emergence-of-induction] This near-orthogonality co-occurs with a scale-invariant response regime [luick-2024-universal-response-emergence-of-induction] Both properties are tied to the induction mechanism, strongest for perturbations at token positions that induction heads copy from [luick-2024-universal-response-emergence-of-induction]

Context

induction

Papers

Universal Response and Emergence of Induction in LLMs — Luick, Niclas2024 · arXiv:2411.07071