MATH · IN · MODELS

A per-tool-pair mean-difference direction reads and switches tool choice

measured in 1 paper

Wu et al. show tool selection in tool-calling agents is carried by a single mean-difference direction per tool-pair in the residual stream across Gemma 3 (270M-27B), Qwen 3 (0.6B-14B) and Llama 3.1 (8B) [wu-etal-2026-tool-calling-is-linearly-readable-and-steerable-in-language-models] Adding the direction switches the chosen tool at 83-100% on a 15-tool synthetic benchmark and 77-94% on tau-bench-airline, versus 0% for a random-direction control [wu-etal-2026-tool-calling-is-linearly-readable-and-steerable-in-language-models] PCA over per-tool mean activations puts ~91% of variance in ~10 components for 15 tools, far below a random-Gaussian control [wu-etal-2026-tool-calling-is-linearly-readable-and-steerable-in-language-models] Base (non-instruction-tuned) models already carry the correct tool internally (cosine readout 61-82% on BFCL vs 2-10% from base generation) [wu-etal-2026-tool-calling-is-linearly-readable-and-steerable-in-language-models] SAEs and cross-layer transcoders trace a three-stage circuit: early tool-selective features, mid-layer attention heads, late-layer JSON-formatting features [wu-etal-2026-tool-calling-is-linearly-readable-and-steerable-in-language-models]

Context

tool-use, agentic-llms

Papers

Tool Calling Is Linearly Readable and Steerable in Language Models — Wu, Yuxuan, Wang, Xin, Cho, Jaemin, Yang, Yi, Koshiyama, Adriano, Bulathwela, Sahan, Perez-Ortiz, Maria2026 · arXiv:2605.07990