MATH · IN · MODELS

A probe predicts tool-call necessity even when generation fails to

measured in 1 paper

Sun et al. show prompting and Reason-then-Act give unreliable control over tool-call decisions, sometimes collapsing accuracy (Llama-3.1-8B-Instruct 79.5% to 31.2%) [sun-etal-2026-llm-agents-already-know-when-to-call-tools-even-without-reasoning] A logistic-regression probe on pre-generation all-layer last-token hidden states predicts binary tool necessity at AUROC 0.89-0.96 across six LLMs [sun-etal-2026-llm-agents-already-know-when-to-call-tools-even-without-reasoning] This includes the two Llama models whose own generation fails to express the knowledge, dissociating what a model knows internally from what it outputs [sun-etal-2026-llm-agents-already-know-when-to-call-tools-even-without-reasoning]

Context

tool-use, agentic-llms

Papers

LLM Agents Already Know When to Call Tools — Even Without Reasoning — Sun, Chung-En, Liu, Linbo, Yan, Ge, Wang, Zimo, Weng, Tsui-Wei2026 · arXiv:2605.09252