MATH · IN · MODELS

A value-like success direction is recoverable from frozen VLAs

measured in 1 paper

Zhang et al. linearly probe frozen OpenVLA, Pi0.5, DINOv2, and CLIP features on LIBERO-Goal trajectories to recover Monte-Carlo task-success targets [zhang-etal-2026-what-frozen-vlas-already-know-about-success] A linear, value-like task-success direction is consistently recoverable across all four frozen backbones, none of which was trained to estimate reward [zhang-etal-2026-what-frozen-vlas-already-know-about-success] Deploying the probe as a test-time selector over Pi0.5 action-prefix candidates raises push-plate success from 26.7% (greedy) to 44.3% (p=0.003) [zhang-etal-2026-what-frozen-vlas-already-know-about-success] The direction is thus not merely decodable but behaviorally useful without any additional policy training [zhang-etal-2026-what-frozen-vlas-already-know-about-success]

Context

vision-language-action models, robot policies, value-like structure, test-time selection, frozen backbone probing

Papers

What Frozen VLAs Already Know About Success: A Probing Study of Value-Like Structure in Foundation Robot Policies — Zhang, Jiachen, Nie, Junnan, Lao, Junyi, Cheng, Wei, Liu, Chenghao, Jiang, Jiaxin, Huang, Songfang2026 · arXiv:2605.28527