MATH · IN · MODELS

Symbolic object-relation and action states are linearly decodable above 90% accuracy from nearly every layer of a real LIBERO-spatial-finetuned OpenVLA-7B checkpoint

measured in 1 paper

33 per-layer linear probes (one affine layer plus sigmoid, trained separately per layer) decode 9 binary object-relation predicates (behind, in-front-of, inside, left-of, on, on-table, right-of, open, turned-on) and 2 action-state predicates (grasped, should-move-towards) from frozen OpenVLA-7B hidden states with accuracy above 0.90 for most of the 33 layers [lu-etal-2025-probing-openvla-symbolic-states] Probed model is a real LIBERO-spatial-finetuned OpenVLA-7B checkpoint (Llama-2-7B backbone, 32 transformer blocks, 4096-dimensional hidden states per layer) [lu-etal-2025-probing-openvla-symbolic-states] The paper's own hypothesized layer-ordering (object states decoded earlier than action states) was not confirmed by the probe-accuracy pattern across layers [lu-etal-2025-probing-openvla-symbolic-states]

Context

symbolic state probing, vision-language-action models, cognitive architecture integration

Confirmed in models

Papers

Probing a Vision-Language-Action Model for Symbolic States and Integration into a Cognitive Architecture — Lu, Hong, Li, Hengxu, Shahani, Prithviraj Singh, Herbers, Stephanie, Scheutz, Matthias2025 · arXiv:2502.04558