From Abstraction to Instantiation: Learning Behavioral Representation for Vision-Language-Action Model
Bing Hu, Zaijing Li, Rui Shao, Junda Chen, April Hua Liu, Wei-Shi Zheng, Liqiang Nie
Abstract
Vision-Language-Action (VLA) models often suffer from performance degradation under distribution shifts, as they struggle to learn generalized behavior representations across varying environments. While existing approaches attempt to construct behavior representations through action-centric latent variables, they are often limited by short-horizon temporal fragmentation and static execution-alignment, leading to inconsistent behaviors in complex scenarios. To address these limitations, we propose BehaviorVLA, a framework that facilitates robust manipulation through the learning of a temporally coherent behavioral representations. Our approach features two symmetric components: (1) the Visuomotor Behavior Encoder (VBE), which utilizes a causal Mamba-based architecture to aggregate long-horizon trajectory information into a unified behavior representation; and (2) the Phase-conditioned Behavior Decoder (PBD), which decodes this representation into precise actions by dynamically aligning task-level priors with real-time execution progress. Experiments on RoboTwin 2.0, LIBERO, and CALVIN demonstrate state-of-the-art success rates of 58%, 98%, and 4.36 (Avg. Len), respectively. Notably, in real-world sim-to-real transfer, BehaviorVLA matches the performance of OpenVLA-OFT using only 50% of the demonstration data, showcasing its superior data efficiency and generalization.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b33f3136-23a7-4616-bbd7-e1daffedec18Builds on23
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
- Behavior Transformers: Cloning modes with one stoneNur Muhammad Shafiullah, Zichen Jeff Cui, Ariuntuya Altanzaya, Lerrel PintoNeurIPS 2022 · 470 citations
- RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic ManipulationTianxing Chen, Zanxin Chen, Baijun Chen, Zijian Cai et al.ICML 2026 · 394 citations
- Prismatic VLMs: Investigating the Design Space of Visually-Conditioned Language ModelsSiddharth Karamcheti, Suraj Nair, Ashwin Balakrishna, Percy Liang et al.ICML 2024 · 306 citations
Related papers
- SimpleVLA-RL: Scaling VLA Training via Reinforcement LearningHaozhan Li, Yuxin Zuo, Jiale Yu, Yuhao Zhang et al.ICLR 2026 · 170 citations
- Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic ManipulationZaijing Li, Bing Hu, Rui Shao, Gongwei Chen et al.CVPR 2026 · 23 citations
- FOCA: Future-Oriented Conditioning for Data-Efficient Vision-Language-Action AdaptationDuc Nguyen, Nghiem Diep, Binh Nguyen Gia, Trong-Bao Ho et al.ICML 2026 · 3 citations
- Unified Vision-Language-Action ModelYuqi Wang, Xinghang Li, Wenxuan Wang, Junbo Zhang et al.ICLR 2026 · 144 citations
- LARA: Latent Action Representation Alignment for Vision-Language-Action ModelsMengya Liu, Baoxiong Jia, Jiangyong Huang, Jingze Zhang et al.ICML 2026 · 3 citations
