ForceVLA: Enhancing VLA Models with a Force-aware MoE for Contact-rich Manipulation
Jiawen Yu, Hairuo Liu, Qiaojun Yu, Jieji Ren, Ce Hao, Haitong Ding, Guangyu Huang, Guofan Huang, Yan Song, Panpan Cai, Wenqiang Zhang, Cewu Lu
Abstract
Vision-Language-Action (VLA) models have advanced general-purpose robotic manipulation by leveraging pretrained visual and linguistic representations. However, they struggle with contact-rich tasks that require fine-grained control involving force, especially under visual occlusion or dynamic uncertainty. To address these limitations, we propose ForceVLA, a novel end-to-end manipulation framework that treats external force sensing as a first-class modality within VLA systems. ForceVLA introduces FVLMoE, a force-aware Mixture-of-Experts fusion module that dynamically integrates pretrained visual-language embeddings with real-time 6-axis force feedback during action decoding. This enables context-aware routing across modality-specific experts, enhancing the robot's ability to adapt to subtle contact dynamics. We also introduce ForceVLA-Data, a new dataset comprising synchronized vision, proprioception, and force-torque signals across five contact-rich manipulation tasks. ForceVLA improves average task success by 23.2% over strong pi_0-based baselines, achieving up to 80% success in tasks such as plug insertion. Our approach highlights the importance of multimodal integration for dexterous manipulation and sets a new benchmark for physically intelligent robotic control. Code and data will be released at https://sites.google.com/view/forcevla2025.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- ForceVLA2: Unleashing Hybrid Force-Position Control with Force Awareness for Contact-Rich ManipulationYang Li, Zhaxizhuoma, Hongru Jiang, Junjie Xia et al.CVPR 2026 · 31 citations
- Adaptive Action Chunking at Inference-time for Vision-Language-Action ModelsYuanchang Liang, Xiaobo Wang, Kai Wang, Shuo Wang et al.CVPR 2026 · 30 citations
- AT-VLA: Adaptive Tactile Injection for Enhanced Feedback Reaction in Vision-Language-Action ModelsXiaoqi Li, Muhe Cai, Jiadong Xu, Juan Zhu et al.CVPR 2026 · 18 citations
- AtomicVLA: Unlocking the Potential of Atomic Skill Learning in RobotsLikui Zhang, Tao Tang, Zhihao Zhan, Xiuwei Chen et al.CVPR 2026 · 18 citations
- Cross-Hand Latent Representation for Vision-Language-Action ModelsGuangqi Jiang, Yutong Liang, Jianglong Ye, Jia-Yang Huang et al.CVPR 2026 · 14 citations
Builds on15
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 2,932 citations
- GShard: Scaling Giant Models with Conditional Computation and Automatic ShardingDmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen et al.ICLR 2021 · 1,954 citations
- Scaling Vision with Sparse Mixture of ExpertsCarlos Riquelme, Joan Puigcerver, Basil Mustafa, Maxim Neumann et al.NeurIPS 2021 · 1,213 citations
- GLaM: Efficient Scaling of Language Models with Mixture-of-ExpertsNan Du, Yanping Huang, Andrew M. Dai, Simon Tong et al.ICML 2022 · 1,173 citations
Related papers
- Tabero: Learning Gentle Manipulation with Closed-Loop Force Feedback from Vision, Touch, and LanguageQiwei Wu, Rui Zhang, Xin Xiang, Tao Li et al.ICML 2026 · 2 citations
- Spatial Forcing: Implicit Spatial Representation Alignment for Vision-language-action ModelFuhao Li, Wenxuan Song, Han Zhao, Jingbo Wang et al.ICLR 2026 · 145 citations
- PixelVLA: Advancing Pixel-level Understanding in Vision-Language-Action ModelWenqi Liang, Gan Sun, Yao He, Jiahua Dong et al.ICLR 2026 · 20 citations
- AIR-VLA: Vision-Language-Action Systems for Aerial ManipulationJianli Sun, Bin Tian, Qiyao Zhang, Chengxiang Li et al.ICML 2026 · 4 citations
- STOLA: Self-Adaptive Touch-Language Framework for Tactile Commonsense Reasoning in Open-Ended ScenariosNing Cheng, Jinan Xu, Jialing Chen, Bin Fang et al.AAAI 2026
