FOCA: Future-Oriented Conditioning for Data-Efficient Vision-Language-Action Adaptation
Duc Nguyen, Nghiem Diep, Binh Nguyen Gia, Trong-Bao Ho, Doanh Le Thien, Quang Nguyen, Thien-Loc Ha, Tran Van Nhiem, Bao Thach, Tran Nhat, Tuan Tran, Artur Habuda
Abstract
Vision-Language-Action (VLA) models enable general-purpose robotic control via large-scale multimodal pretraining, yet their effectiveness under few-shot imitation learning remains limited. We conduct a systematic stress test of state-ofthe-art VLA models and show that performance degrades sharply as demonstrations are reduced, revealing a key weakness of existing adaptation strategies. To address this, we introduce FOCA, a future-oriented conditioning framework for dataefficient VLA adaptation. FOCA combines explicit prediction of task-grounded future interaction embeddings with implicit alignment to future goal observations, enabling long-horizon reasoning in latent space without pixel-level prediction. This formulation naturally supports actionfree co-training with synthetic videos from video world models and can be interpreted as learning a future-conditioned value-like representation. Extensive experiments demonstrate FOCA achieves 95.7% success with 20 demonstrations on LIBERO, improves 7-12% on RoboCasa, and delivers up to 26% absolute gains on real robots, establishing a new state of the art in few-shot VLA adaptation. Our code is available at this link.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d76ecee5-d4b6-42ef-8a4e-15e7ec660572Builds on13
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 2,932 citations
- Video PreTraining (VPT): Learning to Act by Watching Unlabeled Online VideosBowen Baker, Ilge Akkaya, Peter Zhokhov, Joost Huizinga et al.NeurIPS 2022 · 458 citations
- Contrastive Learning as Goal-Conditioned Reinforcement LearningBenjamin Eysenbach, Tianjun Zhang, Sergey Levine, Ruslan SalakhutdinovNeurIPS 2022 · 331 citations
- DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World KnowledgeWenyao Zhang, Hongsi Liu, Zekun Qi, Yunnan Wang et al.NeurIPS 2025 · 244 citations
Related papers
- Motion Dynamics Learning for Few-Shot Embodied AdaptationSibo He, Weiying Xie, Daixun Li, Junhao Zhong et al.ICML 2026
- Actions as Language: Fine-Tuning VLMs into VLAs Without Catastrophic ForgettingAsher J. Hancock, Xindi Wu, Lihan Zha, Olga Russakovsky et al.ICLR 2026 · 58 citations
- RA-VLA: Retrieval-Augmented VLA for Test-Time AdaptationSanghwan Jang, Minjin Jeon, Minsoo Kim, Seong Jin Choi et al.ICML 2026
- ForceVLA2: Unleashing Hybrid Force-Position Control with Force Awareness for Contact-Rich ManipulationYang Li, Zhaxizhuoma, Hongru Jiang, Junjie Xia et al.CVPR 2026 · 31 citations
- From Abstraction to Instantiation: Learning Behavioral Representation for Vision-Language-Action ModelBing Hu, Zaijing Li, Rui Shao, Junda Chen et al.ICML 2026 · 4 citations
