Seeing Realism from Simulation: Efficient Video Transfer for Vision-Language-Action Data Augmentation
Chenyu Hui, Xiaodi Huang, Siyu Xu, Yunke Wang, Shan You, Fei Wang, Tao Huang, Chang Xu
Abstract
Vision-language-action (VLA) models typically rely on large-scale real-world videos, whereas simulated data, despite being inexpensive and highly parallelizable to collect, often suffers from a substantial visual domain gap and limited environmental diversity, resulting in weak real-world generalization. We present an efficient video augmentation framework that converts simulated VLA videos into realistic training videos while preserving task semantics and action trajectories. Our pipeline extracts structured conditions from simulation via video semantic segmentation and video captioning, rewrites captions to diversify environments, and uses a conditional video transfer model to synthesize realistic videos. To make augmentation practical at scale, we introduce a diffusion feature-reuse mechanism that reuses video tokens across adjacent timesteps to accelerate generation, and a coreset sampling strategy that identifies a compact, non-redundant subset for augmentation under limited computation.Extensive experiments on Robotwin 2.0, LIBERO, LIBERO-Plus, and a real robotic platform demonstrate consistent improvements.For example, our method improves RDT-1B by 8% on Robotwin 2.0, and boosts by 5.1% on the more challenging LIBERO-Plus benchmark. Code is available at: https://github.com/nanfangxiansheng/Seeing-Realism-from-Simulation
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7ebc2790-6f59-4fce-9370-c11e3fa6b37dCited by top-tier papers1
Ask how each one uses itBuilds on14
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Depth Anything V2Lihe Yang, Bingyi Kang, Zilong Huang, Zhen Zhao et al.NeurIPS 2024 · 2,305 citations
- Deduplicating Training Data Makes Language Models BetterKatherine Lee, Daphne Ippolito, Andrew Nystrom, Chiyuan Zhang et al.ACL 2022 · 844 citations
- Beyond neural scaling laws: beating power law scaling via data pruningBen Sorscher, Robert Geirhos, Shashank Shekhar, Surya Ganguli et al.NeurIPS 2022 · 720 citations
- RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic ManipulationTianxing Chen, Zanxin Chen, Baijun Chen, Zijian Cai et al.ICML 2026 · 394 citations
Related papers
- VLAW: Iterative Co-Improvement of Vision-Language-Action Policy and World ModelYanjiang Guo, Tony Lee, Lucy Xiaoyang Shi, Jianyu Chen et al.ICML 2026 · 29 citations
- Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action ModelJohn Won, Kyungmin Lee, Huiwon Jang, Dongyoung Kim et al.ICML 2026 · 22 citations
- Learning a Unified Latent Action Space from Videos with Action-centric Cycle ConsistencyGuangyan Chen, Qi Shao, Te Cui, Zichen Zhou et al.CVPR 2026
- Unified Vision-Language-Action ModelYuqi Wang, Xinghang Li, Wenxuan Wang, Junbo Zhang et al.ICLR 2026 · 144 citations
- Sim2Real VLA: Zero-Shot Generalization of Synthesized Skills to Realistic ManipulationRunyi Zhao, Sheng Xu, Ruixing Jin, Yueci Deng et al.ICLR 2026
