Video2Robo: 3DGS-based Synthetic Data from One Video Enables Scalable Robot Learning
Yinan Deng, Kejia Hu, Ye Chen, Jianyu Dou, Jiahui Wang, Jingyu Zhao, Haojia Ao, Yi Yang, Yufeng Yue
摘要
Scalable robot learning is hindered by the high cost of acquiring diverse, high-quality embodied data. Existing data generation approaches partially mitigate this issue but typically depend on hard-to-access hardware and labor-intensive manual effort, with limited generalization to diverse scene configurations. To overcome these limitations, we propose Video2Robo, a framework that generates high-quality and diverse robot data directly from a single human demonstration video, enabling seamless deployment on physical robots. At its core, Video2Robo leverages 3D Gaussian Splatting (3DGS) as a powerful scene representation, enabling high-fidelity rendering and explicit 3D scene editing. The framework tracks temporally consistent motion trajectories of task-relevant objects from raw video footage and identifies key task skills, guiding robots to execute tasks kinematically plausibly under novel object arrangements. Furthermore, by augmenting backgrounds, textures, lighting, and camera views, Video2Robo further enhances the diversity of generated data. Extensive evaluations in both simulation and real-world environments demonstrate that policies trained on Video2Robo data achieve superior generalization and transfer performance.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper21
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 被引用 6,549 次
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 被引用 5,687 次
- Habitat 2.0: Training Home Assistants to Rearrange their HabitatAndrew Szot, Alexander Clegg, Eric Undersander, Erik Wijmans 等NeurIPS 2021 · 被引用 826 次
- RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic ManipulationTianxing Chen, Zanxin Chen, Baijun Chen, Zijian Cai 等ICML 2026 · 被引用 394 次
相关 Paper
- RoboPearls: Editable Video Simulation for Robot ManipulationTang Tao, Likui Zhang, Youpeng Wen, Kaidong Zhang 等ICCV 2025 · 被引用 4 次
- Generative Gaussian Splatting: Generating 3D Scenes with Video Diffusion PriorsKatja Schwarz, Norman Müller, Peter KontschiederICCV 2025 · 被引用 3 次
- CloDS: Visual-Only Unsupervised Cloth Dynamics Learning in Unknown ConditionsYu-Liang Zhan, Jian Li, Wenbing Huang, Yang Liu 等ICLR 2026
- GaussianVideo: Efficient Video Representation via Hierarchical Gaussian SplattingAndrew Bond, Jui-Hsien Wang, Long Mai, Erkut Erdem 等ICCV 2025 · 被引用 14 次
- Lyra: Generative 3D Scene Reconstruction via Video Diffusion Model Self-DistillationSherwin Bahmani, Tianchang Shen, Jiawei Ren, Jiahui Huang 等ICLR 2026 · 被引用 33 次
