D2E: Scaling Vision-Action Pretraining on Desktop Data for Transfer to Embodied AI
Suhwan Choi, Jaeyoon Jung, Haebin Seong, Minchan Kim, Minyeong Kim, Yongjun Cho, Yoonshik Kim, Yubeen Park, Youngjae Yu, Yunsung Lee
摘要
Large language models leverage internet-scale text data, yet embodied AI remains constrained by the prohibitive costs of physical trajectory collection. Desktop environments---particularly gaming---offer a compelling alternative: they provide rich sensorimotor interactions at scale while maintaining the structured observation-action coupling essential for embodied learning. We present D2E (Desktop to Embodied AI), a framework that demonstrates desktop interactions can serve as an effective pretraining substrate for robotics embodied AI tasks. Unlike prior work that remained domain-specific (e.g., VPT for Minecraft) or kept data proprietary (e.g., SIMA), D2E establishes a complete pipeline from scalable desktop data collection to verified transfer in embodied domains. Our framework comprises three components: (1) the OWA Toolkit that unifies diverse desktop interactions into a standardized format with 152× compression, (2) the Generalist-IDM that achieves strong zero-shot generalization across unseen games through timestamp-based event prediction, enabling internet-scale pseudo-labeling, and (3) VAPT that transfers desktop-pretrained representations to physical manipulation and navigation. Using 1.3K+ hours of data (259 hours of human demonstrations and 1K+ hours of pseudo-labeled gameplay), our 1B-parameter model achieves 96.6% success on LIBERO manipulation and 83.3% on CANVAS navigation, matching or surpassing models up to 7 larger, such as (3.3B) and OpenVLA (7B). These results demonstrate that sensorimotor primitives learned from digital interactions transfer effectively to real-world physical tasks, establishing desktop pretraining as a practical paradigm for embodied AI. All resources are publicly available at https://worv-ai.github.io/d2e.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper7
- Learning Universal Policies via Text-Guided Video GenerationYilun Du, Sherry Yang, Bo Dai, Hanjun Dai 等NeurIPS 2023 · 被引用 742 次
- Video PreTraining (VPT): Learning to Act by Watching Unlabeled Online VideosBowen Baker, Ilge Akkaya, Peter Zhokhov, Joost Huizinga 等NeurIPS 2022 · 被引用 458 次
- Knowledge Insulating Vision-Language-Action Models: Train Fast, Run Fast, Generalize BetterDanny Driess, Jost Tobias Springenberg, Brian Ichter, Lili Yu 等NeurIPS 2025 · 被引用 162 次
- GUI-World: A Video Benchmark and Dataset for Multimodal GUI-oriented UnderstandingDongping Chen, Yue Huang, Siyuan Wu, Jingyu Tang 等ICLR 2025 · 被引用 1 次
- Latent Action Pretraining from VideosSeonghyeon Ye, Joel Jang, Byeongguk Jeon, Se June Joo 等ICLR 2025
相关 Paper
- Where are we in the search for an Artificial Visual Cortex for Embodied Intelligence?Arjun Majumdar, Karmesh Yadav, Sergio Arnaud, Yecheng Jason Ma 等NeurIPS 2023 · 被引用 336 次
- Learning to Move Before Learning to Do: Task-Agnostic pretraining for VLAsJunhao Shi, Siyin Wang, Xiaopeng Yu, Li Ji 等ICML 2026
- From Multimodal LLMs to Generalist Embodied Agents: Methods and LessonsAndrew Szot, Bogdan Mazoure, Omar Attia, Aleksei Timofeev 等CVPR 2025
- Architect: Generating Vivid and Interactive 3D Scenes with Hierarchical 2D InpaintingYian Wang, Xiaowen Qiu, Jiageng Liu, Zhehuan Chen 等NeurIPS 2024 · 被引用 48 次
- InternData-A1: Pioneering High-Fidelity Synthetic Data for Pre-training Generalist PolicyYang Tian, Yuyin Yang, Yiman Xie, Zetao Cai 等CVPR 2026 · 被引用 64 次
