InternData-A1: Pioneering High-Fidelity Synthetic Data for Pre-training Generalist Policy
Yang Tian, Yuyin Yang, Yiman Xie, Zetao Cai, Xu Shi, Ning Gao, Hangxu Liu, Xuekun Jiang, Zherui Qiu, Feng Yuan, Yaping Li, Ping Wang
摘要
Recent works explore how real and synthetic data contribute to Vision-Language-Action (VLA) models' generalization. While current VLA models have shown the strong effectiveness of large-scale real-robot pre-training, synthetic data has not previously demonstrated comparable capability at scale. This paper provides the first evidence that synthetic data alone can match the performance of the strongest 𝜋-dataset in pre-training a VLA model, revealing the substantial value of large-scale simulation. The resulting model also exhibits surprisingly zero-shot sim-to-real transfer on several challenging tasks. Our synthetic dataset, InternData-A1, contains over 630k trajectories and 7,433 hours across 4 embodiments, 18 skills, 70 tasks, and 227 scenes, covering rigid, articulated, deformable, and fluid-object manipulation. It is generated through a highly autonomous, fully decoupled, and compositional simulation pipeline that enables long-horizon skill composition, flexible task assembly, and heterogeneous embodiments with minimal manual tuning. Using the same architecture as 𝜋 0 , we pre-train a model entirely on InternData-A1 and find that it matches the official 𝜋 0 across 49 simulation tasks, 5 real-world tasks, and 4 long-horizon dexterous tasks. We release the dataset and will open-source the generation pipeline to broaden access to large-scale robotic data and to lower the barrier to scalable data creation for embodied AI research.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Rethinking Video Generation Model for the Embodied WorldYufan Deng, Zilin Pan, Hongyu Zhang, Xiaojie Li 等ICML 2026 · 被引用 24 次
- RoboInter: A Holistic Intermediate Representation Suite Towards Robotic ManipulationHao Li, Ziqin Wang, Zi-han Ding, Shuai Yang 等ICLR 2026 · 被引用 17 次
- STABLE: Simulation-Ready Tabletop Layout Generation via a Semantics–Physics Dual SystemZhen Luo, Yixuan Yang, Xudong XU, Jinkun Hao 等ICML 2026
- SynthVerse: A Large-Scale Diverse Synthetic Dataset for Point TrackingWeiguang Zhao, Haoran Xu, Xingyu Miao, Qin Zhao 等SIGGRAPH 2026
它引用的顶会 Paper9
- RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic ManipulationTianxing Chen, Zanxin Chen, Baijun Chen, Zijian Cai 等ICML 2026 · 被引用 394 次
- X-VLA: Soft-Prompted Transformer as Scalable Cross-Embodiment Vision-Language-Action ModelJinliang Zheng, Jianxiong Li, Zhihao Wang, Dongxiu Liu 等ICLR 2026 · 被引用 335 次
- Vertex Block DescentAnka He Chen, Ziheng Liu, Yin Yang, Cem YukselSIGGRAPH 2024 · 被引用 35 次
- ManiSkill2: A Unified Benchmark for Generalizable Manipulation SkillsJiayuan Gu, Fanbo Xiang, Xuanlin Li, Zhan Ling 等ICLR 2023 · 被引用 21 次
- ArtVIP: Articulated Digital Assets of Visual Realism, Modular Interaction, and Physical Fidelity for Robot LearningZhao Jin, Zhengping Che, Tao Li, Zhen Zhao 等ICLR 2026 · 被引用 15 次
相关 Paper
- Sim2Real VLA: Zero-Shot Generalization of Synthesized Skills to Realistic ManipulationRunyi Zhao, Sheng Xu, Ruixing Jin, Yueci Deng 等ICLR 2026
- Joint-Aligned Latent Action: Towards Scalable VLA Pretraining in the WildHao Luo, Ye Wang, Wanpeng Zhang, Haoqi Yuan 等CVPR 2026 · 被引用 15 次
- TwinVLA: Data-Efficient Bimanual Manipulation with Twin Single-Arm Vision-Language-Action ModelsHokyun Im, Euijin Jeong, Andrey Kolobov, Jianlong Fu 等ICLR 2026 · 被引用 11 次
- DexVLG: Dexterous Vision-Language-Grasp Model at ScaleJiawei He, Danshi Li, Xinqiang Yu, Zekun Qi 等ICCV 2025 · 被引用 6 次
- Perceiving and Acting in First-Person: A Dataset and Benchmark for Egocentric Human-Object-Human InteractionsLiang Xu, Chengqun Yang, Zili Lin, Fei Xu 等ICCV 2025 · 被引用 2 次
