Perceiving and Acting in First-Person: A Dataset and Benchmark for Egocentric Human-Object-Human Interactions
Liang Xu, Chengqun Yang, Zili Lin, Fei Xu, Yifan Liu, Congsheng Xu, Yiyi Zhang, Jie Qin, Xingdong Sheng, Yunhui Liu, Xin Jin, Yichao Yan
摘要
Learning action models from real-world human-centric interaction datasets is important towards building generalpurpose intelligent assistants with efficiency. However, most existing datasets only offer specialist interaction category and ignore that AI assistants perceive and act based on first-person acquisition. We urge that both the generalist interaction knowledge and egocentric modality are indispensable. In this paper, we embed the manual-assisted task into a vision-language-action framework, where the assistant provides services to the instructor following egocentric vision and commands. With our hybrid RGB-MoCap system, pairs of assistants and instructors engage with multi-ple objects and the scene following GPT-generated scripts. Under this setting, we accomplish InterVLA, the first largescale human-object-human interaction dataset with 11.4 hours and 1.2M frames of multimodal data, spanning 2 egocentric and 5 exocentric videos, accurate human/object motions and verbal commands. Furthermore, we establish novel benchmarks on egocentric human motion estimation, interaction synthesis, and interaction prediction with comprehensive analysis. We believe that our InterVLA testbed and the benchmarks will foster future works on building AI agents in the physical world.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- InterPrior: Scaling Generative Control for Physics-Based Human-Object InteractionsSirui Xu, Samuel Schulter, Morteza Ziyadi, Xialin He 等CVPR 2026 · 被引用 14 次
- IMU-HOI: A Symbiotic Framework for Coherent Human-Object Interaction and Motion Capture via Contact-Conscious Inertial FusionLizhou Lin, Songpengcheng Xia, Zengyuan Lai, Lan Sun 等CVPR 2026 · 被引用 1 次
- Stability-Driven Motion Generation for Object-Guided Human-Human Co-ManipulationJiahao Xu, Xiaohan Yuan, Xingchen Wu, Chongyang Xu 等CVPR 2026
- Unified Number-Free Text-to-Motion Generation Via Flow MatchingGuanhe Huang, Oya ÇeliktutanCVPR 2026
它引用的顶会 Paper68
- PaLM-E: An Embodied Multimodal Language ModelDanny Driess, Fei Xia, Mehdi S. M. Sajjadi, Corey Lynch 等ICML 2023 · 被引用 2,601 次
- Action-Conditioned 3D Human Motion Synthesis with Transformer VAEMathis Petrovich, Michael J. Black, Gül VarolICCV 2021 · 被引用 672 次
- Ego4D: Around the World in 3, 000 Hours of Egocentric VideoKristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis 等CVPR 2022 · 被引用 525 次
- Generating Diverse and Natural 3D Human Motions from TextChuan Guo, Shihao Zou, Xinxin Zuo, Sen Wang 等CVPR 2022 · 被引用 462 次
- EPIC-Fusion: Audio-Visual Temporal Binding for Egocentric Action RecognitionEvangelos Kazakos, Arsha Nagrani, Andrew Zisserman, Dima DamenICCV 2019 · 被引用 395 次
相关 Paper
- HoloAssist: an Egocentric Human Interaction Dataset for Interactive AI Assistants in the Real WorldXin Wang, Taein Kwon, Mahdi Rad, Bowen Pan 等ICCV 2023 · 被引用 151 次
- Exo2Ego: Exocentric Knowledge Guided MLLM for Egocentric Video UnderstandingHaoyu Zhang, Qiaohui Chu, Meng Liu, Haoxiang Shi 等AAAI 2026 · 被引用 17 次
- SIMMC-VR: A Task-oriented Multimodal Dialog Dataset with Situated and Immersive VR StreamsTe-Lin Wu, Satwik Kottur, Andrea Madotto, Mahmoud Azab 等ACL 2023 · 被引用 4 次
- HumanVLA: Towards Vision-Language Directed Object Rearrangement by Physical HumanoidXinyu Xu, Yizheng Zhang, Yonglu Li, Lei Han 等NeurIPS 2024 · 被引用 29 次
- InternData-A1: Pioneering High-Fidelity Synthetic Data for Pre-training Generalist PolicyYang Tian, Yuyin Yang, Yiman Xie, Zetao Cai 等CVPR 2026 · 被引用 64 次
