EgoWorld: Translating Exocentric View to Egocentric View using Rich Exocentric Observations
Junho Park, Andrew Sangwoo Ye, Taein Kwon
摘要
Egocentric vision is essential for both human and machine visual understanding, particularly in capturing the detailed hand-object interactions needed for manipulation tasks. Translating third-person views into first-person views significantly benefits augmented reality (AR), virtual reality (VR) and robotics applications. However, current exocentric-to-egocentric translation methods are limited by their dependence on 2D cues, synchronized multi-view settings, and unrealistic assumptions such as the necessity of an initial egocentric frame and relative camera poses during inference. To overcome these challenges, we introduce EgoWorld, a novel framework that reconstructs an egocentric view from rich exocentric observations, including point clouds, 3D hand poses, and textual descriptions. Our approach reconstructs a point cloud from estimated exocentric depth maps, reprojects it into the egocentric perspective, and then applies diffusion model to produce dense, semantically coherent egocentric images. Evaluated on four datasets (i.e., H2O, TACO, Assembly101, and Ego-Exo4D), EgoWorld achieves state-of-the-art performance and demonstrates robust generalization to new objects, actions, scenes, and subjects. Moreover, EgoWorld exhibits robustness on in-the-wild examples, underscoring its practical applicability. Project page is available at https://redorangeyellowy.github.io/EgoWorld/.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper29
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
相关 Paper
- UniEgoMotion: A Unified Model for Egocentric Motion Reconstruction, Forecasting, and GenerationChaitanya Patel, Hiroki Nakamura, Yuta Kyuragi, Kazuki Kozuka 等ICCV 2025 · 被引用 3 次
- EgoTactile: Learning Grasp Pressure for Everyday Objects from Egocentric VideoYuan Zeng, Yujia Shi, Tiao Tan, Xingting Li 等ICML 2026 · 被引用 1 次
- EgoRoC: Towards Egocentric Robotic Control via Task-Agnostic Visual AlignmentWei Feng, Chi Zhang, Nan Li, Qian Zhang 等CVPR 2026
- Generating 6DoF Object Manipulation Trajectories from Action Description in Egocentric VisionTomoya Yoshida, Shuhei Kurita, Taichi Nishimura, Shinsuke MoriCVPR 2025
- EgoForce: Forearm-Guided Camera-Space 3D Hand Pose from a Monocular Egocentric CameraChristen Millerdurai, Shaoxiang Wang, Yaxu Xie, Vladislav Golyanik 等SIGGRAPH 2026
