Exocentric-to-Egocentric Video Generation
Jia-Wei Liu, Weijia Mao, Zhongcong Xu, Jussi Keppo, Mike Zheng Shou
摘要
We introduce Exo2Ego-V, a novel exocentric-to-egocentric diffusion-based video generation method for daily-life skilled human activities where sparse 4 -view exo-centric viewpoints are configured 360° around the scene. This task is particularly challenging due to the significant variations between exocentric and egocentric viewpoints and high complexity of dynamic motions and real-world daily-life environments. To address these challenges, we first propose a new diffusion-based multi-view exocentric encoder to extract the dense multi-scale features from multi-view exocentric videos as the appearance conditions for egocentric video generation. Then, we design an exocentric-to-egocentric view translation prior to provide spatially aligned egocentric features as a concatenation guidance for the input of egocentric video diffusion model. Finally, we introduce the temporal attention layers into our egocentric video diffusion pipeline to improve the temporal consistency cross egocentric frames. Extensive experiments demonstrate that Exo2Ego-V significantly outperforms SOTA approaches on 5 categories from the Ego-Exo4D dataset with an average of 35% in terms of LPIPS. Our code and model will be made available on https://github.com/showlab/Exo2Ego-V .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- EgoNight: Towards Egocentric Vision Understanding at Night with a Challenging BenchmarkDeheng Zhang, Yuqian Fu, Runyi Yang, Yang Miao 等ICLR 2026 · 被引用 19 次
- EgoWorld: Translating Exocentric View to Egocentric View using Rich Exocentric ObservationsJunho Park, Andrew Sangwoo Ye, Taein KwonICLR 2026 · 被引用 10 次
- EgoX: Egocentric Video Generation from a Single Exocentric VideoTaewoong Kang, Kinam Kim, Dohyeon Kim, Minho Park 等CVPR 2026 · 被引用 9 次
- Towards Comprehensive Scene Understanding: Integrating First and Third-Person Views for LVLMsInsu Lee, Wooje Park, Jaeyun Jang, Minyoung Noh 等NeurIPS 2025 · 被引用 8 次
- ObjectRelator: Enabling Cross-View Object Relation Understanding Across Ego-Centric and Exo-Centric PerspectivesYuqian Fu, Runze Wang, Bin Ren, Guolei Sun 等ICCV 2025 · 被引用 5 次
它引用的顶会 Paper26
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Video Diffusion ModelsJonathan Ho, Tim Salimans, Alexey A. Gritsenko, William Chan 等NeurIPS 2022 · 被引用 2,948 次
- AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific TuningYuwei Guo, Ceyuan Yang, Anyi Rao, Zhengyang Liang 等ICLR 2024 · 被引用 1,493 次
相关 Paper
- Vivid-ZOO: Multi-View Video Generation with Diffusion ModelBing Li, Cheng Zheng, Wenxuan Zhu, Jinjie Mai 等NeurIPS 2024 · 被引用 48 次
- EgoTwin: Dreaming Body and View in First PersonJingqiao Xiu, Fangzhou Hong, Yicong Li, Mengze Li 等ICLR 2026 · 被引用 14 次
- EgoControl: Controllable Egocentric Video Generation via 3D Full-Body PosesEnrico Pallotta, Sina Mokhtarzadeh Azar, Lars Doorenbos, Serdar Ozsoy 等CVPR 2026 · 被引用 7 次
- UniEgoMotion: A Unified Model for Egocentric Motion Reconstruction, Forecasting, and GenerationChaitanya Patel, Hiroki Nakamura, Yuta Kyuragi, Kazuki Kozuka 等ICCV 2025 · 被引用 3 次
- Cross-View Exocentric to Egocentric Video SynthesisGaowen Liu, Hao Tang, Hugo Latapie, Jason J. Corso 等ACM MM 2021 · 被引用 22 次
