EgoExo-Gen: Ego-centric Video Prediction by Watching Exo-centric Videos
Jilan Xu, Yifei Huang, Baoqi Pei, Junlin Hou, Qingqiu Li, Guo Chen, Yuejie Zhang, Rui Feng, Weidi Xie
2025Year
5Top-tier citations
Abstract
First frame Predicted future frames ⋅⋅⋅ "C stirs the noodles in the pot with the spoon in his right hand." Exo-centric video Ego-centric video Figure 1 : The cross-view video prediction task aims to predict future RGB frames of the ego-centric video, given the first ego-centric frame, a text instruction, and a synchronised exo-centric video.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 115100cc-da82-4fec-91a6-a14ce9a1104bCited by top-tier papers5
- EgoWorld: Translating Exocentric View to Egocentric View using Rich Exocentric ObservationsJunho Park, Andrew Sangwoo Ye, Taein KwonICLR 2026 · 10 citations
- EgoX: Egocentric Video Generation from a Single Exocentric VideoTaewoong Kang, Kinam Kim, Dohyeon Kim, Minho Park et al.CVPR 2026 · 9 citations
- Towards Comprehensive Scene Understanding: Integrating First and Third-Person Views for LVLMsInsu Lee, Wooje Park, Jaeyun Jang, Minyoung Noh et al.NeurIPS 2025 · 8 citations
- EgoControl: Controllable Egocentric Video Generation via 3D Full-Body PosesEnrico Pallotta, Sina Mokhtarzadeh Azar, Lars Doorenbos, Serdar Ozsoy et al.CVPR 2026 · 7 citations
- Walk in Others' Shoes with a Single Glance: Human-Centric Visual Grounding with Top-View Perspective TransformationYuqi Bu, Xin Wu, Zirui Zhao, Yi Cai et al.ACL 2025
Builds on31
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific TuningYuwei Guo, Ceyuan Yang, Anyi Rao, Zhengyang Liang et al.ICLR 2024 · 1,493 citations
Related papers
- Joint Hand Motion and Interaction Hotspots Prediction from Egocentric VideosShaowei Liu, Subarna Tripathi, Somdeb Majumdar, Xiaolong WangCVPR 2022 · 69 citations
- Cross-View Exocentric to Egocentric Video SynthesisGaowen Liu, Hao Tang, Hugo Latapie, Jason J. Corso et al.ACM MM 2021 · 22 citations
- Look Before You Speak: Visually Contextualized UtterancesPaul Hongsuck Seo, Arsha Nagrani, Cordelia SchmidCVPR 2021
- Ego-PMOVE: Prompt-aware Mixture of View Experts Network for Egocentric Gaze PredictionHeqian Qiu, Lanxiao Wang, Taijin Zhao, Zhaofeng Shi et al.AAAI 2026
- Exocentric-to-Egocentric Video GenerationJia-Wei Liu, Weijia Mao, Zhongcong Xu, Jussi Keppo et al.NeurIPS 2024 · 27 citations
