Forecasting 3D Scanpaths in Egocentric Video
Fiona Ryan, Ishwarya Ananthabhotla, Yijun Qian, Judy Hoffman, James M. Rehg, Vamsi Krishna Ithapu, Calvin Murdock
Abstract
Forecasting gaze behavior is an important task for understanding user intent and creating AR/VR systems that can anticipate where users will look and interact next. While prior works have addressed predicting scanpaths in static images, forecasting gaze in egocentric videos presents new challenges due to the dynamic nature of the scene and the camera wearer's continuous movement through the 3D environment. To address these challenges, we formulate the novel task of egocentric scanpath prediction as forecasting a sequence of future fixations in 3D Cartesian coordinates relative to the last observed camera pose, producing a 3D scanpath that is grounded in the environment. We propose a transformer architecture that leverages egocentric video frames, head pose, and past 3D gaze observations to predict future 3D fixation sequences. We evaluate our method on the Aria Digital Twin dataset. Our findings establish a baseline for the novel task of 3D scanpath prediction and highlight important architectural elements for our task.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8a83fd5a-8cfe-44ce-abcd-dd412474d3dcBuilds on19
- MotionGPT: Human Motion as a Foreign LanguageBiao Jiang, Xin Chen, Wen Liu, Jingyi Yu et al.NeurIPS 2023 · 698 citations
- Aria Digital Twin: A New Benchmark Dataset for Egocentric 3D Machine PerceptionXiaqing Pan, Nicholas Charron, Yongqian Yang, Scott Peters et al.ICCV 2023 · 145 citations
- Atari-HEAD: Atari Human Eye-Tracking and Demonstration DatasetRuohan Zhang, Calen Walshe, Zhuode Liu, Lin Guan et al.AAAI 2020 · 77 citations
- Breaking The Limits of Text-conditioned 3D Motion Synthesis with Elaborative DescriptionsYijun Qian, Jack Urbanek, Alexander G. Hauptmann, Jungdam WonICCV 2023 · 15 citations
- EyeFormer: Predicting Personalized Scanpaths with Transformer-Guided Reinforcement LearningYue Jiang, Zixin Guo, Hamed Rezazadegan Tavakoli, Luis A. Leiva et al.UIST 2024 · 15 citations
Related papers
- LookOut: Real-World Humanoid Egocentric NavigationBoxiao Pan, Adam W. Harley, Francis Engelmann, C. Karen Liu et al.ICCV 2025 · 2 citations
- Gaze Beyond the Frame: Forecasting Egocentric 3D Visual SpanHeeseung Yun, Joonil Na, Jaeyeon Kim, Calvin Murdock et al.NeurIPS 2025 · 3 citations
- Uncertainty-aware State Space Transformer for Egocentric 3D Hand Trajectory ForecastingWentao Bao, Lele Chen, Libing Zeng, Zhong Li et al.ICCV 2023 · 34 citations
- Joint Hand Motion and Interaction Hotspots Prediction from Egocentric VideosShaowei Liu, Subarna Tripathi, Somdeb Majumdar, Xiaolong WangCVPR 2022 · 69 citations
- Modeling Human Gaze Behavior with Diffusion Models for Unified Scanpath PredictionGiuseppe Cartella, Vittorio Cuculo, Alessandro D'Amelio, Marcella Cornia et al.ICCV 2025 · 3 citations
