Uncertainty-aware State Space Transformer for Egocentric 3D Hand Trajectory Forecasting
Wentao Bao, Lele Chen, Libing Zeng, Zhong Li, Yi Xu, Junsong Yuan, Yu Kong
摘要
Hand trajectory forecasting from egocentric views is crucial for enabling a prompt understanding of human intentions when interacting with AR/VR systems. However, existing methods handle this problem in a 2D image space which is inadequate for 3D real-world applications. In this paper, we set up an egocentric 3D hand trajectory forecasting task that aims to predict hand trajectories in a 3D space from early observed RGB videos in a first-person view. To fulfill this goal, we propose an uncertainty-aware state space Transformer (USST) that takes the merits of the attention mechanism and aleatoric uncertainty within the framework of the classical state-space model. The model can be further enhanced by the velocity constraint and visual prompt tuning (VPT) on large vision transformers. Moreover, we develop an annotation workflow to collect 3D hand trajectories with high quality. Experimental results on H2O and EgoPAT3D datasets demonstrate the superiority of USST for both 2D and 3D trajectory forecasting. The code and datasets are publicly released: https://actionlab-cv.github.io/EgoHandTrajPred.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Gaze Beyond the Frame: Forecasting Egocentric 3D Visual SpanHeeseung Yun, Joonil Na, Jaeyeon Kim, Calvin Murdock 等NeurIPS 2025 · 被引用 3 次
- Self-Supervised Monocular 4D Scene Reconstruction for Egocentric VideosChengbo Yuan, Geng Chen, Li Yi, Yang GaoICCV 2025 · 被引用 2 次
- Perceiving and Acting in First-Person: A Dataset and Benchmark for Egocentric Human-Object-Human InteractionsLiang Xu, Chengqun Yang, Zili Lin, Fei Xu 等ICCV 2025 · 被引用 2 次
- HandWorld: Hand-Centric Unified Video Action GenerationZhihao Sun, Zhiying Du, Xitong Yang, Zuxuan WuCVPR 2026 · 被引用 1 次
- Int3DNet: Scene-Motion Cross Attention Network for 3D Intention Prediction in Mixed RealityTaewook Ha, Woojin Cho, Dooyoung Kim, Woontack WooIEEE VR 2026
它引用的顶会 Paper17
- AgentFormer: Agent-Aware Transformers for Socio-Temporal Multi-Agent ForecastingYe Yuan, Xinshuo Weng, Yanglan Ou, Kris KitaniICCV 2021 · 被引用 658 次
- Ego4D: Around the World in 3, 000 Hours of Egocentric VideoKristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis 等CVPR 2022 · 被引用 525 次
- H2O: Two Hands Manipulating Objects for First Person Interaction RecognitionTaein Kwon, Bugra Tekin, Jan Stühmer, Federica Bogo 等ICCV 2021 · 被引用 271 次
- Anticipative Video TransformerRohit Girdhar, Kristen GraumanICCV 2021 · 被引用 270 次
- Uncertainty-based Traffic Accident Anticipation with Spatio-Temporal Relational LearningWentao Bao, Qi Yu, Yu KongACM MM 2020 · 被引用 191 次
相关 Paper
- Joint Hand Motion and Interaction Hotspots Prediction from Egocentric VideosShaowei Liu, Subarna Tripathi, Somdeb Majumdar, Xiaolong WangCVPR 2022 · 被引用 69 次
- Forecasting 3D Scanpaths in Egocentric VideoFiona Ryan, Ishwarya Ananthabhotla, Yijun Qian, Judy Hoffman 等CVPR 2026 · 被引用 1 次
- COPILOT: Human-Environment Collision Prediction and Localization from Egocentric VideosBoxiao Pan, Bokui Shen, Davis Rempe, Despoina Paschalidou 等ICCV 2023 · 被引用 3 次
- UniEgoMotion: A Unified Model for Egocentric Motion Reconstruction, Forecasting, and GenerationChaitanya Patel, Hiroki Nakamura, Yuta Kyuragi, Kazuki Kozuka 等ICCV 2025 · 被引用 3 次
- Egocentric Prediction of Action Target in 3DYiming Li, Ziang Cao, Andrew Liang, Benjamin Liang 等CVPR 2022 · 被引用 20 次
