Instance Tracking in 3D Scenes from Egocentric Videos
Yunhan Zhao, Haoyu Ma, Shu Kong, Charless C. Fowlkes
Abstract
Egocentric sensors such as AR/VR devices capture human-object interactions and offer the potential to provide task-assistance by recalling 3D locations of objects of interest in the surrounding environment. This capability requires instance tracking in real-world 3D scenes from egocentric videos (IT3DEgo). We explore this problem by first introducing a new benchmark dataset, consisting of RGB and depth videos, per-frame camera pose, and instance-level annotations in both 2D camera and 3D world coordinates. We present an evaluation protocol which evaluates tracking performance in 3D coordinates with two settings for enrolling instances to track: (1) single-view online enrollment where an instance is specified on-the-fly based on the human wearer's interactions. and (2) multi-view pre-enrollment where images of an instance to be tracked are stored in memory ahead of time. To address IT3DEgo, we first repurpose methods from relevant areas, e.g., single object tracking (SOT) - running SOT methods to track instances in 2D frames and lifting them to 3D using camera pose and depth. We also present a simple method that leverages pre-trained segmentation and detection models to generate proposals from RGB frames and match proposals with enrolled instance images. Our experiments show that our method (with no finetuning) significantly outperforms SOT-based approaches in the egocentric setting. We conclude by arguing that the problem of egocentric instance tracking is made easier by leveraging camera pose and using a 3D allocentric (world) coordinate representation. Dataset and open-source code: https://github.com/IT3DEgo/IT3DEgo.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 31af8562-cea5-42ff-8888-c58faa745f3fCited by top-tier papers4
- Is Tracking Really More Challenging in First Person Egocentric Vision?Matteo Dunnhofer, Zaira Manigrasso, Christian MicheloniICCV 2025 · 1 citation
- Solving Instance Detection from an Open-World PerspectiveQianqian Shen, Yunhan Zhao, Nahyun Kwon, Jeeeun Kim et al.CVPR 2025
- EAGLE: Episodic Appearance- and Geometry-aware Memory for Unified 2D-3D Visual Query Localization in Egocentric VisionYifei Cao, Yu Liu, Guolong Wang, Zhu Liu et al.AAAI 2026
- HD-EPIC: A Highly-Detailed Egocentric Video DatasetToby Perrett, Ahmad Darkhalil, Saptarshi Sinha, Omar Emara et al.CVPR 2025
Builds on30
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- Learning Discriminative Model Prediction for TrackingGoutam Bhat, Martin Danelljan, Luc Van Gool, Radu TimofteICCV 2019 · 1,294 citations
- Learning Spatio-Temporal Transformer for Visual TrackingBin Yan, Houwen Peng, Jianlong Fu, Dong Wang et al.ICCV 2021 · 1,062 citations
- Habitat 2.0: Training Home Assistants to Rearrange their HabitatAndrew Szot, Alexander Clegg, Eric Undersander, Erik Wijmans et al.NeurIPS 2021 · 826 citations
Related papers
- GSOT3D: Towards Generic 3D Single Object Tracking in the WildYifan Jiao, Yunhao Li, Junhua Ding, Qing Yang et al.ICCV 2025
- Scene-Aware Egocentric 3D Human Pose EstimationJian Wang, Diogo C. Luvizon, Weipeng Xu, Lingjie Liu et al.CVPR 2023
- EgoPoseVR: Spatiotemporal Multi-Modal Reasoning for Egocentric Full-Body Pose in Virtual RealityHaojie Cheng, Shaun Jing Heng Ong, Shaoyu Cai, Aiden Tat Yang Koh et al.IEEE VR 2026 · 1 citation
- FRAME: Floor-aligned Representation for Avatar Motion from Egocentric VideoAndrea Boscolo Camiletto, Jian Wang, Eduardo Alvarado, Rishabh Dabral et al.CVPR 2025
- SHOW3D: Capturing Scenes of 3D Hands and Objects in the WildPatrick Rim, Kevin Harris, Braden Copple, Shangchen Han et al.CVPR 2026 · 5 citations
