Instance Tracking in 3D Scenes from Egocentric Videos
Yunhan Zhao, Haoyu Ma, Shu Kong, Charless C. Fowlkes
摘要
Egocentric sensors such as AR/VR devices capture human-object interactions and offer the potential to provide task-assistance by recalling 3D locations of objects of interest in the surrounding environment. This capability requires instance tracking in real-world 3D scenes from egocentric videos (IT3DEgo). We explore this problem by first introducing a new benchmark dataset, consisting of RGB and depth videos, per-frame camera pose, and instance-level annotations in both 2D camera and 3D world coordinates. We present an evaluation protocol which evaluates tracking performance in 3D coordinates with two settings for enrolling instances to track: (1) single-view online enrollment where an instance is specified on-the-fly based on the human wearer's interactions. and (2) multi-view pre-enrollment where images of an instance to be tracked are stored in memory ahead of time. To address IT3DEgo, we first repurpose methods from relevant areas, e.g., single object tracking (SOT) - running SOT methods to track instances in 2D frames and lifting them to 3D using camera pose and depth. We also present a simple method that leverages pre-trained segmentation and detection models to generate proposals from RGB frames and match proposals with enrolled instance images. Our experiments show that our method (with no finetuning) significantly outperforms SOT-based approaches in the egocentric setting. We conclude by arguing that the problem of egocentric instance tracking is made easier by leveraging camera pose and using a 3D allocentric (world) coordinate representation. Dataset and open-source code: https://github.com/IT3DEgo/IT3DEgo.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Is Tracking Really More Challenging in First Person Egocentric Vision?Matteo Dunnhofer, Zaira Manigrasso, Christian MicheloniICCV 2025 · 被引用 1 次
- Solving Instance Detection from an Open-World PerspectiveQianqian Shen, Yunhan Zhao, Nahyun Kwon, Jeeeun Kim 等CVPR 2025
- EAGLE: Episodic Appearance- and Geometry-aware Memory for Unified 2D-3D Visual Query Localization in Egocentric VisionYifei Cao, Yu Liu, Guolong Wang, Zhu Liu 等AAAI 2026
- HD-EPIC: A Highly-Detailed Egocentric Video DatasetToby Perrett, Ahmad Darkhalil, Saptarshi Sinha, Omar Emara 等CVPR 2025
它引用的顶会 Paper30
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- Learning Discriminative Model Prediction for TrackingGoutam Bhat, Martin Danelljan, Luc Van Gool, Radu TimofteICCV 2019 · 被引用 1,294 次
- Learning Spatio-Temporal Transformer for Visual TrackingBin Yan, Houwen Peng, Jianlong Fu, Dong Wang 等ICCV 2021 · 被引用 1,062 次
- Habitat 2.0: Training Home Assistants to Rearrange their HabitatAndrew Szot, Alexander Clegg, Eric Undersander, Erik Wijmans 等NeurIPS 2021 · 被引用 826 次
相关 Paper
- GSOT3D: Towards Generic 3D Single Object Tracking in the WildYifan Jiao, Yunhao Li, Junhua Ding, Qing Yang 等ICCV 2025
- Scene-Aware Egocentric 3D Human Pose EstimationJian Wang, Diogo C. Luvizon, Weipeng Xu, Lingjie Liu 等CVPR 2023
- EgoPoseVR: Spatiotemporal Multi-Modal Reasoning for Egocentric Full-Body Pose in Virtual RealityHaojie Cheng, Shaun Jing Heng Ong, Shaoyu Cai, Aiden Tat Yang Koh 等IEEE VR 2026 · 被引用 1 次
- FRAME: Floor-aligned Representation for Avatar Motion from Egocentric VideoAndrea Boscolo Camiletto, Jian Wang, Eduardo Alvarado, Rishabh Dabral 等CVPR 2025
- SHOW3D: Capturing Scenes of 3D Hands and Objects in the WildPatrick Rim, Kevin Harris, Braden Copple, Shangchen Han 等CVPR 2026 · 被引用 5 次
