Detecting Invisible People
Tarasha Khurana, Achal Dave, Deva Ramanan
Abstract
Monocular object detection and tracking have improved drastically in recent years, but rely on a key assumption: that objects are visible to the camera. Many offline tracking approaches reason about occluded objects post-hoc, by linking together tracklets after the object re-appears, making use of reidentification (ReID). However, online tracking in embodied robotic agents (such as a self-driving vehicle) fundamentally requires object permanence, which is the ability to reason about occluded objects before they re-appear. In this work, we re-purpose tracking benchmarks and propose new metrics for the task of detecting invisible objects, focusing on the illustrative case of people. We demonstrate that current detection and tracking systems perform dramatically worse on this task. We introduce two key innovations to recover much of this performance drop. We treat occluded object detection in temporal sequences as a short-term forecasting challenge, bringing to bear tools from dynamic sequence prediction. Second, we build dynamic models that explicitly reason in 3D from monocular videos without calibration, using observations produced by monocular depth estimators. To our knowledge, ours is the first work to demonstrate the effectiveness of monocular depth estimation for the task of tracking and detecting occluded objects. Our approach strongly improves by 11.4% over the baseline in ablations and by 5.0% over the state-of-the-art in F1 score.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0fcbd0fb-d223-4970-8b10-a731ebb9e6cdCited by top-tier papers9
- UCMCTrack: Multi-Object Tracking with Uniform Camera Motion CompensationKefu Yi, Kai Luo, Xiaolei Luo, Jiangui Huang et al.AAAI 2024 · 119 citations
- GLAMR: Global Occlusion-Aware Human Mesh Recovery with Dynamic CamerasYe Yuan, Umar Iqbal, Pavlo Molchanov, Kris Kitani et al.CVPR 2022 · 111 citations
- Quo Vadis: Is Trajectory Forecasting the Key Towards Long-Term Multi-Object Tracking?Patrick Dendorfer, Vladimir Yugay, Aljosa Osep, Laura Leal-TaixéNeurIPS 2022 · 77 citations
- Tracking People by Predicting 3D Appearance, Location and PoseJathushan Rajasegaran, Georgios Pavlakos, Angjoo Kanazawa, Jitendra MalikCVPR 2022 · 57 citations
- Towards Generalizable Multi-Object TrackingZheng Qin, Le Wang, Sanping Zhou, Panpan Fu et al.CVPR 2024 · 21 citations
Builds on4
- Tracking Without Bells and WhistlesPhilipp Bergmann, Tim Meinhardt, Laura Leal-TaixéICCV 2019 · 1,030 citations
- STGAT: Modeling Spatial-Temporal Interactions for Human Trajectory PredictionYingfan Huang, Huikun Bi, Zhaoxin Li, Tianlu Mao et al.ICCV 2019 · 615 citations
- Visualizing the Invisible: Occluded Vehicle Segmentation and RecoveryXiaosheng Yan, Yuanlong Yu, Feigege Wang, Wenxi Liu et al.ICCV 2019 · 46 citations
- PANDA: A Gigapixel-Level Human-Centric Video DatasetXueyang Wang, Xiya Zhang, Yinheng Zhu, Yuchen Guo et al.CVPR 2020
Related papers
- Learning to Track with Object PermanencePavel Tokmakov, Jie Li, Wolfram Burgard, Adrien GaidonICCV 2021 · 241 citations
- Joint Monocular 3D Vehicle Detection and TrackingHou-Ning Hu, Qi-Zhi Cai, Dequan Wang, Ji Lin et al.ICCV 2019 · 242 citations
- Object Permanence Emerges in a Random Walk along MemoryPavel Tokmakov, Allan Jabri, Jie Li, Adrien GaidonICML 2022 · 28 citations
- Depth From Camera Motion and Object DetectionBrent A. Griffin, Jason J. CorsoCVPR 2021
- PnPNet: End-to-End Perception and Prediction With Tracking in the LoopMing Liang, Bin Yang, Wenyuan Zeng, Yun Chen et al.CVPR 2020
