Detecting Invisible People
Tarasha Khurana, Achal Dave, Deva Ramanan
摘要
Monocular object detection and tracking have improved drastically in recent years, but rely on a key assumption: that objects are visible to the camera. Many offline tracking approaches reason about occluded objects post-hoc, by linking together tracklets after the object re-appears, making use of reidentification (ReID). However, online tracking in embodied robotic agents (such as a self-driving vehicle) fundamentally requires object permanence, which is the ability to reason about occluded objects before they re-appear. In this work, we re-purpose tracking benchmarks and propose new metrics for the task of detecting invisible objects, focusing on the illustrative case of people. We demonstrate that current detection and tracking systems perform dramatically worse on this task. We introduce two key innovations to recover much of this performance drop. We treat occluded object detection in temporal sequences as a short-term forecasting challenge, bringing to bear tools from dynamic sequence prediction. Second, we build dynamic models that explicitly reason in 3D from monocular videos without calibration, using observations produced by monocular depth estimators. To our knowledge, ours is the first work to demonstrate the effectiveness of monocular depth estimation for the task of tracking and detecting occluded objects. Our approach strongly improves by 11.4% over the baseline in ablations and by 5.0% over the state-of-the-art in F1 score.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- UCMCTrack: Multi-Object Tracking with Uniform Camera Motion CompensationKefu Yi, Kai Luo, Xiaolei Luo, Jiangui Huang 等AAAI 2024 · 被引用 119 次
- GLAMR: Global Occlusion-Aware Human Mesh Recovery with Dynamic CamerasYe Yuan, Umar Iqbal, Pavlo Molchanov, Kris Kitani 等CVPR 2022 · 被引用 111 次
- Quo Vadis: Is Trajectory Forecasting the Key Towards Long-Term Multi-Object Tracking?Patrick Dendorfer, Vladimir Yugay, Aljosa Osep, Laura Leal-TaixéNeurIPS 2022 · 被引用 77 次
- Tracking People by Predicting 3D Appearance, Location and PoseJathushan Rajasegaran, Georgios Pavlakos, Angjoo Kanazawa, Jitendra MalikCVPR 2022 · 被引用 57 次
- Towards Generalizable Multi-Object TrackingZheng Qin, Le Wang, Sanping Zhou, Panpan Fu 等CVPR 2024 · 被引用 21 次
它引用的顶会 Paper4
- Tracking Without Bells and WhistlesPhilipp Bergmann, Tim Meinhardt, Laura Leal-TaixéICCV 2019 · 被引用 1,030 次
- STGAT: Modeling Spatial-Temporal Interactions for Human Trajectory PredictionYingfan Huang, Huikun Bi, Zhaoxin Li, Tianlu Mao 等ICCV 2019 · 被引用 615 次
- Visualizing the Invisible: Occluded Vehicle Segmentation and RecoveryXiaosheng Yan, Yuanlong Yu, Feigege Wang, Wenxi Liu 等ICCV 2019 · 被引用 46 次
- PANDA: A Gigapixel-Level Human-Centric Video DatasetXueyang Wang, Xiya Zhang, Yinheng Zhu, Yuchen Guo 等CVPR 2020
相关 Paper
- Learning to Track with Object PermanencePavel Tokmakov, Jie Li, Wolfram Burgard, Adrien GaidonICCV 2021 · 被引用 241 次
- Joint Monocular 3D Vehicle Detection and TrackingHou-Ning Hu, Qi-Zhi Cai, Dequan Wang, Ji Lin 等ICCV 2019 · 被引用 242 次
- Object Permanence Emerges in a Random Walk along MemoryPavel Tokmakov, Allan Jabri, Jie Li, Adrien GaidonICML 2022 · 被引用 28 次
- Depth From Camera Motion and Object DetectionBrent A. Griffin, Jason J. CorsoCVPR 2021
- PnPNet: End-to-End Perception and Prediction With Tracking in the LoopMing Liang, Bin Yang, Wenyuan Zeng, Yun Chen 等CVPR 2020
