RetinaTrack: Online Single Stage Joint Detection and Tracking
Zhichao Lu, Vivek Rathod, Ronny Votel, Jonathan Huang
Abstract
Traditionally multi-object tracking and object detection are performed using separate systems with most prior works focusing exclusively on one of these aspects over the other. Tracking systems clearly benefit from having access to accurate detections, however and there is ample evidence in literature that detectors can benefit from tracking which, for example, can help to smooth predictions over time. In this paper we focus on the tracking-by-detection paradigm for autonomous driving where both tasks are mission critical. We propose a conceptually simple and efficient joint model of detection and tracking, called RetinaTrack , which modifies the popular single stage RetinaNet approach such that it is amenable to instance-level embedding training. We show, via evaluations on the Waymo Open Dataset, that we outperform a recent state of the art tracking algorithm while requiring significantly less computation. We believe that our simple yet effective approach can serve as a strong baseline for future work in this area.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1bc4435f-aa53-4add-ba27-ff17a2af2129Cited by top-tier papers12
- Exploring Simple 3D Multi-Object Tracking for Autonomous DrivingChenxu Luo, Xiaodong Yang, Alan L. YuilleICCV 2021 · 122 citations
- TrackFlow: Multi-Object Tracking with Normalizing FlowsGianluca Mancusi, Aniello Panariello, Angelo Porrello, Matteo Fabbri et al.ICCV 2023 · 23 citations
- Joint 3D Object Detection and Tracking Using Spatio-Temporal Representation of Camera Image and LiDAR Point CloudsJunho Koh, Jaekyum Kim, Jin Hyeok Yoo, Yecheol Kim et al.AAAI 2022 · 19 citations
- RT-MOT: Confidence-Aware Real-Time Scheduling Framework for Multi-Object Tracking TasksDonghwa Kang, Seunghoon Lee, Hoon Sung Chwa, Seung-Hwan Bae et al.RTSS 2022 · 10 citations
- Cannot See the Forest for the Trees: Aggregating Multiple Viewpoints to Better Classify Objects in VideosSukjun Hwang, Miran Heo, Seoung Wug Oh, Seon Joo KimCVPR 2022 · 4 citations
Builds on5
- FCOS: Fully Convolutional One-Stage Object DetectionZhi Tian, Chunhua Shen, Hao Chen, Tong HeICCV 2019 · 6,042 citations
- Tracking Without Bells and WhistlesPhilipp Bergmann, Tim Meinhardt, Laura Leal-TaixéICCV 2019 · 1,030 citations
- Sequence Level Semantics Aggregation for Video Object DetectionHaiping Wu, Yuntao Chen, Naiyan Wang, Zhaoxiang ZhangICCV 2019 · 236 citations
- Object Guided External Memory Network for Video Object DetectionHanming Deng, Yang Hua, Tao Song, Zongpu Zhang et al.ICCV 2019 · 109 citations
- Leveraging Long-Range Temporal Relationships Between Proposals for Video Object DetectionMykhailo Shvets, Wei Liu, Alexander C. BergICCV 2019 · 91 citations
Related papers
- Tracking by Instance Detection: A Meta-Learning ApproachGuangting Wang, Chong Luo, Xiaoyan Sun, Zhiwei Xiong et al.CVPR 2020
- SRNet: Spatial Relation Network for Efficient Single-stage Instance Segmentation in VideosXiaowen Ying, Xin Li, Mooi Choo ChuahACM MM 2021 · 4 citations
- Track To Detect and Segment: An Online Multi-Object TrackerJialian Wu, Jiale Cao, Liangchen Song, Yu Wang et al.CVPR 2021
- PnPNet: End-to-End Perception and Prediction With Tracking in the LoopMing Liang, Bin Yang, Wenyuan Zeng, Yun Chen et al.CVPR 2020
- TubeTK: Adopting Tubes to Track Multi-Object in a One-Step Training ModelBo Pang, Yizhuo Li, Yifan Zhang, Muchen Li et al.CVPR 2020
