Robust Multi-Modality Multi-Object Tracking
Wenwei Zhang, Hui Zhou, Shuyang Sun, Zhe Wang, Jianping Shi, Chen Change Loy
Abstract
Multi-sensor perception is crucial to ensure the reliability and accuracy in autonomous driving system, while multi-object tracking (MOT) improves that by tracing sequential movement of dynamic objects. Most current approaches for multi-sensor multi-object tracking are either lack of reliability by tightly relying on a single input source (e.g., center camera), or not accurate enough by fusing the results from multiple sensors in post processing without fully exploiting the inherent information. In this study, we design a generic sensor-agnostic multi-modality MOT framework (mmMOT), where each modality (i.e., sensors) is capable of performing its role independently to preserve reliability, and could further improving its accuracy through a novel multi-modality fusion module. Our mmMOT can be trained in an end-to-end manner, enables joint optimization for the base feature extractor of each modality and an adjacency estimator for cross modality. Our mmMOT also makes the first attempt to encode deep representation of point cloud in data association process in MOT. We conduct extensive experiments to evaluate the effectiveness of the proposed framework on the challenging KITTI benchmark and report state-of-the-art performance. Code and models are available at https://github.com/ZwwWayne/mmMOT.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e8d2c23e-cadf-4789-a76f-06fac7a6b042Cited by top-tier papers19
- Deep Multimodal Fusion by Channel ExchangingYikai Wang, Wenbing Huang, Fuchun Sun, Tingyang Xu et al.NeurIPS 2020 · 321 citations
- EMP: edge-assisted multi-vehicle perceptionXumiao Zhang, Anlan Zhang, Jiachen Sun, Xiao Zhu et al.MobiCom 2021 · 137 citations
- Box-Aware Feature Enhancement for Single Object Tracking on Point CloudsChaoda Zheng, Xu Yan, Jiantao Gao, Weibing Zhao et al.ICCV 2021 · 116 citations
- 3D Siamese Voxel-to-BEV Tracker for Sparse Point CloudsLe Hui, Lingpeng Wang, Mingmei Cheng, Jin Xie et al.NeurIPS 2021 · 105 citations
- Single-Model and Any-Modality for Video Object TrackingZongwei Wu, Jilai Zheng, Xiangxuan Ren, Florin-Alexandru Vasluianu et al.CVPR 2024 · 78 citations
Related papers
- TrajectoryFormer: 3D Object Tracking Transformer with Predictive Trajectory HypothesesXuesong Chen, Shaoshuai Shi, Chao Zhang, Benjin Zhu et al.ICCV 2023 · 25 citations
- Unifying Voxel-based Representation with Transformer for 3D Object DetectionYanwei Li, Yilun Chen, Xiaojuan Qi, Zeming Li et al.NeurIPS 2022 · 401 citations
- End-to-End Pseudo-LiDAR for Image-Based 3D Object DetectionRui Qian, Divyansh Garg, Yan Wang, Yurong You et al.CVPR 2020
- A Novel Object Re-Track Framework for 3D Point CloudsTuo Feng, Licheng Jiao, Hao Zhu, Long SunACM MM 2020 · 22 citations
- CAT-Det: Contrastively Augmented Transformer for Multimodal 3D Object DetectionYanan Zhang, Jiaxin Chen, Di HuangCVPR 2022 · 138 citations
