End-to-end 3D Tracking with Decoupled Queries
Yanwei Li, Zhiding Yu, Jonah Philion, Anima Anandkumar, Sanja Fidler, Jiaya Jia, Jose Alvarez
Abstract
In this work, we present an end-to-end framework for camera-based 3D multi-object tracking, called DQTrack. To avoid heuristic design in detection-based trackers, recent query-based approaches deal with identity-agnostic detection and identity-aware tracking in a single embedding. However, it brings inferior performance because of the inherent representation conflict. To address this issue, we decouple the single embedding into separated queries, i.e., object query and track query. Unlike previous detection-based and query-based methods, the decoupled-query paradigm utilizes task-specific queries and still maintains the compact pipeline without complex post-processing. Moreover, the learnable association and temporal update are designed to provide differentiable trajectory association and frameby-frame query update, respectively. The proposed DQ-Track is demonstrated to achieve consistent gains in various benchmarks, outperforming previous tracking-by-detection and learning-based methods on the nuScenes dataset. 1 * Work done during an internship at NVIDIA.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b61ba901-aebf-428e-a8a6-bba2362c064eCited by top-tier papers11
- Language Prompt for Autonomous DrivingDongming Wu, Wencheng Han, Yingfei Liu, Tiancai Wang et al.AAAI 2025 · 150 citations
- Visual Point Cloud Forecasting Enables Scalable Autonomous DrivingZetong Yang, Li Chen, Yanan Sun, Hongyang LiCVPR 2024 · 40 citations
- RefAV: Towards Planning-Centric Scenario MiningCainan Davidson, Deva Ramanan, Neehar PeriCVPR 2026 · 17 citations
- Cooptrack: Exploring End-to-End Learning for Efficient Cooperative Sequential PerceptionJiaru Zhong, Jiahao Wang, Jiahui Xu, Xiaofan Li et al.ICCV 2025 · 5 citations
- IC-Mapper: Instance-Centric Spatio-Temporal Modeling for Online Vectorized Map ConstructionJiangtong Zhu, Zhao Yang, Yinan Shi, Jianwu Fang et al.ACM MM 2024 · 3 citations
Builds on16
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- TrackFormer: Multi-Object Tracking with TransformersTim Meinhardt, Alexander Kirillov, Laura Leal-Taixé, Christoph FeichtenhoferCVPR 2022 · 927 citations
- TransFusion: Robust LiDAR-Camera Fusion for 3D Object Detection with TransformersXuyang Bai, Zeyu Hu, Xinge Zhu, Qingqiu Huang et al.CVPR 2022 · 794 citations
Related papers
- ADA-Track: End-to-End Multi-Camera 3D Multi-Object Tracking with Alternating Detection and AssociationShuxiao Ding, Lukas Schneider, Marius Cordts, Juergen GallCVPR 2024
- S2-Track: A Simple yet Strong Approach for End-to-End 3D Multi-Object TrackingTao Tang, Lijun Zhou, Pengkun Hao, Zihang He et al.ICML 2025
- Standing Between Past and Future: Spatio-Temporal Modeling for Multi-Camera 3D Multi-Object TrackingZiqi Pang, Jie Li, Pavel Tokmakov, Dian Chen et al.CVPR 2023
- SynCL: A Synergistic Training Strategy with Instance-Aware Contrastive Learning for End-to-End Multi-Camera 3D TrackingShubo Lin, Yutong Kou, Zirui Wu, Shaoru Wang et al.NeurIPS 2025 · 2 citations
- LA-MOTR: End-to-End Multi-Object Tracking by Learnable AssociationPeng Wang, Yongcai Wang, Hualong Cao, Wang Chen et al.ICCV 2025 · 9 citations
