SparseCoop: Cooperative Perception with Kinematic-Grounded Queries
Jiahao Wang, Zhongwei Jiang, Wenchao Sun, Jiaru Zhong, Haibao Yu, Yuner Zhang, Chenyang Lu, Chuang Zhang, Lei He, Shaobing Xu, Jianqiang Wang
Abstract
Cooperative perception is critical for autonomous driving, overcoming the inherent limitations of a single vehicle, such as occlusions and constrained fields-of-view. However, current approaches sharing dense Bird's-Eye-View (BEV) features are constrained by quadratically-scaling communication costs and the lack of flexibility and interpretability for precise alignment across asynchronous or disparate viewpoints. While emerging sparse query-based methods offer an alternative, they often suffer from inadequate geometric representations, suboptimal fusion strategies, and training instability. In this paper, we propose SparseCoop, a fully sparse cooperative perception framework for 3D detection and tracking that completely discards intermediate BEV representations. Our framework features a trio of innovations: a kinematic grounded instance query that uses an explicit state vector with 3D geometry and velocity for precise spatio-temporal alignment; a coarse-to-fine aggregation module that effectively integrates information from both matched and unmatched instances; and a cooperative instance denoising task that provides stable, abundant supervision to accelerate and stabilize training. Experiments on V2X-Seq and Griffin datasets show SparseCoop achieves state-of-the-art performance. Notably, it delivers this performance with superior computational efficiency and a highly competitive transmission cost, while showing remarkable robustness to real-world challenges like communication latency.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7151b2a5-9898-4baa-9b8d-d6e8678c2874Builds on13
- DN-DETR: Accelerate DETR Training by Introducing Query DeNoisingFeng Li, Hao Zhang, Shilong Liu, Jian Guo et al.CVPR 2022 · 879 citations
- Where2comm: Communication-Efficient Collaborative Perception via Spatial Confidence MapsYue Hu, Shaoheng Fang, Zixing Lei, Yiqi Zhong et al.NeurIPS 2022 · 537 citations
- PETRv2: A Unified Framework for 3D Perception from Multi-Camera ImagesYingfei Liu, Junjie Yan, Fan Jia, Shuailin Li et al.ICCV 2023 · 513 citations
- Learning Distilled Collaboration Graph for Multi-Agent PerceptionYiming Li, Shunli Ren, Pengxiang Wu, Siheng Chen et al.NeurIPS 2021 · 464 citations
- End-to-End Autonomous Driving Through V2X CooperationHaibao Yu, Wenxian Yang, Jiaru Zhong, Zhenwei Yang et al.AAAI 2025 · 56 citations
Related papers
- SparseAlign: a Fully Sparse Framework for Cooperative Object DetectionYunshuang Yuan, Yan Xia, Daniel Cremers, Monika SesterCVPR 2025
- Long-SCOPE: Fully Sparse Long-Range Cooperative 3D PerceptionJiahao Wang, Zikun Xu, Yuner Zhang, Zhongwei Jiang et al.CVPR 2026 · 3 citations
- Cooptrack: Exploring End-to-End Learning for Efficient Cooperative Sequential PerceptionJiaru Zhong, Jiahao Wang, Jiahui Xu, Xiaofan Li et al.ICCV 2025 · 5 citations
- SQS: Enhancing Sparse Perception Models via Query-based Splatting in Autonomous DrivingHaiming Zhang, Yiyao Zhu, Wending Zhou, Xu Yan et al.NeurIPS 2025 · 5 citations
- Communication-Efficient Multi-Vehicle Collaborative Semantic Segmentation via Sparse 3D Gaussian SharingTianyu Hong, Xiaobo Zhou, Wenkai Hu, Qi Xie et al.ICCV 2025 · 2 citations
