From Discriminative to Generative: A Diffusion-Based Paradigm for Multi-Agent Collaborative Perception
Kexin Gong, Puyi Yao, Guiyang Luo, Quan Yuan, Tiange Fu, Hui Zhang, Jinglin Li
摘要
Collaborative perception leveraging intermediate feature fusion has emerged as a leading paradigm to significantly enhance the environmental perception capabilities of autonomous driving systems. However, existing methods typically rely on discriminative supervision guided by downstream tasks. This paradigm compels models to learn minimal, task-specific representations, which conflicts with the goal of cooperative perception to capture comprehensive information, thereby limiting generalization. To address this issue, we propose DiGS-CP, a novel two-stage generative supervised collaborative perception framework. Specifically, we introduce a diffusion-based generative task that conditions on fused object-level features to generate representations of object-level point clouds. The proposed generative supervision provides fine-grained, task-agnostic signals that encourages the fusion module to learn comprehensive representations beyond task-specific requirements. By preserving and integrating complementary information from collaborative agents, our approach overcomes the limitations of task-specific learning and enhances the generalizability of the learned features. Furthermore, our two-stage architecture requires agents to transmit only object-level features, significantly reducing communication overhead. Extensive experiments on three benchmark datasets demonstrate that DiGS-CP achieves state-of-the-art performance in 3D object detection, while maintaining low bandwidth requirements and exhibiting excellent generalization ability.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper23
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Directly Denoising Diffusion ModelsDan Zhang, Jingjing Wang, Feng LuoICML 2024 · 被引用 11,724 次
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 被引用 6,759 次
- Where2comm: Communication-Efficient Collaborative Perception via Spatial Confidence MapsYue Hu, Shaoheng Fang, Zixing Lei, Yiqi Zhong 等NeurIPS 2022 · 被引用 537 次
相关 Paper
- Generative Map Priors for Collaborative BEV Semantic SegmentationJiahui Fu, Yue Gong, Luting Wang, Shifeng Zhang 等CVPR 2025
- Core: Cooperative Reconstruction for Multi-Agent PerceptionBinglu Wang, Lei Zhang, Zhaozhong Wang, Yongqiang Zhao 等ICCV 2023 · 被引用 73 次
- Collaborative Semantic Occupancy Prediction with Hybrid Feature Fusion in Connected Automated VehiclesRui Song, Chenwei Liang, Hu Cao, Zhiran Yan 等CVPR 2024
- Point Cloud Reconstruction Is Insufficient to Learn 3D RepresentationsWeichen Xu, Jian Cao, Tianhao Fu, Ruilong Ren 等ACM MM 2024 · 被引用 1 次
- SparseCoop: Cooperative Perception with Kinematic-Grounded QueriesJiahao Wang, Zhongwei Jiang, Wenchao Sun, Jiaru Zhong 等AAAI 2026 · 被引用 1 次
