TransIFF: An Instance-Level Feature Fusion Framework for Vehicle-Infrastructure Cooperative 3D Detection with Transformers
Ziming Chen, Yifeng Shi, Jinrang Jia
摘要
Cooperation between vehicles and infrastructure is vital to enhancing the safety of autonomous driving. Two significant and contradictory challenges now stand in the collaborative perception: fusion accuracy and communication bandwidth. Previous intermediate fusion methods that transmit features balance the accuracy and bandwidth compared with early fusion and late fusion, but usually have problems with feature alignment and domain gaps, and the bandwidth usage still falls short of the industrial application standard to our best knowledge. In this paper, we propose TransIFF, an instance-level feature fusion framework with transformers that can effectively reduce bandwidth usage. Furthermore, it can align the domain gaps between vehicle and infrastructure features, and improve the robustness of feature fusion, leading to a high cooperative perception accuracy. TransIFF is composed of three components: a vehicle-side network, an infrastructure-side network, and a vehicle-infrastructure fusion network. Initially, the vehicle-side and infrastructureside networks independently generate instance-level features. Subsequently, the infrastructure-side instance-level features are transmitted to the vehicles, significantly reducing the communication bandwidth usage. Finally, in the vehicle-infrastructure fusion network, Cross-Domain Adaptation (CDA) module is designed to align the feature domains, followed by Feature Magnet (FM) module which can adaptively fuse the instance features and achieve a robust feature fusion. TransIFF yields state-of-the-art performance on the widely used real-world vehicle-infrastructure cooperative benchmark DAIR-V2X, achieving 59.62% AP with only 2 12 bytes bandwidth consumption.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper15
- TUMTraf V2X Cooperative Perception DatasetWalter Zimmer, Gerhard Arya Wardana, Suren Sritharan, Xingcheng Zhou 等CVPR 2024 · 被引用 76 次
- MonoUNI: A Unified Vehicle and Infrastructure-side Monocular 3D Object Detection Network with Sufficient Depth CluesJinrang Jia, Zhenjia Li, Yifeng ShiNeurIPS 2023 · 被引用 69 次
- End-to-End Autonomous Driving Through V2X CooperationHaibao Yu, Wenxian Yang, Jiaru Zhong, Zhenwei Yang 等AAAI 2025 · 被引用 56 次
- Learning Cooperative Trajectory Representations for Motion ForecastingHongzhi Ruan, Haibao Yu, Wenxian Yang, Siqi Fan 等NeurIPS 2024 · 被引用 36 次
- mmCooper: A Multi-Agent Multi-Stage Communication-Efficient and Collaboration-Robust Cooperative Perception FrameworkBingyi Liu, Jian Teng, Hongfei Xue, Enshu Wang 等ICCV 2025 · 被引用 14 次
它引用的顶会 Paper13
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar 等NeurIPS 2021 · 被引用 9,661 次
- ViViT: A Video Vision TransformerAnurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun 等ICCV 2021 · 被引用 2,947 次
- CrossViT: Cross-Attention Multi-Scale Vision Transformer for Image ClassificationChun-Fu (Richard) Chen, Quanfu Fan, Rameswar PandaICCV 2021 · 被引用 2,072 次
- TransFusion: Robust LiDAR-Camera Fusion for 3D Object Detection with TransformersXuyang Bai, Zeyu Hu, Xinge Zhu, Qingqiu Huang 等CVPR 2022 · 被引用 794 次
相关 Paper
- DI-V2X: Learning Domain-Invariant Representation for Vehicle-Infrastructure Collaborative 3D Object DetectionXiang Li, Junbo Yin, Wei Li, Chengzhong Xu 等AAAI 2024 · 被引用 56 次
- Privacy-Preserving V2X Collaborative Perception Integrating Unknown CollaboratorsBin Lu, Xinyu Xiao, Changzhou Zhang, Yang Zhou 等AAAI 2025
- BEVSync: Asynchronous Data Alignment for Camera-based Vehicle-Infrastructure Cooperative Perception Under Uncertain DelaysWentao Wang, Jiaqian Wang, Yuxin Deng, Guang TanAAAI 2025 · 被引用 2 次
- DAIR-V2X: A Large-Scale Dataset for Vehicle-Infrastructure Cooperative 3D Object DetectionHaibao Yu, Yizhen Luo, Mao Shu, Yiyi Huo 等CVPR 2022 · 被引用 475 次
- DATA: Domain-And-Time Alignment for High-Quality Feature Fusion in Collaborative PerceptionChengchang Tian, Jianwei Ma, Yan Huang, Zhanye Chen 等ICCV 2025 · 被引用 1 次
