TransIFF: An Instance-Level Feature Fusion Framework for Vehicle-Infrastructure Cooperative 3D Detection with Transformers
Ziming Chen, Yifeng Shi, Jinrang Jia
Abstract
Cooperation between vehicles and infrastructure is vital to enhancing the safety of autonomous driving. Two significant and contradictory challenges now stand in the collaborative perception: fusion accuracy and communication bandwidth. Previous intermediate fusion methods that transmit features balance the accuracy and bandwidth compared with early fusion and late fusion, but usually have problems with feature alignment and domain gaps, and the bandwidth usage still falls short of the industrial application standard to our best knowledge. In this paper, we propose TransIFF, an instance-level feature fusion framework with transformers that can effectively reduce bandwidth usage. Furthermore, it can align the domain gaps between vehicle and infrastructure features, and improve the robustness of feature fusion, leading to a high cooperative perception accuracy. TransIFF is composed of three components: a vehicle-side network, an infrastructure-side network, and a vehicle-infrastructure fusion network. Initially, the vehicle-side and infrastructureside networks independently generate instance-level features. Subsequently, the infrastructure-side instance-level features are transmitted to the vehicles, significantly reducing the communication bandwidth usage. Finally, in the vehicle-infrastructure fusion network, Cross-Domain Adaptation (CDA) module is designed to align the feature domains, followed by Feature Magnet (FM) module which can adaptively fuse the instance features and achieve a robust feature fusion. TransIFF yields state-of-the-art performance on the widely used real-world vehicle-infrastructure cooperative benchmark DAIR-V2X, achieving 59.62% AP with only 2 12 bytes bandwidth consumption.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c15554dd-4bf3-4190-9e4e-d6d7b57f1ee6Cited by top-tier papers15
- TUMTraf V2X Cooperative Perception DatasetWalter Zimmer, Gerhard Arya Wardana, Suren Sritharan, Xingcheng Zhou et al.CVPR 2024 · 76 citations
- MonoUNI: A Unified Vehicle and Infrastructure-side Monocular 3D Object Detection Network with Sufficient Depth CluesJinrang Jia, Zhenjia Li, Yifeng ShiNeurIPS 2023 · 69 citations
- End-to-End Autonomous Driving Through V2X CooperationHaibao Yu, Wenxian Yang, Jiaru Zhong, Zhenwei Yang et al.AAAI 2025 · 56 citations
- Learning Cooperative Trajectory Representations for Motion ForecastingHongzhi Ruan, Haibao Yu, Wenxian Yang, Siqi Fan et al.NeurIPS 2024 · 36 citations
- mmCooper: A Multi-Agent Multi-Stage Communication-Efficient and Collaboration-Robust Cooperative Perception FrameworkBingyi Liu, Jian Teng, Hongfei Xue, Enshu Wang et al.ICCV 2025 · 14 citations
Builds on13
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar et al.NeurIPS 2021 · 9,661 citations
- ViViT: A Video Vision TransformerAnurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun et al.ICCV 2021 · 2,947 citations
- CrossViT: Cross-Attention Multi-Scale Vision Transformer for Image ClassificationChun-Fu (Richard) Chen, Quanfu Fan, Rameswar PandaICCV 2021 · 2,072 citations
- TransFusion: Robust LiDAR-Camera Fusion for 3D Object Detection with TransformersXuyang Bai, Zeyu Hu, Xinge Zhu, Qingqiu Huang et al.CVPR 2022 · 794 citations
Related papers
- DI-V2X: Learning Domain-Invariant Representation for Vehicle-Infrastructure Collaborative 3D Object DetectionXiang Li, Junbo Yin, Wei Li, Chengzhong Xu et al.AAAI 2024 · 56 citations
- Privacy-Preserving V2X Collaborative Perception Integrating Unknown CollaboratorsBin Lu, Xinyu Xiao, Changzhou Zhang, Yang Zhou et al.AAAI 2025
- BEVSync: Asynchronous Data Alignment for Camera-based Vehicle-Infrastructure Cooperative Perception Under Uncertain DelaysWentao Wang, Jiaqian Wang, Yuxin Deng, Guang TanAAAI 2025 · 2 citations
- DAIR-V2X: A Large-Scale Dataset for Vehicle-Infrastructure Cooperative 3D Object DetectionHaibao Yu, Yizhen Luo, Mao Shu, Yiyi Huo et al.CVPR 2022 · 475 citations
- DATA: Domain-And-Time Alignment for High-Quality Feature Fusion in Collaborative PerceptionChengchang Tian, Jianwei Ma, Yan Huang, Zhanye Chen et al.ICCV 2025 · 1 citation
