HM-ViT: Hetero-modal Vehicle-to-Vehicle Cooperative Perception with Vision Transformer
Hao Xiang, Runsheng Xu, Jiaqi Ma
Abstract
Vehicle-to-Vehicle technologies have enabled autonomous vehicles to share information to see through occlusions, greatly enhancing perception performance. Nevertheless, existing works all focused on homogeneous traffic where vehicles are equipped with the same type of sensors, which significantly hampers the scale of collaboration and benefit of cross-modality interactions. In this paper, we investigate the multi-agent hetero-modal cooperative perception problem where agents may have distinct sensor modalities. We present HM-ViT, the first unified multi-agent hetero-modal cooperative perception framework that can collaboratively predict 3D objects for highly dynamic Vehicle-to-Vehicle (V2V) collaborations with varying numbers and types of agents. To effectively fuse features from multi-view images and LiDAR point clouds, we design a novel heterogeneous 3D graph transformer to jointly reason inter-agent and intra-agent interactions. The extensive experiments on the V2V perception dataset OPV2V demonstrate that the HM-ViT outperforms SOTA cooperative perception methods for V2V hetero-modal cooperative perception. Our code will be released at https://github.com/XHwind/HM-ViT.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers26
- An Extensible Framework for Open Heterogeneous Collaborative PerceptionYifan Lu, Yue Hu, Yiqi Zhong, Dequan Wang et al.ICLR 2024 · 116 citations
- Asynchrony-Robust Collaborative Perception via Bird's Eye View FlowSizhe Wei, Yuxi Wei, Yue Hu, Yifan Lu et al.NeurIPS 2023 · 102 citations
- V2XPnP: Vehicle-to-Everything Spatio-Temporal Fusion for Multi-Agent Perception and PredictionZewei Zhou, Hao Xiang, Zhaoliang Zheng, Seth Z. Zhao et al.ICCV 2025 · 15 citations
- RCDN: Towards Robust Camera-Insensitivity Collaborative Perception via Dynamic Feature-based 3D Neural ModelingTianhang Wang, Fan Lu, Zehan Zheng, Zhijun Li et al.NeurIPS 2024 · 11 citations
- Pragmatic Heterogeneous Collaborative Perception via Generative Communication MechanismJunfei Zhou, Penglin Dai, Quanmin Wei, Bingyi Liu et al.NeurIPS 2025 · 11 citations
Builds on12
- TransFusion: Robust LiDAR-Camera Fusion for 3D Object Detection with TransformersXuyang Bai, Zeyu Hu, Xinge Zhu, Qingqiu Huang et al.CVPR 2022 · 794 citations
- BEVFusion: A Simple and Robust LiDAR-Camera Fusion FrameworkTingting Liang, Hongwei Xie, Kaicheng Yu, Zhongyu Xia et al.NeurIPS 2022 · 762 citations
- MAXIM: Multi-Axis MLP for Image ProcessingZhengzhong Tu, Hossein Talebi, Han Zhang, Feng Yang et al.CVPR 2022 · 550 citations
- M3D-RPN: Monocular 3D Region Proposal Network for Object DetectionGarrick Brazil, Xiaoming LiuICCV 2019 · 542 citations
- Where2comm: Communication-Efficient Collaborative Perception via Spatial Confidence MapsYue Hu, Shaoheng Fang, Zixing Lei, Yiqi Zhong et al.NeurIPS 2022 · 537 citations
Related papers
- GT-Space: Enhancing Heterogeneous Collaborative Perception with Ground Truth Feature SpaceWentao Wang, Haoran Xu, Guang TanICLR 2026 · 2 citations
- One is Plenty: A Polymorphic Feature Interpreter for Immutable Heterogeneous Collaborative PerceptionYuchen Xia, Quan Yuan, Guiyang Luo, Xiaoyuan Fu et al.CVPR 2025
- V2V4Real: A Real-World Large-Scale Dataset for Vehicle-to-Vehicle Cooperative PerceptionRunsheng Xu, Xin Xia, Jinlong Li, Hanzhao Li et al.CVPR 2023
- V2U4Real: A Real-world Large-scale Dataset for Vehicle-to-UAV Cooperative PerceptionWeijia Li, Haoen Xiang, Tianxu Wang, Shuaibing Wu et al.CVPR 2026 · 4 citations
- Shared Cross-Modal Trajectory Prediction for Autonomous DrivingChiho Choi, Joon Hee Choi, Jiachen Li, Srikanth MallaCVPR 2021
