V2XPnP: Vehicle-to-Everything Spatio-Temporal Fusion for Multi-Agent Perception and Prediction
Zewei Zhou, Hao Xiang, Zhaoliang Zheng, Seth Z. Zhao, Mingyue Lei, Yun Zhang, Tianhui Cai, Xinyi Liu, Johnson Liu, Maheswari Bajji, Xin Xia, Zhiyu Huang
摘要
Vehicle-to-everything (V2X) technologies offer a promising paradigm to mitigate the limitations of constrained observability in single-vehicle systems. Prior work primarily focuses on single-frame cooperative perception, which fuses agents' information across different spatial locations but ignores temporal cues and temporal tasks (e.g., temporal perception and prediction). In this paper, we focus on the spatio-temporal fusion in V2X scenarios and design one-step and multi-step communication strategies (when to transmit) as well as examine their integration with three fusion strategies - early, late, and intermediate (what to transmit), providing comprehensive benchmarks with 11 fusion models (how to fuse). Furthermore, we propose V2XPnP, a novel intermediate fusion framework within one-step communication for end-to-end perception and prediction. Our framework employs a unified Transformer-based architecture to effectively model complex spatio-temporal relationships across multiple agents, frames, and high-definition maps. Moreover, we introduce the V2XPnP Sequential Dataset that supports all V2X collaboration modes and addresses the limitations of existing real-world datasets, which are restricted to single-frame or single-mode cooperation. Extensive experiments demonstrate that our framework outperforms state-of-the-art methods in both perception and prediction tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- AutoVLA: A Vision-Language-Action Model for End-to-End Autonomous Driving with Adaptive Reasoning and Reinforcement Fine-TuningZewei Zhou, Tianhui Cai, Seth Z. Zhao, Yun Zhang 等NeurIPS 2025 · 被引用 310 次
- Cooptrack: Exploring End-to-End Learning for Efficient Cooperative Sequential PerceptionJiaru Zhong, Jiahao Wang, Jiahui Xu, Xiaofan Li 等ICCV 2025 · 被引用 5 次
- TurboTrain: Towards Efficient and Balanced Multi-Task Learning for Multi-Agent Perception and PredictionZewei Zhou, Seth Z. Zhao, Tianhui Cai, Zhiyu Huang 等ICCV 2025 · 被引用 2 次
- EnerGS: Energy-Based Gaussian Splatting under Partial Geometric PriorsRui Song, Tianhui Cai, Markus Gross, Yun Zhang 等ICML 2026 · 被引用 2 次
- ZeRCP: Towards Communication-Efficient Collaborative Perception and Future Scene Prediction via Request-Free Spatial FilteringYijie Chen, Yuzhe Ji, Haotian Wang, Xiaoyun Qiu 等AAAI 2026
它引用的顶会 Paper27
- Large Scale Interactive Motion Forecasting for Autonomous Driving : The Waymo Open Motion DatasetScott Ettinger, Shuyang Cheng, Benjamin Caine, Chenxi Liu 等ICCV 2021 · 被引用 817 次
- VAD: Vectorized Scene Representation for Efficient Autonomous DrivingBo Jiang, Shaoyu Chen, Qing Xu, Bencheng Liao 等ICCV 2023 · 被引用 602 次
- DenseTNT: End-to-end Trajectory Prediction from Dense Goal SetsJunru Gu, Chen Sun, Hang ZhaoICCV 2021 · 被引用 563 次
- Where2comm: Communication-Efficient Collaborative Perception via Spatial Confidence MapsYue Hu, Shaoheng Fang, Zixing Lei, Yiqi Zhong 等NeurIPS 2022 · 被引用 537 次
- DAIR-V2X: A Large-Scale Dataset for Vehicle-Infrastructure Cooperative 3D Object DetectionHaibao Yu, Yizhen Luo, Mao Shu, Yiyi Huo 等CVPR 2022 · 被引用 475 次
相关 Paper
- End-to-End Autonomous Driving Through V2X CooperationHaibao Yu, Wenxian Yang, Jiaru Zhong, Zhenwei Yang 等AAAI 2025 · 被引用 56 次
- Long-SCOPE: Fully Sparse Long-Range Cooperative 3D PerceptionJiahao Wang, Zikun Xu, Yuner Zhang, Zhongwei Jiang 等CVPR 2026 · 被引用 3 次
- TraF-Align: Trajectory-aware Feature Alignment for Asynchronous Multi-agent PerceptionZhiying Song, Lei Yang, Fuxi Wen, Jun LiCVPR 2025
- When Autonomous Vehicle Meets V2X Cooperative Perception: How Far Are We?An Guo, Shuoxiao Zhang, Enyi Tang, Xinyu Gao 等ASE 2025 · 被引用 1 次
- Learning Cooperative Trajectory Representations for Motion ForecastingHongzhi Ruan, Haibao Yu, Wenxian Yang, Siqi Fan 等NeurIPS 2024 · 被引用 36 次
