TurboTrain: Towards Efficient and Balanced Multi-Task Learning for Multi-Agent Perception and Prediction
Zewei Zhou, Seth Z. Zhao, Tianhui Cai, Zhiyu Huang, Bolei Zhou, Jiaqi Ma
Abstract
End-to-end training of multi-agent systems offers significant advantages in improving multi-task performance. However, training such models remains challenging and requires extensive manual design and monitoring. In this work, we introduce TurboTrain, a novel and efficient training framework for multi-agent perception and prediction. TurboTrain comprises two key components: a multi-agent spatiotemporal pretraining scheme based on masked reconstruction learning and a balanced multi-task learning strategy based on gradient conflict suppression. By streamlining the training process, our framework eliminates the need for manually designing and tuning complex multi-stage training pipelines, substantially reducing training time and improving performance. We evaluate TurboTrain on a real-world cooperative driving dataset, V2XPnP-Seq, and demonstrate that it further improves the performance of state-of-the-art multi-agent perception and prediction models. Our results highlight that pretraining effectively captures spatiotemporal multi-agent features and significantly benefits downstream tasks. Moreover, the proposed balanced multi-task learning strategy enhances detection and prediction.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7eb4a574-981a-4c7a-8227-6a72546f0b40Cited by top-tier papers1
Ask how each one uses itBuilds on43
- Gradient Surgery for Multi-Task LearningTianhe Yu, Saurabh Kumar, Abhishek Gupta, Sergey Levine et al.NeurIPS 2020 · 2,261 citations
- Masked Autoencoders As Spatiotemporal LearnersChristoph Feichtenhofer, Haoqi Fan, Yanghao Li, Kaiming HeNeurIPS 2022 · 690 citations
- Conflict-Averse Gradient Descent for Multi-task learningBo Liu, Xingchao Liu, Xiaojie Jin, Peter Stone et al.NeurIPS 2021 · 686 citations
- VAD: Vectorized Scene Representation for Efficient Autonomous DrivingBo Jiang, Shaoyu Chen, Qing Xu, Bencheng Liao et al.ICCV 2023 · 602 citations
- Where2comm: Communication-Efficient Collaborative Perception via Spatial Confidence MapsYue Hu, Shaoheng Fang, Zixing Lei, Yiqi Zhong et al.NeurIPS 2022 · 537 citations
Related papers
- Core: Cooperative Reconstruction for Multi-Agent PerceptionBinglu Wang, Lei Zhang, Zhaozhong Wang, Yongqiang Zhao et al.ICCV 2023 · 73 citations
- Navigation-Guided Sparse Scene Representation for End-to-End Autonomous DrivingPeidong Li, Dixiao CuiICLR 2025
- GPT-ST: Generative Pre-Training of Spatio-Temporal Graph Neural NetworksZhonghang Li, Lianghao Xia, Yong Xu, Chao HuangNeurIPS 2023 · 55 citations
- Visual Exemplar Driven Task-Prompting for Unified Perception in Autonomous DrivingXiwen Liang, Minzhe Niu, Jianhua Han, Hang Xu et al.CVPR 2023
- GT-Space: Enhancing Heterogeneous Collaborative Perception with Ground Truth Feature SpaceWentao Wang, Haoran Xu, Guang TanICLR 2026 · 2 citations
