Video Task Decathlon: Unifying Image and Video Tasks in Autonomous Driving
Thomas E. Huang, Yifan Liu, Luc Van Gool, Fisher Yu
Abstract
Performing multiple heterogeneous visual tasks in dynamic scenes is a hallmark of human perception capability. Despite remarkable progress in image and video recognition via representation learning, current research still focuses on designing specialized networks for singular, homogeneous, or simple combination of tasks. We instead explore the construction of a unified model for major image and video recognition tasks in autonomous driving with diverse input and output structures. To enable such an investigation, we design a new challenge, Video Task Decathlon (VTD), which includes ten representative image and video tasks spanning classification, segmentation, localization, and association of objects and pixels. On VTD, we develop our unified network, VTDNet, that uses a single structure and a single set of weights for all ten tasks. VTDNet groups similar tasks and employs task interaction stages to exchange information within and between task groups. Given the impracticality of labeling all tasks on all frames and the performance degradation associated with joint training of many tasks, we design a Curriculum training, Pseudo-labeling, and Fine-tuning (CPF) scheme to successfully train VTDNet on all tasks and mitigate performance loss. Armed with CPF, VTDNet significantly outperforms its single-task counterparts on most tasks with only 20% overall computations. VTD is a promising new direction for exploring the unification of perception tasks in autonomous driving.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b857b9c6-8799-4ae2-b2da-d24dd4d98082Cited by top-tier papers3
- FusionSAM: Visual Multi-Modal Learning with Segment Anything ModelDaixun Li, Weiying Xie, Mingxiang Cao, Yunke Wang et al.KDD 2025 · 2 citations
- A Backpack Full of Skills: Egocentric Video Understanding with Diverse Task PerspectivesSimone Alberto Peirone, Francesca Pistilli, Antonio Alliegro, Giuseppe AvertaCVPR 2024 · 1 citation
- Learning with Preserving for Continual Multitask LearningHanchen David Wang, Siwoo Bae, Zirong Chen, Meiyi MaAAAI 2026
Builds on21
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer et al.CVPR 2022 · 6,782 citations
- Gradient Surgery for Multi-Task LearningTianhe Yu, Saurabh Kumar, Abhishek Gupta, Sergey Levine et al.NeurIPS 2020 · 2,261 citations
- Per-Pixel Classification is Not All You Need for Semantic SegmentationBowen Cheng, Alexander G. Schwing, Alexander KirillovNeurIPS 2021 · 2,196 citations
Related papers
- BDD100K: A Diverse Driving Dataset for Heterogeneous Multitask LearningFisher Yu, Haofeng Chen, Xin Wang, Wenqi Xian et al.CVPR 2020
- TarViS: A Unified Approach for Target-Based Video SegmentationAli Athar, Alexander Hermans, Jonathon Luiten, Deva Ramanan et al.CVPR 2023
- Planning-oriented Autonomous DrivingYihan Hu, Jiazhi Yang, Li Chen, Keyu Li et al.CVPR 2023
- Many Task Learning With Task RoutingGjorgji Strezoski, Nanne van Noord, Marcel WorringICCV 2019 · 112 citations
- Incremental Multi-Domain Learning with Network Latent Tensor FactorizationAdrian Bulat, Jean Kossaifi, Georgios Tzimiropoulos, Maja PanticAAAI 2020 · 33 citations
