Generative Modeling for Multi-task Visual Learning
Zhipeng Bao, Martial Hebert, Yu-Xiong Wang
Abstract
Generative modeling has recently shown great promise in computer vision, but it has mostly focused on synthesizing visually realistic images. In this paper, motivated by multi-task learning of shareable feature representations, we consider a novel problem of learning a shared generative model that is useful across various visual perception tasks. Correspondingly, we propose a general multi-task oriented generative modeling (MGM) framework, by coupling a discriminative multi-task network with a generative network. While it is challenging to synthesize both RGB images and pixel-level annotations in multi-task scenarios, our framework enables us to use synthesized images paired with only weak annotations (i.e., image-level scene labels) to facilitate multiple visual tasks. Experimental evaluation on challenging multi-task benchmarks, including NYUv2 and Taskonomy, demonstrates that our MGM framework improves the performance of all the tasks by large margins, consistently outperforming state-of-the-art multi-task approaches.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers7
- StyleGAN knows Normal, Depth, Albedo, and MoreAnand Bhattad, Daniel McKee, Derek Hoiem, David A. ForsythNeurIPS 2023 · 61 citations
- Diffusion Models for Multi-Task Generative ModelingChangyou Chen, Han Ding, Bunyamin Sisman, Yi Xu et al.ICLR 2024 · 11 citations
- ReferevErything: Towards Segmenting Everything we can Speak of in VideosAnurag Bagchi, Zhipeng Bao, Yu-Xiong Wang, Pavel Tokmakov et al.ICCV 2025 · 11 citations
- Multi-task View Synthesis with Neural Radiance FieldsShuhong Zheng, Zhipeng Bao, Martial Hebert, Yu-Xiong WangICCV 2023 · 7 citations
- Swiss Army Knife: Synergizing Biases in Knowledge from Vision Foundation Models for Multi-Task LearningYuxiang Lu, Shengcao Cao, Yu-Xiong WangICLR 2025
Builds on8
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Which Tasks Should Be Learned Together in Multi-task Learning?Trevor Standley, Amir Zamir, Dawn Chen, Leonidas J. Guibas et al.ICML 2020 · 651 citations
- 3D Scene Graph: A Structure for Unified Semantics, 3D Space, and CameraIro Armeni, Zhi-Yang He, Amir Zamir, JunYoung Gwak et al.ICCV 2019 · 474 citations
- AdaShare: Learning What To Share For Efficient Deep Multi-Task LearningXimeng Sun, Rameswar Panda, Rogério Feris, Kate SaenkoNeurIPS 2020 · 337 citations
- Self-Ensembling With GAN-Based Data Augmentation for Domain Adaptation in Semantic SegmentationJaehoon Choi, Taekyung Kim, Changick KimICCV 2019 · 264 citations
Related papers
- Bowtie Networks: Generative Modeling for Joint Few-Shot Recognition and Novel-View SynthesisZhipeng Bao, Yu-Xiong Wang, Martial HebertICLR 2021 · 6 citations
- TaskPrompter: Spatial-Channel Multi-Task Prompting for Dense Scene UnderstandingHanrong Ye, Dan XuICLR 2023
- TaskExpert: Dynamically Assembling Multi-Task Representations with Memorial Mixture-of-ExpertsHanrong Ye, Dan XuICCV 2023 · 60 citations
- MAESTRO: Task-Relevant Optimization Via Adaptive Feature Enhancement and Suppression for Multi-Task 3D PerceptionChangwon Kang, Jisong Kim, Hongjae Shin, Junseo Park et al.ICCV 2025
- Reconciling Visual Perception and Generation in Diffusion ModelsLiulei Li, Yi Yang, Wenguan WangICLR 2026
