DeMT: Deformable Mixer Transformer for Multi-Task Learning of Dense Prediction
Yangyang Xu, Yibo Yang, Lefei Zhang
Abstract
Convolution neural networks (CNNs) and Transformers have their own advantages and both have been widely used for dense prediction in multi-task learning (MTL). Most of the current studies on MTL solely rely on CNN or Transformer. In this work, we present a novel MTL model by combining both merits of deformable CNN and query-based Transformer for multi-task learning of dense prediction. Our method, named DeMT, is based on a simple and effective encoder-decoder architecture (i.e., deformable mixer encoder and task-aware transformer decoder). First, the deformable mixer encoder contains two types of operators: the channel-aware mixing operator leveraged to allow communication among different channels (i.e., efficient channel location mixing), and the spatial-aware deformable operator with deformable convolution applied to efficiently sample more informative spatial locations (i.e., deformed features). Second, the task-aware transformer decoder consists of the task interaction block and task query block. The former is applied to capture task interaction features via self-attention. The latter leverages the deformed features and task-interacted features to generate the corresponding task-specific feature through a query-based Transformer for corresponding task predictions. Extensive experiments on two dense image prediction datasets, NYUD-v2 and PASCAL-Context, demonstrate that our model uses fewer GFLOPs and significantly outperforms current Transformer-and CNN-based competitive models on a variety of metrics. The code are available at https://github.com/yangyangxu0/DeMT .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9f853b7f-ffa2-4b61-b9ea-5c5a4a36119cCited by top-tier papers17
- MmAP: Multi-Modal Alignment Prompt for Cross-Domain Multi-Task LearningYi Xin, Junlong Du, Qiang Wang, Ke Yan et al.AAAI 2024 · 102 citations
- VMT-Adapter: Parameter-Efficient Transfer Learning for Multi-Task Dense Scene UnderstandingYi Xin, Junlong Du, Qiang Wang, Zhiwen Lin et al.AAAI 2024 · 94 citations
- Multi-Task Learning with Knowledge Distillation for Dense PredictionYangyang Xu, Yibo Yang, Lefei ZhangICCV 2023 · 18 citations
- Video Task Decathlon: Unifying Image and Video Tasks in Autonomous DrivingThomas E. Huang, Yifan Liu, Luc Van Gool, Fisher YuICCV 2023 · 10 citations
- StableMTL: Repurposing Latent Diffusion Models for Multi-Task Learning from Partially Annotated Synthetic DatasetsAnh-Quan Cao, Ivan Lopes, Raoul de CharetteCVPR 2026 · 2 citations
Builds on11
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without ConvolutionsWenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan et al.ICCV 2021 · 4,909 citations
- Vision Transformers for Dense PredictionRené Ranftl, Alexey Bochkovskiy, Vladlen KoltunICCV 2021 · 2,647 citations
- Vision Transformer with Deformable AttentionZhuofan Xia, Xuran Pan, Shiji Song, Li Erran Li et al.CVPR 2022 · 835 citations
- Learning Affinity from Attention: End-to-End Weakly-Supervised Semantic Segmentation with TransformersLixiang Ru, Yibing Zhan, Baosheng Yu, Bo DuCVPR 2022 · 257 citations
Related papers
- Multi-Task Dense Prediction via Mixture of Low-Rank ExpertsYuqi Yang, Peng-Tao Jiang, Qibin Hou, Hao Zhang et al.CVPR 2024 · 32 citations
- Multi-Task Dense Predictions via Unleashing the Power of DiffusionYuqi Yang, Peng-Tao Jiang, Qibin Hou, Hao Zhang et al.ICLR 2025
- TaskPrompter: Spatial-Channel Multi-Task Prompting for Dense Scene UnderstandingHanrong Ye, Dan XuICLR 2023
- Task-Conditional Adapter for Multi-Task Dense PredictionFengze Jiang, Shuling Wang, Xiaojin GongACM MM 2024 · 6 citations
- Active Token MixerGuoqiang Wei, Zhizheng Zhang, Cuiling Lan, Yan Lu et al.AAAI 2023 · 25 citations
