DeMT: Deformable Mixer Transformer for Multi-Task Learning of Dense Prediction
Yangyang Xu, Yibo Yang, Lefei Zhang
摘要
Convolution neural networks (CNNs) and Transformers have their own advantages and both have been widely used for dense prediction in multi-task learning (MTL). Most of the current studies on MTL solely rely on CNN or Transformer. In this work, we present a novel MTL model by combining both merits of deformable CNN and query-based Transformer for multi-task learning of dense prediction. Our method, named DeMT, is based on a simple and effective encoder-decoder architecture (i.e., deformable mixer encoder and task-aware transformer decoder). First, the deformable mixer encoder contains two types of operators: the channel-aware mixing operator leveraged to allow communication among different channels (i.e., efficient channel location mixing), and the spatial-aware deformable operator with deformable convolution applied to efficiently sample more informative spatial locations (i.e., deformed features). Second, the task-aware transformer decoder consists of the task interaction block and task query block. The former is applied to capture task interaction features via self-attention. The latter leverages the deformed features and task-interacted features to generate the corresponding task-specific feature through a query-based Transformer for corresponding task predictions. Extensive experiments on two dense image prediction datasets, NYUD-v2 and PASCAL-Context, demonstrate that our model uses fewer GFLOPs and significantly outperforms current Transformer-and CNN-based competitive models on a variety of metrics. The code are available at https://github.com/yangyangxu0/DeMT .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper17
- MmAP: Multi-Modal Alignment Prompt for Cross-Domain Multi-Task LearningYi Xin, Junlong Du, Qiang Wang, Ke Yan 等AAAI 2024 · 被引用 102 次
- VMT-Adapter: Parameter-Efficient Transfer Learning for Multi-Task Dense Scene UnderstandingYi Xin, Junlong Du, Qiang Wang, Zhiwen Lin 等AAAI 2024 · 被引用 94 次
- Multi-Task Learning with Knowledge Distillation for Dense PredictionYangyang Xu, Yibo Yang, Lefei ZhangICCV 2023 · 被引用 18 次
- Video Task Decathlon: Unifying Image and Video Tasks in Autonomous DrivingThomas E. Huang, Yifan Liu, Luc Van Gool, Fisher YuICCV 2023 · 被引用 10 次
- StableMTL: Repurposing Latent Diffusion Models for Multi-Task Learning from Partially Annotated Synthetic DatasetsAnh-Quan Cao, Ivan Lopes, Raoul de CharetteCVPR 2026 · 被引用 2 次
它引用的顶会 Paper11
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li 等ICLR 2021 · 被引用 7,353 次
- Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without ConvolutionsWenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan 等ICCV 2021 · 被引用 4,909 次
- Vision Transformers for Dense PredictionRené Ranftl, Alexey Bochkovskiy, Vladlen KoltunICCV 2021 · 被引用 2,647 次
- Vision Transformer with Deformable AttentionZhuofan Xia, Xuran Pan, Shiji Song, Li Erran Li 等CVPR 2022 · 被引用 835 次
- Learning Affinity from Attention: End-to-End Weakly-Supervised Semantic Segmentation with TransformersLixiang Ru, Yibing Zhan, Baosheng Yu, Bo DuCVPR 2022 · 被引用 257 次
相关 Paper
- Multi-Task Dense Prediction via Mixture of Low-Rank ExpertsYuqi Yang, Peng-Tao Jiang, Qibin Hou, Hao Zhang 等CVPR 2024 · 被引用 32 次
- Multi-Task Dense Predictions via Unleashing the Power of DiffusionYuqi Yang, Peng-Tao Jiang, Qibin Hou, Hao Zhang 等ICLR 2025
- TaskPrompter: Spatial-Channel Multi-Task Prompting for Dense Scene UnderstandingHanrong Ye, Dan XuICLR 2023
- Task-Conditional Adapter for Multi-Task Dense PredictionFengze Jiang, Shuling Wang, Xiaojin GongACM MM 2024 · 被引用 6 次
- Active Token MixerGuoqiang Wei, Zhizheng Zhang, Cuiling Lan, Yan Lu 等AAAI 2023 · 被引用 25 次
