Enhancing Mamba Decoder with Bidirectional Interaction in Multi-Task Dense Prediction
Mang Cao, Sanping Zhou, Yizhe Li, Ye Deng, Wenli Huang, Le Wang
Abstract
Sufficient cross-task interaction is crucial for success in multi-task dense prediction. However, sufficient interaction often results in high computational complexity, forcing existing methods to face the trade-off between interaction completeness and computational efficiency. To address this limitation, this work proposes a Bidirectional Interaction Mamba (BIM), which incorporates novel scanning mechanisms to adapt the Mamba modeling approach for multitask dense prediction. On the one hand, we introduce a novel Bidirectional Interaction Scan (BI-Scan) mechanism, which constructs task-specific representations as bidirectional sequences during interaction. By integrating taskfirst and position-first scanning modes within a unified linear complexity architecture, BI-Scan efficiently preserves critical cross-task information. On the other hand, we employ a Multi-Scale Scan (MS-Scan) mechanism to achieve multi-granularity scene modeling. This design not only meets the diverse granularity requirements of various tasks but also enhances nuanced cross-task feature interactions. Extensive experiments on two challenging benchmarks, i.e., NYUD-V2 and PASCAL-Context, show the superiority of our BIM vs its state-of-the-art competitors. Codes are available online: https://github.com/mmm-cc/BIM_ for_MTL.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext da6cd46d-b43d-4d8c-97ac-e5ea42444a37Builds on19
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 3,482 citations
- VMamba: Visual State Space ModelYue Liu, Yunjie Tian, Yuzhong Zhao, Hongtian Yu et al.NeurIPS 2024 · 3,199 citations
- Gradient Surgery for Multi-Task LearningTianhe Yu, Saurabh Kumar, Abhishek Gupta, Sergey Levine et al.NeurIPS 2020 · 2,261 citations
- Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space ModelLianghui Zhu, Bencheng Liao, Qian Zhang, Xinlong Wang et al.ICML 2024 · 1,725 citations
Related papers
- HOIMamba: Efficient Mamba-based Disentangled Progressive Learning for HOI DetectionYongchao Xu, Jiawei Liu, Sen Tao, Qiang Zhang et al.AAAI 2025 · 1 citation
- Going Beyond Multi-Task Dense Prediction with Synergy Embedding ModelsHuimin Huang, Yawen Huang, Lanfen Lin, Ruofeng Tong et al.CVPR 2024
- Contrastive Multi-Task Dense PredictionSiwei Yang, Hanrong Ye, Dan XuAAAI 2023 · 13 citations
- 2D-CrossScan Mamba: Enhancing State Space Models with Spatially Consistent Multi-Path 2D Information PropagationLonglong Yu, Wenxi Li, Yaoqi Sun, Hang Xu et al.AAAI 2026
- SF-Mamba: Rethinking State Space Model for VisionMasakazu Yoshimura, Teruaki Hayashi, Yuki Hoshino, Wei-Yao Wang et al.ICML 2026
