UniRGB-IR: A Unified Framework for Visible-Infrared Semantic Tasks via Adapter Tuning
Maoxun Yuan, Bo Cui, Tianyi Zhao, Jiayi Wang, Shan Fu, Xue Yang, Xingxing Wei
Abstract
Semantic analysis on visible (RGB) and infrared (IR) images has gained significant attention due to their enhanced accuracy and robustness under challenging conditions including low-illumination and adverse weather. However, due to the lack of pre-trained foundation models on the large-scale infrared image datasets, existing methods prefer to design task-specific frameworks and directly fine-tune them with pre-trained foundation models on their RGB-IR semantic relevance datasets, which results in poor scalability and limited generalization. To address these limitations, we propose UniRGB-IR, a scalable and efficient framework for RGB-IR semantic tasks that introduces a novel adapter mechanism to effectively incorporate rich multi-modal features into pre-trained RGB-based foundation models. Our framework comprises three key components: a vision transformer (ViT) foundation model, a Multi-modal Feature Pool (MFP) module, and a Supplementary Feature Injector (SFI) module. The MFP and SFI modules cooperate with each other as an adpater to effectively complement the ViT features with the contextual multi-scale features. During training process, we freeze the entire foundation model to inherit prior knowledge and only optimize the MFP and SFI modules. Furthermore, to verify the effectiveness of our framework, we utilize the ViT-Base as the pre-trained foundation model to perform extensive experiments. Experimental results on various RGB-IR semantic tasks demonstrate that our method can achieve state-of-the-art performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f3516ee9-d705-44e9-80ca-813ad6db6cc7Cited by top-tier papers6
- Rethinking Multi-Modal Object Detection From the Perspective of Mono-Modality Feature LearningTianyi Zhao, Boyang Liu, Yanglei Gao, Yiming Sun et al.ICCV 2025 · 15 citations
- Seeing Through the Noise: Improving Infrared Small Target Detection and Segmentation from Noise Suppression PerspectiveMaoxun Yuan, Duanni Meng, Ziteng Xi, Tianyi Zhao et al.CVPR 2026 · 11 citations
- M-SpecGene: Generalized Foundation Model for RGBT Multispectral VisionKailai Zhou, Fuqiang Yang, Shixian Wang, Bihan Wen et al.ICCV 2025 · 4 citations
- LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMsHaoran Lou, Chunxiao Fan, Ziyan Liu, Yuexin Wu et al.ICCV 2025 · 1 citation
- MFH-NAS:A Hybrid Neural Architecture Search Framework for Multimodal Fusion Object DetectionQuanWei Gao, Shuqi Zhao, Ruyu Wang, Shuyin Zhang et al.ICML 2026
Builds on20
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without ConvolutionsWenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan et al.ICCV 2021 · 4,909 citations
- CCNet: Criss-Cross Attention for Semantic SegmentationZilong Huang, Xinggang Wang, Lichao Huang, Chang Huang et al.ICCV 2019 · 2,972 citations
Related papers
- UNIP: Rethinking Pre-trained Attention Patterns for Infrared Semantic SegmentationTao Zhang, Jinyong Wen, Zhen Chen, Kun Ding et al.ICLR 2025
- Bi-directional Adapter for Multimodal TrackingBing Cao, Junliang Guo, Pengfei Zhu, Qinghua HuAAAI 2024 · 153 citations
- PanAdapter: Two-Stage Fine-Tuning with Spatial-Spectral Priors Injecting for PansharpeningRuoCheng Wu, Zien Zhang, Shangqi Deng, Yule Duan et al.AAAI 2025 · 7 citations
- Simplifying Cross-modal Interaction via Modality-Shared Features for RGBT TrackingLiqiu Chen, Yuqing Huang, Hengyu Li, Zikun Zhou et al.ACM MM 2024 · 2 citations
- ABMDRNet: Adaptive-Weighted Bi-Directional Modality Difference Reduction Network for RGB-T Semantic SegmentationQiang Zhang, Shenlu Zhao, Yongjiang Luo, Dingwen Zhang et al.CVPR 2021
