DuGI-MAE: Improving Infrared Mask Autoencoders via Dual-Domain Guidance
Yinghui Xing, Xiaoting Su, Shizhou Zhang, Donghao Chu, Di Xu
Abstract
Infrared imaging plays a critical role in low-light and adverse weather conditions. However, due to the distinct characteristics of infrared images, existing foundation models such as Masked Autoencoder (MAE) trained on visible data perform suboptimal in infrared image interpretation tasks. To bridge this gap, an infrared foundation model known as Inf-MAE (Liu et al. 2024a) was developed and pre-trained on large-scale infrared datasets. Despite its effectiveness, Inf-MAE still faces several limitations, including the omission of informative tokens, insufficient modeling of global associations, and neglect of non-uniform noise. In this paper, we propose a Dual-domain Guided Infrared foundation model based on MAE (DuGI-MAE). First, we design a deterministic masking strategy based on token entropy, preserving only high-entropy tokens for reconstruction to enhance informativeness. Next, we introduce a Dual-Domain Guidance (DDG) module, which simultaneously captures global token relationships and adaptively filters non-uniform background noise commonly present in infrared imagery. To facilitate large-scale pretraining, we construct Inf-590K, a comprehensive infrared image dataset encompassing diverse scenes, various target types, and multiple spatial resolutions. Pretrained on Inf-590K, DuGI-MAE demonstrates strong generalization capabilities across various downstream tasks, including infrared object detection, semantic segmentation, and small target detection. Experimental results validate the superiority of the proposed method over both supervised and selfsupervised comparison methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 587559ee-1b1e-41ff-b786-7ace1feddf33Builds on10
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- BEiT: BERT Pre-Training of Image TransformersHangbo Bao, Li Dong, Songhao Piao, Furu WeiICLR 2022 · 3,632 citations
- Target-aware Dual Adversarial Learning and a Multi-scenario Multi-Modality Benchmark to Fuse Infrared and Visible for Object DetectionJinyuan Liu, Xin Fan, Zhanbo Huang, Guanyao Wu et al.CVPR 2022 · 929 citations
- FusionDN: A Unified Densely Connected Network for Image FusionHan Xu, Jiayi Ma, Zhuliang Le, Junjun Jiang et al.AAAI 2020 · 559 citations
- Masked Feature Prediction for Self-Supervised Visual Pre-TrainingChen Wei, Haoqi Fan, Saining Xie, Chao-Yuan Wu et al.CVPR 2022 · 524 citations
Related papers
- UniRGB-IR: A Unified Framework for Visible-Infrared Semantic Tasks via Adapter TuningMaoxun Yuan, Bo Cui, Tianyi Zhao, Jiayi Wang et al.ACM MM 2025 · 21 citations
- UNIP: Rethinking Pre-trained Attention Patterns for Infrared Semantic SegmentationTao Zhang, Jinyong Wen, Zhen Chen, Kun Ding et al.ICLR 2025
- RetroMAE-2: Duplex Masked Auto-Encoder For Pre-Training Retrieval-Oriented Language ModelsZheng Liu, Shitao Xiao, Yingxia Shao, Zhao CaoACL 2023 · 7 citations
- Empowering Visible-Infrared Person Re-Identification with Large Foundation ModelsZhangyi Hu, Bin Yang, Mang YeNeurIPS 2024 · 45 citations
- Temporal-Emerged Prompting for Segment Anything in Multiframe Infrared Small Target DetectionYinghui Xing, Donghao Chu, Shizhou Zhang, di xuICML 2026
