IRGPT: Understanding Real-World Infrared Image with Bi-Cross-Modal Curriculum on Large-Scale Benchmark
Zhe Cao, Jin Zhang, Ruiheng Zhang
摘要
Real-world infrared imagery presents unique challenges for vision-language models due to the scarcity of aligned text data and domain-specific characteristics. Although existing methods have advanced the field, their reliance on synthetic infrared images generated through style transfer from visible images, which limits their ability to capture the unique characteristics of the infrared modality. To address this, we propose IRGPT, the first multi-modal large language model for real-world infrared images, built upon a large-scale InfraRed-Text Dataset (IR-TD) comprising over 260K authentic image-text pairs. The proposed IR-TD dataset contains real infrared images paired with meticulously handcrafted texts, where the initial drafts originated from two complementary processes: (1) LLM-generated descriptions of visible images, and (2) rule-based descriptions of annotations. Furthermore, we introduce a bi-cross-modal curriculum transfer learning strategy that systematically transfers knowledge from visible to infrared domains by considering the difficulty scores of both infrared-visible and infrared-text. Evaluated on a benchmark of 9 tasks (e.g., recognition, grounding), IRGPT achieves state-of-the-art performance even compared with larger-scale models.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper30
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
相关 Paper
- Empowering Visible-Infrared Person Re-Identification with Large Foundation ModelsZhangyi Hu, Bin Yang, Mang YeNeurIPS 2024 · 被引用 45 次
- IMUGPT 2.0: Language-Based Cross Modality Transfer for Sensor-Based Human Activity RecognitionZikang Leng, Amitrajit Bhattacharjee, Hrudhai Rajasekhar, Lizhe Zhang 等UbiComp 2024 · 被引用 59 次
- Thermal-Det: Language-Guided Cross-Modal Distillation for Open-Vocabulary Thermal Object DetectionYasiru Ranasinghe, Elim Schenck, Florence Yellin, Shuowen Hu 等CVPR 2026
- Cross-Modal Semantic Decoupling and Transfer for Text-to-Visible-Infrared Person Re-IdentificationZiang Zhang, Bin Yang, Mang YeICML 2026
- SynthRGB-T: Language-Vision Guided Image Translation for Diversity SynthesisJiangang Ding, Yiquan Du, Pengxiang Li, Lili Pei 等CVPR 2026 · 被引用 1 次
