SynthRGB-T: Language-Vision Guided Image Translation for Diversity Synthesis
Jiangang Ding, Yiquan Du, Pengxiang Li, Lili Pei, Yuanlin Zhao, Wei Li
Abstract
Bridging the modality gap between infrared and visible imagery is critical for cross-modal understanding and for enriching multimodal benchmarks. However, existing approaches remain confined to one-to-one mappings and are typically evaluated on unidirectional or closed-set scenarios. To address this challenge, we present SynthRGB-T, a unified framework for diverse and bidirectional image translation. Specifically, we formulate image translation as a vision-language guided denoising diffusion process, enabling flexible conditioning and open-world generalization. To enhance semantic alignment, a Visual Grounding Pipeline (VGP) is introduced to exploit the world knowledge of foundation models for fine-grained translation guidance. During the diffusion process, we propose to adopt a decoupling injection strategy to alleviate interference among multiple guidance. In addition, a Dual Conditional Cross-Attention (DCCA) module is designed to facilitate collaborative representation learning in latent space. SynthRGB-T is simple and versatile-capable of synthesizing diverse, high-fidelity data that substantially extends multimodal resources within the community. Comprehensive evaluations on multiple real-world benchmarks confirm that SynthRGB-T delivers superior performance and enhanced visual diversity over existing approaches.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bfe4367c-0d6c-426d-ae7d-d801002c643fBuilds on16
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Directly Denoising Diffusion ModelsDan Zhang, Jingjing Wang, Feng LuoICML 2024 · 11,724 citations
- Target-aware Dual Adversarial Learning and a Multi-scenario Multi-Modality Benchmark to Fuse Infrared and Visible for Object DetectionJinyuan Liu, Xin Fan, Zhanbo Huang, Guanyao Wu et al.CVPR 2022 · 929 citations
Related papers
- Towards General Modality Translation with Contrastive and Predictive Latent Diffusion BridgeNimrod Berman, Omkar Joglekar, Eitan Kosman, Dotan Di Castro et al.NeurIPS 2025 · 5 citations
- Real-World Image Variation by Aligning Diffusion Inversion ChainYuechen Zhang, Jinbo Xing, Eric Lo, Jiaya JiaNeurIPS 2023 · 56 citations
- DiffTV: Identity-Preserved Thermal-to-Visible Face Translation via Feature Alignment and Dual-Stage ConditionsJingyu Lin, Guiqin Zhao, Jing Xu, Guoli Wang et al.ACM MM 2024 · 9 citations
- UniAlignment: Semantic Alignment for Unified Image Generation, Understanding, Manipulation and PerceptionXinyang Song, Libin Wang, Weining Wang, Shaozhen Liu et al.AAAI 2026
- TherA: Thermal-Aware Visual-Language Prompting for Controllable RGB-to-Thermal Infrared TranslationDong-Guw Lee, Tai Hyoung Rhee, Hyunsoo Jang, Young-Sik Shin et al.CVPR 2026 · 4 citations
