ThermalGen: Style-Disentangled Flow-Based Generative Models for RGB-to-Thermal Image Translation
Jiuhong Xiao, Roshan Nayak, Ning Zhang, Daniel Tortei, Giuseppe Loianno
Abstract
Paired RGB-thermal data is crucial for visual-thermal sensor fusion and cross-modality tasks, including important applications such as multi-modal image alignment and retrieval. However, the scarcity of synchronized and calibrated RGB-thermal image pairs presents a major obstacle to progress in these areas. To overcome this challenge, RGB-to-Thermal (RGB-T) image translation has emerged as a promising solution, enabling the synthesis of thermal images from abundant RGB datasets for training purposes. In this study, we propose ThermalGen, an adaptive flow-based generative model for RGB-T image translation, incorporating an RGB image conditioning architecture and a style-disentangled mechanism. To support large-scale training, we curated eight public satellite-aerial, aerial, and ground RGB-T paired datasets, and introduced three new large-scale satellite-aerial RGB-T datasets--DJI-day, Bosonplus-day, and Bosonplus-night--captured across diverse times, sensor types, and geographic regions. Extensive evaluations across multiple RGB-T benchmarks demonstrate that ThermalGen achieves comparable or superior translation performance compared to existing GAN-based and diffusion-based methods. To our knowledge, ThermalGen is the first RGB-T image translation model capable of synthesizing thermal images that reflect significant variations in viewpoints, sensor characteristics, and environmental conditions. Project page: http://xjh19971.github.io/ThermalGen
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c8681c16-1847-481a-819b-fbfbcb3f17d4Cited by top-tier papers5
- TherA: Thermal-Aware Visual-Language Prompting for Controllable RGB-to-Thermal Infrared TranslationDong-Guw Lee, Tai Hyoung Rhee, Hyunsoo Jang, Young-Sik Shin et al.CVPR 2026 · 4 citations
- AnyMod-LLVE: Low-Light Video Enhancement with Modality-Agnostic InferenceHangfeng Liang, Yutao Hu, Yanhan Hu, Xiaohan Wu et al.ICML 2026
- MMVIP: A Visible-infrared Paired Dataset for Multi-weather Marine VisionYunpeng Yin, Lihan Wang, Zhaoshen He, Xinqiang He et al.CVPR 2026
- Thermal-Det: Language-Guided Cross-Modal Distillation for Open-Vocabulary Thermal Object DetectionYasiru Ranasinghe, Elim Schenck, Florence Yellin, Shuowen Hu et al.CVPR 2026
- No Calibration, No Depth, No Problem: Cross-Sensor View Synthesis with 3D ConsistencyCho-Ying Wu, Zixun Huang, Xinyu Huang, Liu RenCVPR 2026
Builds on18
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- Palette: Image-to-Image Diffusion ModelsChitwan Saharia, William Chan, Huiwen Chang, Chris A. Lee et al.SIGGRAPH 2022 · 1,638 citations
Related papers
- M-SpecGene: Generalized Foundation Model for RGBT Multispectral VisionKailai Zhou, Fuqiang Yang, Shixian Wang, Bihan Wen et al.ICCV 2025 · 4 citations
- SemanticRT: A Large-Scale Dataset and Method for Robust Semantic Segmentation in Multispectral ImagesWei Ji, Jingjing Li, Cheng Bian, Zhicheng Zhang et al.ACM MM 2023 · 22 citations
- Visible-Thermal UAV Tracking: A Large-Scale Benchmark and New BaselinePengyu Zhang, Jie Zhao, Dong Wang, Huchuan Lu et al.CVPR 2022 · 225 citations
- Beyond Strict Pairing: Arbitrarily Paired Training for High-Performance Infrared and Visible Image FusionYanglin Deng, Tianyang Xu, Chunyang Cheng, Hui Li et al.CVPR 2026
- ABMDRNet: Adaptive-Weighted Bi-Directional Modality Difference Reduction Network for RGB-T Semantic SegmentationQiang Zhang, Shenlu Zhao, Yongjiang Luo, Dingwen Zhang et al.CVPR 2021
