I2V-GAN: Unpaired Infrared-to-Visible Video Translation
Shuang Li, Bingfeng Han, Zhenjie Yu, Chi Harold Liu, Kai Chen, Shuigen Wang
Abstract
Human vision is often adversely affected by complex environmental factors, especially in night vision scenarios. Thus, infrared cameras are often leveraged to help enhance the visual effects via detecting infrared radiation in the surrounding environment, but the infrared videos are undesirable due to the lack of detailed semantic information. In such a case, an effective video-to-video translation method from the infrared domain to the visible light counterpart is strongly needed by overcoming the intrinsic huge gap between infrared and visible fields. To address this challenging problem, we propose an infrared-to-visible (I2V) video translation method I2V-GAN to generate fine-grained and spatial-temporal consistent visible light videos by given unpaired infrared videos. Technically, our model capitalizes on three types of constraints: 1) adversarial constraint to generate synthetic frames that are similar to the real ones, 2) cyclic consistency with the introduced perceptual loss for effective content conversion as well as style preservation, and 3) similarity constraints across and within domains to enhance the content and motion consistency in both spatial and temporal spaces at a fine-grained level. Furthermore, the current public available infrared and visible light datasets are mainly used for object detection or tracking, and some are composed of discontinuous images which are not suitable for video tasks. Thus, we provide a new dataset for infrared-to-visible video translation, which is named IRVI. Specifically, it has 12 consecutive video clips of vehicle and monitoring scenes, and both infrared and visible light videos could be apart into 24352 frames. Comprehensive experiments on IRVI validate that I2V-GAN is superior to the compared state-of-the-art methods in the translation of infrared-to-visible videos with higher fluency and finer semantic details. Moreover, additional experimental results on the flower-to-flower dataset indicate I2V-GAN is also applicable to other video translation tasks. The code and IRVI dataset are available at https://github.com/BIT-DA/I2V-GAN.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 35ba12a5-ca2c-40ca-bce8-315c07879d1fCited by top-tier papers3
- 3M-TI: High-Quality Mobile Thermal Imaging via Calibration-free Multi-Camera Cross-Modal DiffusionMinchong Chen, Xiaoyun Yuan, Junzhe Wan, Jianing Zhang et al.CVPR 2026 · 2 citations
- SynthRGB-T: Language-Vision Guided Image Translation for Diversity SynthesisJiangang Ding, Yiquan Du, Pengxiang Li, Lili Pei et al.CVPR 2026 · 1 citation
- On the Difficulty of Unpaired Infrared-to-Visible Video Translation: Fine-Grained Content-Rich Patches TransferZhenjie Yu, Shuang Li, Yirui Shen, Chi Harold Liu et al.CVPR 2023
Builds on3
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Simultaneous Semantic Alignment Network for Heterogeneous Domain AdaptationShuang Li, Binhui Xie, Jiashu Wu, Ying Zhao et al.ACM MM 2020 · 42 citations
- StarGAN v2: Diverse Image Synthesis for Multiple DomainsYunjey Choi, Youngjung Uh, Jaejun Yoo, Jung-Woo HaCVPR 2020
Related papers
- ROMA: Cross-Domain Region Similarity Matching for Unpaired Nighttime Infrared to Daytime Visible Video TranslationZhenjie Yu, Kai Chen, Shuang Li, Bingfeng Han et al.ACM MM 2022 · 20 citations
- Style Transfer Meets Super-Resolution: Advancing Unpaired Infrared-to-Visible Image Translation with Detail EnhancementYirui Shen, Jingxuan Kang, Shuang Li, Zhenjie Yu et al.ACM MM 2023 · 11 citations
- NIR-assisted Video Enhancement via Unpaired 24-hour DataMuyao Niu, Zhihang Zhong, Yinqiang ZhengICCV 2023 · 4 citations
- Beyond Strict Pairing: Arbitrarily Paired Training for High-Performance Infrared and Visible Image FusionYanglin Deng, Tianyang Xu, Chunyang Cheng, Hui Li et al.CVPR 2026
- Dispel Darkness for Better Fusion: A Controllable Visual Enhancer Based on Cross-Modal Conditional Adversarial LearningHao Zhang, Linfeng Tang, Xinyu Xiang, Xuhui Zuo et al.CVPR 2024 · 21 citations
