ROMA: Cross-Domain Region Similarity Matching for Unpaired Nighttime Infrared to Daytime Visible Video Translation
Zhenjie Yu, Kai Chen, Shuang Li, Bingfeng Han, Chi Harold Liu, Shuigen Wang
Abstract
Infrared cameras are often utilized to enhance the night vision since the visible light cameras exhibit inferior efficacy without sufficient illumination. However, infrared data possesses inadequate color contrast and representation ability attributed to its intrinsic heat-related imaging principle, which hinders its application. Although, the domain gaps between unpaired nighttime infrared and daytime visible videos are even huger than paired ones that captured at the same time, establishing an effective translation mapping will greatly contribute to various fields. In this case, the structural knowledge within nighttime infrared videos and semantic information contained in the translated daytime visible pairs could be utilized simultaneously. To this end, we propose a tailored framework ROMA that couples with our introduced cRoss-domain regiOn siMilarity mAtching technique for bridging the huge gaps. To be specific, ROMA could efficiently translate the unpaired nighttime infrared videos into fine-grained daytime visible ones, meanwhile maintain the spatiotemporal consistency via matching the cross-domain region similarity. Furthermore, we design a multiscale region-wise discriminator to distinguish the details from synthesized visible results and real references. Moreover, we provide a new and challenging dataset encouraging further research for unpaired nighttime infrared and daytime visible video translation, named InfraredCity, which is times larger than the recently released infrared-related dataset IRVI. Codes and datasets are available https://github.com/BIT-DA/ROMA here.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- Thermal-Physics Guided Infrared Image Super-Resolution with Dynamic High-Frequency AmplificationMingxuan Zhou, Yirui Shen, Shuang Li, Jing Geng et al.AAAI 2026
- UNIP: Rethinking Pre-trained Attention Patterns for Infrared Semantic SegmentationTao Zhang, Jinyong Wen, Zhen Chen, Kun Ding et al.ICLR 2025
- On the Difficulty of Unpaired Infrared-to-Visible Video Translation: Fine-Grained Content-Rich Patches TransferZhenjie Yu, Shuang Li, Yirui Shen, Chi Harold Liu et al.CVPR 2023
- Thermal Diffusion Matters: Infrared Spatial-Temporal Video Super-Resolution through Heat Conduction PriorsMingxuan Zhou, Shuang Li, Yutang Zhang, Jing Geng et al.CVPR 2026
Builds on2
Related papers
- I2V-GAN: Unpaired Infrared-to-Visible Video TranslationShuang Li, Bingfeng Han, Zhenjie Yu, Chi Harold Liu et al.ACM MM 2021 · 58 citations
- Style Transfer Meets Super-Resolution: Advancing Unpaired Infrared-to-Visible Image Translation with Detail EnhancementYirui Shen, Jingxuan Kang, Shuang Li, Zhenjie Yu et al.ACM MM 2023 · 11 citations
- NIR-assisted Video Enhancement via Unpaired 24-hour DataMuyao Niu, Zhihang Zhong, Yinqiang ZhengICCV 2023 · 4 citations
- NightReID: A Large-Scale Nighttime Person Re-Identification BenchmarkYuxuan Zhao, Weijian Ruan, He Li, Mang YeAAAI 2025 · 5 citations
- CMDA: Cross-Modality Domain Adaptation for Nighttime Semantic SegmentationRuihao Xia, Chaoqiang Zhao, Meng Zheng, Ziyan Wu et al.ICCV 2023 · 54 citations
