DiffTV: Identity-Preserved Thermal-to-Visible Face Translation via Feature Alignment and Dual-Stage Conditions
Jingyu Lin, Guiqin Zhao, Jing Xu, Guoli Wang, Zejin Wang, Antitza Dantcheva, Lan Du, Cunjian Chen
Abstract
The thermal-to-visible (T2V) face translation task is essential for enabling face verification in low-light or dark conditions by converting thermal infrared faces into their visible counterparts. However, this task faces two primary challenges. First, the inherent differences between the modalities hinder the effective use of thermal information to guide RGB face reconstruction. Second, translated RGB faces often lack the identity details of the corresponding visible faces, such as skin color. To tackle these challenges, we introduce DiffTV, the first Latent Diffusion Model (LDM) specifically designed for T2V facial image translation with a focus on preserving identity. Our approach proposes a novel heterogeneous feature alignment strategy that bridges the modal gap and extracts both coarse-and fine-grained identity features consistent with visible images. Furthermore, a dual-stage condition injection strategy introduces control information to guide identity-preserved translation. Experimental results demonstrate the superior performance of DiffTV, particularly in scenarios where maintaining identity integrity is critical.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 77ecdc7d-c6fa-4a07-9a94-b46c42c66bb3Cited by top-tier papers7
- SceneDecorator: Towards Scene-Oriented Story Generation with Scene Planning and Scene ConsistencyQuanjian Song, Donghao Zhou, Jingyu Lin, Fei Shen et al.NeurIPS 2025 · 9 citations
- HiFi-Inpaint: Towards High-Fidelity Reference-Based Inpainting for Generating Detail-Preserving Human-Product ImagesYi Chen Liu, Donghao Zhou, Jie Wang, Xin Gao et al.CVPR 2026 · 5 citations
- Pan-LUT: Efficient Pan-sharpening via Learnable Look-Up TablesZhongnan Cai, Yingying Wang, Hui Zheng, Panwang Pan et al.NeurIPS 2025 · 2 citations
- Semi-Supervised Semantic Segmentation via Derivative Label PropagationYuanbin Fu, Xiaojie GuoAAAI 2026 · 1 citation
- Adaptive-Learngene: Continual Expansion and Task-Aware Selection of Learngenes for Dynamic EnvironmentsShuxia Lin, Qiufeng Wang, Chang Liu, Xu Yang et al.AAAI 2026
Builds on15
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
Related papers
- TherA: Thermal-Aware Visual-Language Prompting for Controllable RGB-to-Thermal Infrared TranslationDong-Guw Lee, Tai Hyoung Rhee, Hyunsoo Jang, Young-Sik Shin et al.CVPR 2026 · 4 citations
- SynthRGB-T: Language-Vision Guided Image Translation for Diversity SynthesisJiangang Ding, Yiquan Du, Pengxiang Li, Lili Pei et al.CVPR 2026 · 1 citation
- UV-IDM: Identity-Conditioned Latent Diffusion Model for Face UV-Texture GenerationHong Li, Yutang Feng, Song Xue, Xuhui Liu et al.CVPR 2024
- Data Generation Scheme for Thermal Modality with Edge-Guided Adversarial Conditional Diffusion ModelGuoqing Zhu, Honghu Pan, Qiang Wang, Chao Tian et al.ACM MM 2024 · 7 citations
- Cross-Modal Semantic Decoupling and Transfer for Text-to-Visible-Infrared Person Re-IdentificationZiang Zhang, Bin Yang, Mang YeICML 2026
