TherA: Thermal-Aware Visual-Language Prompting for Controllable RGB-to-Thermal Infrared Translation
Dong-Guw Lee, Tai Hyoung Rhee, Hyunsoo Jang, Young-Sik Shin, Ukcheol Shin, Ayoung Kim
摘要
Despite the inherent advantages of thermal infrared (TIR) imaging, large-scale data collection and annotation remain a major bottleneck for TIR-based perception. A practical alternative is to synthesize pseudo-TIR data via image translation; however, most RGB-to-TIR approaches heavily rely on RGB-centric priors that overlook thermal physics, yielding implausible heat distributions. In this paper, we introduce TherA, a controllable RGB-to-TIR translation framework that produces diverse and thermally plausible images at both scene and object level. TherA couples TherA-VLM with a latent-diffusion-based translator. Given a single RGB image and a user-prompted condition pair, TherA-VLM yields a thermal-aware embedding that encodes scene, object, material, and heat-emission context reflecting the input scene-condition pair. Conditioning the diffusion model on this embedding enables realistic TIR synthesis and fine-grained control across time of day, weather, and object state. Compared to other baselines, TherA achieves state-of-the-art translation performance, demonstrating improved zero-shot translation performance up to 33% increase averaged across all metrics.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper26
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar 等NeurIPS 2021 · 被引用 9,661 次
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 被引用 6,759 次
相关 Paper
- DiffTV: Identity-Preserved Thermal-to-Visible Face Translation via Feature Alignment and Dual-Stage ConditionsJingyu Lin, Guiqin Zhao, Jing Xu, Guoli Wang 等ACM MM 2024 · 被引用 9 次
- ThermalGen: Style-Disentangled Flow-Based Generative Models for RGB-to-Thermal Image TranslationJiuhong Xiao, Roshan Nayak, Ning Zhang, Daniel Tortei 等NeurIPS 2025 · 被引用 18 次
- SynthRGB-T: Language-Vision Guided Image Translation for Diversity SynthesisJiangang Ding, Yiquan Du, Pengxiang Li, Lili Pei 等CVPR 2026 · 被引用 1 次
- Data Generation Scheme for Thermal Modality with Edge-Guided Adversarial Conditional Diffusion ModelGuoqing Zhu, Honghu Pan, Qiang Wang, Chao Tian 等ACM MM 2024 · 被引用 7 次
- PhysIR-Splat: Physically Consistent Thermal Infrared Radiative Transfer in 3D Gaussian SplattingJingyuan Gao, Yumeng Hu, Fei Gao, Mingjin ZhangCVPR 2026
