Thermal Diffusion Matters: Infrared Spatial-Temporal Video Super-Resolution through Heat Conduction Priors
Mingxuan Zhou, Shuang Li, Yutang Zhang, Jing Geng, Yirui Shen, Jingxuan Kang, Fuzhen Zhuang, Shuigen Wang
Abstract
Infrared video acquisition inherently suffers from low spatial resolution and limited frame rates due to the physical constraints of thermal imaging sensors. These limitations make infrared video enhancement uniquely challenging, as it requires restoring spatial details and temporal continuity from highly undersampled thermal signals. To address this challenge, we propose THERIS, a unified THERmal-physics inspired framework for Infrared spatial-temporal video Super-resolution. Grounded in the physical principles of thermal diffusion, THERIS leverages heat conduction dynamics that govern the spatiotemporal evolution of infrared pixel intensities. Specifically, the proposed Thermal Diffusion Interpolation Module (TDIM) treats temporal feature sequences as one-dimensional heat fields and performs frequency-domain diffusion to synthesize temporally coherent intermediate frames. Building on this foundation, the Thermo-Aware State Space Module (TSSM) refines spatiotemporal representations through learnable spectral filtering and selective state-space modeling, while maintaining consistency guided by the thermodynamic prior inherited from TDIM. Additionally, a Temperature Field Modeling Loss is introduced to enforce adherence to the heat conduction equation, promoting temporal coherence and spatial stability in the generated results. Extensive experiments demonstrate that THERIS achieves state-of-the-art performance while producing visually coherent results. To facilitate further research in the infrared video processing domain, we also introduce IRVAL, a high-resolution dataset comprising 108,512 video frames at 512512 resolution.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext eda3731e-92ec-419f-bda0-a8f65f23d6baBuilds on23
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 3,482 citations
- Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space ModelLianghui Zhu, Bencheng Liao, Qian Zhang, Xinlong Wang et al.ICML 2024 · 1,725 citations
- MUSIQ: Multi-scale Image Quality TransformerJunjie Ke, Qifei Wang, Yilin Wang, Peyman Milanfar et al.ICCV 2021 · 1,325 citations
Related papers
- Thermal-Physics Guided Infrared Image Super-Resolution with Dynamic High-Frequency AmplificationMingxuan Zhou, Yirui Shen, Shuang Li, Jing Geng et al.AAAI 2026
- Streaming Diffusion Model for Fast Infrared and Visible Video FusionJinyuan Liu, Ludan Sun, Tengyu Ma, Chunyan Yang et al.CVPR 2026 · 2 citations
- HATIR: Heat-Aware Diffusion for Turbulent Infrared Video Super-ResolutionYang Zou, Xingyue Zhu, Kaiqi Han, Jun Ma et al.AAAI 2026 · 3 citations
- Toward Real-world Infrared Image Super-Resolution: A Unified Autoregressive Framework and Benchmark DatasetYang Zou, Jun Ma, Zhidong Jiao, Xingyuan Li et al.CVPR 2026 · 4 citations
- Learning Spatial-Temporal Implicit Neural Representations for Event-Guided Video Super-ResolutionYunfan Lu, Zipeng Wang, Minjie Liu, Hongjian Wang et al.CVPR 2023
