FRAMER: Frequency-Aligned Self-Distillation with Adaptive Modulation Leveraging Diffusion Priors for Real-World Image Super-Resolution
Seungho Choi, Jeahun Sung, Jihyong Oh
Abstract
Real-image super-resolution (Real-ISR) seeks to recover HR images from LR inputs with mixed, unknown degradations. While diffusion models surpass GANs in perceptual quality, they under-reconstruct high-frequency (HF) details due to a low-frequency (LF) bias and a depth-wise"low-first, high-later"hierarchy. We introduce FRAMER, a plug-and-play training scheme that exploits diffusion priors without changing the backbone or inference. At each denoising step, the final-layer feature map teaches all intermediate layers. Teacher and student feature maps are decomposed into LF/HF bands via FFT masks to align supervision with the model's internal frequency hierarchy. For LF, an Intra Contrastive Loss (IntraCL) stabilizes globally shared structure. For HF, an Inter Contrastive Loss (InterCL) sharpens instance-specific details using random-layer and in-batch negatives. Two adaptive modulators, Frequency-based Adaptive Weight (FAW) and Frequency-based Alignment Modulation (FAM), reweight per-layer LF/HF signals and gate distillation by current similarity. Across U-Net and DiT backbones (e.g., Stable Diffusion 2, 3), FRAMER consistently improves PSNR/SSIM and perceptual metrics (LPIPS, NIQE, MANIQA, MUSIQ). Ablations validate the final-layer teacher and random-layer negatives.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fff66f4b-0d38-46ce-a810-33f0ecbad057Builds on19
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari et al.ICML 2024 · 3,620 citations
Related papers
- FiDeSR: High-Fidelity and Detail-Preserving One-Step Diffusion Super-ResolutionAro Kim, Myeongjin Jang, Chaewon Moon, Youngjin Shin et al.CVPR 2026 · 3 citations
- One-Step Residual Shifting Diffusion for Image Super-Resolution via DistillationDaniil Selikhanovych, David Li, Aleksei Leonov, Nikita Gushchin et al.ICML 2026 · 3 citations
- SSL: A Self-similarity Loss for Improving Generative Image Super-resolutionDu Chen, Zhengqiang Zhang, Jie Liang, Lei ZhangACM MM 2024 · 7 citations
- Disentangled Textual Priors for Diffusion-based Image Super-ResolutionLei Jiang, Xin Liu, Xinze Tong, Zhiliang Li et al.CVPR 2026 · 2 citations
- Beyond the Ground Truth: Enhanced Supervision for Image RestorationDonghun Ryou, Inju Ha, Sanghyeok Chu, Bohyung HanCVPR 2026 · 4 citations
