A Simple Yet Mighty Hartley Diffusion Versatilist for Generalizable Dense Vision Tasks
Qi Bi, Jingjun Yi, Huimin Huang, Hao Zheng, Haolan Zhan, Wei Ji, Yawen Huang, Yuexiang Li, Yefeng Zheng
Abstract
Diffusion models have demonstrated powerful capability as a versatilist for dense vision tasks, yet the generalization ability to unseen domains remains rarely explored. This paper presents HarDiff, an efficient frequency learning scheme, so as to advance generalizable paradigms for diffusion based dense prediction. It draws inspiration from a fine-grained analysis of Discrete Hartley Transform, where some low-frequency features activate the broader content of an image, while some high-frequency features maintain sufficient details for dense pixels. Consequently, HarDiff consists of two key components. The low-frequency training process extracts structural priors from the source domain, to enhance understanding of task-related content. The high-frequency sampling process utilizes detail-oriented guidance from the unseen target domain, to infer precise dense predictions with target-related details. Extensive empirical evidence shows that HarDiff can be easily plugged into various dense vision tasks, e.g., semantic segmentation, depth estimation and haze removal, yielding improvements over the state-of-the-art methods in twelve public benchmarks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext db46ddbf-1f37-4001-80e5-7626b7787437Cited by top-tier papers1
Ask how each one uses itBuilds on61
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Directly Denoising Diffusion ModelsDan Zhang, Jingjing Wang, Feng LuoICML 2024 · 11,724 citations
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar et al.NeurIPS 2021 · 9,661 citations
Related papers
- DDP: Diffusion Model for Dense Visual PredictionYuanfeng Ji, Zhe Chen, Enze Xie, Lanqing Hong et al.ICCV 2023 · 223 citations
- Frequency Domain-Based Diffusion Model for Unpaired Image DehazingChengxu Liu, Lu Qi, Jinshan Pan, Xueming Qian et al.ICCV 2025 · 13 citations
- Generating Content for HDR Deghosting from Frequency ViewTao Hu, Qingsen Yan, Yuankai Qi, Yanning ZhangCVPR 2024
- Exploiting Diffusion Prior for Generalizable Dense PredictionHsin-Ying Lee, Hung-Yu Tseng, Hsin-Ying Lee, Ming-Hsuan YangCVPR 2024
- Diffusion Priors for Variational Likelihood Estimation and Image DenoisingJun Cheng, Shan TanNeurIPS 2024 · 5 citations
