SimpleDiffusion: A Lightweight and Efficient Conditional Diffusion Model for Multi-Modal Salient Object Detection
Shuo Zhang, Jiaming Huang, Wenbing Tang, Jing Liu, Li Han, Jiandun Li, Hongchun Yuan, Zizhu Fan
Abstract
Multi-modal salient object detection (SOD) through the integration of additional data such as depth or thermal information has become a significant task in computer vision during recent years. Traditionally, the challenges of identifying salient objects in RGB, RGB-D (Depth), and RGB-T (Thermal) images are tackled separately, which often leads to issues like poorly defined object edges or overconfident inaccurate predictions. Recent studies have shown that designing a unified end-to-end framework to handle all these three types of SOD tasks simultaneously is both necessary and difficult. To address this need, we propose a novel approach that treats multi-modal SOD as a conditional mask generation task utilizing diffusion models. Specifically, we introduce DiM-SOD, which enables the concurrent use of local (depth maps, thermal maps) and global controls (images) within a unified model for progressive denoising and refined prediction. DiMSOD is efficient, only requiring fine-tuning of local control adapter on the existing stable diffusion model, which not only reduces the fine-tuning cost and model size, making it more viable for real-world applications, but also enhances the integration of multi-modal conditional controls. Additionally, we have developed modules including SOD-ControlNet, Feature Adaptive Network (FAN), and Feature Injection Attention Network (FIAN) to further enhance the model's performance. Extensive experiments demonstrate that DiMSOD efficiently detects salient objects across RGB, RGB-D, and RGB-T datasets, achieving superior performance compared to previous methods. Our code and datasets are accessible at: https://anonymous.4open.science/r/DiMSOD-0B47/ .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 84eb663d-8a6e-4cef-969a-ffc9ecbaac52Builds on19
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
Related papers
- DiMSOD: A Diffusion-Based Framework for Multi-Modal Salient Object DetectionShuo Zhang, Jiaming Huang, Wenbing Tang, Yan Wu et al.AAAI 2025 · 3 citations
- Specificity-preserving RGB-D Saliency DetectionTao Zhou, Huazhu Fu, Geng Chen, Yi Zhou et al.ICCV 2021 · 210 citations
- Cross-modality Discrepant Interaction Network for RGB-D Salient Object DetectionChen Zhang, Runmin Cong, Qinwei Lin, Lin Ma et al.ACM MM 2021 · 116 citations
- MMNet: Multi-Stage and Multi-Scale Fusion Network for RGB-D Salient Object DetectionGuibiao Liao, Wei Gao, Qiuping Jiang, Ronggang Wang et al.ACM MM 2020 · 53 citations
- Deep RGB-D Saliency Detection With Depth-Sensitive Attention and Automatic Multi-Modal FusionPeng Sun, Wenhu Zhang, Huanyu Wang, Songyuan Li et al.CVPR 2021
