SnowMaster: Comprehensive Real-world Image Desnowing via MLLM with Multi-Model Feedback Optimization
Jianyu Lai, Sixiang Chen, Yunlong Lin, Tian Ye, Yun Liu, Song Fei, Zhaohu Xing, Hongtao Wu, Weiming Wang, Lei Zhu
Abstract
Snowfall presents significant challenges for visual data processing, necessitating specialized desnowing algorithms. However, existing models often fail to generalize effectively due to their heavy reliance on synthetic datasets. Furthermore, current real-world snowfall datasets are limited in scale and lack dedicated evaluation metrics designed specifically for snowfall degradation, thus hindering the effective integration of real snowy images into model training to reduce domain gaps. To address these challenges, we first introduce RealSnow10K, a large-scale, high-quality dataset consisting of over 10,000 annotated real-world snowy images. In addition, we curate a preference dataset comprising 36,000 expert-ranked image pairs, enabling the adaptation of multimodal large language models (MLLMs) to better perceive snowy image quality through our innovative Multi-Model Preference Optimization (MMPO). Finally, we propose the SnowMaster, which employs MMPO-enhanced MLLM to perform accurate snowy image evaluation and pseudo-label filtering for semi-supervised training. Experiments demonstrate that SnowMaster delivers superior desnowing performance under real-world conditions.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3d03b087-c960-469a-9130-bfa6a21789d1Cited by top-tier papers2
- PosterReward: Unlocking Accurate Evaluation for High-Quality Graphic Design GenerationJianyu LAI, Sixiang Chen, Jialin Gao, Hengyu Shi et al.CVPR 2026 · 5 citations
- Genhaze: Pioneering Controllable One-Step Realistic Haze Generation for Real-World DehazingSixiang Chen, Tian Ye, Yunlong Lin, Yeying Jin et al.ICCV 2025 · 3 citations
Builds on30
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar et al.NeurIPS 2021 · 9,661 citations
- CrossViT: Cross-Attention Multi-Scale Vision Transformer for Image ClassificationChun-Fu (Richard) Chen, Quanfu Fan, Rameswar PandaICCV 2021 · 2,072 citations
- Exploring CLIP for Assessing the Look and Feel of ImagesJianyi Wang, Kelvin C. K. Chan, Chen Change LoyAAAI 2023 · 1,208 citations
Related papers
- Snow Removal in Video: A New Dataset and A Novel MethodHaoyu Chen, Jingjing Ren, Jinjin Gu, Hongtao Wu et al.ICCV 2023 · 40 citations
- MotionMaster: Generalizable Text-Driven Motion Generation and EditingNan Jiang, Yunhao Li, Lexi Pang, Zimo He et al.CVPR 2026
- Multimodal Large Language Model-Guided ISP Hyperparameter Optimization with Dynamic Preference LearningXinyu Sun, Zhikun Zhao, Congyan Lang, Bing Li et al.ICCV 2025 · 1 citation
- Multi-modal Preference Alignment Remedies Degradation of Visual Instruction Tuning on Language ModelsShengzhi Li, Rongyu Lin, Shichao PeiACL 2024 · 4 citations
- Adaptive Hallucination Alleviation in Multimodal Large Language Models: From Strategic Data Selection to Severity-Guided TrainingYuanyi Xu, Xiangru Zhu, Sihang Jiang, Zhixu Li et al.AAAI 2026
