Beyond Sequential Tools: A Unified VLM Agent System for Photographic Post-Processing via Dynamic Multi-Expert Fusion
Honglin Xiong, Chenjie Zhu, Jianbiao Ding, Zixuan Ni, Wei Li, Zhenpeng Mi, Qian Wang
Abstract
Real-world photographic post-processing is a formidable challenge due to the frequent co-occurrence of multiple, coupled image degradations. Current paradigms, such as monolithic "all-in-one" models, often face generalization bottlenecks, while recent agent-based systems suffer from time-consuming, sequential tool invocation and suboptimal coordination of isolated, single-task tools. To overcome these limitations, we propose a novel and efficient framework: a vision-language agent system for universal photographic post-processing. Our system employs a powerful Vision-Language Model (VLM) as an orchestrator agent to perform nuanced user intent understanding and in-depth degradation analysis. Based on its assessment, the VLM generates a structured plan, dynamically allocating weights to a suite of specialized expert LoRA modules. These experts, which adapt only the Key (K) and Value (V) matrices for enhanced composability, are then simultaneously merged into a pretrained diffusion backbone to execute a tailored restoration. To ensure perceptually optimal weights, we introduce a lightweight allocation branch trained on the VLM's features using Direct Preference Optimization (DPO) from human feedback. This dynamic fusion paradigm enables a synergistic, context-aware restoration in a single, efficient forward pass. Our method demonstrates state-of-the-art performance across a wide range of synthetic and real-world datasets with diverse degradations. Crucially, it exhibits remarkable zero-shot generalization, achieving excellent results on real-world data. Our code and weights will be made publicly available.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1fb43e9c-0845-4f44-a76b-6a5404fd04b7Builds on15
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Restormer: Efficient Transformer for High-Resolution Image RestorationSyed Waqas Zamir, Aditya Arora, Salman Khan, Munawar Hayat et al.CVPR 2022 · 3,348 citations
- Uformer: A General U-Shaped Transformer for Image RestorationZhendong Wang, Xiaodong Cun, Jianmin Bao, Wengang Zhou et al.CVPR 2022 · 1,970 citations
- MUSIQ: Multi-scale Image Quality TransformerJunjie Ke, Qifei Wang, Yilin Wang, Peyman Milanfar et al.ICCV 2021 · 1,325 citations
Related papers
- An Intelligent Agentic System for Complex Image Restoration ProblemsKaiwen Zhu, Jinjin Gu, Zhiyuan You, Yu Qiao et al.ICLR 2025
- Hybrid Agents for Image RestorationBingchen Li, Xin Li, Yiting Lu, Zhibo ChenCVPR 2026 · 17 citations
- FAPE-IR: Frequency-Aware Planning and Execution Framework for All-in-One Image RestorationJingren Liu, Shuning Xu, Qirui Yang, Yun Wang et al.CVPR 2026 · 4 citations
- RestoreAgent: Autonomous Image Restoration Agent via Multimodal Large Language ModelsHaoyu Chen, Wenbo Li, Jinjin Gu, Jingjing Ren et al.NeurIPS 2024 · 49 citations
- Language-driven All-in-one Adverse Weather RemovalHao Yang, Liyuan Pan, Yan Yang, Wei LiangCVPR 2024 · 28 citations
