RetouchAgent: Towards Interactive and Explainable Image Retouching with MLLM Agents
Shuo Zhang, Xinyu Yang
摘要
Although deep learning-based image retouching has made significant progress, its inherent subjectivity renders current black-box methods limited in interactivity and explainability. Among existing efforts, parameter-controlled methods aim to improve interactivity, but often suffer from ambiguous semantics and lack support for natural language control. Reinforcement learning–based explainability methods are constrained by low-dimensional and limited action spaces, which result in suboptimal performance. To address the above issues, we propose RetouchAgent, a novel framework that leverages collaboration among multiple MLLM agents for image retouching. Our method consists of the following key steps: (1) Retrieval: By constructing a multimodal retouching database, we enable an ICL sample retrieval mechanism guided by retouching intent. (2) Engine: Leveraging the vision-language understanding capabilities of MLLM, a carefully designed prompting strategy, and a dedicated operation library, we enable precise and controllable image retouching. (3) Reflection: We evaluate each retouching interaction and optimize the retouching process for progressive result refinement. Finally, through multiple rounds of collaboration among MLLM agents, RetouchAgent achieves state-of-the-art performance in quantitative and qualitative evaluations.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper20
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- MasaCtrl: Tuning-Free Mutual Self-Attention Control for Consistent Image Synthesis and EditingMingdeng Cao, Xintao Wang, Zhongang Qi, Ying Shan 等ICCV 2023 · 被引用 770 次
- A Survey on In-context LearningQingxiu Dong, Lei Li, Damai Dai, Ce Zheng 等EMNLP 2024 · 被引用 479 次
相关 Paper
- PerTouch: VLM-Driven Agent for Personalized and Semantic Image RetouchingZewei Chang, Zheng-Peng Duan, Jianxing Zhang, Chun-Le Guo 等AAAI 2026 · 被引用 3 次
- JarvisArt: Liberating Human Artistic Creativity via an Intelligent Photo Retouching AgentYunlong Lin, Zixu Lin, Kunjie Lin, Jinbin Bai 等NeurIPS 2025 · 被引用 42 次
- Hybrid Agents for Image RestorationBingchen Li, Xin Li, Yiting Lu, Zhibo ChenCVPR 2026 · 被引用 17 次
- MonetGPT: Solving Puzzles Enhances MLLMs' Image Retouching SkillsNiladri Shekhar Dutt, Duygu Ceylan, Niloy J. MitraSIGGRAPH 2025 · 被引用 2 次
- VeraRetouch: A Lightweight Fully Differentiable Framework for Multi-Task Reasoning Photo RetouchingYihong Guo, Youwei Lyu, Jiajun Tang, Yizhuo Zhou 等SIGGRAPH 2026
