RetouchIQ: MLLM Agents for Instruction-Based Image Retouching with Generalist Reward
Qiucheng Wu, Jing Shi, Simon Jenni, Kushal Kafle, Tianyu Wang, Shiyu Chang, Handong Zhao
Abstract
Convey the vast, serene grandeur of the night sky-a quiet sense of infinity beneath the luminous Milky Way.
Create an epic, moody seascape with a cool teal-blue tone. Bring richness to the sky and waves, give a dramatic, cinematic intensity.
Let's give this image a warmer feel and shift the tones to enhance the ambient glow and overall richness.
I want the Chicago Theater to stand out with a bold cinematic glow, radiant and glorious against the night.
Make the image feel more vivid and alive, softening heavy shadows and bringing warmth and depth to the scene. Make the flowers look more vibrant and lively, with richer tones and a crisp, refreshing feel that brings out their natural beauty.
Give this photo a warm vintage feel, adding depth and a timeless sense of nostalgia for this character.
Make the scene feel warmer and more inviting, with gentle golden light that adds comfort and a calm, cozy atmosphere.
- This work was completed during Qiucheng's internship at Adobe.
goals with precise parameter control. To move beyond conventional, rule-based rewards that compute similarity against a fixed reference image using handcrafted metrics, we propose a generalist reward model-an RL fine-tuned MLLM that evaluates edited results through a set of generated metrics on a case-by-case basis. Then, the reward model provides scalar feedback through multimodal reasoning, enabling reinforcement learning with high-quality, instruction-consistent gradients. We curate an extended dataset with 190k instruction-reasoning pairs and establish a new benchmark for instruction-based image editing. Experiments show that RETOUCHIQ substantially improves both semantic consistency and perceptual quality over previous MLLM-based and diffusion-based editing systems. Our findings demonstrate the potential of generalist reward-driven MLLM agents as flexible, explainable, and executable assistants for professional image editing.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2801ace5-49fc-49a6-a90c-ef8b273aebdbBuilds on22
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Score-Based Generative Modeling through Stochastic Differential EquationsYang Song, Jascha Sohl-Dickstein, Diederik P. Kingma, Abhishek Kumar et al.ICLR 2021 · 1,270 citations
- Prompt-to-Prompt Image Editing with Cross-Attention ControlAmir Hertz, Ron Mokady, Jay Tenenbaum, Kfir Aberman et al.ICLR 2023 · 361 citations
- VL-Rethinker: Incentivizing Self-Reflection of Vision-Language Models with Reinforcement LearningHaozhe Wang, Chao Qu, Zuming Huang, Wei Chu et al.NeurIPS 2025 · 356 citations
- Mastering Text-to-Image Diffusion: Recaptioning, Planning, and Generating with Multimodal LLMsLing Yang, Zhaochen Yu, Chenlin Meng, Minkai Xu et al.ICML 2024 · 231 citations
Related papers
- Refine-IQA: Multi-Stage Reinforcement Finetuning for Perceptual Image Quality AssessmentZiheng Jia, Jiaying Qian, Zicheng Zhang, Zijian Chen et al.AAAI 2026 · 3 citations
- EditScore: Unlocking Online RL for Image Editing via High-Fidelity Reward ModelingXin Luo, Jiahao Wang, Chenyuan Wu, Shitao Xiao et al.ICLR 2026 · 63 citations
- MonetGPT: Solving Puzzles Enhances MLLMs' Image Retouching SkillsNiladri Shekhar Dutt, Duygu Ceylan, Niloy J. MitraSIGGRAPH 2025 · 2 citations
- PerTouch: VLM-Driven Agent for Personalized and Semantic Image RetouchingZewei Chang, Zheng-Peng Duan, Jianxing Zhang, Chun-Le Guo et al.AAAI 2026 · 3 citations
- VeraRetouch: A Lightweight Fully Differentiable Framework for Multi-Task Reasoning Photo RetouchingYihong Guo, Youwei Lyu, Jiajun Tang, Yizhuo Zhou et al.SIGGRAPH 2026
