Lune

CVPR2026Top-tier venue

RetouchIQ: MLLM Agents for Instruction-Based Image Retouching with Generalist Reward

Qiucheng Wu, Jing Shi, Simon Jenni, Kushal Kafle, Tianyu Wang, Shiyu Chang, Handong Zhao

2026Year
4Citations

Abstract

Convey the vast, serene grandeur of the night sky-a quiet sense of infinity beneath the luminous Milky Way.

Create an epic, moody seascape with a cool teal-blue tone. Bring richness to the sky and waves, give a dramatic, cinematic intensity.

Let's give this image a warmer feel and shift the tones to enhance the ambient glow and overall richness.

I want the Chicago Theater to stand out with a bold cinematic glow, radiant and glorious against the night.

Make the image feel more vivid and alive, softening heavy shadows and bringing warmth and depth to the scene. Make the flowers look more vibrant and lively, with richer tones and a crisp, refreshing feel that brings out their natural beauty.

Give this photo a warm vintage feel, adding depth and a timeless sense of nostalgia for this character.

Make the scene feel warmer and more inviting, with gentle golden light that adds comfort and a calm, cozy atmosphere.

  • This work was completed during Qiucheng's internship at Adobe.

goals with precise parameter control. To move beyond conventional, rule-based rewards that compute similarity against a fixed reference image using handcrafted metrics, we propose a generalist reward model-an RL fine-tuned MLLM that evaluates edited results through a set of generated metrics on a case-by-case basis. Then, the reward model provides scalar feedback through multimodal reasoning, enabling reinforcement learning with high-quality, instruction-consistent gradients. We curate an extended dataset with 190k instruction-reasoning pairs and establish a new benchmark for instruction-based image editing. Experiments show that RETOUCHIQ substantially improves both semantic consistency and perceptual quality over previous MLLM-based and diffusion-based editing systems. Our findings demonstrate the potential of generalist reward-driven MLLM agents as flexible, explainable, and executable assistants for professional image editing.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 2801ace5-49fc-49a6-a90c-ef8b273aebdb

Builds on22

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines