Lune

CVPR2026顶会

RetouchIQ: MLLM Agents for Instruction-Based Image Retouching with Generalist Reward

Qiucheng Wu, Jing Shi, Simon Jenni, Kushal Kafle, Tianyu Wang, Shiyu Chang, Handong Zhao

2026年份
4被引次数

摘要

Convey the vast, serene grandeur of the night sky-a quiet sense of infinity beneath the luminous Milky Way.

Create an epic, moody seascape with a cool teal-blue tone. Bring richness to the sky and waves, give a dramatic, cinematic intensity.

Let's give this image a warmer feel and shift the tones to enhance the ambient glow and overall richness.

I want the Chicago Theater to stand out with a bold cinematic glow, radiant and glorious against the night.

Make the image feel more vivid and alive, softening heavy shadows and bringing warmth and depth to the scene. Make the flowers look more vibrant and lively, with richer tones and a crisp, refreshing feel that brings out their natural beauty.

Give this photo a warm vintage feel, adding depth and a timeless sense of nostalgia for this character.

Make the scene feel warmer and more inviting, with gentle golden light that adds comfort and a calm, cozy atmosphere.

  • This work was completed during Qiucheng's internship at Adobe.

goals with precise parameter control. To move beyond conventional, rule-based rewards that compute similarity against a fixed reference image using handcrafted metrics, we propose a generalist reward model-an RL fine-tuned MLLM that evaluates edited results through a set of generated metrics on a case-by-case basis. Then, the reward model provides scalar feedback through multimodal reasoning, enabling reinforcement learning with high-quality, instruction-consistent gradients. We curate an extended dataset with 190k instruction-reasoning pairs and establish a new benchmark for instruction-based image editing. Experiments show that RETOUCHIQ substantially improves both semantic consistency and perceptual quality over previous MLLM-based and diffusion-based editing systems. Our findings demonstrate the potential of generalist reward-driven MLLM agents as flexible, explainable, and executable assistants for professional image editing.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

它引用的顶会 Paper22

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖