The Promise of RL for Autoregressive Image Editing
Saba Ahmadi, Rabiul Awal, Ankur Sikarwar, Amirhossein Kazemnejad, Ge Ya Luo, Juan A. Rodríguez, Sai Rajeswar Mudumba, Siva Reddy, Chris Pal, Benno Krojer, Aishwarya Agrawal
Abstract
While image generation techniques are now capable of producing high-quality images that respect prompts which span multiple sentences, the task of text-guided image editing remains a challenge. Even edit requests that consist of only a few words often fail to be executed correctly. We explore three strategies to enhance performance on a wide range of image editing tasks: supervised fine-tuning (SFT), reinforcement learning (RL), and Chain-of-Thought (CoT) reasoning. In order to study all these components in one consistent framework, we adopt an autoregressive multimodal model that processes textual and visual tokens in a unified manner. We find RL combined with a large multi-modal LLM verifier to be the most effective of these strategies. As a result, we release EARL: Editing with Autoregression and RL, a strong RL-based image editing model that performs competitively on a diverse range of edits compared to strong baselines, despite using much less training data. Thus, EARL pushes the frontier of autoregressive multimodal models on image editing. We release our code, training data, and trained models at https://github.com/mair-lab/EARL.
Turn the color of patient chart to be brown. Remove one vase.
Pour all the orange content from the bag into the big bowl.
Remove the birds from the image. Replace the beaded curtain with window curtains.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext effd400c-bf0c-4eeb-b30e-266c1647ffb5Cited by top-tier papers3
- EditScore: Unlocking Online RL for Image Editing via High-Fidelity Reward ModelingXin Luo, Jiahao Wang, Chenyuan Wu, Shitao Xiao et al.ICLR 2026 · 63 citations
- Rendering-Aware Reinforcement Learning for Vector Graphics GenerationJuan A. Rodríguez, Haotian Zhang, Abhay Puri, Rishav Pramanik et al.NeurIPS 2025 · 42 citations
- Learning an Image Editing Model without Image Editing PairsNupur Kumari, Sheng-Yu Wang, Nanxuan Zhao, Yotam Nitzan et al.ICLR 2026 · 14 citations
Builds on32
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann et al.ICLR 2024 · 4,569 citations
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari et al.ICML 2024 · 3,620 citations
Related papers
- Breaking the SFT Plateau: Multimodal Structured Reinforcement Learning for Chart-to-Code GenerationLei Chen, Xuanle Zhao, Zhixiong Zeng, Jing Huang et al.ICLR 2026 · 16 citations
- VTool-R1: VLMs Learn to Think with Images via Reinforcement Learning on Multimodal Tool UseMingyuan Wu, Jingcheng Yang, Jize Jiang, Meitang Li et al.ICLR 2026 · 87 citations
- BoxCtrl: 3D-Aware Visual Prompting for Geometric Image EditingFeifei Wang, Shiyuan Yang, Xiaoyu Li, Jing LiaoSIGGRAPH 2026
- GoT: Unleashing Reasoning Capability of MLLM for Visual Generation and EditingRongyao Fang, Chengqi Duan, Kun Wang, Linjiang Huang et al.NeurIPS 2025 · 5 citations
- ReFocus: Visual Editing as a Chain of Thought for Structured Image UnderstandingXingyu Fu, Minqian Liu, Zhengyuan Yang, John Corring et al.ICML 2025
