B-Spar: Bayesian Sparse-Reward Modeling for RL-based Image Editing
shusong xu, Peiye Liu, Yongbin Liu, Bangjie Yin, Tianyi Zheng, Zhaomang Sun, Zhenyu Chen, Peng-Tao Jiang, Jian Zhang, Yuzhao Wang, Zhen Gu, Jinwei Chen, Bo Li
Abstract
Autonomous image-editing agents powered by multimodal large language models (MLLMs) improve transparency and controllability by translating high-level instructions into tool-mediated edit sequences, but training such agents with reinforcement learning often relies on dense proxy rewards (e.g., incremental image-quality score gains) to compensate for sparse human feedback. When these proxies overvalue small local changes, the resulting optimization signal can be dominated by numerically measurable yet perceptually negligible edits, biasing policy gradients toward proxy artifacts rather than meaningful progress. We propose B-Spar, a reward-centric Reinforcement Learning framework for perceptually aligned image retouching under sparse feedback that combines prior-guided trajectory sampling to reduce inefficient exploration, Bayesian reward modeling to densify sparse binary feedback into a stable training signal, and anchor-regularized policy optimization to steer updates toward high-reward regions while preventing early mode collapse. Experiments on public benchmarks demonstrate that B-Spar improves perceptual quality and metric alignment with stable training and competitive inference efficiency over strong prompt-based and training-based baselines. Notably, it outperforms AIGC-based baselines by over 95% in perceptual quality, achieving an improvement of approximately 33.5% over the state-of-the-art.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7c3065cb-4b4c-4fb7-8523-5fc5034498aeBuilds on14
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Self-Rewarding Language ModelsWeizhe Yuan, Richard Yuanzhe Pang, Kyunghyun Cho, Xian Li et al.ICML 2024 · 569 citations
- Learning to Utilize Shaping Rewards: A New Approach of Reward ShapingYujing Hu, Weixun Wang, Hangtian Jia, Yixiang Wang et al.NeurIPS 2020 · 256 citations
- Exploration-Guided Reward Shaping for Reinforcement Learning under Sparse RewardsRati Devidze, Parameswaran Kamalaruban, Adish SinglaNeurIPS 2022 · 122 citations
Related papers
- RetouchAgent: Towards Interactive and Explainable Image Retouching with MLLM AgentsShuo Zhang, Xinyu YangAAAI 2026
- RePrompt: Reasoning-Augmented Reprompting for Text-to-Image Generation via Reinforcement LearningMingrui Wu, Lu Wang, Pu Zhao, Fangkai Yang et al.ICLR 2026 · 19 citations
- CogniEdit: Dense Gradient Flow Optimization for Fine-Grained Image EditingYan Li, Lin Liu, Xiaopeng Zhang, Wei Xue et al.CVPR 2026 · 2 citations
- Spatial Preference Rewarding for MLLMs Spatial UnderstandingHan Qiu, Peng Gao, Lewei Lu, Xiaoqin Zhang et al.ICCV 2025 · 3 citations
- Retrospective In-Context Learning for Temporal Credit Assignment with Large Language ModelsWen-Tse Chen, Jiayu Chen, Fahim Tajwar, Hao Zhu et al.NeurIPS 2025 · 4 citations
