B-Spar: Bayesian Sparse-Reward Modeling for RL-based Image Editing
shusong xu, Peiye Liu, Yongbin Liu, Bangjie Yin, Tianyi Zheng, Zhaomang Sun, Zhenyu Chen, Peng-Tao Jiang, Jian Zhang, Yuzhao Wang, Zhen Gu, Jinwei Chen, Bo Li
摘要
Autonomous image-editing agents powered by multimodal large language models (MLLMs) improve transparency and controllability by translating high-level instructions into tool-mediated edit sequences, but training such agents with reinforcement learning often relies on dense proxy rewards (e.g., incremental image-quality score gains) to compensate for sparse human feedback. When these proxies overvalue small local changes, the resulting optimization signal can be dominated by numerically measurable yet perceptually negligible edits, biasing policy gradients toward proxy artifacts rather than meaningful progress. We propose B-Spar, a reward-centric Reinforcement Learning framework for perceptually aligned image retouching under sparse feedback that combines prior-guided trajectory sampling to reduce inefficient exploration, Bayesian reward modeling to densify sparse binary feedback into a stable training signal, and anchor-regularized policy optimization to steer updates toward high-reward regions while preventing early mode collapse. Experiments on public benchmarks demonstrate that B-Spar improves perceptual quality and metric alignment with stable training and competitive inference efficiency over strong prompt-based and training-based baselines. Notably, it outperforms AIGC-based baselines by over 95% in perceptual quality, achieving an improvement of approximately 33.5% over the state-of-the-art.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper14
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Self-Rewarding Language ModelsWeizhe Yuan, Richard Yuanzhe Pang, Kyunghyun Cho, Xian Li 等ICML 2024 · 被引用 569 次
- Learning to Utilize Shaping Rewards: A New Approach of Reward ShapingYujing Hu, Weixun Wang, Hangtian Jia, Yixiang Wang 等NeurIPS 2020 · 被引用 256 次
- Exploration-Guided Reward Shaping for Reinforcement Learning under Sparse RewardsRati Devidze, Parameswaran Kamalaruban, Adish SinglaNeurIPS 2022 · 被引用 122 次
相关 Paper
- RetouchAgent: Towards Interactive and Explainable Image Retouching with MLLM AgentsShuo Zhang, Xinyu YangAAAI 2026
- RePrompt: Reasoning-Augmented Reprompting for Text-to-Image Generation via Reinforcement LearningMingrui Wu, Lu Wang, Pu Zhao, Fangkai Yang 等ICLR 2026 · 被引用 19 次
- CogniEdit: Dense Gradient Flow Optimization for Fine-Grained Image EditingYan Li, Lin Liu, Xiaopeng Zhang, Wei Xue 等CVPR 2026 · 被引用 2 次
- Spatial Preference Rewarding for MLLMs Spatial UnderstandingHan Qiu, Peng Gao, Lewei Lu, Xiaoqin Zhang 等ICCV 2025 · 被引用 3 次
- Retrospective In-Context Learning for Temporal Credit Assignment with Large Language ModelsWen-Tse Chen, Jiayu Chen, Fahim Tajwar, Hao Zhu 等NeurIPS 2025 · 被引用 4 次
