Lune

ICML2026顶会

Benchmarking and Improving Fine-Grained Text-to-Image Alignment via Paired Reinforcement Learning

Kaihang Pan, Wendong Bu, Yuruo Wu, Kai Shen, Yang Wu, Yun Zhu, Zehan Wang, liyunfei, ZhaoHang, Juncheng Li, Siliang Tang

出版方
2026年份

摘要

While recent autoregressive models have achieved text-to-image generation performance comparable to diffusion models, they significantly struggle with fine-grained semantic alignment. To rigorously evaluate this limitation, we introduce DeltaBench, a benchmark featuring paired prompts with subtle fine-grained differences, which reveals that existing models fail to achieve precise control over visual tokens. To bridge this gap, we propose FocusDiff, a comprehensive framework that enhances alignment by learning from subtle differences in similar text-image pairs. Specifically, we construct FocusDiff-Data, a large-scale dataset of paired samples derived from image editing tasks to capture localized semantic shifts. Furthermore, we introduce Pair-GRPO, an improved reinforcement learning algorithm that extends GRPO to paired samples. Extensive experiments demonstrate that our approach outperforms most prior prominent methods on both DeltaBench and existing benchmarks.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

它引用的顶会 Paper22

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖