Lune

CVPR2026顶会

Aligning Multi-Character Narrative Image Generation with Multi-Aspect Human Preferences

Ziyi Gao, Zhipeng Wei, Jingjing Chen, Zhiyu Tan, Hao Li, Yi-Ping Phoebe Chen

出版方
2026年份

摘要

Narrative image generation aims to create images featuring multiple distinct characters while capturing their interrelationships, posing significant challenges for current text-to-image diffusion models. As a result, general personalized methods often suffer from poor semantic alignment, identity blending, and aesthetic implausibility. These issues are inadequately captured by existing evaluation metrics such as CLIP, ArcFace, and conventional reward models, which fundamentally fail to align with human perceptual preferences. To align with human preferences, we first construct a fine-grained human preference dataset, NI-RLHF, by collecting both detailed human critiques and preference judgments across three core dimensions: prompt following, identity consistency, and visual quality. This comprehensive dataset facilitates the training of NIReward, a critiquebased reward model capable of generating interpretable image evaluations. Building upon the interpretable reward signal from NIReward, we propose Adaptive Dominancebased Preference Optimization (ADPO) to balance learning across diverse preference dimensions while dynamically adapting to reward margins. Experimental results indicate that NIReward significantly outperforms existing evaluation models and reward models, and ADPO yields a significant improvement across the three key preference dimensions. By introducing NIReward and ADPO, our work paves the way for generating narrative images aligned with actual human preferences.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

它引用的顶会 Paper12

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖