Preference Adaptive and Sequential Text-to-Image Generation
Ofir Nabati, Guy Tennenholtz, Chih-Wei Hsu, Moonkyung Ryu, Deepak Ramachandran, Yinlam Chow, Xiang Li, Craig Boutilier
摘要
We address the problem of interactive text-toimage (T2I) generation, designing a reinforcement learning (RL) agent which iteratively improves a set of generated images for a user through a sequence of prompt expansions. Using human raters, we create a novel dataset of sequential preferences, which we leverage, together with largescale open-source (non-sequential) datasets. We construct user-preference and user-choice models using an EM strategy and identify varying user preference types. We then leverage a large multimodal language model (LMM) and a valuebased RL approach to suggest an adaptive and diverse slate of prompt expansions to the user. Our Preference Adaptive and Sequential Text-toimage Agent (PASTA) extends T2I models with adaptive multi-turn capabilities, fostering collaborative co-creation and addressing uncertainty or underspecification in a user's intent. We evaluate PASTA using human raters, showing significant improvement compared to baseline methods. We also open-source our sequential rater dataset and simulated user-rater interactions to support future research in user-centric multi-turn T2I systems.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Premier: Personalized Preference Modulation with Learnable User Embedding in Text-to-Image GenerationZihao Wang, Yuxiang Wei, Xinpeng Zhou, Tianyu Zhang 等CVPR 2026 · 被引用 1 次
- Visual Personalization Turing TestRameen Abdal, James Burgess, Sergey Tulyakov, Kuan-Chieh Jackson WangCVPR 2026
它引用的顶会 Paper22
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann 等ICLR 2024 · 被引用 4,569 次
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes 等ICLR 2020 · 被引用 4,112 次
相关 Paper
- RetouchAgent: Towards Interactive and Explainable Image Retouching with MLLM AgentsShuo Zhang, Xinyu YangAAAI 2026
- T2I-Copilot: A Training-Free Multi-Agent Text-to-Image System for Enhanced Prompt Interpretation and Interactive GenerationChieh-Yun Chen, Min Shi, Gong Zhang, Humphrey ShiICCV 2025 · 被引用 1 次
- Taming Text-to-Image Synthesis for Novices: User-centric Prompt Generation via Multi-turn GuidanceYilun Liu, Minggui He, Feiyu Yao, Yuhe Ji 等EMNLP 2025
- CTR-Driven Advertising Image Generation with Multimodal Large Language ModelsXingye Chen, Wei Feng, Zhenbang Du, Weizhen Wang 等WWW 2025 · 被引用 15 次
- Boosting Chart-to-Code Generation in MLLM via Dual Preference-Guided RefinementZhihan Zhang, Yixin Cao, Lizi LiaoACM MM 2025
