Human Learning by Model Feedback: The Dynamics of Iterative Prompting with Midjourney
Shachar Don-Yehiya, Leshem Choshen, Omri Abend
摘要
Generating images with a Text-to-Image model often requires multiple trials, where human users iteratively update their prompt based on feedback, namely the output image. Taking inspiration from cognitive work on reference games and dialogue alignment, this paper analyzes the dynamics of the user prompts along such iterations. We compile a dataset of iterative interactions of human users with Midjourney. 1 Our analysis then reveals that prompts predictably converge toward specific traits along these iterations. We further study whether this convergence is due to human users, realizing they missed important details, or due to adaptation to the model's "preferences", producing better images for a specific language style. We show initial evidence that both possibilities are at play. The possibility that users adapt to the model's preference raises concerns about reusing user data for further training. The prompts may be biased towards the preferences of a specific model, rather than align with human intentions and natural manner of expression.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Exploring the Evolvement of User Engagement in Online Creative Community under the Surge of Generative AI: A Case Study of DeviantArtQingyu Guo, Yuqi Zhang, Kangyu Yuan, Changyang He 等CSCW 2025
- LLMs Behind the Scenes: Enabling Narrative Scene IllustrationMelissa Roemmele, John Joon Young Chung, Taewook Kim, Yuqian Sun 等EMNLP 2025
它引用的顶会 Paper14
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger 等ICLR 2020 · 被引用 8,443 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- Pick-a-Pic: An Open Dataset of User Preferences for Text-to-Image GenerationYuval Kirstain, Adam Polyak, Uriel Singer, Shahbuland Matiana 等NeurIPS 2023 · 被引用 1,192 次
- CLIPScore: A Reference-free Evaluation Metric for Image CaptioningJack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras 等EMNLP 2021 · 被引用 937 次
相关 Paper
- An Exploration of Default Images in Text-to-Image GenerationHannu Simonen, Atte Kiviniemi, Hannah Johnston, Helena Barranha 等CHI 2026 · 被引用 2 次
- Is It AI or Is It Me? Understanding Users' Prompt Journey with Text-to-Image Generative AI ToolsAtefeh Mahdavi Goloujeh, Anne Sullivan, Brian MagerkoCHI 2024 · 被引用 88 次
- Self-Consuming Generative Models with Curated Data Provably Optimize Human PreferencesDamien Ferbach, Quentin Bertrand, Avishek Joey Bose, Gauthier GidelNeurIPS 2024 · 被引用 41 次
- MPPR: Memory-Prior-based Prompt Refinement in Continuous Space for Advanced Text-to-Image GenerationZhibing Zhang, Jiantao Lin, Cangqi Zhou, Rui XiaACM MM 2025
- The Chosen One: Consistent Characters in Text-to-Image Diffusion ModelsOmri Avrahami, Amir Hertz, Yael Vinker, Moab Arar 等SIGGRAPH 2024 · 被引用 26 次
