Tailored Visions: Enhancing Text-to-Image Generation with Personalized Prompt Rewriting
Zijie Chen, Lichao Zhang, Fangsheng Weng, Lili Pan, Zhenzhong Lan
摘要
Despite significant progress in the field, it is still challenging to create personalized visual representations that align closely with the desires and preferences of individual users. This process requires users to articulate their ideas in words that are both comprehensible to the models and accurately capture their vision, posing difficulties for many users. In this paper, we tackle this challenge by leveraging historical user interactions with the system to enhance user prompts. We propose a novel approach that involves rewriting user prompts based on a newly collected large-scale text-to-image dataset with over 300k prompts from 3115 users. Our rewriting model enhances the expressiveness and alignment of user prompts with their intended visual outputs. Experimental results demonstrate the superiority of our methods over baseline approaches, as evidenced in our new offline evaluation method and online tests. Our code and dataset are available at https://github.com/zzjchen/Tailored-Visions
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- SnapMoGen: Human Motion Generation from Expressive TextsChuan Guo, Inwoo Hwang, Jian Wang, Bing ZhouNeurIPS 2025 · 被引用 50 次
- Personalized Generation In Large Model Era: A SurveyYiyan Xu, Jinghao Zhang, Alireza Salemi, Xinting Hu 等ACL 2025 · 被引用 45 次
- POET: Supporting Prompting Creativity and Personalization with Automated Expansion of Text-to-Image GenerationEvans Xu Han, Alice Qian Zhang, Haiyi Zhu, Hong Shen 等UIST 2025 · 被引用 5 次
- Draw Your Mind: Personalized Generation via Condition-Level Modeling in Text-to-Image Diffusion ModelsHyungjin Kim, Seokho Ahn, Young-Duk SeoICCV 2025 · 被引用 4 次
- DRC: Enhancing Personalized Image Generation via Disentangled Representation CompositionYiyan Xu, Wuqiang Zheng, Wenjie Wang, Fengbin Zhu 等ACM MM 2025 · 被引用 2 次
它引用的顶会 Paper19
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 被引用 6,759 次
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray 等ICML 2021 · 被引用 6,356 次
相关 Paper
- A User-Friendly Framework for Generating Model-Preferred Prompts in Text-to-Image SynthesisNailei Hei, Qianyu Guo, Zihao Wang, Yan Wang 等AAAI 2024 · 被引用 11 次
- Capability-aware Prompt Reformulation Learning for Text-to-Image GenerationJingtao Zhan, Qingyao Ai, Yiqun Liu, Jia Chen 等SIGIR 2024 · 被引用 7 次
- Promptify: Text-to-Image Generation through Interactive Prompt Exploration with Large Language ModelsStephen Brade, Bryan Wang, Maurício Sousa, Sageev Oore 等UIST 2023 · 被引用 179 次
- Is It AI or Is It Me? Understanding Users' Prompt Journey with Text-to-Image Generative AI ToolsAtefeh Mahdavi Goloujeh, Anne Sullivan, Brian MagerkoCHI 2024 · 被引用 88 次
- Learning to Rewrite Prompts for Personalized Text GenerationCheng Li, Mingyang Zhang, Qiaozhu Mei, Weize Kong 等WWW 2024 · 被引用 54 次
