Personalized Image Generation with Large Multimodal Models
Yiyan Xu, Wenjie Wang, Yang Zhang, Biao Tang, Peng Yan, Fuli Feng, Xiangnan He
摘要
Personalized content filtering, such as recommender systems, has become a critical infrastructure to alleviate information overload. However, these systems merely filter existing content and are constrained by its limited diversity, making it difficult to meet users' varied content needs. To address this limitation, personalized content generation has emerged as a promising direction with broad applications. Nevertheless, most existing research focuses on personalized text generation, with relatively little attention given to personalized image generation. The limited work in personalized image generation faces challenges in accurately capturing users' visual preferences and needs from noisy user-interacted images and complex multimodal instructions. Worse still, there is a lack of supervised data for training personalized image generation models. To overcome the challenges, we propose a Personalized Image Generation Framework named Pigeon, which adopts exceptional large multimodal models with three dedicated modules to capture users' visual preferences and needs from noisy user history and multimodal instructions. To alleviate the data scarcity, we introduce a two-stage preference alignment scheme, comprising masked preference reconstruction and pairwise preference alignment, to align Pigeon with the personalized image generation task. We apply Pigeon to personalized sticker and movie poster generation, where extensive quantitative results and human evaluation highlight its superiority over various generative baselines.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- Easier Painting Than Thinking: Can Text-to-Image Models Set the Stage, but Not Direct the Play?Ouxiang Li, Yuan Wang, Xinting Hu, Huijuan Huang 等ICLR 2026 · 被引用 39 次
- SPEED: Scalable, Precise, and Efficient Concept Erasure for Diffusion ModelsOuxiang Li, Yuan Wang, Xinting Hu, Houcheng Jiang 等ICLR 2026 · 被引用 37 次
- Personalized Image Editing in Text-to-Image Diffusion Models via Collaborative Direct Preference OptimizationConnor Dunlop, Matthew Zheng, Kavana Venkatesh, Pinar YanardagNeurIPS 2025 · 被引用 8 次
- Thinking with Frames: Generative Video Distortion Evaluation via Frame Reward ModelYuan Wang, Borui Liao, Huijuan Huang, Jinda Lu 等CVPR 2026 · 被引用 5 次
- Personalized Visual Content Generation in Conversational SystemsXianquan Wang, Zhaocheng Du, Huibo Xu, Shukang Yin 等NeurIPS 2025 · 被引用 4 次
它引用的顶会 Paper27
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
相关 Paper
- Premier: Personalized Preference Modulation with Learnable User Embedding in Text-to-Image GenerationZihao Wang, Yuxiang Wei, Xinpeng Zhou, Tianyu Zhang 等CVPR 2026 · 被引用 1 次
- PMG : Personalized Multimodal Generation with Large Language ModelsXiaoteng Shen, Rui Zhang, Xiaoyan Zhao, Jieming Zhu 等WWW 2024 · 被引用 40 次
- DRC: Enhancing Personalized Image Generation via Disentangled Representation CompositionYiyan Xu, Wuqiang Zheng, Wenjie Wang, Fengbin Zhu 等ACM MM 2025 · 被引用 2 次
- I-AM-G: Interest Augmented Multimodal Generator for Item PersonalizationXianquan Wang, Likang Wu, Shukang Yin, Zhi Li 等EMNLP 2024 · 被引用 1 次
- Prompt2Poster: Automatically Artistic Chinese Poster Creation from Prompt OnlyShaodong Wang, Yunyang Ge, Liuhan Chen, Haiyang Zhou 等ACM MM 2024 · 被引用 5 次
