I-AM-G: Interest Augmented Multimodal Generator for Item Personalization
Xianquan Wang, Likang Wu, Shukang Yin, Zhi Li, Yanjiang Chen, Feng Hu, Yu Su, Qi Liu
Abstract
The emergence of personalized generation has made it possible to create texts or images that meet the unique needs of users. Recent advances mainly focus on style or scene transfer based on given keywords. However, in e-commerce and recommender systems, it is almost an untouched area to explore user historical interactions, automatically mine user interests with semantic associations, and create item representations that closely align with user individual interests. In this paper, we propose a brand new framework called Interest Augmented Multimodal Generator (I-AM-G). The framework first extracts tags from the multimodal information of items that the user has interacted with, and the most frequently occurred ones are extracted to rewrite the text description of the item. Then, the framework uses a decoupled text-to-text and image-to-image retriever to search for the top-K similar item text and image embeddings from the item pool. Finally, the Attention module for user interests fuses the retrieved information in a cross-modal manner and further guides the personalized generation process collaborating with the rewritten text. We conducted extensive and comprehensive experiments to demonstrate that our framework can effectively generate results aligned with user preferences, which potentially provides a new paradigm of Rewrite and Retrieve for personalized generation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f006dab6-7010-45f7-824c-9e47370f59e4Cited by top-tier papers3
- Personalized Visual Content Generation in Conversational SystemsXianquan Wang, Zhaocheng Du, Huibo Xu, Shukang Yin et al.NeurIPS 2025 · 4 citations
- Fewer Battles, More Gain: An Information-Efficient Framework for Arena-based LLM EvaluationZirui Liu, Xianquan Wang, Yan Zhuang, Jiatong Li et al.ICLR 2026
- Foundation Encoders Are All You Need for Preference-Aware PersonalizationHyungjin Kim, Seokho Ahn, Young-Duk SeoCVPR 2026
Builds on16
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
Related papers
- RAGAR: Retrieval Augmented Personalized Image Generation Guided by RecommendationRun Ling, Wenji Wang, Yuting Liu, Guibing Guo et al.AAAI 2026 · 5 citations
- Disentangling and Generating Modalities for Recommendation in Missing Modality ScenariosJiwan Kim, Hongseok Kang, Sein Kim, Kibum Kim et al.SIGIR 2025 · 8 citations
- PMG : Personalized Multimodal Generation with Large Language ModelsXiaoteng Shen, Rui Zhang, Xiaoyan Zhao, Jieming Zhu et al.WWW 2024 · 40 citations
- Design Your Ad: Personalized Advertising Image and Text Generation with Unified Autoregressive ModelsYexing Xu, Wei Feng, Shen Zhang, Haohan Wang et al.CVPR 2026 · 1 citation
- Towards Unified Multi-Modal Personalization: Large Vision-Language Models for Generative Recommendation and BeyondTianxin Wei, Bowen Jin, Ruirui Li, Hansi Zeng et al.ICLR 2024 · 46 citations
