CTR-Driven Advertising Image Generation with Multimodal Large Language Models
Xingye Chen, Wei Feng, Zhenbang Du, Weizhen Wang, Yanyin Chen, Haohan Wang, Linkai Liu, Yaoyu Li, Jinyuan Zhao, Yu Li, Zheng Zhang, Jingjing Lv
Abstract
In web data, advertising images are crucial for capturing user attention and improving advertising effectiveness. Most existing methods generate background for products primarily focus on the aesthetic quality, which may fail to achieve satisfactory online performance. To address this limitation, we explore the use of Multimodal Large Language Models (MLLMs) for generating advertising images by optimizing for Click-Through Rate (CTR) as the primary objective. Firstly, we build targeted pre-training tasks, and leverage a large-scale e-commerce multimodal dataset to equip MLLMs with initial capabilities for advertising image generation tasks. To further improve the CTR of generated images, we propose a novel reward model to fine-tune pre-trained MLLMs through Reinforcement Learning (RL), which can jointly utilize multimodal features and accurately reflect user click preferences. Meanwhile, a product-centric preference optimization strategy is developed to ensure that the generated background content aligns with the product characteristics after fine-tuning, enhancing the overall relevance and effectiveness of the advertising images. Extensive experiments have demonstrated that our method achieves state-of-the-art performance in both online and offline metrics. Our code and pre-trained models are publicly available at: https://github.com/Chenguoz/CAIG.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 576b640d-b8a5-47c0-bab7-8b157bc9394dCited by top-tier papers9
- RelaCtrl: Relevance-Guided Efficient Control for Diffusion TransformersKe Cao, Jing Wang, Ao Ma, Jiasong Feng et al.AAAI 2026 · 15 citations
- Lay2Story: Extending Diffusion Transformers for Layout-Togglable Story GenerationAo Ma, Jiasong Feng, Ke Cao, Jing Wang et al.ICCV 2025 · 13 citations
- AutoPP: Towards Automated Product Poster Generation and OptimizationJiahao Fan, Yuxin Qin, Wei Feng, Yanyin Chen et al.AAAI 2026 · 2 citations
- Customization under Fire: Plugin Poisoning in Text-to-Image EcosystemJiahao Chen, Xing He, Yong Yang, Xinfeng Li et al.CCS 2026 · 2 citations
- Design Your Ad: Personalized Advertising Image and Text Generation with Unified Autoregressive ModelsYexing Xu, Wei Feng, Shen Zhang, Haohan Wang et al.CVPR 2026 · 1 citation
Builds on15
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
Related papers
- Open-World Attribute Mining for E-Commerce Products with Multimodal Self-Correction Instruction TuningJiaqi Li, Yanming Li, Xiaoli Shen, Chuanyi Zhang et al.ACL 2025 · 2 citations
- RePIC: Reinforced Post-Training for Personalizing Multi-Modal Language ModelsYeongtak Oh, Dohyun Chung, Juhyeon Shin, Sangha Park et al.NeurIPS 2025 · 12 citations
- Boosting Chart-to-Code Generation in MLLM via Dual Preference-Guided RefinementZhihan Zhang, Yixin Cao, Lizi LiaoACM MM 2025
- Harnessing Multimodal Large Language Models for Multimodal Sequential RecommendationYuyang Ye, Zhi Zheng, Yishan Shen, Tianshu Wang et al.AAAI 2025 · 68 citations
- ICG: Improving Cover Image Generation via MLLM-based Prompting and Personalized Preference AlignmentZhipeng Bian, Jieming Zhu, Qijiong Liu, Wang Lin et al.EMNLP 2025
