Multi-Modal Experience Inspired AI Creation
Qian Cao, Xu Chen, Ruihua Song, Hao Jiang, Guang Yang, Zhao Cao
Abstract
AI creation, such as poem or lyrics generation, has attracted increasing attention from both industry and academic communities, with many promising models proposed in the past few years. Existing methods usually estimate the outputs based on single and independent visual or textual information. However, in reality, humans usually make creations according to their experiences, which may involve different modalities and be sequentially correlated. To model such human capabilities, in this paper, we define and solve a novel AI creation problem based on human experiences. More specifically, we study how to generate texts based on sequential multi-modal information. Compared with the previous works, this task is much more difficult because the designed model has to well understand and adapt the semantics among different modalities and effectively convert them into the output in a sequential manner. To alleviate these difficulties, we firstly design a multi-channel sequence-to-sequence architecture equipped with a multi-modal attention network. For more effective optimization, we then propose a curriculum negative sampling strategy tailored for the sequential inputs. To benchmark this problem and demonstrate the effectiveness of our model, we manually labeled a new multi-modal experience dataset. With this dataset, we conduct extensive experiments by comparing our model with a series of representative baselines, where we can demonstrate significant improvements in our model based on both automatic and human-centered metrics. The code and data are available at: https://github.com/Aman-4-Real/MMTG.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8dc50f8b-d99c-460f-96bd-0b586b9a72fbCited by top-tier papers2
- Evaluating Text Creativity across Diverse Domains: a Dataset and Large Language Model EvaluatorQian Cao, Xiting Wang, Yuzhuo Yuan, Yahui Liu et al.ICLR 2026 · 11 citations
- See or Guess: Counterfactually Regularized Image CaptioningQian Cao, Xu Chen, Ruihua Song, Xiting Wang et al.ACM MM 2024 · 7 citations
Builds on11
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- Unicoder-VL: A Universal Encoder for Vision and Language by Cross-Modal Pre-TrainingGen Li, Nan Duan, Yuejian Fang, Ming Gong et al.AAAI 2020 · 966 citations
- Contrastive Learning with Adversarial Perturbations for Conditional Text GenerationSeanie Lee, Dong Bok Lee, Sung Ju HwangICLR 2021 · 117 citations
Related papers
- AI-Lyricist: Generating Music and Vocabulary Constrained LyricsXichu Ma, Ye Wang, Min-Yen Kan, Wee Sun LeeACM MM 2021 · 22 citations
- Preference Adaptive and Sequential Text-to-Image GenerationOfir Nabati, Guy Tennenholtz, Chih-Wei Hsu, Moonkyung Ryu et al.ICML 2025
- Sequential Attention GAN for Interactive Image EditingYu Cheng, Zhe Gan, Yitong Li, Jingjing Liu et al.ACM MM 2020 · 74 citations
- Unified Discrete Diffusion for Simultaneous Vision-Language GenerationMinghui Hu, Chuanxia Zheng, Zuopeng Yang, Tat-Jen Cham et al.ICLR 2023 · 8 citations
- AtHom: Two Divergent Attentions Stimulated By Homomorphic Training in Text-to-Image SynthesisZhenbo Shi, Zhi Chen, Zhenbo Xu, Wei Yang et al.ACM MM 2022 · 3 citations
