Multimodal Persona Based Generation of Comic Dialogs
Harsh Agrawal, Aditya Mishra, Manish Gupta, Mausam
Abstract
We focus on the novel problem of persona based dialogue generation for comic strips. Dialogs in comic strips is a unique and unexplored area where every strip contains utterances from various characters with each one building upon the previous utterances and the associated visual scene. Previous works like Di-aloGPT, PersonaGPT and other dialog generation models encode two-party dialogues and do not account for the visual information. To the best of our knowledge we are the first to propose the paradigm of multimodal persona based dialogue generation. We contribute a novel dataset, COMSET, consisting of 54K strips, harvested from 13 popular comics available online. Further, we propose a multimodal persona-based architecture, MPDIALOG, to generate dialogues for the next panel in the strip which decreases the perplexity score by ∼10 points over strong dialogue generation baseline models. We demonstrate that there is still ample opportunity for improvement, highlighting the importance of building stronger dialogue systems that are able to generate persona-consistent dialogues and understand the context through various modalities.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e77354cf-74bb-401c-aa2b-3f9aec560a3fCited by top-tier papers2
- Sparse Activation Editing for Reliable Instruction Following in NarrativesRuncong Zhao, Chengyu Cao, Qinglin Zhu, Xiucheng Lyu et al.EMNLP 2025
- Aligning VLM Assistants with Personalized Situated CognitionYongqi Li, Shen Zhou, Xiaohu Li, Xin Miao et al.ACL 2025
Builds on15
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes et al.ICLR 2020 · 4,112 citations
- Multimodal Few-Shot Learning with Frozen Language ModelsMaria Tsimpoukelli, Jacob Menick, Serkan Cabi, S. M. Ali Eslami et al.NeurIPS 2021 · 1,020 citations
- SimVLM: Simple Visual Language Model Pretraining with Weak SupervisionZirui Wang, Jiahui Yu, Adams Wei Yu, Zihang Dai et al.ICLR 2022 · 950 citations
Related papers
- MPCHAT: Towards Multimodal Persona-Grounded ConversationJaewoo Ahn, Yeda Song, Sangdoo Yun, Gunhee KimACL 2023 · 4 citations
- Let's Go Real Talk: Spoken Dialogue Model for Face-to-Face ConversationSe Jin Park, Chae Won Kim, Hyeongseop Rha, Minsu Kim et al.ACL 2024
- Zero-Shot Character Identification and Speaker Prediction in Comics via Iterative Multimodal FusionYingxuan Li, Ryota Hinami, Kiyoharu Aizawa, Yusuke MatsuiACM MM 2024 · 3 citations
- MMDialog: A Large-scale Multi-turn Dialogue Dataset Towards Multi-modal Open-domain ConversationJiazhan Feng, Qingfeng Sun, Can Xu, Pu Zhao et al.ACL 2023 · 20 citations
- Structure-Aware Multimodal Sequential Learning for Visual DialogYoung-Jin Kim, Min-Jun Kim, Kyunghwan An, Jinwoo Ahn et al.AAAI 2024 · 3 citations
