MMLSCU: A Dataset for Multi-modal Multi-domain Live Streaming Comment Understanding
Zixiang Meng, Qiang Gao, Di Guo, Yunlong Li, Bobo Li, Hao Fei, Shengqiong Wu, Fei Li, Chong Teng, Donghong Ji
Abstract
With the increasing popularity of live streaming, the interactions from viewers during a live streaming can provide more specific and constructive feedback for both the streamer and platform. In such scenario, the primary and most direct feedback method from the audience is through comments. Thus, mining these live streaming comments to unearth the intentions behind them and, in turn, aiding streamers to enhance their live streaming quality is significant for the well development of live streaming ecosystem. To this end, we introduce the MMLSCU dataset, containing 50,129 intention-annotated comments across multiple modalities (text, images, vi-deos, audio) from eight streaming domains. Using multimodal pretrained large model and drawing inspiration from the Chain of Thoughts (CoT) concept, we implement an end-to-end model to sequentially perform the following tasks: viewer comment intent detection ➛ intent cause mining ➛ viewer comment explanation ➛ streamer policy suggestion. We employ distinct branches for video and audio to process their respective modalities. After obtaining the video and audio representations, we conduct a multimodal fusion with the comment. This integrated data is then fed into the large language model to perform inference across the four tasks following the CoT framework. Experimental results indicate that our model outperforms three multimodal classification baselines on comment intent detection and streamer policy suggestion, and one multimodal generation baselines on intent cause mining and viewer comment explanation. Compared to the models using only text, our multimodal setting yields superior outcomes. Moreover, incorporating CoT allows our model to enhance comment interpretation and more precise suggestions for the streamers. Our proposed dataset and model will bring new research attention on multimodal live streaming comment understanding.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Cited by top-tier papers3
- SpeechEE: A Novel Benchmark for Speech Event ExtractionBin Wang, Meishan Zhang, Hao Fei, Yu Zhao et al.ACM MM 2024 · 1 citation
- Peeking Ahead of the Field Study: Exploring VLM Personas as Support Tools for Embodied Studies in HCIXinyue Gui, Ding Xia, Mark Colley, Yuan Li et al.CHI 2026 · 1 citation
- COSMMIC: Comment-Sensitive Multimodal Multilingual Indian Corpus for Summarization and Headline GenerationRaghvendra Kumar, Mohammed Salman S. A, Aryan Sahu, Tridib Nandi et al.ACL 2025
Related papers
- VideoIC: A Video Interactive Comments Dataset and Multimodal Multitask Learning for Comments GenerationWeiying Wang, Jieting Chen, Qin JinACM MM 2020 · 26 citations
- LiViBench: An Omnimodal Benchmark for Interactive Livestream Video UnderstandingXiaodong Wang, Langling Huang, Zhirong Wu, Xu Zhao et al.AAAI 2026 · 1 citation
- CogStream: Context-guided Streaming Video Question AnsweringZicheng Zhao, Kangyu Wang, Shijie Li, Rui Qian et al.AAAI 2026 · 3 citations
- ChartSketcher: Reasoning with Multimodal Feedback and Reflection for Chart UnderstandingMuye Huang, Lingling Zhang, Jie Ma, Han Lai et al.NeurIPS 2025 · 13 citations
- ThinkSound: Chain-of-Thought Reasoning in Multimodal LLMs for Audio Generation and EditingHuadai Liu, Kaicheng Luo, Jialei Wang, Wen Wang et al.NeurIPS 2025 · 5 citations
