VCMaster: Generating Diverse and Fluent Live Video Comments Based on Multimodal Contexts
Manman Zhang, Ge Luo, Yuchen Ma, Sheng Li, Zhenxing Qian, Xinpeng Zhang
摘要
Live video commenting, or "bullet screen," is a popular social style on video platforms. Automatic live commenting has been explored as a promising approach to enhance the appeal of videos. However, existing methods neglect the diversity of generated sentences, limiting the potential to obtain human-like comments. In this paper, we introduce a novel framework called "VCMaster" for multimodal live video comments generation, which balances the diversity and quality of generated comments to create human-like sentences. We involve images, subtitles, and contextual comments as inputs to better understand complex video contexts. Then, we propose an effective Hierarchical Cross-Fusion Decoder to integrate high-quality trimodal feature representations by cross-fusing critical information from previous layers. Additionally, we develop a Sentence-Level Contrastive Loss to enlarge the distance between generated and contextual comments by contrastive learning. It helps the model to avoid the pitfall of simply imitating provided contextual comments and losing creativity, encouraging the model to achieve more diverse comments while maintaining high quality. We also construct a large-scale multimodal live video comments dataset with 292,507 comments and three sub-datasets that cover nine general categories. Extensive experiments demonstrate that our model achieves a level of human-like language expression and remarkably fluent, diverse, and engaging generated comments compared to baselines.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- VideoIC: A Video Interactive Comments Dataset and Multimodal Multitask Learning for Comments GenerationWeiying Wang, Jieting Chen, Qin JinACM MM 2020 · 被引用 26 次
- Enhancing Multimodal Affective Analysis with Learned Live Comment FeaturesZhaoyuan Deng, Amith Ananthram, Kathleen McKeownAAAI 2025 · 被引用 4 次
- Open-domain Video Commentary GenerationEdison Marrese-Taylor, Yumi Hamazono, Tatsuya Ishigaki, Goran Topic 等EMNLP 2022 · 被引用 2 次
- Understanding Human Preferences: Towards More Personalized Video to Text GenerationYihan Wu, Ruihua Song, Xu Chen, Hao Jiang 等WWW 2024 · 被引用 6 次
- ChinaOpen: A Dataset for Open-world Multimodal LearningAozhu Chen, Ziyuan Wang, Chengbo Dong, Kaibin Tian 等ACM MM 2023 · 被引用 8 次
