Open-domain Video Commentary Generation
Edison Marrese-Taylor, Yumi Hamazono, Tatsuya Ishigaki, Goran Topic, Yusuke Miyao, Ichiro Kobayashi, Hiroya Takamura
Abstract
Live commentary plays an important role in sports broadcasts and video games, making spectators more excited and immersed. In this context, though approaches for automatically generating such commentary have been proposed in the past, they have been generally concerned with specific fields, where it is possible to leverage domain-specific information. In light of this, we propose the task of generating video commentary in an open-domain fashion. We detail the construction of a new large-scale dataset of transcribed commentary aligned with videos containing various human actions in a variety of domains, and propose approaches based on well-known neural architectures to tackle the task. To understand the strengths and limitations of current approaches, we present an in-depth empirical study based on our data. Our results suggest clear trade-offs between textual and visual inputs for the models and highlight the importance of relying on external knowledge in this open-domain setting, resulting in a set of robust baselines for our task.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 58eeecd0-527a-4df8-a40e-ca35fa59a667Builds on5
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- End-to-End Dense Video Captioning with Parallel DecodingTeng Wang, Ruimao Zhang, Zhichao Lu, Feng Zheng et al.ICCV 2021 · 238 citations
- Fast Learning of Temporal Action Proposal via Dense Boundary GeneratorChuming Lin, Jian Li, Yabiao Wang, Ying Tai et al.AAAI 2020 · 226 citations
- VLN BERT: A Recurrent Vision-and-Language BERT for NavigationYicong Hong, Qi Wu, Yuankai Qi, Cristian Rodriguez Opazo et al.CVPR 2021
- Sketch, Ground, and Refine: Top-Down Dense Video CaptioningChaorui Deng, Shizhe Chen, Da Chen, Yuan He et al.CVPR 2021
Related papers
- A Descriptive Basketball Highlight Dataset for Automatic Commentary GenerationBenhui Zhang, Junyu Gao, Yuan YuanACM MM 2024 · 14 citations
- Synchronized Video Storytelling: Generating Video Narrations with Structured StorylineDingyi Yang, Chunru Zhan, Ziheng Wang, Biao Wang et al.ACL 2024
- MatchTime: Towards Automatic Soccer Game Commentary GenerationJiayuan Rao, Haoning Wu, Chang Liu, Yanfeng Wang et al.EMNLP 2024 · 14 citations
- VideoIC: A Video Interactive Comments Dataset and Multimodal Multitask Learning for Comments GenerationWeiying Wang, Jieting Chen, Qin JinACM MM 2020 · 26 citations
- VCMaster: Generating Diverse and Fluent Live Video Comments Based on Multimodal ContextsManman Zhang, Ge Luo, Yuchen Ma, Sheng Li et al.ACM MM 2023 · 3 citations
