StreamHover: Livestream Transcript Summarization and Annotation
Sangwoo Cho, Franck Dernoncourt, Tim Ganter, Trung Bui, Nedim Lipka, Walter Chang, Hailin Jin, Jonathan Brandt, Hassan Foroosh, Fei Liu
摘要
With the explosive growth of livestream broadcasting, there is an urgent need for new summarization technology that enables us to create a preview of streamed content and tap into this wealth of knowledge. However, the problem is nontrivial due to the informal nature of spoken language. Further, there has been a shortage of annotated datasets that are necessary for transcript summarization. In this paper, we present StreamHover, a framework for annotating and summarizing livestream transcripts. With a total of over 500 hours of videos annotated with both extractive and abstractive summaries, our benchmark dataset is significantly larger than currently existing annotated corpora. We explore a neural extractive summarization model that leverages vector-quantized variational autoencoder to learn latent vector representations of spoken utterances and identify salient utterances from the transcripts to form summaries. We show that our model generalizes better and improves performance over strong baselines. The results of this study provide an avenue for future research to improve summarization solutions for efficient browsing of livestreams.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- MeetingBank: A Benchmark Dataset for Meeting SummarizationYebowen Hu, Timothy Ganter, Hanieh Deilamsalehy, Franck Dernoncourt 等ACL 2023 · 被引用 19 次
- Record Once, Post Everywhere: Automatic Shortening of Audio Stories for Social MediaBryan Wang, Zeyu Jin, Gautham J. MysoreUIST 2022 · 被引用 13 次
- Reproducibility in Computational Linguistics: Is Source Code Enough?Mohammad Arvan, Luís Pina, Natalie PardeEMNLP 2022 · 被引用 12 次
- Towards Abstractive Grounded Summarization of Podcast TranscriptsKaiqiang Song, Chen Li, Xiaoyang Wang, Dong Yu 等ACL 2022 · 被引用 11 次
- Moment Detection in Long Tutorial VideosIoana Croitoru, Simion-Vlad Bogolin, Samuel Albanie, Yang Liu 等ICCV 2023 · 被引用 7 次
它引用的顶会 Paper11
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger 等ICLR 2020 · 被引用 8,443 次
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad 等ACL 2020 · 被引用 1,224 次
- Heterogeneous Graph Neural Networks for Extractive Document SummarizationDanqing Wang, Pengfei Liu, Yining Zheng, Xipeng Qiu 等ACL 2020 · 被引用 275 次
- Transformer-XH: Multi-Evidence Reasoning with eXtra Hop AttentionChen Zhao, Chenyan Xiong, Corby Rosset, Xia Song 等ICLR 2020 · 被引用 120 次
- VMSMO: Learning to Generate Multimodal Summary for Video-based News ArticlesMingzhe Li, Xiuying Chen, Shen Gao, Zhangming Chan 等EMNLP 2020 · 被引用 65 次
相关 Paper
- Toward Unifying Text Segmentation and Long Document SummarizationSangwoo Cho, Kaiqiang Song, Xiaoyang Wang, Fei Liu 等EMNLP 2022 · 被引用 19 次
- CatchLive: Real-time Summarization of Live Streams with Stream Content and Interaction DataSaelyne Yang, Jisu Yim, Juho Kim, Hijung Valentina ShinCHI 2022 · 被引用 29 次
- Unsupervised Abstractive Dialogue Summarization for Tete-a-TetesXinyuan Zhang, Ruiyi Zhang, Manzil Zaheer, Amr AhmedAAAI 2021 · 被引用 27 次
- CLIP-It! Language-Guided Video SummarizationMedhini Narasimhan, Anna Rohrbach, Trevor DarrellNeurIPS 2021 · 被引用 196 次
- Streaming Video Understanding and Multi-round Interaction with Memory-enhanced KnowledgeHaomiao Xiong, Zongxin Yang, Jiazuo Yu, Yunzhi Zhuge 等ICLR 2025
