StreamHover: Livestream Transcript Summarization and Annotation
Sangwoo Cho, Franck Dernoncourt, Tim Ganter, Trung Bui, Nedim Lipka, Walter Chang, Hailin Jin, Jonathan Brandt, Hassan Foroosh, Fei Liu
Abstract
With the explosive growth of livestream broadcasting, there is an urgent need for new summarization technology that enables us to create a preview of streamed content and tap into this wealth of knowledge. However, the problem is nontrivial due to the informal nature of spoken language. Further, there has been a shortage of annotated datasets that are necessary for transcript summarization. In this paper, we present StreamHover, a framework for annotating and summarizing livestream transcripts. With a total of over 500 hours of videos annotated with both extractive and abstractive summaries, our benchmark dataset is significantly larger than currently existing annotated corpora. We explore a neural extractive summarization model that leverages vector-quantized variational autoencoder to learn latent vector representations of spoken utterances and identify salient utterances from the transcripts to form summaries. We show that our model generalizes better and improves performance over strong baselines. The results of this study provide an avenue for future research to improve summarization solutions for efficient browsing of livestreams.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9b90c780-b20e-40f1-9b6e-da5ab504997eCited by top-tier papers8
- MeetingBank: A Benchmark Dataset for Meeting SummarizationYebowen Hu, Timothy Ganter, Hanieh Deilamsalehy, Franck Dernoncourt et al.ACL 2023 · 19 citations
- Record Once, Post Everywhere: Automatic Shortening of Audio Stories for Social MediaBryan Wang, Zeyu Jin, Gautham J. MysoreUIST 2022 · 13 citations
- Reproducibility in Computational Linguistics: Is Source Code Enough?Mohammad Arvan, Luís Pina, Natalie PardeEMNLP 2022 · 12 citations
- Towards Abstractive Grounded Summarization of Podcast TranscriptsKaiqiang Song, Chen Li, Xiaoyang Wang, Dong Yu et al.ACL 2022 · 11 citations
- Moment Detection in Long Tutorial VideosIoana Croitoru, Simion-Vlad Bogolin, Samuel Albanie, Yang Liu et al.ICCV 2023 · 7 citations
Builds on11
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- Heterogeneous Graph Neural Networks for Extractive Document SummarizationDanqing Wang, Pengfei Liu, Yining Zheng, Xipeng Qiu et al.ACL 2020 · 275 citations
- Transformer-XH: Multi-Evidence Reasoning with eXtra Hop AttentionChen Zhao, Chenyan Xiong, Corby Rosset, Xia Song et al.ICLR 2020 · 120 citations
- VMSMO: Learning to Generate Multimodal Summary for Video-based News ArticlesMingzhe Li, Xiuying Chen, Shen Gao, Zhangming Chan et al.EMNLP 2020 · 65 citations
Related papers
- Toward Unifying Text Segmentation and Long Document SummarizationSangwoo Cho, Kaiqiang Song, Xiaoyang Wang, Fei Liu et al.EMNLP 2022 · 19 citations
- CatchLive: Real-time Summarization of Live Streams with Stream Content and Interaction DataSaelyne Yang, Jisu Yim, Juho Kim, Hijung Valentina ShinCHI 2022 · 29 citations
- Unsupervised Abstractive Dialogue Summarization for Tete-a-TetesXinyuan Zhang, Ruiyi Zhang, Manzil Zaheer, Amr AhmedAAAI 2021 · 27 citations
- CLIP-It! Language-Guided Video SummarizationMedhini Narasimhan, Anna Rohrbach, Trevor DarrellNeurIPS 2021 · 196 citations
- Streaming Video Understanding and Multi-round Interaction with Memory-enhanced KnowledgeHaomiao Xiong, Zongxin Yang, Jiazuo Yu, Yunzhi Zhuge et al.ICLR 2025
