Toward Unifying Text Segmentation and Long Document Summarization
Sangwoo Cho, Kaiqiang Song, Xiaoyang Wang, Fei Liu, Dong Yu
Abstract
Text segmentation is important for signaling a document’s structure. Without segmenting a long document into topically coherent sections, it is difficult for readers to comprehend the text, let alone find important information. The problem is only exacerbated by a lack of segmentation in transcripts of audio/video recordings. In this paper, we explore the role that section segmentation plays in extractive summarization of written and spoken documents. Our approach learns robust sentence representations by performing summarization and segmentation simultaneously, which is further enhanced by an optimization-based regularizer to promote selection of diverse summary sentences. We conduct experiments on multiple datasets ranging from scientific articles to spoken transcripts to evaluate the model’s performance. Our findings suggest that the model can not only achieve state-of-the-art performance on publicly available benchmarks, but demonstrate better cross-genre transferability when equipped with text segmentation. We perform a series of analyses to quantify the impact of section segmentation on summarizing written and spoken documents of substantial length and complexity.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7535e7ed-7e57-4a57-922a-91928c642da0Cited by top-tier papers6
- MoC: Mixtures of Text Chunking Learners for Retrieval-Augmented Generation SystemJihao Zhao, Zhiyuan Ji, Zhaoxin Fan, Hanyu Wang et al.ACL 2025 · 21 citations
- MeetingBank: A Benchmark Dataset for Meeting SummarizationYebowen Hu, Timothy Ganter, Hanieh Deilamsalehy, Franck Dernoncourt et al.ACL 2023 · 19 citations
- HiChunk: Evaluating and Enhancing Retrieval Augmented Generation with Hierarchical ChunkingWensheng Lu, Keyu Chen, Zhifeng Shen, Ruizhi Qiao et al.ACL 2026 · 10 citations
- GreedyCAS: Unsupervised Scientific Abstract Segmentation with Normalized Mutual InformationYingqiang Gao, Jessica Lam, Nianlong Gu, Richard H. R. HahnloserEMNLP 2023
- QChunker: Learning Question-Aware Text Chunking for Domain RAG via Multi-Agent DebateJihao Zhao, Daixuan Li, Pengfei Li, Shuaishuai Zu et al.WWW 2026
Builds on8
- Big Bird: Transformers for Longer SequencesManzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie et al.NeurIPS 2020 · 3,159 citations
- BRIO: Bringing Order to Abstractive SummarizationYixin Liu, Pengfei Liu, Dragomir R. Radev, Graham NeubigACL 2022 · 329 citations
- On Extractive and Abstractive Neural Document Summarization with Transformer Language ModelsJonathan Pilault, Raymond Li, Sandeep Subramanian, Chris PalEMNLP 2020 · 186 citations
- Evaluating the Factual Consistency of Abstractive Text SummarizationWojciech Kryscinski, Bryan McCann, Caiming Xiong, Richard SocherEMNLP 2020 · 67 citations
- On Faithfulness and Factuality in Abstractive SummarizationJoshua Maynez, Shashi Narayan, Bernd Bohnet, Ryan T. McDonaldACL 2020 · 54 citations
Related papers
- Multi-Granularity Interaction Network for Extractive and Abstractive Multi-Document SummarizationHanqi Jin, Tianming Wang, Xiaojun WanACL 2020 · 92 citations
- Weakly supervised discourse segmentation for multiparty oral conversationsLila Gravellier, Julie Hunter, Philippe Muller, Thomas Pellegrini et al.EMNLP 2021 · 1 citation
- PEGASUS: Pre-training with Extracted Gap-sentences for Abstractive SummarizationJingqing Zhang, Yao Zhao, Mohammad Saleh, Peter J. LiuICML 2020 · 2,453 citations
- Multi-View Sequence-to-Sequence Models with Conversational Structure for Abstractive Dialogue SummarizationJiaao Chen, Diyi YangEMNLP 2020 · 121 citations
- Two-Level Transformer and Auxiliary Coherence Modeling for Improved Text SegmentationGoran Glavas, Swapna SomasundaranAAAI 2020 · 71 citations
