A Joint Model for Document Segmentation and Segment Labeling
Joe Barrow, Rajiv Jain, Vlad I. Morariu, Varun Manjunatha, Douglas W. Oard, Philip Resnik
摘要
Text segmentation aims to uncover latent structure by dividing text from a document into coherent sections. Where previous work on text segmentation considers the tasks of document segmentation and segment labeling separately, we show that the tasks contain complementary information and are best addressed jointly. We introduce the Segment Pooling LSTM (S-LSTM) model, which is capable of jointly segmenting a document and labeling segments. In support of joint training, we develop a method for teaching the model to recover from errors by aligning the predicted and ground truth segments. We show that S-LSTM reduces segmentation error by 30% on average, while also improving segment labeling.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- SegFormer: A Topic Segmentation Model with Controllable Range of AttentionHaitao Bai, Pinghui Wang, Ruofei Zhang, Zhou SuAAAI 2023 · 被引用 18 次
- Hierarchical Macro Discourse Parsing Based on Topic SegmentationFeng Jiang, Yaxin Fan, Xiaomin Chu, Peifeng Li 等AAAI 2021 · 被引用 14 次
- Improving Long Document Topic Segmentation Models With Enhanced Coherence ModelingHai Yu, Chong Deng, Qinglin Zhang, Jiaqing Liu 等EMNLP 2023 · 被引用 7 次
- SuperDialseg: A Large-scale Dataset for Supervised Dialogue SegmentationJunfeng Jiang, Chengzhang Dong, Sadao Kurohashi, Akiko AizawaEMNLP 2023 · 被引用 2 次
- GreedyCAS: Unsupervised Scientific Abstract Segmentation with Normalized Mutual InformationYingqiang Gao, Jessica Lam, Nianlong Gu, Richard H. R. HahnloserEMNLP 2023
相关 Paper
- Two-Level Transformer and Auxiliary Coherence Modeling for Improved Text SegmentationGoran Glavas, Swapna SomasundaranAAAI 2020 · 被引用 71 次
- Efficient Transformers with Dynamic Token PoolingPiotr Nawrot, Jan Chorowski, Adrian Lancucki, Edoardo Maria PontiACL 2023 · 被引用 14 次
- A Simple yet Effective Layout Token in Large Language Models for Document UnderstandingZhaoqing Zhu, Chuwei Luo, Zirui Shao, Feiyu Gao 等CVPR 2025
- A Sequence-to-Sequence Approach with Mixed Pointers to Topic Segmentation and Segment LabelingJinxiong Xia, Houfeng WangKDD 2023 · 被引用 3 次
- Toward Unifying Text Segmentation and Long Document SummarizationSangwoo Cho, Kaiqiang Song, Xiaoyang Wang, Fei Liu 等EMNLP 2022 · 被引用 19 次
