Two-Level Transformer and Auxiliary Coherence Modeling for Improved Text Segmentation
Goran Glavas, Swapna Somasundaran
摘要
Breaking down the structure of long texts into semantically coherent segments makes the texts more readable and supports downstream applications like summarization and retrieval. Starting from an apparent link between text coherence and segmentation, we introduce a novel supervised model for text segmentation with simple but explicit coherence modeling. Our model – a neural architecture consisting of two hierarchically connected Transformer networks – is a multi-task learning model that couples the sentence-level segmentation objective with the coherence objective that differentiates correct sequences of sentences from corrupt ones. The proposed model, dubbed Coherence-Aware Text Segmentation (CATS), yields state-of-the-art segmentation performance on a collection of benchmark datasets. Furthermore, by coupling CATS with cross-lingual word embeddings, we demonstrate its effectiveness in zero-shot language transfer: it can successfully segment texts in languages unseen in training.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Hierarchical Macro Discourse Parsing Based on Topic SegmentationFeng Jiang, Yaxin Fan, Xiaomin Chu, Peifeng Li 等AAAI 2021 · 被引用 14 次
- Improving Long Document Topic Segmentation Models With Enhanced Coherence ModelingHai Yu, Chong Deng, Qinglin Zhang, Jiaqing Liu 等EMNLP 2023 · 被引用 7 次
- Human Guided Exploitation of Interpretable Attention Patterns in Summarization and Topic SegmentationRaymond Li, Wen Xiao, Linzi Xing, Lanjun Wang 等EMNLP 2022 · 被引用 4 次
- Topical Segmentation of Spoken Narratives: A Test Case on Holocaust Survivor TestimoniesEitan Wagner, Renana Keydar, Amit Pinchevski, Omri AbendEMNLP 2022 · 被引用 3 次
- SuperDialseg: A Large-scale Dataset for Supervised Dialogue SegmentationJunfeng Jiang, Chengzhang Dong, Sadao Kurohashi, Akiko AizawaEMNLP 2023 · 被引用 2 次
相关 Paper
- A Joint Model for Document Segmentation and Segment LabelingJoe Barrow, Rajiv Jain, Vlad I. Morariu, Varun Manjunatha 等ACL 2020 · 被引用 47 次
- Towards Global Video Scene Segmentation with Context-Aware TransformerYang Yang, Yurui Huang, Weili Guo, Baohua Xu 等AAAI 2023 · 被引用 34 次
- ReSTR: Convolution-free Referring Image Segmentation Using TransformersNamyup Kim, Dongwon Kim, Suha Kwak, Cuiling Lan 等CVPR 2022 · 被引用 149 次
- Language-driven Semantic SegmentationBoyi Li, Kilian Q. Weinberger, Serge J. Belongie, Vladlen Koltun 等ICLR 2022 · 被引用 885 次
- Toward Unifying Text Segmentation and Long Document SummarizationSangwoo Cho, Kaiqiang Song, Xiaoyang Wang, Fei Liu 等EMNLP 2022 · 被引用 19 次
