A unified approach to sentence segmentation of punctuated text in many languages
Rachel Wicks, Matt Post
摘要
The sentence is a fundamental unit of text processing. Yet sentences in the wild are commonly encountered not in isolation, but unsegmented within larger paragraphs and documents. Therefore, the first step in many NLP pipelines is sentence segmentation. Despite its importance, this step is the subject of relatively little research. There are no standard test sets or even methods for evaluation, leaving researchers and engineers without a clear footing for evaluating and selecting models for the task. Existing tools have relatively small language coverage, and efforts to extend them to other languages are often ad hoc. We introduce a modern context-based modeling approach that provides a solution to the problem of segmenting punctuated text in many languages, and show how it can be trained on noisily-annotated data. We also establish a new 23-language multilingual evaluation set. Our approach exceeds high baselines set by existing methods on prior English corpora (WSJ and Brown corpora), and also performs well on average on our new evaluation set. We release our tool, ERSATZ, as open source.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Segment Any Text: A Universal Approach for Robust, Efficient and Adaptable Sentence SegmentationMarkus Frohmann, Igor Sterner, Ivan Vulic, Benjamin Minixhofer 等EMNLP 2024 · 被引用 10 次
- Where's the Point? Self-Supervised Multilingual Punctuation-Agnostic Sentence SegmentationBenjamin Minixhofer, Jonas Pfeiffer, Ivan VulicACL 2023 · 被引用 8 次
- Measuring, Modeling, and Helping People Account for Privacy Risks in Online Self-Disclosures with AIIsadora Krsek, Anubha Kabra, Yao Dou, Tarek Naous 等CSCW 2025 · 被引用 6 次
- Automatic sentence segmentation of clinical record narratives in real-world dataDongfang Xu, Davy Weissenbacher, Karen O'Connor, Siddharth Rawal 等EMNLP 2024 · 被引用 1 次
- Extending Automatic Machine Translation Evaluation to Book-Length DocumentsKuang-Da Wang, Shuoyang Ding, Chao-Han Huck Yang, Ping-Chun Hsieh 等EMNLP 2025
它引用的顶会 Paper1
相关 Paper
- Improving Segmentation for Technical Support ProblemsKushal Chauhan, Abhirut GuptaACL 2020
- SegFormer: A Topic Segmentation Model with Controllable Range of AttentionHaitao Bai, Pinghui Wang, Ruofei Zhang, Zhou SuAAAI 2023 · 被引用 18 次
- CommonLID: Re-evaluating State-of-the-Art Language Identification Performance on Web DataPedro Ortiz Suarez, Laurie Burchell, Catherine Arnett, Rafael Mosquera 等ACL 2026 · 被引用 5 次
- XL-WiC: A Multilingual Benchmark for Evaluating Semantic ContextualizationAlessandro Raganato, Tommaso Pasini, José Camacho-Collados, Mohammad Taher PilehvarEMNLP 2020 · 被引用 2 次
- PAUSE: Positive and Annealed Unlabeled Sentence EmbeddingLele Cao, Emil Larsson, Vilhelm von Ehrenheim, Dhiana Deva Cavalcanti Rocha 等EMNLP 2021
