Where's the Point? Self-Supervised Multilingual Punctuation-Agnostic Sentence Segmentation
Benjamin Minixhofer, Jonas Pfeiffer, Ivan Vulic
Abstract
Many NLP pipelines split text into sentences as one of the crucial preprocessing steps. Prior sentence segmentation tools either rely on punctuation or require a considerable amount of sentence-segmented training data: both central assumptions might fail when porting sentence segmenters to diverse languages on a massive scale. In this work, we thus introduce a multilingual punctuation-agnostic sentence segmentation method, currently covering 85 languages, trained in a self-supervised fashion on unsegmented text, by making use of newline characters which implicitly perform segmentation into paragraphs. We further propose an approach that adapts our method to the segmentation in a given corpus by using only a small number (64-256) of sentence-segmented examples. The main results indicate that our method outperforms all the prior best sentence-segmentation tools by an average of 6.1% F1 points. Furthermore, we demonstrate that proper sentence segmentation has a point: the use of a (powerful) sentence segmenter makes a considerable difference for a downstream application such as machine translation (MT). By using our method to match sentence segmentation to the segmentation used during training of MT models, we achieve an average improvement of 2.3 BLEU points over the best prior segmentation tool, as well as massive gains over a trivial segmenter that splits text into equally sized blocks. * Work done during the time BM interned at Cohere. † Equal senior authorship. Collection Sentences UD This is the high season for tourism; | between December and April few people visit and many tour companies and restaurants close down. OPUS100 'I couldn't help it,' said Five, in a sulky tone; 'Seven jogged my elbow.' | On which Seven looked up and said, 'That's right, Five! Always lay the blame (...)!' Ersatz "A lot of people would like to go back to 1970," before program trading, he said. | "I would like to go back to 1970. | But we're not going back (...
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers7
- Segment Any Text: A Universal Approach for Robust, Efficient and Adaptable Sentence SegmentationMarkus Frohmann, Igor Sterner, Ivan Vulic, Benjamin Minixhofer et al.EMNLP 2024 · 10 citations
- The Harmonic Structure of Information ContoursEleftheria Tsipidi, Samuel Kiegeland, Franz Nowak, Tianyang Xu et al.ACL 2025 · 6 citations
- CompoundPiece: Evaluating and Improving Decompounding Performance of Language ModelsBenjamin Minixhofer, Jonas Pfeiffer, Ivan VulicEMNLP 2023 · 3 citations
- Beyond Transcripts: A Renewed Perspective on Audio ChapteringFabian Retkowski, Maike Züfle, Thai-Binh Nguyen, Jan Niehues et al.ACL 2026 · 2 citations
- Scaling Low-Resource MT via Synthetic Data Generation with LLMsOna de Gibert, Joseph Attieh, Teemu Vahtola, Mikko Aulamo et al.EMNLP 2025 · 2 citations
Builds on6
- Improving Massively Multilingual Neural Machine Translation and Zero-Shot TranslationBiao Zhang, Philip Williams, Ivan Titov, Rico SennrichACL 2020 · 213 citations
- Charformer: Fast Character Transformers via Gradient-based Subword TokenizationYi Tay, Vinh Q. Tran, Sebastian Ruder, Jai Prakash Gupta et al.ICLR 2022 · 198 citations
- Fast, Effective, and Self-Supervised: Transforming Masked Language Models into Universal Lexical and Sentence EncodersFangyu Liu, Ivan Vulic, Anna Korhonen, Nigel CollierEMNLP 2021 · 85 citations
- Making Monolingual Sentence Embeddings Multilingual using Knowledge DistillationNils Reimers, Iryna GurevychEMNLP 2020 · 54 citations
- How Good is Your Tokenizer? On the Monolingual Performance of Multilingual Language ModelsPhillip Rust, Jonas Pfeiffer, Ivan Vulic, Sebastian Ruder et al.ACL 2021
Related papers
- A unified approach to sentence segmentation of punctuated text in many languagesRachel Wicks, Matt PostACL 2021
- Not Low-Resource Anymore: Aligner Ensembling, Batch Filtering, and New Datasets for Bengali-English Machine TranslationTahmid Hasan, Abhik Bhattacharjee, Kazi Samin, Masum Hasan et al.EMNLP 2020 · 7 citations
- Language-agnostic Representation from Multilingual Sentence Encoders for Cross-lingual Similarity EstimationNattapong Tiyajamorn, Tomoyuki Kajiwara, Yuki Arase, Makoto OnizukaEMNLP 2021 · 16 citations
- Two-Level Transformer and Auxiliary Coherence Modeling for Improved Text SegmentationGoran Glavas, Swapna SomasundaranAAAI 2020 · 71 citations
- Language-agnostic BERT Sentence EmbeddingFangxiaoyu Feng, Yinfei Yang, Daniel Cer, Naveen Arivazhagan et al.ACL 2022
