Segment First or Comprehend First? Explore the Limit of Unsupervised Word Segmentation with Large Language Models
Zihong Zhang, Liqi He, Zuchao Li, Lefei Zhang, Hai Zhao, Bo Du
Abstract
Word segmentation stands as a cornerstone of Natural Language Processing (NLP). Based on the concept of"comprehend first, segment later", we propose a new framework to explore the limit of unsupervised word segmentation with Large Language Models (LLMs) and evaluate the semantic understanding capabilities of LLMs based on word segmentation. We employ current mainstream LLMs to perform word segmentation across multiple languages to assess LLMs'"comprehension". Our findings reveal that LLMs are capable of following simple prompts to segment raw text into words. There is a trend suggesting that models with more parameters tend to perform better on multiple languages. Additionally, we introduce a novel unsupervised method, termed LLACA (arge anguage Model-Inspired ho-orasick utomaton). Leveraging the advanced pattern recognition capabilities of Aho-Corasick automata, LLACA innovatively combines these with the deep insights of well-pretrained LLMs. This approach not only enables the construction of a dynamic -gram model that adjusts based on contextual information but also integrates the nuanced understanding of LLMs, offering significant improvements over traditional methods. Our source code is available at https://github.com/hkr04/LLACA
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d9d0fb49-b0a7-4dcb-aac9-1c1c1f3b5dbeBuilds on3
- Multi-Modal Latent Space Learning for Chain-of-Thought Reasoning in Language ModelsLiqi He, Zuchao Li, Xiantao Cai, Ping WangAAAI 2024 · 38 citations
- Explicit Sentence Compression for Neural Machine TranslationZuchao Li, Rui Wang, Kehai Chen, Masao Utiyama et al.AAAI 2020 · 31 citations
- Seeking Common but Distinguishing Difference, A Joint Aspect-based Sentiment Analysis ModelHongjiang Jing, Zuchao Li, Hai Zhao, Shu JiangEMNLP 2021 · 20 citations
Related papers
- Training-Free Long-Context Scaling of Large Language ModelsChenxin An, Fei Huang, Jun Zhang, Shansan Gong et al.ICML 2024 · 68 citations
- LangHOPS: Language Grounded Hierarchical Open-Vocabulary Part SegmentationYang Miao, Jan-Nico Zaech, Xi Wang, Fabien Despinoy et al.NeurIPS 2025 · 3 citations
- Do Large Language Models Understand Word Senses?Domenico Meconi, Simone Stirpe, Federico Martelli, Leonardo Lavalle et al.EMNLP 2025 · 7 citations
- HyperSeg: Hybrid Segmentation Assistant with Fine-grained Visual PerceiverCong Wei, Yujie Zhong, Haoxian Tan, Yong Liu et al.CVPR 2025
- Probing LLMs for Multilingual Discourse Generalization Through a Unified Label SetFlorian Eichin, Yang Janet Liu, Barbara Plank, Michael A. HedderichACL 2025
