TopWORDS-Seg: Simultaneous Text Segmentation and Word Discovery for Open-Domain Chinese Texts via Bayesian Inference
Changzai Pan, Maosong Sun, Ke Deng
2022年份
6被引次数
1顶会引用
摘要
Processing open-domain Chinese texts has been a critical bottleneck in computational linguistics for decades, partially because text segmentation and word discovery often entangle with each other in this challenging scenario. No existing methods yet can achieve effective text segmentation and word discovery simultaneously in open domain. This study fills in this gap by proposing a novel method called TopWORDS-Seg based on Bayesian inference, which enjoys robust performance and transparent interpretation when no training corpus and domain vocabulary are available. Advantages of TopWORDS-Seg are demonstrated by a series of experimental studies.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper1
相关 Paper
- Coupling Distant Annotation and Adversarial Training for Cross-Domain Chinese Word SegmentationNing Ding, Dingkun Long, Guangwei Xu, Muhua Zhu 等ACL 2020 · 被引用 12 次
- Advancing Multi-Criteria Chinese Word Segmentation Through Criterion Classification and DenoisingTzu-Hsuan Chou, Chun-Yi Lin, Hung-Yu KaoACL 2023 · 被引用 2 次
- A Joint Multiple Criteria Model in Transfer Learning for Cross-domain Chinese Word SegmentationKaiyu Huang, Degen Huang, Zhuang Liu, Fengran MoEMNLP 2020 · 被引用 22 次
- Shatter and Gather: Learning Referring Image Segmentation with Text SupervisionDongwon Kim, Namyup Kim, Cuiling Lan, Suha KwakICCV 2023 · 被引用 29 次
- SegFormer: A Topic Segmentation Model with Controllable Range of AttentionHaitao Bai, Pinghui Wang, Ruofei Zhang, Zhou SuAAAI 2023 · 被引用 18 次
