Attention Is All You Need for Chinese Word Segmentation
Sufeng Duan, Hai Zhao
摘要
Taking greedy decoding algorithm as it should be, this work focuses on further strengthening the model itself for Chinese word segmentation (CWS), which results in an even more fast and more accurate CWS model. Our model consists of an attention only stacked encoder and a light enough decoder for the greedy segmentation plus two highway connections for smoother training, in which the encoder is composed of a newly proposed Transformer variant, Gaussian-masked Directional (GD) Transformer, and a biaffine attention scorer. With the effective encoder design, our model only needs to take unigram features for scoring. Our model is evaluated on SIGHAN Bakeoff benchmark datasets. The experimental results show that with the highest segmentation speed, the proposed model achieves new state-of-the-art or comparable performance against strong baselines in terms of strict closed test setting.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper2
相关 Paper
- A Joint Multiple Criteria Model in Transfer Learning for Cross-domain Chinese Word SegmentationKaiyu Huang, Degen Huang, Zhuang Liu, Fengran MoEMNLP 2020 · 被引用 22 次
- Advancing Multi-Criteria Chinese Word Segmentation Through Criterion Classification and DenoisingTzu-Hsuan Chou, Chun-Yi Lin, Hung-Yu KaoACL 2023 · 被引用 2 次
- Joint Chinese Word Segmentation and Part-of-speech Tagging via Two-way Attentions of Auto-analyzed KnowledgeYuanhe Tian, Yan Song, Xiang Ao, Fei Xia 等ACL 2020 · 被引用 50 次
- SCTNet: Single-Branch CNN with Transformer Semantic Information for Real-Time SegmentationZhengze Xu, Dongyue Wu, Changqian Yu, Xiangxiang Chu 等AAAI 2024 · 被引用 166 次
- WB-DETR: Transformer-Based Detector without BackboneFanfan Liu, Haoran Wei, Wenzhe Zhao, Guozhen Li 等ICCV 2021 · 被引用 45 次
