Optimizing Chinese Lexical Simplification Across Word Types: A Hybrid Approach
Zihao Xiao, Jiefu Gong, Shijin Wang, Wei Song
Abstract
This paper addresses the task of Chinese Lexical Simplification (CLS). A key challenge in CLS is the scarcity of data resources. We begin by evaluating the performance of various language models at different scales in unsupervised and few-shot settings, finding that their effectiveness is sensitive to word types. Expensive large language models (LLMs), such as GPT-4, outperform small models in simplifying complex content words and Chinese idioms from the dictionary. To take advantage of this, we propose an automatic knowledge distillation framework called PivotKD for generating training data to fine-tune small models. In addition, all models face difficulties with out-ofdictionary (OOD) words such as internet slang. To address this, we implement a retrieval-based interpretation augmentation (RIA) strategy, injecting word interpretations from external resources into the context. Experimental results demonstrate that fine-tuned small models outperform GPT-4 in simplifying complex content words and Chinese idioms. Additionally, the RIA strategy enhances the performance of most models, particularly in handling OOD words. Our findings suggest that a hybrid approach could optimize CLS performance while managing inference costs. This would involve configuring choices such as model scale, linguistic resources, and the use of RIA based on specific word types to strike an ideal balance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on3
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Lexical Simplification with Pretrained EncodersJipeng Qiang, Yun Li, Yi Zhu, Yunhao Yuan et al.AAAI 2020 · 86 citations
Related papers
- A New Dataset and Empirical Study for Sentence Simplification in ChineseShiping Yang, Renliang Sun, Xiaojun WanACL 2023 · 4 citations
- In-context Mixing (ICM): Code-mixed Prompts for Multilingual LLMsBhavani Shankar, Preethi Jyothi, Pushpak BhattacharyyaACL 2024
- Investigating Large Language Models for Complex Word Identification in Multilingual and Multidomain SetupsRazvan-Alexandru Smadu, David-Gabriel Ion, Dumitru-Clementin Cercel, Florin Pop et al.EMNLP 2024 · 4 citations
- A Systematic Study of Knowledge Distillation for Natural Language Generation with Pseudo-Target TrainingNitay Calderon, Subhabrata Mukherjee, Roi Reichart, Amir KantorACL 2023 · 5 citations
- MiniPLM: Knowledge Distillation for Pre-training Language ModelsYuxian Gu, Hao Zhou, Fandong Meng, Jie Zhou et al.ICLR 2025
