LLM-SLM Collaborative Framework of Idiomatic Expression Generation
Hui Gao, Changhao Song, Peng Zhang, Jing Zhang, Chang Yang, Liuxian Ge
Abstract
Idiomatic Expression Generation, which aims to produce idiomatic text from plain text, is a valuable yet challenging NLP task. However, existing methods suffer from the scarcity of parallel data and dependence on high-quality manual annotations. To address this, we propose an iterative LLM-SLM (Large Language Model-Small Language Model) collaborative framework-Auto-IDEA, that replaces human supervision for idiomatic expression data generation. In this self-improving cycle, the LLM constructs parallel corpora (idiomatic and plain text) via bidirectional semantic reconstruction, automatically generating "Locate-Then-Polish" (LTP) annotations; the SLM filters low-quality corpora while continuously enhancing its verification ability through incremental learning. We instantiate Auto-IDEA for Chinese Idiom Polishing (CIP), constructing CIP-200K, a largescale dataset of 206K parallel sentences with LTP annotations. The Qwen3-8B fine-tuned on CIP-200K achieves a 25.2% absolute Idiom Polishing Accuracy (IPA) improvement over a supervised fine-tuning (SFT) baseline, outperforming DeepSeek-R1 by 6.2%. Extensive experiments (e.g., Chinese idiom cloze tests and English idiom generation tasks) and human evaluations verify the generalization and effectiveness of Auto-IDEA, demonstrating a new pathway for high-quality, annotation-free data generation through LLM-SLM collaboration.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b257df33-a91b-4583-927f-d31a0b7a45ddBuilds on7
- WizardCoder: Empowering Code Large Language Models with Evol-InstructZiyang Luo, Can Xu, Pu Zhao, Qingfeng Sun et al.ICLR 2024 · 945 citations
- Large Language Models for Data Annotation and Synthesis: A SurveyZhen Tan, Dawei Li, Song Wang, Alimohammad Beigi et al.EMNLP 2024 · 119 citations
- NaturalConv: A Chinese Dialogue Dataset Towards Multi-turn Topic-driven ConversationXiaoyang Wang, Chen Li, Jianqiao Zhao, Dong YuAAAI 2021 · 54 citations
- Translate Meanings, Not Just Words: IdiomKB's Role in Optimizing Idiomatic Translation with Language ModelsShuang Li, Jiangjie Chen, Siyu Yuan, Xinyi Wu et al.AAAI 2024 · 44 citations
- FreeAL: Towards Human-Free Active Learning in the Era of Large Language ModelsRuixuan Xiao, Yiwen Dong, Junbo Zhao, Runze Wu et al.EMNLP 2023 · 9 citations
Related papers
- Idiomatic Expression Paraphrasing without Strong SupervisionJianing Zhou, Ziheng Zeng, Hongyu Gong, Suma BhatAAAI 2022 · 12 citations
- An Iterative Polishing Framework Based on Quality Aware Masked Language Model for Chinese Poetry GenerationLiming Deng, Jie Wang, Hang-Ming Liang, Hui Chen et al.AAAI 2020 · 26 citations
- Writing Polishment with Simile: Task, Dataset and A Neural ApproachJiayi Zhang, Zhi Cui, Xiaoqiang Xia, Yalong Guo et al.AAAI 2021 · 20 citations
- ScholarGEC: Enhancing Controllability of Large Language Model for Chinese Academic Grammatical Error CorrectionZixiao Kong, Xianquan Wang, Shuanghong Shen, Keyu Zhu et al.AAAI 2025 · 2 citations
- Inverse IFEval: Can LLMs Unlearn Stubborn Training Conventions to Follow Real Instructions?Qinyan Zhang, Xinping Lei, Ruijie Miao, FU YU et al.ICLR 2026 · 10 citations
