LLM-SLM Collaborative Framework of Idiomatic Expression Generation
Hui Gao, Changhao Song, Peng Zhang, Jing Zhang, Chang Yang, Liuxian Ge
摘要
Idiomatic Expression Generation, which aims to produce idiomatic text from plain text, is a valuable yet challenging NLP task. However, existing methods suffer from the scarcity of parallel data and dependence on high-quality manual annotations. To address this, we propose an iterative LLM-SLM (Large Language Model-Small Language Model) collaborative framework-Auto-IDEA, that replaces human supervision for idiomatic expression data generation. In this self-improving cycle, the LLM constructs parallel corpora (idiomatic and plain text) via bidirectional semantic reconstruction, automatically generating "Locate-Then-Polish" (LTP) annotations; the SLM filters low-quality corpora while continuously enhancing its verification ability through incremental learning. We instantiate Auto-IDEA for Chinese Idiom Polishing (CIP), constructing CIP-200K, a largescale dataset of 206K parallel sentences with LTP annotations. The Qwen3-8B fine-tuned on CIP-200K achieves a 25.2% absolute Idiom Polishing Accuracy (IPA) improvement over a supervised fine-tuning (SFT) baseline, outperforming DeepSeek-R1 by 6.2%. Extensive experiments (e.g., Chinese idiom cloze tests and English idiom generation tasks) and human evaluations verify the generalization and effectiveness of Auto-IDEA, demonstrating a new pathway for high-quality, annotation-free data generation through LLM-SLM collaboration.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper7
- WizardCoder: Empowering Code Large Language Models with Evol-InstructZiyang Luo, Can Xu, Pu Zhao, Qingfeng Sun 等ICLR 2024 · 被引用 945 次
- Large Language Models for Data Annotation and Synthesis: A SurveyZhen Tan, Dawei Li, Song Wang, Alimohammad Beigi 等EMNLP 2024 · 被引用 119 次
- NaturalConv: A Chinese Dialogue Dataset Towards Multi-turn Topic-driven ConversationXiaoyang Wang, Chen Li, Jianqiao Zhao, Dong YuAAAI 2021 · 被引用 54 次
- Translate Meanings, Not Just Words: IdiomKB's Role in Optimizing Idiomatic Translation with Language ModelsShuang Li, Jiangjie Chen, Siyu Yuan, Xinyi Wu 等AAAI 2024 · 被引用 44 次
- FreeAL: Towards Human-Free Active Learning in the Era of Large Language ModelsRuixuan Xiao, Yiwen Dong, Junbo Zhao, Runze Wu 等EMNLP 2023 · 被引用 9 次
相关 Paper
- Idiomatic Expression Paraphrasing without Strong SupervisionJianing Zhou, Ziheng Zeng, Hongyu Gong, Suma BhatAAAI 2022 · 被引用 12 次
- An Iterative Polishing Framework Based on Quality Aware Masked Language Model for Chinese Poetry GenerationLiming Deng, Jie Wang, Hang-Ming Liang, Hui Chen 等AAAI 2020 · 被引用 26 次
- Writing Polishment with Simile: Task, Dataset and A Neural ApproachJiayi Zhang, Zhi Cui, Xiaoqiang Xia, Yalong Guo 等AAAI 2021 · 被引用 20 次
- ScholarGEC: Enhancing Controllability of Large Language Model for Chinese Academic Grammatical Error CorrectionZixiao Kong, Xianquan Wang, Shuanghong Shen, Keyu Zhu 等AAAI 2025 · 被引用 2 次
- Inverse IFEval: Can LLMs Unlearn Stubborn Training Conventions to Follow Real Instructions?Qinyan Zhang, Xinping Lei, Ruijie Miao, FU YU 等ICLR 2026 · 被引用 10 次
