AugNLG: Few-shot Natural Language Generation using Self-trained Data Augmentation
Xinnuo Xu, Guoyin Wang, Young-Bum Kim, Sungjin Lee
Abstract
Natural Language Generation (NLG) is a key component in a task-oriented dialogue system, which converts the structured meaning representation (MR) to the natural language. For large-scale conversational systems, where it is common to have over hundreds of intents and thousands of slots, neither template-based approaches nor model-based approaches are scalable. Recently, neural NLGs started leveraging transfer learning and showed promising results in few-shot settings. This paper proposes AUGNLG, a novel data augmentation approach that combines a self-trained neural retrieval model with a few-shot learned NLU model, to automatically create MR-to-Text data from open-domain texts. The proposed system mostly outperforms the state-ofthe-art methods on the FEWSHOTWOZ data in both BLEU and Slot Error Rate. We further confirm improved results on the FEW-SHOTSGD data and provide comprehensive analysis results on key components of our system. Our code and data are available at https: //github.com/XinnuoXu/AugNLG.
MR: inform_no_match (kidsallowed = yes) TEXT: I cannot find restaurants with kids allowed.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 86eaedd2-9078-4dcf-9fee-25d8c2ef79acCited by top-tier papers6
- Turning Dust into Gold: Distilling Complex Reasoning Capabilities from LLMs by Leveraging Negative DataYiwei Li, Peiwen Yuan, Shaoxiong Feng, Boyuan Pan et al.AAAI 2024 · 31 citations
- Aspect-Based Sentiment Analysis with Explicit Sentiment AugmentationsJihong Ouyang, Zhiyao Yang, Silong Liang, Bing Wang et al.AAAI 2024 · 21 citations
- Dual-level Mixup for Graph Few-shot Learning with Fewer TasksYonghao Liu, Mengyu Li, Fausto Giunchiglia, Lan Huang et al.WWW 2025 · 8 citations
- Towards Zero-Shot Multilingual Transfer for Code-Switched ResponsesTing-Wei Wu, Changsheng Zhao, Ernie Chang, Yangyang Shi et al.ACL 2023 · 2 citations
- PromDA: Prompt-based Data Augmentation for Low-Resource NLU TasksYufei Wang, Can Xu, Qingfeng Sun, Huang Hu et al.ACL 2022
Builds on6
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Unsupervised Data Augmentation for Consistency TrainingQizhe Xie, Zihang Dai, Eduard H. Hovy, Thang Luong et al.NeurIPS 2020 · 2,774 citations
- MixText: Linguistically-Informed Interpolation of Hidden Space for Semi-Supervised Text ClassificationJiaao Chen, Zichao Yang, Diyi YangACL 2020 · 340 citations
- Paraphrase Augmented Task-Oriented Dialog GenerationSilin Gao, Yichi Zhang, Zhijian Ou, Zhou YuACL 2020 · 78 citations
- Continual Learning in Task-Oriented Dialogue SystemsAndrea Madotto, Zhaojiang Lin, Zhenpeng Zhou, Seungwhan Moon et al.EMNLP 2021 · 68 citations
Related papers
- A Generative Model for Joint Natural Language Understanding and GenerationBo-Hsiang Tseng, Jianpeng Cheng, Yimai Fang, David VandykeACL 2020 · 25 citations
- Self-training Improves Pre-training for Few-shot Learning in Task-oriented Dialog SystemsFei Mi, Wanhao Zhou, Lingjing Kong, Fengyu Cai et al.EMNLP 2021 · 18 citations
- Attention-Informed Mixed-Language Training for Zero-Shot Cross-Lingual Task-Oriented Dialogue SystemsZihan Liu, Genta Indra Winata, Zhaojiang Lin, Peng Xu et al.AAAI 2020 · 105 citations
- Improving Compositional Generalization with Self-Training for Data-to-Text GenerationSanket Vaibhav Mehta, Jinfeng Rao, Yi Tay, Mihir Kale et al.ACL 2022 · 34 citations
- Expand, Highlight, Generate: RL-driven Document Generation for Passage RerankingArian Askari, Mohammad Aliannejadi, Chuan Meng, Evangelos Kanoulas et al.EMNLP 2023 · 7 citations
