C2C-GenDA: Cluster-to-Cluster Generation for Data Augmentation of Slot Filling
Yutai Hou, Sanyuan Chen, Wanxiang Che, Cheng Chen, Ting Liu
Abstract
Slot filling, a fundamental module of spoken language understanding, often suffers from insufficient quantity and diversity of training data. To remedy this, we propose a novel Cluster-to-Cluster generation framework for Data Augmentation (DA), named C2C-GenDA. It enlarges the training set by reconstructing existing utterances into alternative expressions while keeping semantic. Different from previous DA works that reconstruct utterances one by one independently, C2C-GenDA jointly encodes multiple existing utterances of the same semantics and simultaneously decodes multiple unseen expressions. Jointly generating multiple new utterances allows to consider the relations between generated instances and encourages diversity. Besides, encoding multiple existing utterances endows C2C with a wider view of existing expressions, helping to reduce generation that duplicates existing data. Experiments on ATIS and Snips datasets show that instances augmented by C2C-GenDA improve slot filling by 7.99 (11.9%↑) and 5.76 (13.6%↑) F-scores respectively, when there are only hundreds of training utterances. Code: https://github.com/Sanyuan-Chen/C2C-DA.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on1
Related papers
- Simple Conversational Data Augmentation for Semi-supervised Abstractive Dialogue SummarizationJiaao Chen, Diyi YangEMNLP 2021 · 31 citations
- SNIDA: Unlocking Few-Shot Object Detection with Non-Linear Semantic Decoupling AugmentationYanjie Wang, Xu Zou, Luxin Yan, Sheng Zhong et al.CVPR 2024 · 22 citations
- RepCodec: A Speech Representation Codec for Speech TokenizationZhichao Huang, Chutong Meng, Tom KoACL 2024 · 16 citations
- Unsupervised Multimodal Clustering for Semantics Discovery in Multimodal UtterancesHanlei Zhang, Hua Xu, Fei Long, Xin Wang et al.ACL 2024
- Linguistically-Enriched and Context-AwareZero-shot Slot FillingA. B. Siddique, Fuad T. Jamour, Vagelis HristidisWWW 2021 · 29 citations
