Learning to Perturb Word Embeddings for Out-of-distribution QA
Seanie Lee, Minki Kang, Juho Lee, Sung Ju Hwang
Abstract
QA models based on pretrained language models have achieved remarkable performance on various benchmark datasets. However, QA models do not generalize well to unseen data that falls outside the training distribution, due to distributional shifts. Data augmentation (DA) techniques which drop/replace words have shown to be effective in regularizing the model from overfitting to the training data. Yet, they may adversely affect the QA tasks since they incur semantic changes that may lead to wrong answers for the QA task. To tackle this problem, we propose a simple yet effective DA method based on a stochastic noise generator, which learns to perturb the word embedding of the input questions and context without changing their semantics. We validate the performance of the QA models trained with our word embedding perturbation on a single source dataset, on five different target domains. The results show that our method significantly outperforms the baseline DA methods. Notably, the model trained with ours outperforms the model trained with more than 240K artificially generated QA pairs. Q: In what year was the Theodore m. Hesburgh library at Notre Dame finished? C: (…) the main building is the 14 -story Theodore m. Hesburgh library, completed in 1963, (…) this mural is popularly known as "touchdown jesus" because of its proximity … Q: In what year was the Theodore m. Hesburgh library at Notre Dame finished? Q: each last year was the Theodore m. Vanroth library at Notre Dame finished. C: (…) the first building is the 14 -story Theodore p von Hesburgh library, completed in 1963 ; (…) this mural is popularly known as our confession jesus christ because all its … C: (…) the main building is the 14 -story Theodore m. Hesburgh library, completed in 1963, (…) this mural is popularly known as "touchdown jesus" because of its proximity …
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- Sequential Reptile: Inter-Task Gradient Alignment for Multilingual LearningSeanie Lee, Haebeom Lee, Juho Lee, Sung Ju HwangICLR 2022 · 20 citations
- A Positive-Unlabeled Metric Learning Framework for Document-Level Relation Extraction with Incomplete LabelingYe Wang, Huazheng Pan, Tao Zhang, Wen Wu et al.AAAI 2024 · 11 citations
- Latent Paraphrasing: Perturbation on Layers Improves Knowledge Injection in Language ModelsMinki Kang, Sung Ju Hwang, Gibbeum Lee, Jaewoong ChoNeurIPS 2024 · 3 citations
- Improving Cooperation in Language Games with Bayesian Inference and the Cognitive HierarchyJoseph Bills, Christopher Archibald, Diego BlaylockAAAI 2025 · 1 citation
- HarmAug: Effective Data Augmentation for Knowledge Distillation of Safety Guard ModelsSeanie Lee, Haebin Seong, Dong Bok Lee, Minki Kang et al.ICLR 2025
Builds on8
- CutMix: Regularization Strategy to Train Strong Classifiers With Localizable FeaturesSangdoo Yun, Dongyoon Han, Sanghyuk Chun, Seong Joon Oh et al.ICCV 2019 · 5,843 citations
- Unsupervised Data Augmentation for Consistency TrainingQizhe Xie, Zihang Dai, Eduard H. Hovy, Thang Luong et al.NeurIPS 2020 · 2,774 citations
- ELECTRA: Pre-training Text Encoders as Discriminators Rather Than GeneratorsKevin Clark, Minh-Thang Luong, Quoc V. Le, Christopher D. ManningICLR 2020 · 541 citations
- On the Sentence Embeddings from Pre-trained Language ModelsBohan Li, Hao Zhou, Junxian He, Mingxuan Wang et al.EMNLP 2020 · 538 citations
- Cross-Domain Few-Shot Classification via Learned Feature-Wise TransformationHung-Yu Tseng, Hsin-Ying Lee, Jia-Bin Huang, Ming-Hsuan YangICLR 2020 · 467 citations
Related papers
- Virtual Data Augmentation: A Robust and General Framework for Fine-tuning Pre-trained ModelsKun Zhou, Wayne Xin Zhao, Sirui Wang, Fuzheng Zhang et al.EMNLP 2021 · 6 citations
- Domain Gap Embeddings for Generative Dataset AugmentationYinong Oliver Wang, Younjoon Chung, Chen Henry Wu, Fernando De la TorreCVPR 2024 · 8 citations
- Q: How to Specialize Large Vision-Language Models to Data-Scarce VQA Tasks? A: Self-Train on Unlabeled Images!Zaid Khan, B. G. Vijay Kumar, Samuel Schulter, Xiang Yu et al.CVPR 2023
- Tell Me How to Ask Again: Question Data Augmentation with Controllable Rewriting in Continuous SpaceDayiheng Liu, Yeyun Gong, Jie Fu, Yu Yan et al.EMNLP 2020 · 36 citations
- A Simple Feature Augmentation for Domain GeneralizationPan Li, Da Li, Wei Li, Shaogang Gong et al.ICCV 2021 · 242 citations
