HiddenCut: Simple Data Augmentation for Natural Language Understanding with Better Generalizability
Jiaao Chen, Dinghan Shen, Weizhu Chen, Diyi Yang
Abstract
Fine-tuning large pre-trained models with taskspecific data has achieved great success in NLP. However, it has been demonstrated that the majority of information within the selfattention networks are redundant and not utilized effectively during the fine-tuning stage. This leads to inferior results when generalizing the obtained models to out-of-domain distributions. To this end, we propose a simple yet effective data augmentation technique, Hidden-Cut, to better regularize the model and encourage it to learn more generalizable features. Specifically, contiguous spans within the hidden space are dynamically and strategically dropped during training. Experiments show that our HiddenCut method outperforms the state-of-the-art augmentation methods on the GLUE benchmark, and consistently exhibit superior generalization performances on out-ofdistribution and challenging counterexamples. We have publicly released our code at https: //github.com/GT-SALT/HiddenCut .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1ffb50bb-5278-41b3-9c4e-a490a37c1cdaCited by top-tier papers7
- ConsistTL: Modeling Consistency in Transfer Learning for Low-Resource Neural Machine TranslationZhaocong Li, Xuebo Liu, Derek F. Wong, Lidia S. Chao et al.EMNLP 2022 · 20 citations
- AD-DROP: Attribution-Driven Dropout for Robust Language Model Fine-TuningTao Yang, Jinghao Deng, Xiaojun Quan, Qifan Wang et al.NeurIPS 2022 · 7 citations
- Beware of Model Collapse! Fast and Stable Test-time Adaptation for Robust Question AnsweringYi Su, Yixin Ji, Juntao Li, Hai Ye et al.EMNLP 2023 · 2 citations
- Curriculum Consistency Learning for Conditional Sentence GenerationLiangxin Liu, Xuebo Liu, Lian Lian, Shengjun Cheng et al.EMNLP 2024 · 1 citation
- QA Domain Adaptation using Hidden Space Augmentation and Self-Supervised Contrastive AdaptationZhenrui Yue, Huimin Zeng, Bernhard Kratzwald, Stefan Feuerriegel et al.EMNLP 2022 · 1 citation
Builds on14
- Deberta: decoding-Enhanced Bert with Disentangled AttentionPengcheng He, Xiaodong Liu, Jianfeng Gao, Weizhu ChenICLR 2021 · 3,729 citations
- Unsupervised Data Augmentation for Consistency TrainingQizhe Xie, Zihang Dai, Eduard H. Hovy, Thang Luong et al.NeurIPS 2020 · 2,774 citations
- Learning The Difference That Makes A Difference With Counterfactually-Augmented DataDivyansh Kaushik, Eduard H. Hovy, Zachary Chase LiptonICLR 2020 · 625 citations
- Adversarial NLI: A New Benchmark for Natural Language UnderstandingYixin Nie, Adina Williams, Emily Dinan, Mohit Bansal et al.ACL 2020 · 602 citations
- ELECTRA: Pre-training Text Encoders as Discriminators Rather Than GeneratorsKevin Clark, Minh-Thang Luong, Quoc V. Le, Christopher D. ManningICLR 2020 · 541 citations
Related papers
- Raise a Child in Large Language Model: Towards Effective and Generalizable Fine-tuningRunxin Xu, Fuli Luo, Zhiyuan Zhang, Chuanqi Tan et al.EMNLP 2021 · 129 citations
- HyPe: Better Pre-trained Language Model Fine-tuning with Hidden Representation PerturbationHongyi Yuan, Zheng Yuan, Chuanqi Tan, Fei Huang et al.ACL 2023 · 5 citations
- Fine-Tuning Pre-Trained Language Models Effectively by Optimizing Subnetworks AdaptivelyHaojie Zhang, Ge Li, Jia Li, Zhongjin Zhang et al.NeurIPS 2022 · 43 citations
- Mixout: Effective Regularization to Finetune Large-scale Pretrained Language ModelsCheolhyoung Lee, Kyunghyun Cho, Wanmo KangICLR 2020 · 233 citations
- APrompt: Attention Prompt Tuning for Efficient Adaptation of Pre-trained Language ModelsQifan Wang, Yuning Mao, Jingang Wang, Hanchao Yu et al.EMNLP 2023 · 25 citations
