HiddenCut: Simple Data Augmentation for Natural Language Understanding with Better Generalizability
Jiaao Chen, Dinghan Shen, Weizhu Chen, Diyi Yang
摘要
Fine-tuning large pre-trained models with taskspecific data has achieved great success in NLP. However, it has been demonstrated that the majority of information within the selfattention networks are redundant and not utilized effectively during the fine-tuning stage. This leads to inferior results when generalizing the obtained models to out-of-domain distributions. To this end, we propose a simple yet effective data augmentation technique, Hidden-Cut, to better regularize the model and encourage it to learn more generalizable features. Specifically, contiguous spans within the hidden space are dynamically and strategically dropped during training. Experiments show that our HiddenCut method outperforms the state-of-the-art augmentation methods on the GLUE benchmark, and consistently exhibit superior generalization performances on out-ofdistribution and challenging counterexamples. We have publicly released our code at https: //github.com/GT-SALT/HiddenCut .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- ConsistTL: Modeling Consistency in Transfer Learning for Low-Resource Neural Machine TranslationZhaocong Li, Xuebo Liu, Derek F. Wong, Lidia S. Chao 等EMNLP 2022 · 被引用 20 次
- AD-DROP: Attribution-Driven Dropout for Robust Language Model Fine-TuningTao Yang, Jinghao Deng, Xiaojun Quan, Qifan Wang 等NeurIPS 2022 · 被引用 7 次
- Beware of Model Collapse! Fast and Stable Test-time Adaptation for Robust Question AnsweringYi Su, Yixin Ji, Juntao Li, Hai Ye 等EMNLP 2023 · 被引用 2 次
- Curriculum Consistency Learning for Conditional Sentence GenerationLiangxin Liu, Xuebo Liu, Lian Lian, Shengjun Cheng 等EMNLP 2024 · 被引用 1 次
- QA Domain Adaptation using Hidden Space Augmentation and Self-Supervised Contrastive AdaptationZhenrui Yue, Huimin Zeng, Bernhard Kratzwald, Stefan Feuerriegel 等EMNLP 2022 · 被引用 1 次
它引用的顶会 Paper14
- Deberta: decoding-Enhanced Bert with Disentangled AttentionPengcheng He, Xiaodong Liu, Jianfeng Gao, Weizhu ChenICLR 2021 · 被引用 3,729 次
- Unsupervised Data Augmentation for Consistency TrainingQizhe Xie, Zihang Dai, Eduard H. Hovy, Thang Luong 等NeurIPS 2020 · 被引用 2,774 次
- Learning The Difference That Makes A Difference With Counterfactually-Augmented DataDivyansh Kaushik, Eduard H. Hovy, Zachary Chase LiptonICLR 2020 · 被引用 625 次
- Adversarial NLI: A New Benchmark for Natural Language UnderstandingYixin Nie, Adina Williams, Emily Dinan, Mohit Bansal 等ACL 2020 · 被引用 602 次
- ELECTRA: Pre-training Text Encoders as Discriminators Rather Than GeneratorsKevin Clark, Minh-Thang Luong, Quoc V. Le, Christopher D. ManningICLR 2020 · 被引用 541 次
相关 Paper
- Raise a Child in Large Language Model: Towards Effective and Generalizable Fine-tuningRunxin Xu, Fuli Luo, Zhiyuan Zhang, Chuanqi Tan 等EMNLP 2021 · 被引用 129 次
- HyPe: Better Pre-trained Language Model Fine-tuning with Hidden Representation PerturbationHongyi Yuan, Zheng Yuan, Chuanqi Tan, Fei Huang 等ACL 2023 · 被引用 5 次
- Fine-Tuning Pre-Trained Language Models Effectively by Optimizing Subnetworks AdaptivelyHaojie Zhang, Ge Li, Jia Li, Zhongjin Zhang 等NeurIPS 2022 · 被引用 43 次
- Mixout: Effective Regularization to Finetune Large-scale Pretrained Language ModelsCheolhyoung Lee, Kyunghyun Cho, Wanmo KangICLR 2020 · 被引用 233 次
- APrompt: Attention Prompt Tuning for Efficient Adaptation of Pre-trained Language ModelsQifan Wang, Yuning Mao, Jingang Wang, Hanchao Yu 等EMNLP 2023 · 被引用 25 次
