Distributionally Robust Finetuning BERT for Covariate Drift in Spoken Language Understanding
Samuel Broscheit, Quynh Do, Judith Gaspers
Abstract
In this study, we investigate robustness against covariate drift in spoken language understanding (SLU). Covariate drift can occur in SLUwhen there is a drift between training and testing regarding what users request or how they request it. To study this we propose a method that exploits natural variations in data to create a covariate drift in SLU datasets. Experiments show that a state-of-the-art BERT-based model suffers performance loss under this drift. To mitigate the performance loss, we investigate distributionally robust optimization (DRO) for finetuning BERT-based models. We discuss some recent DRO methods, propose two new variants and empirically show that DRO improves robustness under drift.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- Adaptive Preference Scaling for Reinforcement Learning with Human FeedbackIlgee Hong, Zichong Li, Alexander Bukharin, Yixiao Li et al.NeurIPS 2024 · 23 citations
- Out-of-Distribution Generalization in Natural Language Processing: Past, Present, and FutureLinyi Yang, Yaoxian Song, Xuan Ren, Chenyang Lyu et al.EMNLP 2023 · 12 citations
- Characterizing and Measuring Linguistic Dataset DriftTyler A. Chang, Kishaloy Halder, Neha Anna John, Yogarshi Vyas et al.ACL 2023 · 2 citations
- Interventional Speech Noise Injection for ASR Generalizable Spoken Language UnderstandingYeonJoon Jung, Jaeseong Lee, Seungtaek Choi, Dohyeon Lee et al.EMNLP 2024
Builds on5
- WILDS: A Benchmark of in-the-Wild Distribution ShiftsPang Wei Koh, Shiori Sagawa, Henrik Marklund, Sang Michael Xie et al.ICML 2021 · 1,773 citations
- Distributionally Robust Neural NetworksShiori Sagawa, Pang Wei Koh, Tatsunori B. Hashimoto, Percy LiangICLR 2020 · 1,578 citations
- Just Train Twice: Improving Group Robustness without Training Group InformationEvan Zheran Liu, Behzad Haghgoo, Annie S. Chen, Aditi Raghunathan et al.ICML 2021 · 683 citations
- Large-Scale Methods for Distributionally Robust OptimizationDaniel Levy, Yair Carmon, John C. Duchi, Aaron SidfordNeurIPS 2020 · 281 citations
- Modeling the Second Player in Distributionally Robust OptimizationPaul Michel, Tatsunori Hashimoto, Graham NeubigICLR 2021 · 39 citations
Related papers
- COCO-DR: Combating the Distribution Shift in Zero-Shot Dense Retrieval with Contrastive and Distributionally Robust LearningYue Yu, Chenyan Xiong, Si Sun, Chao Zhang et al.EMNLP 2022 · 21 citations
- Examining and Combating Spurious Features under Distribution ShiftChunting Zhou, Xuezhe Ma, Paul Michel, Graham NeubigICML 2021 · 78 citations
- Learning Distributionally Robust Models at Scale via Composite OptimizationFarzin Haddadpour, Mohammad Mahdi Kamani, Mehrdad Mahdavi, Amin KarbasiICLR 2022 · 5 citations
- Towards Robust Alignment of Language Models: Distributionally Robustifying Direct Preference OptimizationJunkang Wu, Yuexiang Xie, Zhengyi Yang, Jiancan Wu et al.ICLR 2025
- Task-level Distributionally Robust Optimization for Large Language Model-based Dense RetrievalGuangyuan Ma, Yongliang Ma, Xing Wu, Zhenpeng Su et al.AAAI 2025 · 6 citations
