MASKER: Masked Keyword Regularization for Reliable Text Classification
Seung Jun Moon, Sangwoo Mo, Kimin Lee, Jaeho Lee, Jinwoo Shin
Abstract
Pre-trained language models have achieved state-of-the-art accuracies on various text classification tasks, e.g., sentiment analysis, natural language inference, and semantic textual similarity. However, the reliability of the fine-tuned text classifiers is an often underlooked performance criterion. For instance, one may desire a model that can detect out-of-distribution (OOD) samples (drawn far from training distribution) or be robust against domain shifts. We claim that one central obstacle to the reliability is the over-reliance of the model on a limited number of keywords, instead of looking at the whole context. In particular, we find that (a) OOD samples often contain in-distribution keywords, while (b) cross-domain samples may not always contain keywords; over-relying on the keywords can be problematic for both cases. In light of this observation, we propose a simple yet effective fine-tuning method, coined masked keyword regularization (MASKER), that facilitates context-based prediction. MASKER regularizes the model to reconstruct the keywords from the rest of the words and make low-confidence predictions without enough context. When applied to various pre-trained language models (e.g., BERT, RoBERTa, and ALBERT), we demonstrate that MASKER improves OOD detection and cross-domain generalization without degrading classification accuracy. Code is available at https://github.com/alinlab/MASKER.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- Object-aware Contrastive Learning for Debiased Scene RepresentationSangwoo Mo, Hyunwoo Kang, Kihyuk Sohn, Chun-Liang Li et al.NeurIPS 2021 · 57 citations
- TACIT: A Target-Agnostic Feature Disentanglement Framework for Cross-Domain Text ClassificationRui Song, Fausto Giunchiglia, Yingji Li, Mingjie Tian et al.AAAI 2024 · 10 citations
- MaskDroid: Robust Android Malware Detection with Masked Graph RepresentationsJingnan Zheng, Jiahao Liu, An Zhang, Jun Zeng et al.ASE 2024 · 6 citations
- Discovering and Mitigating Visual Biases Through Keyword ExplanationYounghyun Kim, Sangwoo Mo, Minkyu Kim, Kyungmin Lee et al.CVPR 2024
- STINMatch: Semi-Supervised Semantic-Topological Iteration Network for Financial Risk Detection via News Label DiffusionXurui Li, Yue Qin, Rui Zhu, Tianqianjin Lin et al.EMNLP 2023
Builds on6
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel et al.ICLR 2020 · 7,418 citations
- CSI: Novelty Detection via Contrastive Learning on Distributionally Shifted InstancesJihoon Tack, Sangwoo Mo, Jongheon Jeong, Jinwoo ShinNeurIPS 2020 · 755 citations
- ELECTRA: Pre-training Text Encoders as Discriminators Rather Than GeneratorsKevin Clark, Minh-Thang Luong, Quoc V. Le, Christopher D. ManningICLR 2020 · 541 citations
- On the Stability of Fine-tuning BERT: Misconceptions, Explanations, and Strong BaselinesMarius Mosbach, Maksym Andriushchenko, Dietrich KlakowICLR 2021 · 448 citations
- Revisiting Few-sample BERT Fine-tuningTianyi Zhang, Felix Wu, Arzoo Katiyar, Kilian Q. Weinberger et al.ICLR 2021 · 172 citations
Related papers
- Is Fine-tuning Needed? Pre-trained Language Models Are Near Perfect for Out-of-Domain DetectionRheeya Uppaal, Junjie Hu, Yixuan LiACL 2023 · 9 citations
- Calibrated Language Model Fine-Tuning for In- and Out-of-Distribution DataLingkai Kong, Haoming Jiang, Yuchen Zhuang, Jie Lyu et al.EMNLP 2020 · 47 citations
- Preserving Pre-trained Features Helps Calibrate Fine-tuned Language ModelsGuande He, Jianfei Chen, Jun ZhuICLR 2023 · 1 citation
- Debiased Fine-Tuning for Vision-Language Models by Prompt RegularizationBeier Zhu, Yulei Niu, Saeil Lee, Minhoe Hur et al.AAAI 2023 · 34 citations
- Masked Images Are Counterfactual Samples for Robust Fine-TuningYao Xiao, Ziyi Tang, Pengxu Wei, Cong Liu et al.CVPR 2023
