Does Your Model Classify Entities Reasonably? Diagnosing and Mitigating Spurious Correlations in Entity Typing
Nan Xu, Fei Wang, Bangzheng Li, Mingtao Dong, Muhao Chen
摘要
Entity typing aims at predicting one or more words that describe the type(s) of a specific mention in a sentence. Due to shortcuts from surface patterns to annotated entity labels and biased training, existing entity typing models are subject to the problem of spurious correlations. To comprehensively investigate the faithfulness and reliability of entity typing methods, we first systematically define distinct kinds of model biases that are reflected mainly from spurious correlations. Particularly, we identify six types of existing model biases, including mention-context bias, lexical overlapping bias, named entity bias, pronoun bias, dependency bias, and overgeneralization bias. To mitigate model biases, we then introduce a counterfactual data augmentation method. By augmenting the original training set with their debiased counterparts, models are forced to fully comprehend sentences and discover the fundamental cues for entity typing, rather than relying on spurious correlations for shortcuts. Experimental results on the UFET dataset show our counterfactual data augmentation approach helps improve generalization of different entity typing models with consistently better performance on both the original and debiased test sets 1 . PLM Prompts Entity Typing Instances Mention-Context: Prompt I: <Mention> is a type of <mask>. S1: fire is a type of <mask>. RoBERTa: energy, heat, explosion, fire, gas S2: the war is a type of <mask>. True labels: war, battle, conflict RoBERTa: war, battle, conflict, violence, warfare T1: A teacher who survived the shooting said he would never forgive the police for taking an hour to arrive after the gunman opened fire.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Knowledge Conflicts for LLMs: A SurveyRongwu Xu, Zehan Qi, Zhijiang Guo, Cunxiang Wang 等EMNLP 2024 · 被引用 38 次
- Compositional Generalization from Learned Skills via CoT Training: A Theoretical and Structural Analysis for ReasoningXinhao Yao, Ruifeng Ren, Yun Liao, Lizhong Ding 等ICLR 2026 · 被引用 6 次
- Hopping Too Late: Exploring the Limitations of Large Language Models on Multi-Hop QueriesEden Biran, Daniela Gottesman, Sohee Yang, Mor Geva 等EMNLP 2024 · 被引用 3 次
- The Curious Case of Analogies: Investigating Analogical Reasoning in Large Language ModelsTaewhoo Lee, Minju Song, Chanwoong Yoon, Jungwoo Park 等AAAI 2026 · 被引用 1 次
- Controllable Context Sensitivity and the Knob Behind ItJulian Minder, Kevin Du, Niklas Stoehr, Giovanni Monea 等ICLR 2025
它引用的顶会 Paper12
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
- Learning The Difference That Makes A Difference With Counterfactually-Augmented DataDivyansh Kaushik, Eduard H. Hovy, Zachary Chase LiptonICLR 2020 · 被引用 625 次
- Adversarial Filters of Dataset BiasesRonan Le Bras, Swabha Swayamdipta, Chandra Bhagavatula, Rowan Zellers 等ICML 2020 · 被引用 242 次
- End-to-End Bias Mitigation by Modelling Biases in CorporaRabeeh Karimi Mahabadi, Yonatan Belinkov, James HendersonACL 2020 · 被引用 136 次
- Contrastive Out-of-Distribution Detection for Pretrained TransformersWenxuan Zhou, Fangyu Liu, Muhao ChenEMNLP 2021 · 被引用 63 次
相关 Paper
- CoBA: Counterbias Text Augmentation for Mitigating Various Spurious Correlations via Semantic TriplesKyohoon Jin, Juhwan Choi, Jungmin Yun, Junho Lee 等EMNLP 2025
- Counterfactual Generator: A Weakly-Supervised Method for Named Entity RecognitionXiangji Zeng, Yunliang Li, Yuchen Zhai, Yin ZhangEMNLP 2020 · 被引用 55 次
- C2L: Causally Contrastive Learning for Robust Text ClassificationSeungtaek Choi, Myeongho Jeong, Hojae Han, Seung-won HwangAAAI 2022 · 被引用 52 次
- BiasAdv: Bias-Adversarial Augmentation for Model DebiasingJongin Lim, Youngdong Kim, Byungjai Kim, Chanho Ahn 等CVPR 2023
- Ultra-Fine Entity Typing with Weak Supervision from a Masked Language ModelHongliang Dai, Yangqiu Song, Haixun WangACL 2021
