SCENE: Self-Labeled Counterfactuals for Extrapolating to Negative Examples
Deqing Fu, Ameya Godbole, Robin Jia
摘要
Detecting negatives (such as non-entailment relationships, unanswerable questions, and false claims) is an important and challenging aspect of many natural language understanding tasks. Though manually collecting challenging negative examples can help models detect them, it is both costly and domain-specific. In this work, we propose Self-labeled Counterfactuals for Extrapolating to Negative Examples (SCENE), an automatic method for synthesizing training data that greatly improves models’ ability to detect challenging negative examples. In contrast with standard data augmentation, which synthesizes new examples for existing labels, SCENE can synthesize negative examples zero-shot from only positive ones. Given a positive example, SCENE perturbs it with a mask infilling model, then determines whether the resulting example is negative based on a self-training heuristic. With access to only answerable training examples, SCENE can close 69.6% of the performance gap on SQuAD 2.0, a dataset where half of the evaluation examples are unanswerable, compared to a model trained on SQuAD 2.0. Our method also extends to boolean question answering and recognizing textual entailment, and improves generalization from SQuAD to ACE-whQA, an out-of-domain extractive QA benchmark.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Cancer-Myth: Evaluating Large Language Models on Patient Questions with False PresuppositionsWang Zhu, Tianqi Chen, Xinyan Yu, Ching Ying Lin 等ICLR 2026 · 被引用 15 次
- Don't Just Say "I don't know"! Self-aligning Large Language Models for Responding to Unknown Questions with ExplanationsYang Deng, Yong Zhao, Moxin Li, See-Kiong Ng 等EMNLP 2024 · 被引用 7 次
- TLDR: Token-Level Detective Reward Model for Large Vision Language ModelsDeqing Fu, Tong Xiao, Rui Wang, Wang Zhu 等ICLR 2025
它引用的顶会 Paper15
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad 等ACL 2020 · 被引用 1,224 次
- Learning The Difference That Makes A Difference With Counterfactually-Augmented DataDivyansh Kaushik, Eduard H. Hovy, Zachary Chase LiptonICLR 2020 · 被引用 625 次
- Understanding Self-Training for Gradual Domain AdaptationAnanya Kumar, Tengyu Ma, Percy LiangICML 2020 · 被引用 266 次
- Dense Passage Retrieval for Open-Domain Question AnsweringVladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis 等EMNLP 2020 · 被引用 142 次
- Robustness to Spurious Correlations in Text Classification via Automatically Generated CounterfactualsZhao Wang, Aron CulottaAAAI 2021 · 被引用 114 次
相关 Paper
- Retrieval-guided Counterfactual Generation for QABhargavi Paranjape, Matthew Lamm, Ian TenneyACL 2022 · 被引用 39 次
- Language Model Pre-training on True NegativesZhuosheng Zhang, Hai Zhao, Masao Utiyama, Eiichiro SumitaAAAI 2023 · 被引用 3 次
- Generating Information-Seeking Conversations from Unlabeled DocumentsGangwoo Kim, Sungdong Kim, Kang Min Yoo, Jaewoo KangEMNLP 2022 · 被引用 4 次
- Training Question Answering Models From Synthetic DataRaul Puri, Ryan Spring, Mohammad Shoeybi, Mostofa Patwary 等EMNLP 2020 · 被引用 15 次
- Counterfactual Active Learning for Out-of-Distribution GeneralizationXun Deng, Wenjie Wang, Fuli Feng, Hanwang Zhang 等ACL 2023 · 被引用 10 次
