Transitive self-consistency evaluation of NLI models without gold labels
Wei Wu, Mark Last
摘要
Natural Language Inference (NLI) is an important task in natural language processing. NLI models are aimed at automatically determining logical relationships between pairs of sentences. However, recent studies based on gold labels assigned to sentence pairs by human experts have provided some evidence that NLI models tend to make inconsistent model decisions during inference. Previous studies have used existing NLI datasets to test the transitive consistency of language models. However, they test only variations of two transitive consistency rules out of four. To further evaluate the transitive consistency of NLI models, we propose a novel evaluation approach that allows us to test all four rules automatically by generating adversarial examples via antonym replacements. Since we are testing self-consistency, human labeling of generated adversarial examples is unnecessary. Our experiments on several benchmark datasets indicate that the examples generated by the proposed antonym replacement methodology can reveal transitive inconsistencies in the state-of-the-art NLI models.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper7
- Deberta: decoding-Enhanced Bert with Disentangled AttentionPengcheng He, Xiaodong Liu, Jianfeng Gao, Weizhu ChenICLR 2021 · 被引用 3,729 次
- Adversarial NLI: A New Benchmark for Natural Language UnderstandingYixin Nie, Adina Williams, Emily Dinan, Mohit Bansal 等ACL 2020 · 被引用 602 次
- Smarter, Better, Faster, Longer: A Modern Bidirectional Encoder for Fast, Memory Efficient, and Long Context Finetuning and InferenceBenjamin Warner, Antoine Chaffin, Benjamin Clavié, Orion Weller 等ACL 2025 · 被引用 552 次
- Generating Persona Consistent Dialogues by Exploiting Natural Language InferenceHaoyu Song, Wei-Nan Zhang, Jingwen Hu, Ting LiuAAAI 2020 · 被引用 85 次
- Consistency Analysis of ChatGPTMyeongjun Jang, Thomas LukasiewiczEMNLP 2023 · 被引用 55 次
相关 Paper
- LogiConBench: Benchmarking Logical Consistencies of LLMsZheng Chen, Chuan Zhou, Fengxiang Cheng, Tin Po Yip 等ICLR 2026
- Introducing Verification Task of Set Consistency with Set-Consistency Energy NetworksMooho Song, Hye Ryung Son, Jay-Yoon LeeACL 2025 · 被引用 1 次
- NILE : Natural Language Inference with Faithful Natural Language ExplanationsSawan Kumar, Partha P. TalukdarACL 2020 · 被引用 15 次
- How Hard is this Test Set? NLI Characterization by Exploiting Training DynamicsAdrian Cosma, Stefan Ruseti, Mihai Dascalu, Cornelia CarageaEMNLP 2024
- A synthetic data approach for domain generalization of NLI modelsMohammad Javad Hosseini, Andrey Petrov, Alex Fabrikant, Annie LouisACL 2024 · 被引用 3 次
