Natural Language Inference in Context - Investigating Contextual Reasoning over Long Texts
Hanmeng Liu, Leyang Cui, Jian Liu, Yue Zhang
摘要
Natural language inference (NLI) is a fundamental NLP task, investigating the entailment relationship between two texts. Popular NLI datasets present the task at sentence-level. While adequate for testing semantic representations, they fall short for testing contextual reasoning over long texts, which is a natural part of the human inference process. We introduce ConTRoL, a new dataset for ConTextual Reasoning over Long texts. Consisting of 8,325 expert-designed "context-hypothesis" pairs with gold labels, ConTRoL is a passage-level NLI dataset with a focus on complex contextual reasoning types such as logical reasoning. It is derived from competitive selection and recruitment test (verbal reasoning test) for police recruitment, with expert level quality. Compared with previous NLI benchmarks, the materials in ConTRoL are much more challenging, involving a range of reasoning types. Empirical results show that state-of-the-art language models perform by far worse than educated humans. Our dataset can also serve as a testing-set for downstream tasks like checking the factual correctness of summaries.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Prismatic Synthesis: Gradient-based Data Diversification Boosts Generalization in LLM ReasoningJaehun Jung, Seungju Han, Ximing Lu, Skyler Hallinan 等NeurIPS 2025 · 被引用 50 次
- DocInfer: Document-level Natural Language Inference using Optimal Evidence SelectionPuneet Mathur, Gautam Kunapuli, Riyaz A. Bhat, Manish Shrivastava 等EMNLP 2022 · 被引用 5 次
- Flexible Generation of Natural Language DeductionsKaj Bostrom, Xinyu Zhao, Swarat Chaudhuri, Greg DurrettEMNLP 2021 · 被引用 1 次
- Extractive Fact Decomposition for Interpretable Natural Language Inference in one Forward PassNicholas Popovic, Michael FärberEMNLP 2025 · 被引用 1 次
- EMO: Embedding Model Distillation via Intra-Model Relation and Optimal Transport AlignmentsMinh-Phuc Truong, Hai An Vu, Tu Vu, Nguyen Thi Ngoc Diep 等EMNLP 2025
它引用的顶会 Paper6
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel 等ICLR 2020 · 被引用 7,418 次
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad 等ACL 2020 · 被引用 1,224 次
- Adversarial NLI: A New Benchmark for Natural Language UnderstandingYixin Nie, Adina Williams, Emily Dinan, Mohit Bansal 等ACL 2020 · 被引用 602 次
- ReClor: A Reading Comprehension Dataset Requiring Logical ReasoningWeihao Yu, Zihang Jiang, Yanfei Dong, Jiashi FengICLR 2020 · 被引用 325 次
- MuTual: A Dataset for Multi-Turn Dialogue ReasoningLeyang Cui, Yu Wu, Shujie Liu, Yue Zhang 等ACL 2020 · 被引用 115 次
相关 Paper
- Can Large Language Models Infer Causal Relationships from Real-World Text?Ryan Saklad, Aman Chadha, Oleg V. Pavlov, Raha MoraffahACL 2026 · 被引用 4 次
- SciNLI: A Corpus for Natural Language Inference on Scientific TextMobashir Sadat, Cornelia CarageaACL 2022 · 被引用 41 次
- Entailed Between the Lines: Incorporating Implication into NLIShreya Havaldar, Hamidreza Alvari, John Palowitch, Mohammad Javad Hosseini 等ACL 2025
- Inferential Question AnsweringJamshid Mozafari, Hamed Zamani, Guido Zuccon, Adam JatowtWWW 2026
- CRIT: Graph-Based Automatic Data Synthesis to Enhance Cross-Modal Multi-Hop ReasoningJunyoung Sung, Seungwoo Lyu, Minjun Kim, Sumin An 等CVPR 2026 · 被引用 2 次
