Natural Language Inference in Context - Investigating Contextual Reasoning over Long Texts
Hanmeng Liu, Leyang Cui, Jian Liu, Yue Zhang
Abstract
Natural language inference (NLI) is a fundamental NLP task, investigating the entailment relationship between two texts. Popular NLI datasets present the task at sentence-level. While adequate for testing semantic representations, they fall short for testing contextual reasoning over long texts, which is a natural part of the human inference process. We introduce ConTRoL, a new dataset for ConTextual Reasoning over Long texts. Consisting of 8,325 expert-designed "context-hypothesis" pairs with gold labels, ConTRoL is a passage-level NLI dataset with a focus on complex contextual reasoning types such as logical reasoning. It is derived from competitive selection and recruitment test (verbal reasoning test) for police recruitment, with expert level quality. Compared with previous NLI benchmarks, the materials in ConTRoL are much more challenging, involving a range of reasoning types. Empirical results show that state-of-the-art language models perform by far worse than educated humans. Our dataset can also serve as a testing-set for downstream tasks like checking the factual correctness of summaries.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8ffc767c-6783-4850-813f-0e3d34951db7Cited by top-tier papers8
- Prismatic Synthesis: Gradient-based Data Diversification Boosts Generalization in LLM ReasoningJaehun Jung, Seungju Han, Ximing Lu, Skyler Hallinan et al.NeurIPS 2025 · 50 citations
- DocInfer: Document-level Natural Language Inference using Optimal Evidence SelectionPuneet Mathur, Gautam Kunapuli, Riyaz A. Bhat, Manish Shrivastava et al.EMNLP 2022 · 5 citations
- Flexible Generation of Natural Language DeductionsKaj Bostrom, Xinyu Zhao, Swarat Chaudhuri, Greg DurrettEMNLP 2021 · 1 citation
- Extractive Fact Decomposition for Interpretable Natural Language Inference in one Forward PassNicholas Popovic, Michael FärberEMNLP 2025 · 1 citation
- EMO: Embedding Model Distillation via Intra-Model Relation and Optimal Transport AlignmentsMinh-Phuc Truong, Hai An Vu, Tu Vu, Nguyen Thi Ngoc Diep et al.EMNLP 2025
Builds on6
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel et al.ICLR 2020 · 7,418 citations
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- Adversarial NLI: A New Benchmark for Natural Language UnderstandingYixin Nie, Adina Williams, Emily Dinan, Mohit Bansal et al.ACL 2020 · 602 citations
- ReClor: A Reading Comprehension Dataset Requiring Logical ReasoningWeihao Yu, Zihang Jiang, Yanfei Dong, Jiashi FengICLR 2020 · 325 citations
- MuTual: A Dataset for Multi-Turn Dialogue ReasoningLeyang Cui, Yu Wu, Shujie Liu, Yue Zhang et al.ACL 2020 · 115 citations
Related papers
- Can Large Language Models Infer Causal Relationships from Real-World Text?Ryan Saklad, Aman Chadha, Oleg V. Pavlov, Raha MoraffahACL 2026 · 4 citations
- SciNLI: A Corpus for Natural Language Inference on Scientific TextMobashir Sadat, Cornelia CarageaACL 2022 · 41 citations
- Entailed Between the Lines: Incorporating Implication into NLIShreya Havaldar, Hamidreza Alvari, John Palowitch, Mohammad Javad Hosseini et al.ACL 2025
- Inferential Question AnsweringJamshid Mozafari, Hamed Zamani, Guido Zuccon, Adam JatowtWWW 2026
- CRIT: Graph-Based Automatic Data Synthesis to Enhance Cross-Modal Multi-Hop ReasoningJunyoung Sung, Seungwoo Lyu, Minjun Kim, Sumin An et al.CVPR 2026 · 2 citations
