What does the Failure to Reason with "Respectively" in Zero/Few-Shot Settings Tell Us about Language Models?
Ruixiang Cui, Seolhwa Lee, Daniel Hershcovich, Anders Søgaard
Abstract
Humans can effortlessly understand the coordinate structure of sentences such as "Niels Bohr and Kurt Cobain were born in Copenhagen and Seattle, respectively". In the context of natural language inference (NLI), we examine how language models (LMs) reason with respective readings (Gawron and Kehler, 2004) from two perspectives: syntactic-semantic and commonsense-world knowledge. We propose a controlled synthetic dataset WikiResNLI and a naturally occurring dataset NatResNLI to encompass various explicit and implicit realizations of "respectively". We show that finetuned NLI models struggle with understanding such readings without explicit supervision. While few-shot learning is easy in the presence of explicit cues, longer training is required when the reading is evoked implicitly, leaving models to rely on common sense inferences. Furthermore, our fine-grained analysis indicates models fail to generalize across different constructions. To conclude, we demonstrate that LMs still lag behind humans in generalizing to the long tail of linguistic constructions. Denotation Natural Language Example Premise: w1 and w3 p w2 and w4 , respectively. Emiliano Zapata and Gerhart Münch died in Morelos and Michoacán , respectively Hypotheses: Entailment (1), 1S1O w1 p w2 . Emiliano Zapata died in Morelos . Entailment (2), 1S1O w3 p w4 . Gerhart Münch died in Michoacán . Contradiction (1), 1S1O w1 p w4 . Emiliano Zapata died in Michoacán . Contradiction (2), 1S1O w3 p w2 . Gerhart Münch died in Morelos . Contradiction (3), 1S2O w1 p w2 and w4 . Emiliano Zapata died in Morelos and Michoacán . Contradiction (4), 1S2O w3 p w2 and w4 . Gerhart Münch died in Morelos and Michoacán . Contradiction (5), 2S1O w1 and w3 p w2 . Emiliano Zapata and Gerhart Münch died in Morelos . Contradiction (6), 2S1O w1 and w3 p w4 . Emiliano Zapata and Gerhart Münch died in Michoacán .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7c66cc06-a3ee-416c-9625-927f6ffc76dcCited by top-tier papers1
Ask how each one uses itBuilds on7
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel et al.ICLR 2020 · 7,418 citations
- Deberta: decoding-Enhanced Bert with Disentangled AttentionPengcheng He, Xiaodong Liu, Jianfeng Gao, Weizhu ChenICLR 2021 · 3,729 citations
- Adversarial NLI: A New Benchmark for Natural Language UnderstandingYixin Nie, Adina Williams, Emily Dinan, Mohit Bansal et al.ACL 2020 · 602 citations
- Large Language Models Can Self-ImproveJiaxin Huang, Shixiang Gu, Le Hou, Yuexin Wu et al.EMNLP 2023 · 184 citations
Related papers
- This is not a Dataset: A Large Negation Benchmark to Challenge Large Language ModelsIker García-Ferrero, Begoña Altuna, Javier Álvez, Itziar Gonzalez-Dios et al.EMNLP 2023 · 8 citations
- Entailed Between the Lines: Incorporating Implication into NLIShreya Havaldar, Hamidreza Alvari, John Palowitch, Mohammad Javad Hosseini et al.ACL 2025
- Natural Language Inference in Context - Investigating Contextual Reasoning over Long TextsHanmeng Liu, Leyang Cui, Jian Liu, Yue ZhangAAAI 2021 · 57 citations
- IMPLI: Investigating NLI Models' Performance on Figurative LanguageKevin Stowe, Prasetya Ajie Utama, Iryna GurevychACL 2022 · 52 citations
- Are Natural Language Inference Models IMPPRESsive? Learning IMPlicature and PRESuppositionPaloma Jeretic, Alex Warstadt, Suvrat Bhooshan, Adina WilliamsACL 2020 · 2 citations
