An Investigation of LLMs' Inefficacy in Understanding Converse Relations
Chengwen Qi, Bowen Li, Binyuan Hui, Bailin Wang, Jinyang Li, Jinwang Wu, Yuanjun Laili
Abstract
Large Language Models (LLMs) have achieved remarkable success in many formal language oriented tasks, such as structural data-to-text and semantic parsing. However current benchmarks mostly follow the data distribution of the pre-training data of LLMs. Therefore, a natural question rises that do LLMs really understand the structured semantics of formal languages. In this paper, we investigate this problem on a special case, converse binary relation. We introduce a new benchmark ConvRe focusing on converse relations, which contains 17 relations and 1240 triples extracted from popular knowledge graph completion datasets. Our ConvRe features two tasks, Re2Text and Text2Re, which are formulated as multi-choice question answering to evaluate LLMs' ability to determine the matching between relations and associated text. For the evaluation protocol, apart from different prompting methods, we further introduce variants to the test text and few-shot example text. We conduct experiments on three popular LLM families and have observed various scaling trends. The results suggest that LLMs often resort to shortcut learning and still face challenges on our proposed benchmark. Re2Text Task Read the instruction and then answer the question using A or B. Instruction: (x, has part, y) indicates that y has a part called x. Question: (?, has part, hilt) A: Find an entity that has a part called hilt. B: Find an entity that is a part of hilt. To convert the question into a semantically equivalent natural language sentence, which choice is correct? Answer: Re2Text Task (hard) Read the instruction and then answer the question using A or B. Instruction: (x, has part, y) indicates that y has a part called x. Question: (?, has part, hilt) A: Find an entity that has a part called hilt. B: Find an entity that hilt contains. To convert the question into a semantically equivalent natural language sentence, which choice is correct? Answer: Text2Re Task Read the instruction and then answer the question using A or B. Instruction: (x, has part, y) indicates that y has a part called x. Question: Find an entity that possesses a specific component named hilt. A: (?, has part, hilt) B: (hilt, has part, ?) To convert the question into a semantically equivalent triple query, which choice is correct? Answer: Text2Re Task (hard) Read the instruction and then answer the question using A or B. Instruction: (x, has part, y) indicates that y has a part called x. Question: Find an entity that has a part called hilt. A: (?, has part, hilt) B: (hilt, has part, ?) To convert the question into a semantically equivalent triple query, which choice is correct? Answer: Figure 3: Examples of Re2Text and Text2Re tasks on converse relation. We additionally paraphrase the natural language representations (answer candidates for Re2Text, question for Text2Re) to make them differ from the sentences in the Instruction. to the two tasks, which will be evidenced by the 173 empirical results in our experiments (section 4.2). 174 An intuitive explanation is provided in figure 3. De-175 tailed zero-shot prompting methods can be found 176 in table 2. 2 177 Example Variants in Few-shot Prompting Be-178 side the variants on the test text, we additionally 179 introduce variants to the text in examples for the 180 few-shot prompting. Since we have identified the 181 most challenging settings for the two tasks in zero-182 shot, we will employ such settings for the test text 183 and dub them as hard tests in few-shot. Accord-184 ingly, we incorporate text variants to the examples 185 used in the few-shot prompting. Comprehensively, 186 the few-shot prompts used in our benchmark are 187 listed in table 3. Details of arrangement of text
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- Towards a Theoretical Understanding of the 'Reversal Curse' via Training DynamicsHanlin Zhu, Baihe Huang, Shaolun Zhang, Michael I. Jordan et al.NeurIPS 2024 · 37 citations
- Deep sequence models tend to memorize geometrically; it is unclear whyShahriar Noroozizadeh, Vaishnavh Nagarajan, Elan Rosenfeld, Sanjiv KumarICML 2026 · 11 citations
- Generative Subgraph Retrieval for Knowledge Graph-Grounded Dialog GenerationJinyoung Park, Minseok Joo, Joo-Kyung Kim, Hyunwoo J. KimEMNLP 2024 · 3 citations
- Large Language Models Meet Symbolic Provers for Logical Reasoning EvaluationChengwen Qi, Ren Ma, Bowen Li, He Du et al.ICLR 2025
Builds on8
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo et al.NeurIPS 2022 · 8,168 citations
- TruthfulQA: Measuring How Models Mimic Human FalsehoodsStephanie Lin, Jacob Hilton, Owain EvansACL 2022 · 3,228 citations
- True Few-Shot Learning with Language ModelsEthan Perez, Douwe Kiela, Kyunghyun ChoNeurIPS 2021 · 547 citations
Related papers
- FollowBench: A Multi-level Fine-grained Constraints Following Benchmark for Large Language ModelsYuxin Jiang, Yufei Wang, Xingshan Zeng, Wanjun Zhong et al.ACL 2024 · 10 citations
- Do Language Models Have a Common Sense regarding Time? Revisiting Temporal Commonsense Reasoning in the Era of Large Language ModelsRaghav Jain, Daivik Sojitra, Arkadeep Acharya, Sriparna Saha et al.EMNLP 2023 · 17 citations
- Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-JudgeMd. Tahmid Rahman Laskar, Israt Jahan, Elham Dolatabadi, Chun Peng et al.ACL 2025
- BizBench: A Quantitative Reasoning Benchmark for Business and FinanceMichael Krumdick, Rik Koncel-Kedziorski, Viet Dac Lai, Varshini Reddy et al.ACL 2024 · 10 citations
- ALCUNA: Large Language Models Meet New KnowledgeXunjian Yin, Baizhou Huang, Xiaojun WanEMNLP 2023 · 5 citations
