Revisiting a Pain in the Neck: A Semantic Reasoning Benchmark for Language Models
Yang Liu, Hongming Li, Melissa Xiaohui Qin, Chao Huang, Qiankun Liu
Abstract
We present SEMANTICQA, an evaluation suite designed to assess language models (LMs) in semantic phrase processing tasks. The benchmark consolidates existing multiword expression (MWE) resources and reorganizes them into a unified testbed. It covers both general lexical phenomena, such as lexical collocations, and three fine-grained categories: idiomatic expressions, noun compounds, and verbal constructions. Through SEMANTICQA, we assess LMs of diverse architectures and scales in extraction, classification, and interpretation tasks, as well as sequential task compositions. We reveal substantial performance variation, particularly on tasks requiring semantic reasoning, highlighting differences in reasoning efficacy and semantic understanding of LMs, providing insights for pushing LMs with stronger comprehension on non-trivial semantic phrases. The evaluation harness and data of SEMANTICQA are available at https: //github.com/jacklanda/SemanticQA .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on12
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes et al.ICLR 2020 · 4,112 citations
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
- Towards Understanding Factual Knowledge of Large Language ModelsXuming Hu, Junzhe Chen, Xiaochuan Li, Yufei Guo et al.ICLR 2024 · 21 citations
- Idiomatic Expression Paraphrasing without Strong SupervisionJianing Zhou, Ziheng Zeng, Hongyu Gong, Suma BhatAAAI 2022 · 12 citations
Related papers
- MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Long-tail KnowledgeJie He, Nan Hu, Wanqiu Long, Jiaoyan Chen et al.ACL 2026 · 1 citation
- Memorization or Reasoning? Exploring the Idiom Understanding of LLMsJisu Kim, Youngwoo Shin, Uiji Hwang, Jihun Choi et al.EMNLP 2025
- Evaluating the Impact of Verbal Multiword Expressions on Machine TranslationLinfeng Liu, Saptarshi Ghosh, Tianyu JiangACL 2026 · 2 citations
- M³GQA: A Multi-Entity Multi-Hop Multi-Setting Graph Question Answering BenchmarkBoci Peng, Yongchao Liu, Xiaohe Bo, Jiaxin Guo et al.ACL 2025 · 1 citation
- KG-MuLQA: A Framework for KG-based Multi-Level QA Extraction and Long-Context LLM EvaluationNikita Tatarinov, Vidhyakshaya Kannan, Haricharana Srinivasa, Arnav Raj et al.ACL 2026 · 2 citations
