CoAM: Corpus of All-Type Multiword Expressions
Yusuke Ide, Joshua Tanner, Adam Nohejl, Jacob Hoffman, Justin Vasselli, Hidetaka Kamigaito, Taro Watanabe
Abstract
Multiword expressions (MWEs) refer to idiomatic sequences of multiple words. MWE identification, i.e., detecting MWEs in text, can play a key role in downstream tasks such as machine translation, but existing datasets for the task are inconsistently annotated, limited to a single type of MWE, or limited in size.
To enable reliable and comprehensive evaluation, we created CoAM: Corpus of All-Type Multiword Expressions, a dataset of 1.3K sentences constructed through a multi-step process to enhance data quality consisting of human annotation, human review, and automated consistency checking. Additionally, for the first time in a dataset for MWE identification, CoAM's MWEs are tagged with MWE types, such as NOUN and VERB, enabling fine-grained error analysis. 1 Annotations for CoAM were collected using a new interface created with our interface generator, which allows easy and flexible annotation of MWEs in any form. 2 Through experiments using CoAM, we find that a fine-tuned large language model outperforms MWEasWSD, which achieved the state-of-theart performance on the DiMSUM dataset. Furthermore, analysis using our MWE type tagged data reveals that VERB MWEs are easier than NOUN MWEs to identify across approaches.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9fc58d4f-131c-4943-a74b-b31b01b657ddCited by top-tier papers1
Ask how each one uses itBuilds on7
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 5,863 citations
- UniversalNER: Targeted Distillation from Large Language Models for Open Named Entity RecognitionWenxuan Zhou, Sheng Zhang, Yu Gu, Muhao Chen et al.ICLR 2024 · 118 citations
- Translate Meanings, Not Just Words: IdiomKB's Role in Optimizing Idiomatic Translation with Language ModelsShuang Li, Jiangjie Chen, Siyu Yuan, Xinyi Wu et al.AAAI 2024 · 44 citations
- Generationary or "How We Went beyond Word Sense Inventories and Learned to Gloss"Michele Bevilacqua, Marco Maru, Roberto NavigliEMNLP 2020 · 41 citations
- Understanding Jargon: Combining Extraction and Generation for Definition ModelingJie Huang, Hanyin Shao, Kevin Chen-Chuan Chang, Jinjun Xiong et al.EMNLP 2022 · 11 citations
Related papers
- Evaluating the Impact of Verbal Multiword Expressions on Machine TranslationLinfeng Liu, Saptarshi Ghosh, Tianyu JiangACL 2026 · 2 citations
- Revisiting a Pain in the Neck: A Semantic Reasoning Benchmark for Language ModelsYang Liu, Hongming Li, Melissa Xiaohui Qin, Chao Huang et al.ACL 2026
- Pre-tokenization of Multi-word Expressions in Cross-lingual Word EmbeddingsNaoki Otani, Satoru Ozaki, Xingyuan Zhao, Yucen Li et al.EMNLP 2020 · 7 citations
- BABEL: Bodies, Action and Behavior With English LabelsAbhinanda R. Punnakkal, Arjun Chandrasekaran, Nikos Athanasiou, Alejandra Quiros-Ramirez et al.CVPR 2021
- Improving Large-scale Paraphrase Acquisition and GenerationYao Dou, Chao Jiang, Wei XuEMNLP 2022 · 11 citations
