Easy as PIE? Identifying Multi-Word Expressions with LLMs
Kai Golan Hashiloni, Ofri Hefetz, Kfir Bar
Abstract
We investigate the identification of idiomatic expressions-a semantically noncompositional subclass of multiword expressions (MWEs)-in running text using large language models (LLMs) without any fine-tuning. Instead, we adopt a prompt-based approach and evaluate a range of prompting strategies, including zero-shot, few-shot, and chain-of-thought variants, across multiple languages, datasets, and model types. Our experiments show that, with well-crafted prompts, LLMs can perform competitively with supervised models trained on annotated data. These findings highlight the potential of prompt-based LLMs as a flexible and effective alternative for idiomatic expression identification.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 927bda96-e61c-4e75-a96f-a710e03109b5Builds on5
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Self-Consistency Improves Chain of Thought Reasoning in Language ModelsXuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le et al.ICLR 2023 · 681 citations
- Rolling the DICE on Idiomaticity: How LLMs Fail to Grasp ContextMaggie Mi, Aline Villavicencio, Nafise Sadat MoosaviACL 2025 · 9 citations
- Investigating Large Language Models for Complex Word Identification in Multilingual and Multidomain SetupsRazvan-Alexandru Smadu, David-Gabriel Ion, Dumitru-Clementin Cercel, Florin Pop et al.EMNLP 2024 · 4 citations
- CoAM: Corpus of All-Type Multiword ExpressionsYusuke Ide, Joshua Tanner, Adam Nohejl, Jacob Hoffman et al.ACL 2025
Related papers
- Multilingual Large Language Models Are Not (Yet) Code-SwitchersRuochen Zhang, Samuel Cahyawijaya, Jan Christian Blaise Cruz, Genta Indra Winata et al.EMNLP 2023 · 19 citations
- Interactive and Visual Prompt Engineering for Ad-hoc Task Adaptation with Large Language ModelsHendrik Strobelt, Albert Webson, Victor Sanh, Benjamin Hoover et al.IEEE VIS 2022 · 191 citations
- Pre-trained Language Models Can be Fully Zero-Shot LearnersXuandong Zhao, Siqi Ouyang, Zhiguo Yu, Ming Wu et al.ACL 2023 · 22 citations
- RulePrompt: Weakly Supervised Text Classification with Prompting PLMs and Self-Iterative Logical RulesMiaomiao Li, Jiaqi Zhu, Yang Wang, Yi Yang et al.WWW 2024 · 5 citations
- Large Language Models for Anomaly Detection in Computational Workflows: From Supervised Fine-Tuning to In-Context LearningHongwei Jin, George Papadimitriou, Krishnan Raghavan, Pawel Zuk et al.SC 2024 · 14 citations
