Probing Semantic Alignment, Lexical Invariance, and Syntactic Influence in LLM Metaphor Processing
Fengying Ye, Shanshan Wang, Lidia S. Chao, Derek F. Wong
Abstract
Large language models (LLMs) achieve strong performance on metaphor detection and interpretation tasks, yet it remains unclear what such behavioral success reveals about metaphor processing. We present a diagnostic analysis that examines the limits of behavioral evidence by probing three complementary dimensions: semantic attribute alignment, lexical invariance, and syntactic sensitivity. Using geometric probing, we assess whether model-generated interpretations align with reference semantic attributes; through context-varying substitution, we analyze the stability of lexical associations between metaphorical and literal expressions; and via controlled syntactic perturbations, we examine sensitivity in metaphor detection. Our analysis reveals that LLM-generated interpretations can exhibit semantic drift relative to reference attributes; stable lexical anchors persist across contextual conditions, potentially supporting conventional metaphors while biasing novel metaphors requiring contextual integration; and detection performance is sensitive to syntactic irregularities. These findings suggest that strong behavioral performance may reflect heterogeneous underlying signals, highlighting the need for caution when interpreting metaphor benchmarks as evidence of robust, integrated semantic understanding.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 60b7d5b1-5c40-4b98-ae54-bfa389cc5946Cited by top-tier papers1
Ask how each one uses itBuilds on8
- Calibrate Before Use: Improving Few-shot Performance of Language ModelsZihao Zhao, Eric Wallace, Shi Feng, Dan Klein et al.ICML 2021 · 1,843 citations
- Explainable Metaphor Identification Inspired by Conceptual Metaphor TheoryMengshi Ge, Rui Mao, Erik CambriaAAAI 2022 · 67 citations
- Does GPT-3 Grasp Metaphors? Identifying Metaphor Mappings with Generative Language ModelsLennart Wachowiak, Dagmar GromannACL 2023 · 11 citations
- Does It Capture STEL? A Modular, Similarity-based Linguistic Style Evaluation FrameworkAnna Wegmann, Dong NguyenEMNLP 2021 · 7 citations
- G-IdiomAlign: A Gloss-Pivoted Benchmark for Cross-Lingual Idiom AlignmentFengying Ye, Yanming Sun, Runzhe Zhan, Lidia S. Chao et al.ACL 2026 · 1 citation
Related papers
- Metaphors in Pre-Trained Language Models: Probing and Generalization Across Datasets and LanguagesEhsan Aghazadeh, Mohsen Fayyaz, Yadollah YaghoobzadehACL 2022
- Do Large Language Models Truly Grasp Addition? A Rule-Focused Diagnostic Using Two-Integer ArithmeticYang Yan, Yu Lu, Renjun Xu, Zhenzhong LanEMNLP 2025 · 1 citation
- When Language Models Lose Their Mind: The Consequences of Brain MisalignmentGabriele Merlin, Mariya TonevaICLR 2026 · 3 citations
- Mechanisms vs. Outcomes: Probing for Syntax Fails to Explain Performance on Targeted Syntactic EvaluationsAnanth Agarwal, Jasper Jian, Christopher D. Manning, Shikhar MurtyEMNLP 2025 · 5 citations
- Can Large Language Models Interpret Noun-Noun Compounds? A Linguistically-Motivated Study on Lexicalized and Novel CompoundsGiulia Rambelli, Emmanuele Chersoni, Claudia Collacciani, Marianna BolognesiACL 2024
