Shedding Light on Software Engineering-specific Metaphors and Idioms
Mia Mohammad Imran, Preetha Chatterjee, Kostadin Damevski
摘要
Use of figurative language, such as metaphors and idioms, is common in our daily-life communications, and it can also be found in Software Engineering (SE) channels, such as comments on GitHub. Automatically interpreting figurative language is a challenging task, even with modern Large Language Models (LLMs), as it often involves subtle nuances. This is particularly true in the SE domain, where figurative language is frequently used to convey technical concepts, often bearing developer affect (e.g., 'spaghetti code'). Surprisingly, there is a lack of studies on how figurative language in SE communications impacts the performance of automatic tools that focus on understanding developer communications, e.g., bug prioritization, incivility detection. Furthermore, it is an open question to what extent state-of-the-art LLMs interpret figurative expressions in domain-specific communication such as software engineering. To address this gap, we study the prevalence and impact of figurative language in SE communication channels. This study contributes to understanding the role of figurative language in SE, the potential of LLMs in interpreting them, and its impact on automated SE communication analysis. Our results demonstrate the effectiveness of fine-tuning LLMs with figurative language in SE and its potential impact on automated tasks that involve affect. We found that, among three state-of-the-art LLMs, the best improved fine-tuned versions have an average improvement of 6.66% on a GitHub emotion classification dataset, 7.07% on a GitHub incivility classification dataset, and 3.71% on a Bugzilla bug report prioritization dataset.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper9
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel 等ICLR 2020 · 被引用 7,418 次
- Adversarial NLI: A New Benchmark for Natural Language UnderstandingYixin Nie, Adina Williams, Emily Dinan, Mohit Bansal 等ACL 2020 · 被引用 602 次
- Fast Changeset-based Bug Localization with BERTAgnieszka Ciborowska, Kostadin DamevskiICSE 2022 · 被引用 53 次
- IMPLI: Investigating NLI Models' Performance on Figurative LanguageKevin Stowe, Prasetya Ajie Utama, Iryna GurevychACL 2022 · 被引用 52 次
- FLUTE: Figurative Language Understanding through Textual ExplanationsTuhin Chakrabarty, Arkadiy Saakyan, Debanjan Ghosh, Smaranda MuresanEMNLP 2022 · 被引用 35 次
相关 Paper
- Data Augmentation for Improving Emotion Recognition in Software Engineering CommunicationMia Mohammad Imran, Yashasvi Jain, Preetha Chatterjee, Kostadin DamevskiASE 2022 · 被引用 22 次
- Uncovering the Causes of Emotions in Software Developer Communication Using Zero-shot LLMsMia Mohammad Imran, Preetha Chatterjee, Kostadin DamevskiICSE 2024 · 被引用 23 次
- FLUID QA: A Multilingual Benchmark for Figurative Language Usage in Dialogue across English, Chinese, and KoreanSeoyoon Park, Hyeji Choi, Minseon Kim, Subin An 等EMNLP 2025 · 被引用 1 次
- Inside Out: Uncovering How Comment Internalization Steers LLMs for Better or WorseAaron Imani, Mohammad Moshirpour, Iftekhar AhmedICSE 2026
- An Empirical Study on Commit Message Generation Using LLMs via In-Context LearningYifan Wu, Yunpeng Wang, Ying Li, Wei Tao 等ICSE 2025 · 被引用 1 次
