Data Augmentation for Improving Emotion Recognition in Software Engineering Communication
Mia Mohammad Imran, Yashasvi Jain, Preetha Chatterjee, Kostadin Damevski
摘要
Emotions (e.g., Joy, Anger) are prevalent in daily software engineering (SE) activities, and are known to be significant indicators of work productivity (e.g., bug fixing efficiency). Recent studies have shown that directly applying general purpose emotion classification tools to SE corpora is not effective. Even within the SE domain, tool performance degrades significantly when trained on one communication channel and evaluated on another (e.g, StackOverflow vs. GitHub comments). Retraining a tool with channel-specific data takes significant effort since manually annotating a large dataset of ground truth data is expensive. In this paper, we address this data scarcity problem by automatically creating new training data using a data augmentation technique. Based on an analysis of the types of errors made by popular SE-specific emotion recognition tools, we specifically target our data augmentation strategy in order to improve the performance of emotion recognition. Our results show an average improvement of 9.3% in micro F1-Score for three existing emotion classification tools (ESEM-E, EMTk, SEntiMoji) when trained with our best augmentation strategy.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Uncovering the Causes of Emotions in Software Developer Communication Using Zero-shot LLMsMia Mohammad Imran, Preetha Chatterjee, Kostadin DamevskiICSE 2024 · 被引用 23 次
- Shedding Light on Software Engineering-specific Metaphors and IdiomsMia Mohammad Imran, Preetha Chatterjee, Kostadin DamevskiICSE 2024 · 被引用 6 次
- A Weak Supervision-Based Approach to Improve Chatbots for Code RepositoriesFarbod Farhour, Ahmad Abdellatif, Essam Mansour, Emad ShihabFSE 2024 · 被引用 2 次
- Toxicity Ahead: Forecasting Conversational Derailment on GitHubMia Mohammad Imran, Robert Zita, Rahat Rizvi Rahman, Preetha Chatterjee 等ICSE 2026
它引用的顶会 Paper10
- Unsupervised Data Augmentation for Consistency TrainingQizhe Xie, Zihang Dai, Eduard H. Hovy, Thang Luong 等NeurIPS 2020 · 被引用 2,774 次
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad 等ACL 2020 · 被引用 1,224 次
- Do Not Have Enough Data? Deep Learning to the Rescue!Ateret Anaby-Tavor, Boaz Carmeli, Esther Goldbraich, Amir Kantor 等AAAI 2020 · 被引用 398 次
- Recognizing developers' emotions while programmingDaniela Girardi, Nicole Novielli, Davide Fucci, Filippo LanubileICSE 2020 · 被引用 60 次
- Textual Data Augmentation for Efficient Active Learning on Tiny DatasetsHusam Quteineh, Spyridon Samothrakis, Richard F. E. SutcliffeEMNLP 2020 · 被引用 47 次
相关 Paper
- "I Didn't Know I Looked Angry": Characterizing Observed Emotion and Reported Affect at WorkHarmanpreet Kaur, Daniel McDuff, Alex C. Williams, Jaime Teevan 等CHI 2022 · 被引用 68 次
- Multilingual training for Software EngineeringToufique Ahmed, Premkumar T. DevanbuICSE 2022
- The Landscape of Toxicity: An Empirical Investigation of Toxicity on GitHubJaydeb Sarker, Asif Kamal Turzo, Amiangshu BosuFSE 2025 · 被引用 1 次
- AU-assisted Graph Attention Convolutional Network for Micro-Expression RecognitionHong-Xia Xie, Ling Lo, Hong-Han Shuai, Wen-Huang ChengACM MM 2020 · 被引用 189 次
- Assessing Emoji Use in Modern Text Processing ToolsAbu Awal Md Shoeb, Gerard de MeloACL 2021
