Data Augmentation for Improving Emotion Recognition in Software Engineering Communication
Mia Mohammad Imran, Yashasvi Jain, Preetha Chatterjee, Kostadin Damevski
Abstract
Emotions (e.g., Joy, Anger) are prevalent in daily software engineering (SE) activities, and are known to be significant indicators of work productivity (e.g., bug fixing efficiency). Recent studies have shown that directly applying general purpose emotion classification tools to SE corpora is not effective. Even within the SE domain, tool performance degrades significantly when trained on one communication channel and evaluated on another (e.g, StackOverflow vs. GitHub comments). Retraining a tool with channel-specific data takes significant effort since manually annotating a large dataset of ground truth data is expensive. In this paper, we address this data scarcity problem by automatically creating new training data using a data augmentation technique. Based on an analysis of the types of errors made by popular SE-specific emotion recognition tools, we specifically target our data augmentation strategy in order to improve the performance of emotion recognition. Our results show an average improvement of 9.3% in micro F1-Score for three existing emotion classification tools (ESEM-E, EMTk, SEntiMoji) when trained with our best augmentation strategy.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ab61aad4-ef19-4263-bb2c-5571e8e10e09Cited by top-tier papers4
- Uncovering the Causes of Emotions in Software Developer Communication Using Zero-shot LLMsMia Mohammad Imran, Preetha Chatterjee, Kostadin DamevskiICSE 2024 · 23 citations
- Shedding Light on Software Engineering-specific Metaphors and IdiomsMia Mohammad Imran, Preetha Chatterjee, Kostadin DamevskiICSE 2024 · 6 citations
- A Weak Supervision-Based Approach to Improve Chatbots for Code RepositoriesFarbod Farhour, Ahmad Abdellatif, Essam Mansour, Emad ShihabFSE 2024 · 2 citations
- Toxicity Ahead: Forecasting Conversational Derailment on GitHubMia Mohammad Imran, Robert Zita, Rahat Rizvi Rahman, Preetha Chatterjee et al.ICSE 2026
Builds on10
- Unsupervised Data Augmentation for Consistency TrainingQizhe Xie, Zihang Dai, Eduard H. Hovy, Thang Luong et al.NeurIPS 2020 · 2,774 citations
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- Do Not Have Enough Data? Deep Learning to the Rescue!Ateret Anaby-Tavor, Boaz Carmeli, Esther Goldbraich, Amir Kantor et al.AAAI 2020 · 398 citations
- Recognizing developers' emotions while programmingDaniela Girardi, Nicole Novielli, Davide Fucci, Filippo LanubileICSE 2020 · 60 citations
- Textual Data Augmentation for Efficient Active Learning on Tiny DatasetsHusam Quteineh, Spyridon Samothrakis, Richard F. E. SutcliffeEMNLP 2020 · 47 citations
Related papers
- "I Didn't Know I Looked Angry": Characterizing Observed Emotion and Reported Affect at WorkHarmanpreet Kaur, Daniel McDuff, Alex C. Williams, Jaime Teevan et al.CHI 2022 · 68 citations
- Multilingual training for Software EngineeringToufique Ahmed, Premkumar T. DevanbuICSE 2022
- The Landscape of Toxicity: An Empirical Investigation of Toxicity on GitHubJaydeb Sarker, Asif Kamal Turzo, Amiangshu BosuFSE 2025 · 1 citation
- AU-assisted Graph Attention Convolutional Network for Micro-Expression RecognitionHong-Xia Xie, Ling Lo, Hong-Han Shuai, Wen-Huang ChengACM MM 2020 · 189 citations
- Assessing Emoji Use in Modern Text Processing ToolsAbu Awal Md Shoeb, Gerard de MeloACL 2021
