Bridging Fairness and Explainability: Can Input-Based Explanations Promote Fairness in Hate Speech Detection?
Yifan Wang, Mayank Jobanputra, Ji-Ung Lee, Soyoung Oh, Isabel Valera, Vera Demberg
Abstract
Natural language processing (NLP) models often replicate or amplify social bias from training data, raising concerns about fairness. At the same time, their black-box nature makes it difficult for users to recognize biased predictions and for developers to effectively mitigate them. While some studies suggest that input-based explanations can help detect and mitigate bias, others question their reliability in ensuring fairness. Existing research on explainability in fair NLP has been predominantly qualitative, with limited large-scale quantitative analysis. In this work, we conduct the first systematic study of the relationship between explainability and fairness in hate speech detection, focusing on both encoder- and decoder-only models. We examine three key dimensions: (1) identifying biased predictions, (2) selecting fair models, and (3) mitigating bias during model training. Our findings show that input-based explanations can effectively detect biased predictions and serve as useful supervision for reducing bias during training, but they are unreliable for selecting fair models among candidates. Our code is available at https://github.com/Ewanwong/fairness_x_explainability.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f59ce236-19d2-4cad-a242-014afdfbe509Builds on20
- ContextCite: Attributing Model Generation to ContextBenjamin Cohen-Wang, Harshay Shah, Kristian Georgiev, Aleksander MadryNeurIPS 2024 · 118 citations
- Language (Technology) is Power: A Critical Survey of "Bias" in NLPSu Lin Blodgett, Solon Barocas, Hal Daumé III, Hanna M. WallachACL 2020 · 68 citations
- ATMAN: Understanding Transformer Predictions Through Memory Efficient Attention ManipulationBjörn Deiseroth, Mayukh Deb, Samuel Weinbach, Manuel Brack et al.NeurIPS 2023 · 45 citations
- ERASER: A Benchmark to Evaluate Rationalized NLP ModelsJay DeYoung, Sarthak Jain, Nazneen Fatema Rajani, Eric P. Lehman et al.ACL 2020 · 36 citations
- Interpreting Language Models with Contrastive ExplanationsKayo Yin, Graham NeubigEMNLP 2022 · 32 citations
Related papers
- Intrinsic Bias Metrics Do Not Correlate with Application BiasSeraphina Goldfarb-Tarrant, Rebecca Marchant, Ricardo Muñoz Sánchez, Mugdha Pandya et al.ACL 2021
- From Pretraining Data to Language Models to Downstream Tasks: Tracking the Trails of Political Biases Leading to Unfair NLP ModelsShangbin Feng, Chan Young Park, Yuhan Liu, Yulia TsvetkovACL 2023 · 117 citations
- Towards Conceptualization of "Fair Explanation": Disparate Impacts of anti-Asian Hate Speech Explanations on Content ModeratorsTin Nguyen, Jiannan Xu, Aayushi Roy, Hal Daumé III et al.EMNLP 2023 · 1 citation
- Constructing Fair Latent Space for Intersection of Fairness and ExplainabilityHyungjun Joo, Hyeonggeun Han, Sehwan Kim, Sangwoo Hong et al.AAAI 2025 · 2 citations
- BanHADEX: Towards Explainable HAte Speech Detection in Bangla Using Human Annotated EXplanationFaisal Hossain Raquib, Akm Moshiur Rahman Mazumder, Md Fahim, Md. Tahmid Hasan Fuad et al.ACL 2026
