Human Rationales as Attribution Priors for Explainable Stance Detection
Sahil Jayaram, Emily Allaway
Abstract
As NLP systems become better at detecting opinions and beliefs from text, it is important to ensure not only that models are accurate but also that they arrive at their predictions in ways that align with human reasoning. In this work, we present a method for imparting human-like rationalization to a stance detection model using crowdsourced annotations on a small fraction of the training data. We show that in a data-scarce setting, our approach can improve the reasoning of a state-of-the-art classifierparticularly for inputs containing challenging phenomena such as sarcasm-at no cost in predictive performance. Furthermore, we demonstrate that attention weights surpass a leading attribution method in providing faithful explanations of our model's predictions, thus serving as a computationally cheap and reliable source of attributions for our model.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 067e91cd-3eca-413e-a3d1-1adb7a8822b5Cited by top-tier papers4
- Neglected Free Lunch - Learning Image Classifiers Using Annotation ByproductsDongyoon Han, Junsuk Choe, Seonghyeok Chun, John Joon Young Chung et al.ICCV 2023 · 4 citations
- Why Should This Article Be Deleted? Transparent Stance Detection in Multilingual Wikipedia Editor DiscussionsLucie-Aimée Kaffee, Arnav Arora, Isabelle AugensteinEMNLP 2023 · 3 citations
- Leveraging Machine-Generated Rationales to Facilitate Social Meaning Detection in ConversationsRitam Dutt, Zhen Wu, Jiaxin Shi, Divyanshu Sheth et al.ACL 2024 · 2 citations
- SOCIAL SCAFFOLDS: A Generalization Framework for Social Understanding TasksRitam Dutt, Carolyn P. Rosé, Maarten SapEMNLP 2025
Builds on5
- A Diagnostic Study of Explainability Techniques for Text ClassificationPepa Atanasova, Jakob Grue Simonsen, Christina Lioma, Isabelle AugensteinEMNLP 2020 · 158 citations
- Enhancing Cross-target Stance Detection with Transferable Semantic-Emotion KnowledgeBowen Zhang, Min Yang, Xutao Li, Yunming Ye et al.ACL 2020 · 115 citations
- ERASER: A Benchmark to Evaluate Rationalized NLP ModelsJay DeYoung, Sarthak Jain, Nazneen Fatema Rajani, Eric P. Lehman et al.ACL 2020 · 36 citations
- Zero-Shot Stance Detection: A Dataset and Model using Generalized Topic RepresentationsEmily Allaway, Kathleen R. McKeownEMNLP 2020 · 7 citations
- Cross-Domain Label-Adaptive Stance DetectionMomchil Hardalov, Arnav Arora, Preslav Nakov, Isabelle AugensteinEMNLP 2021 · 3 citations
Related papers
- Human Attention Maps for Text Classification: Do Humans and Neural Networks Focus on the Same Words?Cansu Sen, Thomas Hartvigsen, Biao Yin, Xiangnan Kong et al.ACL 2020 · 56 citations
- S³-MSD: Large Vision-Language Model for Explainable and Generalizable Multi-modal Sarcasm DetectionZhihong Zhu, Fan Zhang, Yunyan Zhang, Jinghan Sun et al.AAAI 2026
- Less is More: Attention Supervision with Counterfactuals for Text ClassificationSeungtaek Choi, Haeju Park, Jinyoung Yeo, Seung-won HwangEMNLP 2020 · 16 citations
- Exploring the Efficacy of Automatically Generated Counterfactuals for Sentiment AnalysisLinyi Yang, Jiazheng Li, Padraig Cunningham, Yue Zhang et al.ACL 2021
- Towards Multi-Modal Sarcasm Detection via Hierarchical Congruity Modeling with Knowledge EnhancementHui Liu, Wenya Wang, Haoliang LiEMNLP 2022 · 91 citations
