MARTA: Leveraging Human Rationales for Explainable Text Classification
Ines Arous, Ljiljana Dolamic, Jie Yang, Akansha Bhardwaj, Giuseppe Cuccu, Philippe Cudré-Mauroux
Abstract
Explainability is a key requirement for text classification in many application domains ranging from sentiment analysis to medical diagnosis or legal reviews. Existing methods often rely on "attention" mechanisms for explaining classification results by estimating the relative importance of input units. However, recent studies have shown that such mechanisms tend to mis-identify irrelevant input units in their explanation. In this work, we propose a hybrid human-AI approach that incorporates human rationales into attention-based text classification models to improve the explainability of classification results. Specifically, we ask workers to provide rationales for their annotation by selecting relevant pieces of text. We introduce MARTA, a Bayesian framework that jointly learns an attention-based model and the reliability of workers while injecting human rationales into model training. We derive a principled optimization algorithm based on variational inference with efficient updating rules for learning MARTA parameters. Extensive validation on real-world datasets shows that our framework significantly improves the state of the art both in terms of classification explainability and accuracy.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6dec5422-661d-4314-bb9c-3f4af851ab96Cited by top-tier papers3
- Thinking Like a Developer? Comparing the Attention of Humans with Neural Models of CodeMatteo Paltenghi, Michael PradelASE 2021 · 21 citations
- Compose with Me: Collaborative Music Inpainter for Symbolic Music InfillingZhejing Hu, Yan Liu, Gong Chen, Bruce X. B. YuAAAI 2025 · 2 citations
- Is Attention Explanation? An Introduction to the DebateAdrien Bibal, Rémi Cardon, David Alfter, Rodrigo Wilkens et al.ACL 2022
Builds on2
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel et al.ICLR 2020 · 7,418 citations
- Towards Transparent and Explainable Attention ModelsAkash Kumar Mohankumar, Preksha Nema, Sharan Narasimhan, Mitesh M. Khapra et al.ACL 2020 · 11 citations
Related papers
- Less is More: Attention Supervision with Counterfactuals for Text ClassificationSeungtaek Choi, Haeju Park, Jinyoung Yeo, Seung-won HwangEMNLP 2020 · 16 citations
- SPECTRA: Sparse Structured Text RationalizationNuno Miguel Guerreiro, André F. T. MartinsEMNLP 2021 · 1 citation
- RA3: A Human-in-the-loop Framework for Interpreting and Improving Image Captioning with Relation-Aware Attribution AnalysisLei Chai, Lu Qi, Hailong Sun, Jingzheng LiICDE 2024
- Unifying Model Explainability and Robustness for Joint Text Classification and Rationale ExtractionDongfang Li, Baotian Hu, Qingcai Chen, Tujie Xu et al.AAAI 2022 · 16 citations
- Multi-Dimensional Explanation of Target Variables from DocumentsDiego Antognini, Claudiu Musat, Boi FaltingsAAAI 2021 · 14 citations
