Rationalizing Transformer Predictions via End-To-End Differentiable Self-Training
Marc Felix Brinner, Sina Zarrieß
Abstract
We propose an end-to-end differentiable training paradigm for stable training of a rationalized transformer classifier. Our approach results in a single model that simultaneously classifies a sample and scores input tokens based on their relevance to the classification. To this end, we build on the widely-used three-playergame for training rationalized models, which typically relies on training a rationale selector, a classifier and a complement classifier. We simplify this approach by making a single model fulfill all three roles, leading to a more efficient training paradigm that is not susceptible to the common training instabilities that plague existing approaches. Further, we extend this paradigm to produce class-wise rationales while incorporating recent advances in parameterizing and regularizing the resulting rationales, thus leading to substantially improved and state-of-the-art alignment with human annotations without any explicit supervision.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0bf3d8d8-18ca-46df-b899-6fc5099bc33eBuilds on12
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Deberta: decoding-Enhanced Bert with Disentangled AttentionPengcheng He, Xiaodong Liu, Jianfeng Gao, Weizhu ChenICLR 2021 · 3,729 citations
- Generic Attention-model Explainability for Interpreting Bi-Modal and Encoder-Decoder TransformersHila Chefer, Shir Gur, Lior WolfICCV 2021 · 451 citations
- DeBERTaV3: Improving DeBERTa using ELECTRA-Style Pre-Training with Gradient-Disentangled Embedding SharingPengcheng He, Jianfeng Gao, Weizhu ChenICLR 2023 · 394 citations
- Understanding Interlocking Dynamics of Cooperative RationalizationMo Yu, Yang Zhang, Shiyu Chang, Tommi S. JaakkolaNeurIPS 2021 · 52 citations
Related papers
- Enhancing the Rationale-Input Alignment for Self-explaining RationalizationWei Liu, Haozhao Wang, Jun Wang, Zhiying Deng et al.ICDE 2024 · 6 citations
- Adversarial Cooperative Rationalization: The Risk of Spurious Correlations in Even Clean DatasetsWei Liu, Zhongyu Niu, Lang Gao, Zhiying Deng et al.ICML 2025
- FR: Folded Rationalization with a Unified EncoderWei Liu, Haozhao Wang, Jun Wang, Ruixuan Li et al.NeurIPS 2022 · 33 citations
- SPECTRA: Sparse Structured Text RationalizationNuno Miguel Guerreiro, André F. T. MartinsEMNLP 2021 · 1 citation
- MARTA: Leveraging Human Rationales for Explainable Text ClassificationInes Arous, Ljiljana Dolamic, Jie Yang, Akansha Bhardwaj et al.AAAI 2021 · 47 citations
