An Unsupervised Approach to Achieve Supervised-Level Explainability in Healthcare Records
Joakim Edin, Maria Maistro, Lars Maaløe, Lasse Borgholt, Jakob D. Havtorn, Tuukka Ruotsalo
Abstract
Electronic healthcare records are vital for patient safety as they document conditions, plans, and procedures in both free text and medical codes. Language models have significantly enhanced the processing of such records, streamlining workflows and reducing manual data entry, thereby saving healthcare providers significant resources. However, the black-box nature of these models often leaves healthcare professionals hesitant to trust them. State-ofthe-art explainability methods increase model transparency but rely on human-annotated evidence spans, which are costly. In this study, we propose an approach to produce plausible and faithful explanations without needing such annotations. We demonstrate on the automated medical coding task that adversarial robustness training improves explanation plausibility and introduce AttInGrad, a new explanation method superior to previous ones. By combining both contributions in a fully unsupervised setup, we produce explanations of comparable quality, or better, to that of a supervised approach. We release our code and model weights. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 388bbdb4-538a-44fc-a7c0-8a41f8dd3367Cited by top-tier papers3
- Towards Explainable Diagnosis: A Self-learned Explanatory Knowledge Base ApproachDongqi Huang, Tong Zhou, Zhuoran Jin, Shenghui Shi et al.ACL 2026
- Less is More: Explainable and Efficient ICD Code Prediction with Clinical EntitiesJames C. Douglas, Yidong Gan, Ben Hachey, Jonathan K. KummerfeldACL 2025
- ICDAGENT: Empowering Agentic Large Language Models for Explainable Medical CodingZiyi Yin, Yuanpu Cao, Ting Wang, Jinghui Chen et al.ACL 2026
Builds on6
- ERASER: A Benchmark to Evaluate Rationalized NLP ModelsJay DeYoung, Sarthak Jain, Nazneen Fatema Rajani, Eric P. Lehman et al.ACL 2020 · 36 citations
- Incorporating Residual and Normalization Layers into Analysis of Masked Language ModelsGoro Kobayashi, Tatsuki Kuribayashi, Sho Yokoi, Kentaro InuiEMNLP 2021 · 28 citations
- Discriminative Feature Attributions: Bridging Post Hoc Explainability and Inherent InterpretabilityUsha Bhalla, Suraj Srinivas, Himabindu LakkarajuNeurIPS 2023 · 18 citations
- Ignorance is Bliss: Robust Control via Information GatingManan Tomar, Riashat Islam, Matthew E. Taylor, Sergey Levine et al.NeurIPS 2023 · 14 citations
- MDACE: MIMIC Documents Annotated with Code EvidenceHua Cheng, Rana Jafari, April Russell, Russell Klopfer et al.ACL 2023 · 11 citations
Related papers
- Beyond Label Attention: Transparency in Language Models for Automated Medical Coding via Dictionary LearningJohn Wu, David Wu, Jimeng SunEMNLP 2024 · 4 citations
- Faithful Serum: Mitigating the Faithfulness Gap in Textual Explanations of LLM Decisions via Attribution GuidanceBar Alon, Itamar Zimerman, Lior WolfACL 2026
- Explaining Black-Box Language Models: Learning to Optimize Linguistically-Structured Word SubsetsMinyoung Hwang, Seokhyun Lee, Changhee LeeKDD 2026
- Exploring Accurate and Transparent Domain Adaptation in Predictive Healthcare via Concept-Grounded Orthogonal InferencePengfei Hu, Chang Lu, Feifan Liu, Yue NingICML 2026 · 1 citation
- Robust Explanation for Free or At the Cost of FaithfulnessZeren Tan, Yang TianICML 2023 · 12 citations
