Beyond Label Attention: Transparency in Language Models for Automated Medical Coding via Dictionary Learning
John Wu, David Wu, Jimeng Sun
Abstract
Medical coding, the translation of unstructured clinical text into standardized medical codes, is a crucial but time-consuming healthcare practice. Though large language models (LLM) could automate the coding process and improve the efficiency of such tasks, interpretability remains paramount for maintaining patient trust. Current efforts in interpretability of medical coding applications rely heavily on label attention mechanisms, which often leads to the highlighting of extraneous tokens irrelevant to the ICD code. To facilitate accurate interpretability in medical language models, this paper leverages dictionary learning that can efficiently extract sparsely activated representations from dense language model embeddings in superposition. Compared with common label attention mechanisms, our model goes beyond token-level representations by building an interpretable dictionary which enhances the mechanistic-based explanations for each ICD code prediction, even when the highlighted tokens are medically irrelevant. We show that dictionary features are human interpretable, can elucidate the hidden meanings of upwards of 90% of medically irrelevant tokens, and steer model behavior.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext deeaee6f-8be6-4d82-8e6b-c105e1bb60baCited by top-tier papers1
Ask how each one uses itBuilds on6
- Sparse Autoencoders Find Highly Interpretable Features in Language ModelsRobert Huben, Hoagy Cunningham, Logan Riggs Smith, Aidan Ewart et al.ICLR 2024 · 1,072 citations
- Towards Automated Circuit Discovery for Mechanistic InterpretabilityArthur Conmy, Augustine N. Mavor-Parker, Aengus Lynch, Stefan Heimersheim et al.NeurIPS 2023 · 861 citations
- Generic Attention-model Explainability for Interpreting Bi-Modal and Encoder-Decoder TransformersHila Chefer, Shir Gur, Lior WolfICCV 2021 · 451 citations
- Optimizing Relevance Maps of Vision Transformers Improves RobustnessHila Chefer, Idan Schwartz, Lior WolfNeurIPS 2022 · 55 citations
- ERASER: A Benchmark to Evaluate Rationalized NLP ModelsJay DeYoung, Sarthak Jain, Nazneen Fatema Rajani, Eric P. Lehman et al.ACL 2020 · 36 citations
Related papers
- Temporal Sparse Autoencoders: Leveraging the Sequential Nature of Language for InterpretabilityUsha Bhalla, Alex Oesterling, Claudio Mayrink Verdun, Himabindu Lakkaraju et al.ICLR 2026 · 18 citations
- Towards Principled Evaluations of Sparse Autoencoders for Interpretability and ControlAleksandar Makelov, Georg Lange, Neel NandaICLR 2025
- Efficient Transcoder Adaptation for Fine-Tuned Models: Revealing Medical Reasoning Mechanisms in Large Language ModelsZhouxing Tan, Hanlin Xue, Yulong Wan, Ruochong Xiong et al.AAAI 2026
- A Concept-Based Explainability Framework for Large Multimodal ModelsJayneel Parekh, Pegah Khayatan, Mustafa Shukor, Alasdair Newson et al.NeurIPS 2024 · 48 citations
- Constructing Interpretable Features from Compositional Neuron GroupsOr David Shafran, Atticus Geiger, Mor GevaACL 2026 · 4 citations
