MDACE: MIMIC Documents Annotated with Code Evidence
Hua Cheng, Rana Jafari, April Russell, Russell Klopfer, Edmond Lu, Benjamin Striner, Matthew Gormley
Abstract
We introduce a dataset for evidence/rationale extraction on an extreme multi-label classification task over long medical documents. One such task is Computer-Assisted Coding (CAC) which has improved significantly in recent years thanks to advances in machine learning technologies. However, simply predicting a set of final codes for a patient encounter is insufficient, as CAC systems are required to provide supporting textual evidence to justify the billing codes. A model able to produce accurate and reliable supporting evidence for each code would be a tremendous benefit. However, a human-annotated code evidence corpus is extremely difficult to create because it requires specialized knowledge. In this paper, we introduce MDACE, the first publicly available code evidence dataset, which is built on a subset of the MIMIC-III (English) clinical records. The dataset -annotated by professional medical coders -consists of 302 Inpatient charts with 3,934 evidence spans and 52 Profee charts with 5,563 evidence spans. We implemented several evidence extraction methods based on the EffectiveCAN model (Liu et al., 2021) to establish baseline performance on this dataset. MDACE can be used to evaluate code evidence extraction methods for CAC systems, as well as the accuracy and interpretability of deep learning models for multi-label classification. We believe that the release of MDACE will greatly improve the understanding and application of deep learning technologies for medical coding and document classification.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0687fb91-9701-433f-bb7e-c81ef82b5b6fCited by top-tier papers4
- An Unsupervised Approach to Achieve Supervised-Level Explainability in Healthcare RecordsJoakim Edin, Maria Maistro, Lars Maaløe, Lasse Borgholt et al.EMNLP 2024 · 4 citations
- Less is More: Explainable and Efficient ICD Code Prediction with Clinical EntitiesJames C. Douglas, Yidong Gan, Ben Hachey, Jonathan K. KummerfeldACL 2025
- ICDAGENT: Empowering Agentic Large Language Models for Explainable Medical CodingZiyi Yin, Yuanpu Cao, Ting Wang, Jinghui Chen et al.ACL 2026
- Aligning AI Research with the Needs of Clinical Coding Workflows: Eight Recommendations Based on US Data Analysis and Critical ReviewYidong Gan, Maciej Rybinski, Ben Hachey, Jonathan K. KummerfeldACL 2025
Builds on2
- Understanding Interlocking Dynamics of Cooperative RationalizationMo Yu, Yang Zhang, Shiyu Chang, Tommi S. JaakkolaNeurIPS 2021 · 52 citations
- Effective Convolutional Attention Network for Multi-label Clinical Document ClassificationYang Liu, Hua Cheng, Russell Klopfer, Matthew R. Gormley et al.EMNLP 2021 · 51 citations
Related papers
- ICD Coding from Clinical Text Using Multi-Filter Residual Convolutional Neural NetworkFei Li, Hong YuAAAI 2020 · 201 citations
- Multimodal Medical Code TokenizerXiaorui Su, Shvat Messica, Yepeng Huang, Ruth Johnson et al.ICML 2025 · 2 citations
- RuCCoD: Towards Automated ICD Coding in RussianAlexandr Nesterov, Andrey Sakhovskiy, Ivan Sviridov, Airat Valiev et al.EMNLP 2025 · 1 citation
- Multi-Label Few-Shot ICD Coding as Autoregressive Generation with PromptZhichao Yang, Sunjae Kwon, Zonghai Yao, Hong YuAAAI 2023 · 29 citations
- EviCare: Enhancing Diagnosis Prediction with Deep Model-Guided Evidence for In-Context ReasoningHengyu Zhang, Xuyun Zhang, Pengxiang Zhan, Linhao Luo et al.KDD 2026
