Enhancing Chemical Explainability Through Counterfactual Masking
Lukasz Janisiów, Marek Kochanczyk, Bartosz Michal Zielinski, Tomasz Danel
Abstract
Molecular property prediction is a crucial task that guides the design of new compounds, including drugs and materials. While explainable artificial intelligence methods aim to scrutinize model predictions by identifying influential molecular substructures, many existing approaches rely on masking strategies that remove either atoms or atom-level features to assess importance via fidelity metrics. These methods, however, often fail to adhere to the underlying molecular distribution and thus yield unintuitive explanations. In this work, we propose counterfactual masking, a novel framework that replaces masked substructures with chemically reasonable fragments sampled from generative models trained to complete molecular graphs. Rather than evaluating masked predictions against implausible zeroed-out baselines, we assess them relative to counterfactual molecules drawn from the data distribution. Our method offers two key benefits: (1) molecular realism that underpins robust and distribution-consistent explanations, and (2) meaningful counterfactuals that directly indicate how structural modifications may affect predicted properties. We demonstrate that counterfactual masking is well-suited for benchmarking model explainers and yields more actionable insights across multiple datasets and property prediction tasks. Our approach bridges the gap between explainability and molecular design, offering a principled and generative path toward explainable machine learning in chemistry.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cdf1cac2-1554-45b4-9119-7b5c5af2fddbBuilds on4
- On Explainability of Graph Neural Networks via Subgraph ExplorationsHao Yuan, Haiyang Yu, Jie Wang, Kang Li et al.ICML 2021 · 498 citations
- ProtGNN: Towards Self-Explaining Graph Neural NetworksZaixi Zhang, Qi Liu, Hao Wang, Chengqiang Lu et al.AAAI 2022 · 173 citations
- Global Human-guided Counterfactual Explanations for Molecular Properties via Reinforcement LearningDanqing Wang, Antonis Antoniades, Kha-Dinh Luong, Edwin Zhang et al.KDD 2024
- GenMol: A Drug Discovery Generalist with Discrete DiffusionSeul Lee, Karsten Kreis, Srimukh Prasad Veccham, Meng Liu et al.ICML 2025
Related papers
- LeapFactual: Reliable Visual Counterfactual Explanation Using Conditional Flow MatchingZhuo Cao, Xuan Zhao, Lena Krieger, Hanno Scharr et al.NeurIPS 2025 · 5 citations
- FragFM: Hierarchical Framework for Efficient Molecule Generation via Fragment-Level Discrete Flow MatchingJoongwon Lee, Seonghwan Kim, Seokhyun Moon, Hyunwoo Kim et al.ICLR 2026 · 6 citations
- Multi-Objective Molecule Generation using Interpretable SubstructuresWengong Jin, Regina Barzilay, Tommi S. JaakkolaICML 2020 · 238 citations
- MAGE: Model-Level Graph Neural Networks Explanations via Motif-based Graph GenerationZhaoning Yu, Hongyang GaoICLR 2025
- MAGNet: Motif-Agnostic Generation of Molecules from ScaffoldsLeon Hetzel, Johanna Sommer, Bastian Rieck, Fabian J. Theis et al.ICLR 2025
