Enhancing Chemical Explainability Through Counterfactual Masking
Lukasz Janisiów, Marek Kochanczyk, Bartosz Michal Zielinski, Tomasz Danel
摘要
Molecular property prediction is a crucial task that guides the design of new compounds, including drugs and materials. While explainable artificial intelligence methods aim to scrutinize model predictions by identifying influential molecular substructures, many existing approaches rely on masking strategies that remove either atoms or atom-level features to assess importance via fidelity metrics. These methods, however, often fail to adhere to the underlying molecular distribution and thus yield unintuitive explanations. In this work, we propose counterfactual masking, a novel framework that replaces masked substructures with chemically reasonable fragments sampled from generative models trained to complete molecular graphs. Rather than evaluating masked predictions against implausible zeroed-out baselines, we assess them relative to counterfactual molecules drawn from the data distribution. Our method offers two key benefits: (1) molecular realism that underpins robust and distribution-consistent explanations, and (2) meaningful counterfactuals that directly indicate how structural modifications may affect predicted properties. We demonstrate that counterfactual masking is well-suited for benchmarking model explainers and yields more actionable insights across multiple datasets and property prediction tasks. Our approach bridges the gap between explainability and molecular design, offering a principled and generative path toward explainable machine learning in chemistry.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper4
- On Explainability of Graph Neural Networks via Subgraph ExplorationsHao Yuan, Haiyang Yu, Jie Wang, Kang Li 等ICML 2021 · 被引用 498 次
- ProtGNN: Towards Self-Explaining Graph Neural NetworksZaixi Zhang, Qi Liu, Hao Wang, Chengqiang Lu 等AAAI 2022 · 被引用 173 次
- Global Human-guided Counterfactual Explanations for Molecular Properties via Reinforcement LearningDanqing Wang, Antonis Antoniades, Kha-Dinh Luong, Edwin Zhang 等KDD 2024
- GenMol: A Drug Discovery Generalist with Discrete DiffusionSeul Lee, Karsten Kreis, Srimukh Prasad Veccham, Meng Liu 等ICML 2025
相关 Paper
- LeapFactual: Reliable Visual Counterfactual Explanation Using Conditional Flow MatchingZhuo Cao, Xuan Zhao, Lena Krieger, Hanno Scharr 等NeurIPS 2025 · 被引用 5 次
- FragFM: Hierarchical Framework for Efficient Molecule Generation via Fragment-Level Discrete Flow MatchingJoongwon Lee, Seonghwan Kim, Seokhyun Moon, Hyunwoo Kim 等ICLR 2026 · 被引用 6 次
- Multi-Objective Molecule Generation using Interpretable SubstructuresWengong Jin, Regina Barzilay, Tommi S. JaakkolaICML 2020 · 被引用 238 次
- MAGE: Model-Level Graph Neural Networks Explanations via Motif-based Graph GenerationZhaoning Yu, Hongyang GaoICLR 2025
- MAGNet: Motif-Agnostic Generation of Molecules from ScaffoldsLeon Hetzel, Johanna Sommer, Bastian Rieck, Fabian J. Theis 等ICLR 2025
