Dividing and Conquering a BlackBox to a Mixture of Interpretable Models: Route, Interpret, Repeat
Shantanu Ghosh, Ke Yu, Forough Arabshahi, Kayhan Batmanghelich
Abstract
ML model design either starts with an interpretable model or a Blackbox and explains it post hoc. Blackbox models are flexible but difficult to explain, while interpretable models are inherently explainable. Yet, interpretable models require extensive ML knowledge and tend to be less flexible and underperforming than their Blackbox variants. This paper aims to blur the distinction between a post hoc explanation of a Blackbox and constructing interpretable models. Beginning with a Blackbox, we iteratively carve out a mixture of interpretable experts (MoIE) and a residual network. Each interpretable model specializes in a subset of samples and explains them using First Order Logic (FOL), providing basic reasoning on concepts from the Blackbox. We route the remaining samples through a flexible residual. We repeat the method on the residual network until all the interpretable models explain the desired proportion of data. Our extensive experiments show that our route, interpret, and repeat approach (1) identifies a diverse set of instance-specific concepts with high concept completeness via MoIE without compromising in performance, (2) identifies the relatively "harder" samples to explain via residuals, (3) outperforms the interpretable by-design models by significant margins during test-time interventions, and (4) fixes the shortcut learned by the original Blackbox. The code for MoIE is publicly available at: https://github.com/batmanlab/ ICML-2023-Route-interpret-repeat .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- Energy-Based Concept Bottleneck Models: Unifying Prediction, Concept Intervention, and Probabilistic InterpretationsXinyue Xu, Yi Qin, Lu Mi, Hao Wang et al.ICLR 2024 · 32 citations
- Relational Concept Bottleneck ModelsPietro Barbiero, Francesco Giannini, Gabriele Ciravegna, Michelangelo Diligenti et al.NeurIPS 2024 · 21 citations
- Neural Concept BinderWolfgang Stammer, Antonia Wüst, David Steinmann, Kristian KerstingNeurIPS 2024
- I2MoE: Interpretable Multimodal Interaction-aware Mixture-of-ExpertsJiayi Xin, Sukwon Yun, Jie Peng, Inyoung Choi et al.ICML 2025
Builds on8
- Concept Bottleneck ModelsPang Wei Koh, Thao Nguyen, Yew Siang Tang, Stephen Mussmann et al.ICML 2020 · 1,233 citations
- Addressing Leakage in Concept Bottleneck ModelsMarton Havasi, Sonali Parbhoo, Finale Doshi-VelezNeurIPS 2022 · 163 citations
- Explanation by Progressive ExaggerationSumedha Singla, Brian Pollack, Junxiang Chen, Kayhan BatmanghelichICLR 2020 · 116 citations
- Entropy-Based Logic Explanations of Neural NetworksPietro Barbiero, Gabriele Ciravegna, Francesco Giannini, Pietro Lió et al.AAAI 2022 · 97 citations
- A Framework for Learning Ante-hoc Explainable Models via ConceptsAnirban Sarkar, Deepak Vijaykeerthy, Anindya Sarkar, Vineeth N. BalasubramanianCVPR 2022 · 40 citations
Related papers
- Understanding Cross-layer Contributions to Mixture-of-Experts Routing in LLMsWengang Li, Lingqi Zhang, Toshio Endo, Mohamed WahibICLR 2026
- Self-MoE: Towards Compositional Large Language Models with Self-Specialized ExpertsJunmo Kang, Leonid Karlinsky, Hongyin Luo, Zhen Wang et al.ICLR 2025
- Mixture of Experts Made Intrinsically InterpretableXingyi Yang, Constantin Venhoff, Ashkan Khakzar, Christian Schröder de Witt et al.ICML 2025
- RouterInterp: Understanding Superposed Specialisation in Mixture of Experts RoutingIlya Lasy, Nora Cai, Kola AyonrindeICML 2026
- A Closer Look at the Intervention Procedure of Concept Bottleneck ModelsSungbin Shin, Yohan Jo, Sungsoo Ahn, Namhoon LeeICML 2023 · 59 citations
