Decomposing and Editing Predictions by Modeling Model Computation
Harshay Shah, Andrew Ilyas, Aleksander Madry
Abstract
How does the internal computation of a machine learning model transform inputs into predictions? In this paper, we introduce a task called component modeling that aims to address this question. The goal of component modeling is to decompose an ML model's prediction in terms of its components -- simple functions (e.g., convolution filters, attention heads) that are the"building blocks"of model computation. We focus on a special case of this task, component attribution, where the goal is to estimate the counterfactual impact of individual components on a given prediction. We then present COAR, a scalable algorithm for estimating component attributions; we demonstrate its effectiveness across models, datasets, and modalities. Finally, we show that component attributions estimated with COAR directly enable model editing across five tasks, namely: fixing model errors, ``forgetting'' specific classes, boosting subpopulation robustness, localizing backdoor attacks, and improving robustness to typographic attacks. We provide code for COAR at https://github.com/MadryLab/modelcomponents .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e1859372-e48f-4fce-92cc-52a0f17ca292Cited by top-tier papers18
- ContextCite: Attributing Model Generation to ContextBenjamin Cohen-Wang, Harshay Shah, Kristian Georgiev, Aleksander MadryNeurIPS 2024 · 118 citations
- Optimal ablation for interpretabilityMaximilian Li, Lucas JansonNeurIPS 2024 · 32 citations
- Decomposing and Interpreting Image Representations via Text in ViTs Beyond CLIPSriram Balasubramanian, Samyadeep Basu, Soheil FeiziNeurIPS 2024 · 26 citations
- ProxySPEX: Inference-Efficient Interpretability via Sparse Feature Interactions in LLMsLandon Butler, Abhineet Agarwal, Justin Singh Kang, Yigit Efe Erginbas et al.NeurIPS 2025 · 19 citations
- Unveiling Concept Attribution in Diffusion ModelsNguyen Hung-Quang, Hoang Phan, Khoa D. DoanNeurIPS 2025 · 13 citations
Builds on43
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Locating and Editing Factual Associations in GPTKevin Meng, David Bau, Alex Andonian, Yonatan BelinkovNeurIPS 2022 · 3,415 citations
- Towards Automated Circuit Discovery for Mechanistic InterpretabilityArthur Conmy, Augustine N. Mavor-Parker, Aengus Lynch, Stefan Heimersheim et al.NeurIPS 2023 · 861 citations
- Investigating Gender Bias in Language Models Using Causal Mediation AnalysisJesse Vig, Sebastian Gehrmann, Yonatan Belinkov, Sharon Qian et al.NeurIPS 2020 · 851 citations
Related papers
- Compressed Sensing for Capability Localization in Large Language ModelsAnna Bair, Yixuan Xu, Mingjie Sun, Zico KolterICML 2026 · 1 citation
- DePass: Unified Feature Attributing by Simple Decomposed Forward PassXiangyu Hong, Che Jiang, Kai Tian, Biqing Qi et al.NeurIPS 2025 · 4 citations
- When Parts Are Greater Than Sums: Individual LLM Components Can Outperform Full ModelsTing-Yun Chang, Jesse Thomason, Robin JiaEMNLP 2024
- A Closer Look at Backdoor Attacks on CLIPShuo He, Zhifang Zhang, Feng Liu, Roy Ka-Wei Lee et al.ICML 2025
- Clustering Effect of Adversarial Robust ModelsYang Bai, Xin Yan, Yong Jiang, Shu-Tao Xia et al.NeurIPS 2021 · 13 citations
