Lune

NeurIPS2024Top-tier venue

Measuring Per-Unit Interpretability at Scale Without Humans

Roland S. Zimmermann, David A. Klindt, Wieland Brendel

2024Year
5Citations

Abstract

In today’s era, whatever we can measure at scale, we can optimize. So far, measuring the interpretability of units in deep neural networks (DNNs) for computer vision still requires direct human evaluation and is not scalable. As a result, the inner workings of DNNs remain a mystery despite the remarkable progress we have seen in their applications. In this work, we introduce the first scalable method to measure the per-unit interpretability in vision DNNs. This method does not require any human evaluations, yet its prediction correlates well with existing human interpretability measurements. We validate its predictive power through an interventional human psychophysics study. We demonstrate the usefulness of this measure by performing previously infeasible experiments: (1) A large-scale interpretability analysis across more than 70 million units from 835 computer vision models, and (2) an extensive analysis of how units transform during training. We find an anti-correlation between a model’s downstream classification performance and per-unit interpretability, which is also observable during model training. Furthermore, we see that a layer’s location and width influence its interpretability. Online version, code and interactive visualizations available at brendel-group.github.io/mis.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 091fefa0-4e4d-4c41-b818-7bd455cf1369

Builds on10

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines