Towards Improved Input Masking for Convolutional Neural Networks
Sriram Balasubramanian, Soheil Feizi
摘要
The ability to remove features from the input of machine learning models is very important to understand and interpret model predictions. However, this is non-trivial for vision models since masking out parts of the input image typically causes large distribution shifts. This is because the baseline color used for masking (typically grey or black) is out of distribution. Furthermore, the shape of the mask itself can contain unwanted signals which can be used by the model for its predictions. Recently, there has been some progress in mitigating this issue (called missingness bias) in image masking for vision transformers. In this work, we propose a new masking method for CNNs we call layer masking in which the missingness bias caused by masking is reduced to a large extent. Intuitively, layer masking applies a mask to intermediate activation maps so that the model only processes the unmasked input. We show that our method (i) is able to eliminate or minimize the influence of the mask shape or color on the output of the model, and (ii) is much better than replacing the masked region by black or grey for input perturbation based interpretability techniques like LIME. Thus, layer masking is much less affected by missingness bias than other masking strategies. We also demonstrate how the shape of the mask may leak information about the class, thus affecting estimates of model reliance on class-relevant features derived from input masking. Furthermore, we discuss the role of data augmentation techniques for tackling this problem, and argue that they are not sufficient for preventing model reliance on mask shape.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Decomposing and Interpreting Image Representations via Text in ViTs Beyond CLIPSriram Balasubramanian, Samyadeep Basu, Soheil FeiziNeurIPS 2024 · 被引用 26 次
- Missingness Bias Calibration in Feature Attribution ExplanationsShailesh Sridhar, Anton Xue, Eric WongICLR 2026
- Towards Human-Understandable Multi-Dimensional Concept DiscoveryArne Grobrügge, Niklas Kühl, Gerhard Satzger, Philipp SpitzerCVPR 2025
- ConEx: Human-Interpretable Saliency Maps via Concept-Aware AttributionYehonatan Elisha, Oren Barkan, Ziv Haddad, Noam KoenigsteinICML 2026
它引用的顶会 Paper15
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- RandAugment: Practical Automated Data Augmentation with a Reduced Search SpaceEkin Dogus Cubuk, Barret Zoph, Jonathon Shlens, Quoc LeNeurIPS 2020 · 被引用 4,453 次
- Random Erasing Data AugmentationZhun Zhong, Liang Zheng, Guoliang Kang, Shaozi Li 等AAAI 2020 · 被引用 4,134 次
- Intriguing Properties of Vision TransformersMuzammal Naseer, Kanchana Ranasinghe, Salman Khan, Munawar Hayat 等NeurIPS 2021 · 被引用 863 次
- Balanced Datasets Are Not Enough: Estimating and Mitigating Gender Bias in Deep Image RepresentationsTianlu Wang, Jieyu Zhao, Mark Yatskar, Kai-Wei Chang 等ICCV 2019 · 被引用 469 次
相关 Paper
- Missingness Bias in Model DebuggingSaachi Jain, Hadi Salman, Eric Wong, Pengchuan Zhang 等ICLR 2022 · 被引用 45 次
- What is Missing? Explaining Neurons Activated by Absent ConceptsRobin Hesse, Simone Schaub-Meyer, Janina Hesse, Bernt Schiele 等ICML 2026 · 被引用 1 次
- Masked Image Training for Generalizable Deep Image DenoisingHaoyu Chen, Jinjin Gu, Yihao Liu, Salma Abdel Magid 等CVPR 2023
- Token Activation Map to Visually Explain Multimodal LLMsYi Li, Hualiang Wang, Xinpeng Ding, Haonan Wang 等ICCV 2025 · 被引用 3 次
- Small Transformers Don’t Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and Implications for Mechanistic InterpretabilityLuca Baroni, Galvin Khara, Joachim Schaeffer, Marat Subkhankulov 等ICLR 2026 · 被引用 8 次
