Lune

ICLR2024Top-tier venue

INViTE: INterpret and Control Vision-Language Models with Text Explanations

Haozhe Chen, Junfeng Yang, Carl Vondrick, Chengzhi Mao

2024Year
9Citations
6Top-tier citations

Abstract

Large-scale pre-trained vision foundation models, such as CLIP, have become de facto backbones for various vision tasks. However, due to their black-box nature, understanding the underlying rules behind these models' predictions and controlling model behaviors have remained open challenges. We present INViTE: a framework for INterpreting Vision Transformer's latent tokens with Text Explanations. Given a latent token, INViTE retains its semantic information to the final layer using transformer's local operations and retrieves the closest text for explanation. IN-ViTE enables understanding of model visual reasoning procedure without needing additional model training or data collection. Based on the obtained interpretations, INViTE allows for model editing that controls model reasoning behaviors and improves model robustness against biases and spurious correlations. Our code is available at https://github.com/tonychenxyz/vit-interpret . Recent works seek to interpret models with natural language. MILAN (Hernandez et al., 2022) finds natural language descriptions by maximizing the pointwise mutual information between input regions and human annotations. However, it requires additional data collection and training, thus cannot

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 335574bf-b3fb-47f9-bf8f-34efdd3771b7

Cited by top-tier papers6

Ask how each one uses it

Builds on24

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines