Learning Concise and Descriptive Attributes for Visual Recognition
An Yan, Yu Wang, Yiwu Zhong, Chengyu Dong, Zexue He, Yujie Lu, William Yang Wang, Jingbo Shang, Julian J. McAuley
摘要
Recent advances in foundation models present new opportunities for interpretable visual recognition – one can first query Large Language Models (LLMs) to obtain a set of attributes that describe each class, then apply vision-language models to classify images via these attributes. Pioneering work shows that querying thousands of attributes can achieve performance competitive with image features. However, our further investigation on 8 datasets reveals that LLM-generated attributes in a large quantity perform almost the same as random words. This surprising finding suggests that significant noise may be present in these attributes. We hypothesize that there exist subsets of attributes that can maintain the classification performance with much smaller sizes, and propose a novel learning-to-search method to discover those concise sets of attributes. As a result, on the CUB dataset, our method achieves performance close to that of massive LLM-generated attributes (e.g., 10k attributes for CUB), yet using only 32 attributes in total to distinguish 200 bird species. Furthermore, our new paradigm demonstrates several additional benefits: higher interpretability and interactivity for humans, and the ability to summarize knowledge for a recognition task.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper39
- VLG-CBM: Training Concept Bottleneck Models with Vision-Language GuidanceDivyansh Srivastava, Ge Yan, Lily WengNeurIPS 2024 · 被引用 87 次
- Scene Graph Generation with Role-Playing Large Language ModelsGuikun Chen, Jin Li, Wenguan WangNeurIPS 2024 · 被引用 33 次
- ArGue: Attribute-Guided Prompt Tuning for Vision-Language ModelsXinyu Tian, Shu Zou, Zhaoyuan Yang, Jing ZhangCVPR 2024 · 被引用 29 次
- Democratizing Fine-grained Visual Recognition with Large Language ModelsMingxuan Liu, Subhankar Roy, Wenjing Li, Zhun Zhong 等ICLR 2024 · 被引用 27 次
- ChatTracker: Enhancing Visual Tracking Performance via Chatting with Multimodal Large Language ModelYiming Sun, Fan Yu, Shaoxiang Chen, Yu Zhang 等NeurIPS 2024 · 被引用 21 次
它引用的顶会 Paper14
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech 等NeurIPS 2022 · 被引用 6,707 次
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 被引用 6,549 次
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen 等ICML 2021 · 被引用 5,401 次
- Concept Bottleneck ModelsPang Wei Koh, Thao Nguyen, Yew Siang Tang, Stephen Mussmann 等ICML 2020 · 被引用 1,233 次
相关 Paper
- Visual Classification via Description from Large Language ModelsSachit Menon, Carl VondrickICLR 2023 · 被引用 57 次
- Attribute-formed Class-specific Concept Space: Endowing Language Bottleneck Model with Better Interpretability and ScalabilityJianyang Zhang, Qianli Luo, Guowu Yang, Wenjing Yang 等CVPR 2025
- Learning Interpretable Queries for Explainable Image Classification with Information PursuitStefan Kolek, Aditya Chattopadhyay, Kwan Ho Ryan Chan, Héctor Andrade-Loarca 等ICCV 2025 · 被引用 1 次
- Verbalized Representation Learning for Interpretable Few-Shot GeneralizationCheng-Fu Yang, Da Yin, Wenbo Hu, Heng Ji 等ICCV 2025 · 被引用 1 次
- Language-driven Fine-grained RetrievalShijie Wang, Xin Yu, Yadan Luo, Zijian Wang 等CVPR 2026 · 被引用 2 次
