DesCo: Learning Object Recognition with Rich Language Descriptions
Liunian Harold Li, Zi-Yi Dou, Nanyun Peng, Kai-Wei Chang
Abstract
Recent development in vision-language approaches has instigated a paradigm shift in learning visual recognition models from language supervision. These approaches align objects with language queries (e.g. "a photo of a cat") and improve the models' adaptability to identify novel objects and domains. Recently, several studies have attempted to query these models with complex language expressions that include specifications of fine-grained semantic details, such as attributes, shapes, textures, and relations. However, simply incorporating language descriptions as queries does not guarantee accurate interpretation by the models. In fact, our experiments show that GLIP, the state-of-the-art vision-language model for object detection, often disregards contextual information in the language descriptions and instead relies heavily on detecting objects solely by their names. To tackle the challenges, we propose a new description-conditioned (DesCo) paradigm of learning object recognition models with rich language descriptions consisting of two major innovations: 1) we employ a large language model as a commonsense knowledge engine to generate rich language descriptions of objects based on object names and the raw image-text caption; 2) we design context-sensitive queries to improve the model's ability in deciphering intricate nuances embedded within descriptions and enforce the model to focus on context rather than object names alone. On two novel object detection benchmarks, LVIS and OminiLabel, under the zero-shot detection setting, our approach achieves 34.8 APr minival (+9.1) and 29.3 AP (+3.6), respectively, surpassing the prior state-of-the-art models, GLIP and FIBER, by a large margin. * Equal contribution. Preprint. Under review.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ef9d15d2-3904-4074-99c0-159e9f72a91cCited by top-tier papers3
- PerceptionCLIP: Visual Classification by Inferring and Conditioning on ContextsBang An, Sicheng Zhu, Michael-Andrei Panaitescu-Liess, Chaithanya Kumar Mummadi et al.ICLR 2024 · 15 citations
- If CLIP Could Talk: Understanding Vision-Language Model Representations Through Their Preferred Concept DescriptionsReza Esfandiarpoor, Cristina Menghini, Stephen H. BachEMNLP 2024 · 3 citations
- DoGA: Enhancing Grounded Object Detection via Grouped Pre-Training with AttributesYang Liu, Feng Hou, Yunjie Peng, Gangjian Zhang et al.AAAI 2025 · 1 citation
Builds on27
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
Related papers
- Multi-modal Queried Object Detection in the WildYifan Xu, Mengdan Zhang, Chaoyou Fu, Peixian Chen et al.NeurIPS 2023 · 73 citations
- Simple Image-Level Classification Improves Open-Vocabulary Object DetectionRuohuan Fang, Guansong Pang, Xiao BaiAAAI 2024 · 26 citations
- Hyperbolic Learning with Synthetic Captions for Open-World DetectionFanjie Kong, Yanbei Chen, Jiarui Cai, Davide ModoloCVPR 2024
- Generative Region-Language Pretraining for Open-Ended Object DetectionChuang Lin, Yi Jiang, Lizhen Qu, Zehuan Yuan et al.CVPR 2024
- DetCLIP: Dictionary-Enriched Visual-Concept Paralleled Pre-training for Open-world DetectionLewei Yao, Jianhua Han, Youpeng Wen, Xiaodan Liang et al.NeurIPS 2022 · 285 citations
