OvarNet: Towards Open-Vocabulary Object Attribute Recognition
Keyan Chen, Xiaolong Jiang, Yao Hu, Xu Tang, Yan Gao, Jianqi Chen, Weidi Xie
Abstract
In this paper, we consider the problem of simultaneously detecting objects and inferring their visual attributes in an image, even for those with no manual annotations provided at the training stage, resembling an open-vocabulary scenario. To achieve this goal, we make the following contributions: (i) we start with a naïve two-stage approach for open-vocabulary object detection and attribute classification, termed CLIP-Attr. The candidate objects are first proposed with an offline RPN and later classified for semantic category and attributes; (ii) we combine all available datasets and train with a federated strategy to finetune the CLIP model, aligning the visual representation with attributes, additionally, we investigate the efficacy of leveraging freely available online image-caption pairs under weakly supervised learning; (iii) in pursuit of efficiency, we train a Faster-RCNN type model end-to-end with knowledge distillation, that performs class-agnostic object proposals and classification on semantic categories and attributes with classifiers generated from a text encoder; Finally, (iv) we conduct extensive experiments on VAW, MS-COCO, LSA, and OVAD datasets, and show that recog-⋆ Equal contribution. † Corresponding author. nition of semantic category and attributes is complementary for visual scene understanding, i.e., jointly training object detection and attributes prediction largely outperform existing approaches that treat the two tasks independently, demonstrating strong generalization ability to novel attributes and categories.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 22e35cd3-c083-43ab-8ffb-131b774e0a10Cited by top-tier papers14
- Open-Vocabulary Semantic Segmentation via Attribute Decomposition-AggregationChaofan Ma, Yuhuan Yang, Chen Ju, Fei Zhang et al.NeurIPS 2023 · 40 citations
- ImOV3D: Learning Open Vocabulary Point Clouds 3D Object Detection from Only 2D ImagesTiming Yang, Yuanliang Ju, Li YiNeurIPS 2024 · 22 citations
- Advancing Textual Prompt Learning with Anchored AttributesZheng Li, Yibing Song, Ming-Ming Cheng, Xiang Li et al.ICCV 2025 · 8 citations
- Exploring Region-Word Alignment in Built-in Detector for Open-Vocabulary Object DetectionHeng Zhang, Qiuyu Zhao, Linyu Zheng, Hao Zeng et al.CVPR 2024 · 6 citations
- PixCLIP: Towards Fine-grained Vision-Language Understanding via Any-granularity Pixel-Text AlignmentYicheng Xiao, Yu Chen, Hao-Xuan Ma, Jiale Hong et al.ICML 2026 · 4 citations
Builds on16
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- Open-vocabulary Object Detection via Vision and Language Knowledge DistillationXiuye Gu, Tsung-Yi Lin, Weicheng Kuo, Yin CuiICLR 2022 · 1,274 citations
- RegionCLIP: Region-based Language-Image PretrainingYiwu Zhong, Jianwei Yang, Pengchuan Zhang, Chunyuan Li et al.CVPR 2022 · 481 citations
- Attribute Prototype Network for Zero-Shot LearningWenjia Xu, Yongqin Xian, Jiuniu Wang, Bernt Schiele et al.NeurIPS 2020 · 392 citations
Related papers
- Simple Image-Level Classification Improves Open-Vocabulary Object DetectionRuohuan Fang, Guansong Pang, Xiao BaiAAAI 2024 · 26 citations
- Object-level Semantic and Spatial Distillation for Open Vocabulary DetectionZitong Li, Jinzhuo Wu, Fukang Zhao, Xinyue Wang et al.ICML 2026
- SRA-Det: Learning Omni-Grained Open-Vocabulary Detection Beyond Category NamesLi Yang, Boyu Cai, Wei Liu, Yan Wang et al.CVPR 2026 · 1 citation
- CORA: Adapting CLIP for Open-Vocabulary Detection with Region Prompting and Anchor Pre-MatchingXiaoshi Wu, Feng Zhu, Rui Zhao, Hongsheng LiCVPR 2023
- Bridging the Gap between Object and Image-level Representations for Open-Vocabulary DetectionHanoona Abdul Rasheed, Muhammad Maaz, Muhammad Uzair Khattak, Salman H. Khan et al.NeurIPS 2022 · 215 citations
