Visual Recognition by Request
Chufeng Tang, Lingxi Xie, Xiaopeng Zhang, Xiaolin Hu, Qi Tian
Abstract
Humans have the ability of recognizing visual semantics in an unlimited granularity, but existing visual recognition algorithms cannot achieve this goal. In this paper, we establish a new paradigm named visual recognition by request (ViRReq 1 ) to bridge the gap. The key lies in decomposing visual recognition into atomic tasks named requests and leveraging a knowledge base, a hierarchical and text-based dictionary, to assist task definition. ViRReq allows for (i) learning complicated whole-part hierarchies from highly incomplete annotations and (ii) inserting new concepts with minimal efforts. We also establish a solid baseline by integrating language-driven recognition into recent semantic and instance segmentation methods, and demonstrate its flexible recognition ability on CPP and ADE20K, two datasets with hierarchical whole-part annotations.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0c573143-b613-4925-bc07-0d752dbb4233Cited by top-tier papers8
- VisRL: Intention-Driven Visual Perception via Reinforced ReasoningZhangquan Chen, Xufang Luo, Dongsheng LiICCV 2025 · 2 citations
- Open-Vocabulary Part Segmentation via Progressive and Boundary-Aware StrategyXinlong Li, Di Lin, Shaoyiyi Gao, Jiaxin Li et al.NeurIPS 2025 · 1 citation
- One-shot In-context Part SegmentationZhenqi Dai, Ting Liu, Xingxing Zhang, Yunchao Wei et al.ACM MM 2024 · 1 citation
- SAM-CP: Marrying SAM with Composable Prompts for Versatile SegmentationPengfei Chen, Lingxi Xie, Xinyue Huo, Xuehui Yu et al.ICLR 2025
- HOPS: Hierarchical Open-vocabulary Part Segmentation with Attention-Aware Filtering and Affinity-Guided EnhancementXinlong Li, Di Lin, Shaoyiyi Gao, Yaxuan Liu et al.CVPR 2026
Builds on16
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar et al.NeurIPS 2021 · 9,661 citations
- FCOS: Fully Convolutional One-Stage Object DetectionZhi Tian, Chunhua Shen, Hao Chen, Tong HeICCV 2019 · 6,042 citations
- Swin Transformer V2: Scaling Up Capacity and ResolutionZe Liu, Han Hu, Yutong Lin, Zhuliang Yao et al.CVPR 2022 · 2,138 citations
Related papers
- Open-Vocabulary Semantic Segmentation via Attribute Decomposition-AggregationChaofan Ma, Yuhuan Yang, Chen Ju, Fei Zhang et al.NeurIPS 2023 · 40 citations
- 3D-DRES: Detailed 3D Referring Expression SegmentationQi Chen, Changli Wu, Jiayi Ji, Yiwei Ma et al.AAAI 2026 · 1 citation
- Learning Hierarchical Image Segmentation For Recognition and By RecognitionTsung-Wei Ke, Sangwoo Mo, Stella X. YuICLR 2024 · 20 citations
- Aligning and Prompting Everything All at Once for Universal Visual PerceptionYunhang Shen, Chaoyou Fu, Peixian Chen, Mengdan Zhang et al.CVPR 2024 · 19 citations
- Universal Segmentation at Arbitrary Granularity with Language InstructionYong Liu, Cairong Zhang, Yitong Wang, Jiahao Wang et al.CVPR 2024 · 15 citations
