What Can Human Sketches Do for Object Detection?
Pinaki Nath Chowdhury, Ayan Kumar Bhunia, Aneeshan Sain, Subhadeep Koley, Tao Xiang, Yi-Zhe Song
摘要
Sketches are highly expressive, inherently capturing subjective and fine-grained visual cues. The exploration of such innate properties of human sketches has, however, been limited to that of image retrieval. In this paper, for the first time, we cultivate the expressiveness of sketches but for the fundamental vision task of object detection. The end result is a sketch-enabled object detection framework that detects based on what you sketch -that "zebra" (e.g., one that is eating the grass) in a herd of zebras (instance-aware detection), and only the part (e.g., "head" of a "zebra") that you desire (part-aware detection). We further dictate that our model works without (i) knowing which category to expect at testing (zero-shot) and (ii) not requiring additional bounding boxes (as per fully supervised) and class labels (as per weakly supervised). Instead of devising a model from the ground up, we show an intuitive synergy between foundation models (e.g., CLIP) and existing sketch models build for sketch-based image retrieval (SBIR), which can already elegantly solve the task -CLIP to provide model generalisation, and SBIR to bridge the (sketch→photo) gap. In particular, we first perform independent prompting on both sketch and photo branches of an SBIR model to build highly generalisable sketch and photo encoders on the back of the generalisation ability of CLIP. We then devise a training paradigm to adapt the learned encoders for object detection, such that the region embeddings of detected boxes are aligned with the sketch and photo embeddings from SBIR. Evaluating our framework on standard object detection datasets like PASCAL-VOC and MS-COCO outperforms both supervised (SOD) and weakly-supervised object detectors (WSOD) on zero-shot setups. Project
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper16
- Democratising 2D Sketch to 3D Shape Retrieval Through PivotingPinaki Nath Chowdhury, Ayan Kumar Bhunia, Aneeshan Sain, Subhadeep Koley 等ICCV 2023 · 被引用 10 次
- Composite Sketch+Text Queries for Retrieving Objects with Elusive Names and Complex InteractionsPrajwal Gatti, Kshitij Parikh, Dhriti Prasanna Paul, Manish Gupta 等AAAI 2024 · 被引用 6 次
- Sketch2Saliency: Learning to Detect Salient Objects from Human DrawingsAyan Kumar Bhunia, Subhadeep Koley, Amandeep Kumar, Aneeshan Sain 等CVPR 2023
- Picture that Sketch: Photorealistic Image Generation from Abstract SketchesSubhadeep Koley, Ayan Kumar Bhunia, Aneeshan Sain, Pinaki Nath Chowdhury 等CVPR 2023
- Doodle Your Keypoints: Sketch-Based Few-Shot Keypoint DetectionSubhajit Maity, Ayan Kumar Bhunia, Subhadeep Koley, Pinaki Nath Chowdhury 等ICCV 2025
它引用的顶会 Paper29
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- CutMix: Regularization Strategy to Train Strong Classifiers With Localizable FeaturesSangdoo Yun, Dongyoon Han, Sanghyuk Chun, Seong Joon Oh 等ICCV 2019 · 被引用 5,843 次
- Conditional Prompt Learning for Vision-Language ModelsKaiyang Zhou, Jingkang Yang, Chen Change Loy, Ziwei LiuCVPR 2022 · 被引用 1,438 次
- Open-vocabulary Object Detection via Vision and Language Knowledge DistillationXiuye Gu, Tsung-Yi Lin, Weicheng Kuo, Yin CuiICLR 2022 · 被引用 1,274 次
相关 Paper
- CLIP for All Things Zero-Shot Sketch-Based Image Retrieval, Fine-Grained or NotAneeshan Sain, Ayan Kumar Bhunia, Pinaki Nath Chowdhury, Subhadeep Koley 等CVPR 2023
- Semi-transductive Learning for Generalized Zero-Shot Sketch-Based Image RetrievalCe Ge, Jingyu Wang, Qi Qi, Haifeng Sun 等AAAI 2023 · 被引用 10 次
- Dr. CLIP: CLIP-Driven Universal Framework for Zero-Shot Sketch Image RetrievalXue Li, Jiong Yu, Ziyang Li, Hongchun Lu 等ACM MM 2024 · 被引用 12 次
- Semantic-Aware Knowledge Preservation for Zero-Shot Sketch-Based Image RetrievalQing Liu, Lingxi Xie, Huiyu Wang, Alan L. YuilleICCV 2019 · 被引用 126 次
- Sketch3T: Test-Time Training for Zero-Shot SBIRAneeshan Sain, Ayan Kumar Bhunia, Vaishnav Potlapalli, Pinaki Nath Chowdhury 等CVPR 2022 · 被引用 55 次
