Partially Does It: Towards Scene-Level FG-SBIR with Partial Input
Pinaki Nath Chowdhury, Ayan Kumar Bhunia, Viswanatha Reddy Gajjala, Aneeshan Sain, Tao Xiang, Yi-Zhe Song
Abstract
We scrutinise an important observation plaguing scene-level sketch research - that a significant portion of scene sketches are “partial”. A quick pilot study reveals: (i) a scene sketch does not necessarily contain all objects in the corresponding photo, due to the subjective holistic interpretation of scenes, (ii) there exists significant empty (white) regions as a result of object-level abstraction, and as a result, (iii) existing scene-level fine-grained sketch-based image retrieval methods collapse as scene sketches become more partial. To solve this “partial” problem, we advocate for a simple set-based approach using optimal transport (OT) to model cross-modal region associativity in a partially-aware fashion. Importantly, we improve upon OT to further account for holistic partialness by comparing intra-modal adjacency matrices. Our proposed method is not only robust to partial scene-sketches but also yields state-of-the-art performance on existing datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers16
- Sketch3T: Test-Time Training for Zero-Shot SBIRAneeshan Sain, Ayan Kumar Bhunia, Vaishnav Potlapalli, Pinaki Nath Chowdhury et al.CVPR 2022 · 55 citations
- Sketching without Worrying: Noise-Tolerant Sketch-Based Image RetrievalAyan Kumar Bhunia, Subhadeep Koley, Abdullah Faiz Ur Rahman Khilji, Aneeshan Sain et al.CVPR 2022 · 53 citations
- Democratising 2D Sketch to 3D Shape Retrieval Through PivotingPinaki Nath Chowdhury, Ayan Kumar Bhunia, Aneeshan Sain, Subhadeep Koley et al.ICCV 2023 · 10 citations
- ARNet: Self-Supervised FG-SBIR with Unified Sample Feature Alignment and Multi-Scale Token RecyclingJianan Jiang, Hao Tang, Zhilin Jiang, Weiren Yu et al.AAAI 2025 · 4 citations
- ELViS: Efficient Visual Similarity from Local Descriptors that Generalizes Across DomainsPavel Suma, Giorgos Kordopatis-Zilos, Yannis Kalantidis, Giorgos ToliasICLR 2026 · 1 citation
Builds on22
- CAMP: Cross-Modal Adaptive Message Passing for Text-Image RetrievalZihao Wang, Xihui Liu, Hongsheng Li, Lu Sheng et al.ICCV 2019 · 349 citations
- Graph Optimal Transport for Cross-Domain AlignmentLiqun Chen, Zhe Gan, Yu Cheng, Linjie Li et al.ICML 2020 · 193 citations
- Interactive Sketch & Fill: Multiclass Sketch-to-Image TranslationArnab Ghosh, Richard Zhang, Puneet K. Dokania, Oliver Wang et al.ICCV 2019 · 148 citations
- Sketch3T: Test-Time Training for Zero-Shot SBIRAneeshan Sain, Ayan Kumar Bhunia, Vaishnav Potlapalli, Pinaki Nath Chowdhury et al.CVPR 2022 · 55 citations
- Sketching without Worrying: Noise-Tolerant Sketch-Based Image RetrievalAyan Kumar Bhunia, Subhadeep Koley, Abdullah Faiz Ur Rahman Khilji, Aneeshan Sain et al.CVPR 2022 · 53 citations
Related papers
- Partially Aligned Cross-modal Retrieval via Optimal Transport-based Prototype Alignment LearningJunsheng Wang, Tiantian Gong, Yan YanACM MM 2024 · 3 citations
- Learning to Rematch Mismatched Pairs for Robust Cross-Modal RetrievalHaochen Han, Qinghua Zheng, Guang Dai, Minnan Luo et al.CVPR 2024
- Optimal Transport-based Labor-free Text Prompt Modeling for Sketch Re-identificationRui Li, Tingting Ren, Jie Wen, Jinxing LiNeurIPS 2024 · 3 citations
- Scene-Level Sketch-Based Image Retrieval with Minimal Pairwise SupervisionCe Ge, Jingyu Wang, Qi Qi, Haifeng Sun et al.AAAI 2023 · 6 citations
- Sketch Transformer: Asymmetrical Disentanglement Learning from Dynamic SynthesisCuiqun Chen, Mang Ye, Meibin Qi, Bo DuACM MM 2022 · 29 citations
