Seeing the Whole Through the Parts: Discovering Objects through Semantic Part Mining in Weak Supervision
Shucheng Li, Weixuan Xu, Le Jiang, Hao Wu, Fengyuan Xu, Fan Wu, Feng Lyu
Abstract
Weakly Supervised Object Detection (WSOD) is fundamentally limited by instance ambiguity, manifesting as either part domination (focusing on discriminative fragments) or merged detection (confusing objects with context). Unlike existing approaches that rely on object-level patterns, we draw inspiration from human cognition, where objects are perceived structurally by integrating constituent parts. Building on this perspective, we propose P2WDet (Part-to-Whole Detection), a novel framework that shifts WSOD from traditional instance selection to semantic reconstruction. P2WDet comprises three systematic stages: (1) constructing a comprehensive visually-grounded part vocabulary leveraging large language models; (2) training dedicated part detectors by mining consistent patterns from cross-image proposal clusters; and (3) assembling detected parts into complete objects to enforce structural consistency. Extensive experiments show that P2WDet not only outperforms state-of-the-art methods on standard benchmarks but also demonstrates superior generalization in open-world settings. Furthermore, by decoupling detection from fixed category labels, P2WDet enables the flexible detection of novel or ambiguous objects defined solely by their components.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 2bcb5a8c-3454-4dcf-8167-555df7abf9d7Related papers
- Cap2Det: Learning to Amplify Weak Caption Supervision for Object DetectionKeren Ye, Mingda Zhang, Adriana Kovashka, Wei Li et al.ICCV 2019 · 61 citations
- Weakly Supervised Open-Vocabulary Object DetectionJianghang Lin, Yunhang Shen, Bingquan Wang, Shaohui Lin et al.AAAI 2024 · 18 citations
- UWSOD: Toward Fully-Supervised-Level Capacity Weakly Supervised Object DetectionYunhang Shen, Rongrong Ji, Zhiwei Chen, Yongjian Wu et al.NeurIPS 2020 · 37 citations
- Going Denser with Open-Vocabulary Part SegmentationPeize Sun, Shoufa Chen, Chenchen Zhu, Fanyi Xiao et al.ICCV 2023 · 83 citations
- From Scene to Object: Enhancing Open-Vocabulary Object Detection via Foreground-Background Context ReasoningYanqi Li, Jianwei Niu, Ningbo Gu, Tao RenAAAI 2026
