Seeing the Whole Through the Parts: Discovering Objects through Semantic Part Mining in Weak Supervision
Shucheng Li, Weixuan Xu, Le Jiang, Hao Wu, Fengyuan Xu, Fan Wu, Feng Lyu
摘要
Weakly Supervised Object Detection (WSOD) is fundamentally limited by instance ambiguity, manifesting as either part domination (focusing on discriminative fragments) or merged detection (confusing objects with context). Unlike existing approaches that rely on object-level patterns, we draw inspiration from human cognition, where objects are perceived structurally by integrating constituent parts. Building on this perspective, we propose P2WDet (Part-to-Whole Detection), a novel framework that shifts WSOD from traditional instance selection to semantic reconstruction. P2WDet comprises three systematic stages: (1) constructing a comprehensive visually-grounded part vocabulary leveraging large language models; (2) training dedicated part detectors by mining consistent patterns from cross-image proposal clusters; and (3) assembling detected parts into complete objects to enforce structural consistency. Extensive experiments show that P2WDet not only outperforms state-of-the-art methods on standard benchmarks but also demonstrates superior generalization in open-world settings. Furthermore, by decoupling detection from fixed category labels, P2WDet enables the flexible detection of novel or ambiguous objects defined solely by their components.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- Cap2Det: Learning to Amplify Weak Caption Supervision for Object DetectionKeren Ye, Mingda Zhang, Adriana Kovashka, Wei Li 等ICCV 2019 · 被引用 61 次
- Weakly Supervised Open-Vocabulary Object DetectionJianghang Lin, Yunhang Shen, Bingquan Wang, Shaohui Lin 等AAAI 2024 · 被引用 18 次
- UWSOD: Toward Fully-Supervised-Level Capacity Weakly Supervised Object DetectionYunhang Shen, Rongrong Ji, Zhiwei Chen, Yongjian Wu 等NeurIPS 2020 · 被引用 37 次
- Going Denser with Open-Vocabulary Part SegmentationPeize Sun, Shoufa Chen, Chenchen Zhu, Fanyi Xiao 等ICCV 2023 · 被引用 83 次
- From Scene to Object: Enhancing Open-Vocabulary Object Detection via Foreground-Background Context ReasoningYanqi Li, Jianwei Niu, Ningbo Gu, Tao RenAAAI 2026
