PhaseWin Search Framework Enable Efficient Object-Level Interpretation
Zihan Gu, Ruoyu Chen, Junchi Zhang, Yue Hu, Hua Zhang, Xiaochun Cao
摘要
Attribution is essential for interpreting object-level foundation models. Recent methods based on submodular subset selection have achieved high faithfulness, but their efficiency limitations hinder practical deployment in real-world scenarios. To address this, we propose PhaseWin, a novel phase-window search algorithm that enables faithful region attribution with near-linear complexity. PhaseWin replaces traditional quadratic-cost greedy selection with a phased coarse-to-fine search, combining adaptive pruning, windowed fine-grained selection, and dynamic supervision mechanisms to closely approximate greedy behavior while dramatically reducing model evaluations. Theoretically, PhaseWin retains near-greedy approximation guarantees under mild monotone submodular assumptions. Empirically, PhaseWin achieves over 95% of greedy attribution faithfulness using only 20% of the computational budget, and consistently outperforms other attribution baselines across object detection and visual grounding tasks with Grounding DINO and Florence-2. PhaseWin establishes a new state of the art in scalable, high-faithfulness attribution for object-level multimodal models.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper13
- Interpreting CLIP's Image Representation via Text-Based DecompositionYossi Gandelsman, Alexei A. Efros, Jacob SteinhardtICLR 2024 · 被引用 179 次
- Making Sense of Dependence: Efficient Black-box Explanations Using Dependence MeasurePaul Novello, Thomas Fel, David VigourouxNeurIPS 2022 · 被引用 48 次
- SAFE: Sensitivity-Aware Features for Out-of-Distribution Object DetectionSamuel Wilson, Tobias Fischer, Feras Dayoub, Dimity Miller 等ICCV 2023 · 被引用 46 次
- ThirdEye: Attention Maps for Safe Autonomous Driving SystemsAndrea Stocco, Paulo J. Nunes, Marcelo d'Amorim, Paolo TonellaASE 2022 · 被引用 43 次
- The FAST Algorithm for Submodular MaximizationAdam Breuer, Eric Balkanski, Yaron SingerICML 2020 · 被引用 39 次
相关 Paper
- Interpreting Object-level Foundation Models via Visual Precision SearchRuoyu Chen, Siyuan Liang, Jingzhi Li, Shiming Liu 等CVPR 2025
- Dynamic-DINO: Fine-Grained Mixture of Experts Tuning for Real-Time Open-Vocabulary Object DetectionYehao Lu, Minghe Weng, Zekang Xiao, Rui Jiang 等ICCV 2025 · 被引用 2 次
- Towards Better Understanding Attribution MethodsSukrut Rao, Moritz Böhle, Bernt SchieleCVPR 2022 · 被引用 32 次
- Multi-Object 3D Grounding with Dynamic Modules and Language-Informed Spatial AttentionHaomeng Zhang, Chiao-An Yang, Raymond A. YehNeurIPS 2024 · 被引用 10 次
- Multi-Attribute Interactions Matter for 3D Visual GroundingCan Xu, Yuehui Han, Rui Xu, Le Hui 等CVPR 2024 · 被引用 5 次
