Visual Consensus Prompting for Co-Salient Object Detection
Jie Wang, Nana Yu, Zihao Zhang, Yahong Han
摘要
Existing co-salient object detection (CoSOD) methods generally employ a three-stage architecture (i.e., encoding, consensus extraction & dispersion, and prediction) along with a typical full fine-tuning paradigm. Although they yield certain benefits, they exhibit two notable limitations: 1) This architecture relies on encoded features to facilitate consensus extraction, but the meticulously extracted consensus does not provide timely guidance to the encoding stage. 2) This paradigm involves globally updating all parameters of the model, which is parameter-inefficient and hinders the effective representation of knowledge within the foundation model for this task. Therefore, in this paper, we propose an interaction-effective and parameter-efficient concise architecture for the CoSOD task, addressing two key limitations. It introduces, for the first time, a parameter-efficient prompt tuning paradigm and seamlessly embeds consensus into the prompts to formulate task-specific Visual Consensus Prompts (VCP). Our VCP aims to induce the frozen foundation model to perform better on CoSOD tasks by formulating task-specific visual consensus prompts with minimized tunable parameters. Concretely, the primary insight of the purposeful Consensus Prompt Generator (CPG) is to enforce limited tunable parameters to focus on cosalient representations and generate consensus prompts. The formulated Consensus Prompt Disperser (CPD) leverages consensus prompts to form task-specific visual consensus prompts, thereby arousing the powerful potential of pretrained models in addressing CoSOD tasks. Extensive experiments demonstrate that our concise VCP outperforms 13 cutting-edge full fine-tuning models, achieving the new state of the art (with 6.8% improvement in F m metrics on the most challenging CoCA dataset). Source code has been available at https://github.com/WJ-CV/VCP .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Follow the Saliency: Supervised Saliency for Retrieval-augmented Dense Video CaptioningSeung Hee Choi, MinJu Jeon, Hyunwoo Oh, Jihwan Lee 等CVPR 2026 · 被引用 1 次
- TF-SSD: A Strong Pipeline via Synergic Mask Filter for Training-free Co-salient Object DetectionZhijin He, Shuo Jin, Siyue Yu, Shuwei Wu 等CVPR 2026 · 被引用 1 次
它引用的顶会 Paper18
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar 等NeurIPS 2021 · 被引用 9,661 次
- AdaptFormer: Adapting Vision Transformers for Scalable Visual RecognitionShoufa Chen, Chongjian Ge, Zhan Tong, Jiangliu Wang 等NeurIPS 2022 · 被引用 1,291 次
- AutoPrompt: Eliciting Knowledge from Language Models with Automatically Generated PromptsTaylor Shin, Yasaman Razeghi, Robert L. Logan IV, Eric Wallace 等EMNLP 2020 · 被引用 1,162 次
- VSCode: General Visual Salient and Camouflaged Object Detection with 2D Prompt LearningZiyang Luo, Nian Liu, Wangbo Zhao, Xuguang Yang 等CVPR 2024 · 被引用 96 次
相关 Paper
- Co-Salient Object Detection with Semantic-Level Consensus Extraction and DispersionPeiran Xu, Yadong MuACM MM 2023 · 被引用 7 次
- CoADNet: Collaborative Aggregation-and-Distribution Networks for Co-Salient Object DetectionQijian Zhang, Runmin Cong, Junhui Hou, Chongyi Li 等NeurIPS 2020 · 被引用 70 次
- Memory-Aided Contrastive Consensus Learning for Co-salient Object DetectionPeng Zheng, Jie Qin, Shuo Wang, Tian-Zhu Xiang 等AAAI 2023 · 被引用 31 次
- Visual Instance-aware Prompt TuningXi Xiao, Yunbei Zhang, Xingjian Li, Tianyang Wang 等ACM MM 2025 · 被引用 12 次
- Democracy Does Matter: Comprehensive Feature Mining for Co-Salient Object DetectionSiyue Yu, Jimin Xiao, Bingfeng Zhang, Eng Gee LimCVPR 2022 · 被引用 76 次
