Visual Consensus Prompting for Co-Salient Object Detection
Jie Wang, Nana Yu, Zihao Zhang, Yahong Han
Abstract
Existing co-salient object detection (CoSOD) methods generally employ a three-stage architecture (i.e., encoding, consensus extraction & dispersion, and prediction) along with a typical full fine-tuning paradigm. Although they yield certain benefits, they exhibit two notable limitations: 1) This architecture relies on encoded features to facilitate consensus extraction, but the meticulously extracted consensus does not provide timely guidance to the encoding stage. 2) This paradigm involves globally updating all parameters of the model, which is parameter-inefficient and hinders the effective representation of knowledge within the foundation model for this task. Therefore, in this paper, we propose an interaction-effective and parameter-efficient concise architecture for the CoSOD task, addressing two key limitations. It introduces, for the first time, a parameter-efficient prompt tuning paradigm and seamlessly embeds consensus into the prompts to formulate task-specific Visual Consensus Prompts (VCP). Our VCP aims to induce the frozen foundation model to perform better on CoSOD tasks by formulating task-specific visual consensus prompts with minimized tunable parameters. Concretely, the primary insight of the purposeful Consensus Prompt Generator (CPG) is to enforce limited tunable parameters to focus on cosalient representations and generate consensus prompts. The formulated Consensus Prompt Disperser (CPD) leverages consensus prompts to form task-specific visual consensus prompts, thereby arousing the powerful potential of pretrained models in addressing CoSOD tasks. Extensive experiments demonstrate that our concise VCP outperforms 13 cutting-edge full fine-tuning models, achieving the new state of the art (with 6.8% improvement in F m metrics on the most challenging CoCA dataset). Source code has been available at https://github.com/WJ-CV/VCP .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4d04442b-3d30-4ac3-a4cf-40b638df4ee8Cited by top-tier papers2
- Follow the Saliency: Supervised Saliency for Retrieval-augmented Dense Video CaptioningSeung Hee Choi, MinJu Jeon, Hyunwoo Oh, Jihwan Lee et al.CVPR 2026 · 1 citation
- TF-SSD: A Strong Pipeline via Synergic Mask Filter for Training-free Co-salient Object DetectionZhijin He, Shuo Jin, Siyue Yu, Shuwei Wu et al.CVPR 2026 · 1 citation
Builds on18
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar et al.NeurIPS 2021 · 9,661 citations
- AdaptFormer: Adapting Vision Transformers for Scalable Visual RecognitionShoufa Chen, Chongjian Ge, Zhan Tong, Jiangliu Wang et al.NeurIPS 2022 · 1,291 citations
- AutoPrompt: Eliciting Knowledge from Language Models with Automatically Generated PromptsTaylor Shin, Yasaman Razeghi, Robert L. Logan IV, Eric Wallace et al.EMNLP 2020 · 1,162 citations
- VSCode: General Visual Salient and Camouflaged Object Detection with 2D Prompt LearningZiyang Luo, Nian Liu, Wangbo Zhao, Xuguang Yang et al.CVPR 2024 · 96 citations
Related papers
- Co-Salient Object Detection with Semantic-Level Consensus Extraction and DispersionPeiran Xu, Yadong MuACM MM 2023 · 7 citations
- CoADNet: Collaborative Aggregation-and-Distribution Networks for Co-Salient Object DetectionQijian Zhang, Runmin Cong, Junhui Hou, Chongyi Li et al.NeurIPS 2020 · 70 citations
- Memory-Aided Contrastive Consensus Learning for Co-salient Object DetectionPeng Zheng, Jie Qin, Shuo Wang, Tian-Zhu Xiang et al.AAAI 2023 · 31 citations
- Visual Instance-aware Prompt TuningXi Xiao, Yunbei Zhang, Xingjian Li, Tianyang Wang et al.ACM MM 2025 · 12 citations
- Democracy Does Matter: Comprehensive Feature Mining for Co-Salient Object DetectionSiyue Yu, Jimin Xiao, Bingfeng Zhang, Eng Gee LimCVPR 2022 · 76 citations
