Explicit Visual Prompting for Low-Level Structure Segmentations
Weihuang Liu, Xi Shen, Chi-Man Pun, Xiaodong Cun
摘要
We consider the generic problem of detecting low-level structures in images, which includes segmenting the manipulated parts, identifying out-of-focus pixels, separating shadow regions, and detecting concealed objects. Whereas each such topic has been typically addressed with a domainspecific solution, we show that a unified approach performs well across all of them. We take inspiration from the widelyused pre-training and then prompt tuning protocols in NLP and propose a new visual prompting model, named Explicit Visual Prompting (EVP). Different from the previous visual prompting which is typically a dataset-level implicit embedding, our key insight is to enforce the tunable parameters focusing on the explicit visual content from each individual image, i.e., the features from frozen patch embeddings and the input's high-frequency components. The proposed EVP significantly outperforms other parameter-efficient tuning protocols under the same amount of tunable parameters (5.7% extra trainable parameters of each task). EVP also achieves state-of-the-art performances on diverse lowlevel structure segmentation tasks compared to task-specific solutions. Our code is available at: https://github . com/NiFangBaAGe/Explicit-Visual-Prompt.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper29
- Tuning Multi-mode Token-level Prompt Alignment across ModalitiesDongsheng Wang, Miaoge Li, Xinyang Liu, Mingsheng Xu 等NeurIPS 2023 · 被引用 49 次
- DiffMOT: A Real-time Diffusion-based Multiple Object Tracker with Non-linear PredictionWeiyi Lv, Yuhang Huang, Ning Zhang, Ruei-Sung Lin 等CVPR 2024 · 被引用 36 次
- Visual Prompting Upgrades Neural Network Sparsification: A Data-Model PerspectiveCan Jin, Tianjin Huang, Yihua Zhang, Mykola Pechenizkiy 等AAAI 2025 · 被引用 30 次
- SAFIRE: Segment Any Forged Image RegionMyung-Joon Kwon, Wonjun Lee, Seung-Hun Nam, Minji Son 等AAAI 2025 · 被引用 25 次
- Spider: A Unified Framework for Context-dependent Concept SegmentationXiaoqi Zhao, Youwei Pang, Wei Ji, Baicheng Sheng 等ICML 2024 · 被引用 21 次
它引用的顶会 Paper21
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar 等NeurIPS 2021 · 被引用 9,661 次
相关 Paper
- LoR-VP: Low-Rank Visual Prompting for Efficient Vision Model AdaptationCan Jin, Ying Li, Mingyu Zhao, Shiyu Zhao 等ICLR 2025
- AutoVP: An Automated Visual Prompting Framework and BenchmarkHsi-Ai Tsao, Lei Hsiung, Pin-Yu Chen, Si Liu 等ICLR 2024 · 被引用 29 次
- SA²VP: Spatially Aligned-and-Adapted Visual PromptWenjie Pei, Tongqi Xia, Fanglin Chen, Jinsong Li 等AAAI 2024 · 被引用 33 次
- E2VPT: An Effective and Efficient Approach for Visual Prompt TuningCheng Han, Qifan Wang, Yiming Cui, Zhiwen Cao 等ICCV 2023 · 被引用 108 次
- Exploring Interpretability for Visual Prompt Tuning with Cross-layer ConceptsYubin Wang, Xinyang Jiang, De Cheng, Xiangqian Zhao 等ICLR 2026 · 被引用 1 次
