Can We Get Rid of Handcrafted Feature Extractors? SparseViT: Nonsemantics-Centered, Parameter-Efficient Image Manipulation Localization Through Spare-Coding Transformer
Lei Su, Xiaochen Ma, Xuekang Zhu, Chaoqun Niu, Zeyu Lei, Ji-Zhe Zhou
摘要
Non-semantic features or semantic-agnostic features, which are irrelevant to image context but sensitive to image manipulations, are recognized as evidential to Image Manipulation Localization (IML). Since manual labels are impossible, existing works rely on handcrafted methods to extract non-semantic features. Handcrafted non-semantic features jeopardize IML model's generalization ability in unseen or complex scenarios. Therefore, for IML, the elephant in the room is: How to adaptively extract non-semantic features? Non-semantic features are context-irrelevant and manipulation-sensitive. That is, within an image, they are consistent across patches unless manipulation occurs. Then, spare and discrete interactions among image patches are sufficient for extracting non-semantic features. However, image semantics vary drastically on different patches, requiring dense and continuous interactions among image patches for learning semantic representations. Hence, in this paper, we propose a Sparse Vision Transformer (SparseViT), which reformulates the dense, global self-attention in ViT into a sparse, discrete manner. Such sparse self-attention breaks image semantics and forces SparseViT to adaptively extract non-semantic features for images. Besides, compared with existing IML models, the sparse self-attention mechanism largely reduced the model size (max 80% in FLOPs), achieving stunning parameter efficiency and computation reduction. Extensive experiments demonstrate that, without any handcrafted feature extractors, SparseViT is superior in both generalization and efficiency across benchmark datasets.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- Zooming In on Fakes: A Novel Dataset for Localized AI-Generated Image Detection with Forgery Amplification ApproachLvpan Cai, Haowei Wang, Jiayi Ji, YanShu ZhouMen 等AAAI 2026 · 被引用 8 次
- TextShield-R1: Reinforced Reasoning for Tampered Text DetectionChenfan Qu, Yiwu Zhong, Jian Liu, Xuekang Zhu 等AAAI 2026 · 被引用 4 次
- Beyond Fully Supervised Pixel Annotations: Scribble-Driven Weakly-Supervised Framework for Image Manipulation LocalizationSonglin Li, Guofeng Yu, Zhiqing Guo, Yunfeng Diao 等AAAI 2026 · 被引用 3 次
- From Passive Perception to Active Memory: A Weakly Supervised Image Manipulation Localization Framework Driven by Coarse-Grained AnnotationsZhiqing Guo, Dongdong Xi, Songlin Li, Gaobo YangAAAI 2026 · 被引用 2 次
- EARG-Net: Edge-Aware Reconstruction-Guided Network for Image Manipulation Detection and LocalizationYanpu Yu, Zhaoxin Shi, Hanqing Zhao, Tianyi Wei 等AAAI 2026 · 被引用 1 次
它引用的顶会 Paper11
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar 等NeurIPS 2021 · 被引用 9,661 次
- ViViT: A Video Vision TransformerAnurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun 等ICCV 2021 · 被引用 2,947 次
- Tokens-to-Token ViT: Training Vision Transformers from Scratch on ImageNetLi Yuan, Yunpeng Chen, Tao Wang, Weihao Yu 等ICCV 2021 · 被引用 2,462 次
相关 Paper
- Vision Transformer with Sparse Scan PriorYuguang Zhang, Qihang Fan, Huaibo HuangACM MM 2025 · 被引用 2 次
- Vision Transformers Need More Than RegistersCheng Shi, Yizhou Yu, Sibei YangCVPR 2026 · 被引用 17 次
- Slide-Transformer: Hierarchical Vision Transformer with Local Self-AttentionXuran Pan, Tianzhu Ye, Zhuofan Xia, Shiji Song 等CVPR 2023
- Sparsifiner: Learning Sparse Instance-Dependent Attention for Efficient Vision TransformersCong Wei, Brendan Duke, Ruowei Jiang, Parham Aarabi 等CVPR 2023
- Image Manipulation Detection by Multi-View Multi-Scale SupervisionXinru Chen, Chengbo Dong, Jiaqi Ji, Juan Cao 等ICCV 2021 · 被引用 271 次
