From Passive Perception to Active Memory: A Weakly Supervised Image Manipulation Localization Framework Driven by Coarse-Grained Annotations
Zhiqing Guo, Dongdong Xi, Songlin Li, Gaobo Yang
Abstract
Image manipulation localization (IML) faces a fundamental trade-off between minimizing annotation cost and achieving fine-grained localization accuracy. Existing fully-supervised IML methods depend heavily on dense pixel-level mask annotations, which limits scalability to large datasets or real-world deployment. In contrast, the majority of existing weakly-supervised IML approaches are based on image-level labels, which greatly reduce annotation effort but typically lack precise spatial localization. To address this dilemma, we propose BoxPromptIML, a novel weakly-supervised IML framework that effectively balances annotation cost and localization performance. Specifically, we propose a coarse region annotation strategy, which can generate relatively accurate manipulation masks at lower cost. To improve model efficiency and facilitate deployment, we further design an efficient lightweight student model, which learns to perform fine-grained localization through knowledge distillation from a fixed teacher model based on the Segment Anything Model (SAM). Moreover, inspired by the human subconscious memory mechanism, our feature fusion module employs a dual-guidance strategy that actively contextualizes recalled prototypical patterns with real-time observational cues derived from the input. Instead of passive feature extraction, this strategy enables a dynamic process of knowledge recollection, where long-term memory is adapted to the specific context of the current image, significantly enhancing localization accuracy and robustness. Extensive experiments across both in-distribution and out-of-distribution datasets show that BoxPromptIML outperforms or rivals fully-supervised models, while maintaining strong generalization, low annotation cost, and efficient deployment characteristics.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on6
- GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion ModelsAlexander Quinn Nichol, Prafulla Dhariwal, Aditya Ramesh, Pranav Shyam et al.ICML 2022 · 4,691 citations
- Mesoscopic Insights: Orchestrating Multi-Scale & Hybrid Architecture for Image Manipulation LocalizationXuekang Zhu, Xiaochen Ma, Lei Su, Zhuohang Jiang et al.AAAI 2025 · 44 citations
- Towards Generic Image Manipulation Detection with Weakly-Supervised Self-Consistency LearningYuanhao Zhai, Tianyu Luan, David S. Doermann, Junsong YuanICCV 2023 · 35 citations
- Can We Get Rid of Handcrafted Feature Extractors? SparseViT: Nonsemantics-Centered, Parameter-Efficient Image Manipulation Localization Through Spare-Coding TransformerLei Su, Xiaochen Ma, Xuekang Zhu, Chaoqun Niu et al.AAAI 2025 · 10 citations
- Beyond Fully Supervised Pixel Annotations: Scribble-Driven Weakly-Supervised Framework for Image Manipulation LocalizationSonglin Li, Guofeng Yu, Zhiqing Guo, Yunfeng Diao et al.AAAI 2026 · 3 citations
Related papers
- Segment Anything Model Meets Semi-supervised Medical Image Segmentation: A Novel PerspectiveHaifeng Zhao, Haiyang Li, Lei-Lei Ma, Dengdi SunNeurIPS 2025 · 1 citation
- Improving the Generalization of Segmentation Foundation Model under Distribution Shift via Weakly Supervised AdaptationHaojie Zhang, Yongyi Su, Xun Xu, Kui JiaCVPR 2024 · 26 citations
- AoP-SAM: Automation of Prompts for Efficient SegmentationYi Chen, Muyoung Son, Chuanbo Hua, Joo-Young KimAAAI 2025 · 9 citations
- WISH: Weakly Supervised Instance Segmentation using Heterogeneous LabelsHyeokjun Kweon, Kuk-Jin YoonCVPR 2025
- CSDN: CLIP-Driven Similarity-Aligned Distillation Network for Weakly-Supervised Object LocalizationSifan Zuo, Youfa Liu, Bo DuACM MM 2025
