Mesoscopic Insights: Orchestrating Multi-Scale & Hybrid Architecture for Image Manipulation Localization
Xuekang Zhu, Xiaochen Ma, Lei Su, Zhuohang Jiang, Bo Du, Xiwen Wang, Zeyu Lei, Wentao Feng, Chi-Man Pun, Ji-Zhe Zhou
摘要
The mesoscopic level serves as a bridge between the macroscopic and microscopic worlds, addressing gaps overlooked by both. Image manipulation localization (IML), a crucial technique to pursue truth from fake images, has long relied on low-level (microscopic-level) traces. However, in practice, most tampering aims to deceive the audience by altering image semantics. As a result, manipulation commonly occurs at the object level (macroscopic level), which is equally important as microscopic traces. Therefore, integrating these two levels into the mesoscopic level presents a new perspective for IML research. Inspired by this, our paper explores how to simultaneously construct mesoscopic representations of micro and macro information for IML and introduces the Mesorch architecture to orchestrate both. Specifically, this architecture i) combines Transformers and CNNs in parallel, with Transformers extracting macro information and CNNs capturing micro details, and ii) explores across different scales, assessing micro and macro information seamlessly. Additionally, based on the Mesorch architecture, the paper introduces two baseline models aimed at solving IML tasks through mesoscopic representation. Extensive experiments across four datasets have demonstrated that our models surpass the current state-of-the-art in terms of performance, computational complexity, and robustness.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper14
- RAM: Recover Any 3D Human Motion in-the-WildSen Jia, Ning Zhu, Jinqin Zhong, Jiale Zhou 等CVPR 2026 · 被引用 12 次
- Towards Reliable Identification of Diffusion-based Image ManipulationsAlex Costanzino, Woody Bayliss, Juil Sock, Marc Górriz Blanch 等NeurIPS 2025 · 被引用 4 次
- TextShield-R1: Reinforced Reasoning for Tampered Text DetectionChenfan Qu, Yiwu Zhong, Jian Liu, Xuekang Zhu 等AAAI 2026 · 被引用 4 次
- Beyond Fully Supervised Pixel Annotations: Scribble-Driven Weakly-Supervised Framework for Image Manipulation LocalizationSonglin Li, Guofeng Yu, Zhiqing Guo, Yunfeng Diao 等AAAI 2026 · 被引用 3 次
- From Passive Perception to Active Memory: A Weakly Supervised Image Manipulation Localization Framework Driven by Coarse-Grained AnnotationsZhiqing Guo, Dongdong Xi, Songlin Li, Gaobo YangAAAI 2026 · 被引用 2 次
它引用的顶会 Paper11
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar 等NeurIPS 2021 · 被引用 9,661 次
- Image Manipulation Detection by Multi-View Multi-Scale SupervisionXinru Chen, Chengbo Dong, Jiaqi Ji, Juan Cao 等ICCV 2021 · 被引用 271 次
- ObjectFormer for Image Manipulation Detection and LocalizationJunke Wang, Zuxuan Wu, Jingjing Chen, Xintong Han 等CVPR 2022 · 被引用 190 次
- Localization of Deep Inpainting Using High-Pass Fully Convolutional NetworkHaodong Li, Jiwu HuangICCV 2019 · 被引用 157 次
相关 Paper
- Collaborative Transformers with Multi-Level Forensic Attention for Image Manipulation LocalizationJiwei Zhang, Wenbo Feng, Siwei Wang, Feifei Kou 等AAAI 2026
- DiffForensics: Leveraging Diffusion Prior to Image Forgery Detection and LocalizationZeqin Yu, Jiangqun Ni, Yuzhen Lin, Haoyi Deng 等CVPR 2024 · 被引用 25 次
- Omni-IML: Towards Unified Interpretable Image Manipulation LocalizationChenfan Qu, Yiwu Zhong, Fengjun Guo, Lianwen JinICLR 2026 · 被引用 5 次
- M2sformer: Multi-Spectral and Multi-Scale Attention With Edge-Aware Difficulty Guidance for Image Forgery LocalizationJu-Hyeon Nam, Dong-Hyun Moon, Sang-Chul LeeICCV 2025 · 被引用 4 次
- TransForensics: Image Forgery Localization with Dense Self-AttentionJing Hao, Zhixin Zhang, Shicai Yang, Di Xie 等ICCV 2021 · 被引用 77 次
