ICAS: Detecting Training Data from Autoregressive Image Generative Models
Hongyao Yu, Yixiang Qiu, Yiheng Yang, Hao Fang, Tianqu Zhuang, Jiaxin Hong, Bin Chen, Hao Wu, Shu-Tao Xia
摘要
Autoregressive image generation has witnessed rapid advancements, with prominent models such as scale-wise visual auto-regression pushing the boundaries of visual synthesis. However, these developments also raise significant concerns regarding data privacy and copyright. In response, training data detection has emerged as a critical task for identifying unauthorized data usage in model training. To better understand the vulnerability of autoregressive image generative models to such detection, we conduct the first study that applies membership inference to this domain. Our approach comprises two key components: implicit classification and an adaptive score aggregation strategy. First, we compute the implicit token-wise classification score within the query image. Then we propose an adaptive score aggregation strategy to acquire a final score, which places greater emphasis on the tokens with lower scores. A higher final score indicates that the sample is more likely to be involved in the training set. Extensive experiments demonstrate the superiority of our method over those designed for LLMs, in both class-conditional and text-to-image scenarios. Moreover, our approach exhibits strong robustness and generalization under various data transformations. Furthermore, sufficient experiments suggest two novel key findings: (1) A linear scaling law on membership inference, exposing the vulnerability of large foundation models. (2) Training data from scale-wise visual autoregressive models is easier to detect than other autoregressive paradigms. Our code is available at https://github.com/Chrisqcwx/ImageAR-MIA.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Data Provenance for Image Auto-Regressive GenerationBihe Zhao, Louis Kerner, Michel Meintz, Tameem Bakr 等ICLR 2026 · 被引用 5 次
- Rank Matters: Understanding and Defending Model Inversion Attacks via Low-Rank Feature FilteringHongyao Yu, Yixiang Qiu, Hao Fang, Tianqu Zhuang 等KDD 2026 · 被引用 2 次
- Causally-Grounded Dual-Path Attention Intervention for Object Hallucination Mitigation in LVLMsLiu Yu, Zhonghao Chen, Ping Kuang, Zhikun Feng 等AAAI 2026
它引用的顶会 Paper23
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari 等ICML 2024 · 被引用 3,620 次
- Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale PredictionKeyu Tian, Yi Jiang, Zehuan Yuan, Bingyue Peng 等NeurIPS 2024 · 被引用 1,199 次
- Membership Inference Attacks From First PrinciplesNicholas Carlini, Steve Chien, Milad Nasr, Shuang Song 等S&P 2022 · 被引用 1,049 次
- Detecting Pretraining Data from Large Language ModelsWeijia Shi, Anirudh Ajith, Mengzhou Xia, Yangsibo Huang 等ICLR 2024 · 被引用 365 次
相关 Paper
- Robust Membership Inference for Large Language Models under Adversarial Generative CorruptionYuanhong Huang, Huili Wang, Xueying Bai, Jinrui Wang 等ACL 2026
- Membership Inference Attacks against Large Vision-Language ModelsZhan Li, Yongtao Wu, Yihang Chen, Francesco Tonin 等NeurIPS 2024 · 被引用 43 次
- Privacy Attacks on Image AutoRegressive ModelsAntoni Kowalczuk, Jan Dubinski, Franziska Boenisch, Adam DziedzicICML 2025
- Unveiling Structural Memorization: Structural Membership Inference Attack for Text-to-Image Diffusion ModelsQiao Li, Xiaomeng Fu, Xi Wang, Jin Liu 等ACM MM 2024 · 被引用 6 次
- LLM Dataset Inference: Did you train on my dataset?Pratyush Maini, Hengrui Jia, Nicolas Papernot, Adam DziedzicNeurIPS 2024 · 被引用 162 次
