Membership Inference Attack Against Large Language Model-Based Recommendation Systems: A New Distillation-Based Paradigm
Cuihong Li, Xiaowen Huang, Chuanhuan Yin, Jitao Sang
摘要
Membership Inference Attack (MIA) aims to determine whether a specific data sample was included in the training dataset of a target model. Traditional MIA approaches rely on shadow models to mimic target model behavior, but their effectiveness diminishes for Large Language Model (LLM)-based recommendation systems due to the scale and complexity of training data. This paper introduces a novel knowledge distillation-based MIA paradigm tailored for LLM-based recommendation systems. Our method constructs a reference model via distillation, applying distinct strategies for member and non-member data to enhance discriminative capabilities. The paradigm extracts fused features (e.g., confidence, entropy, loss, and hidden layer vectors) from the reference model to train an attack model, overcoming limitations of individual features. Extensive experiments on extended datasets (Last.FM, MovieLens, Book-Crossing, Delicious) and diverse LLMs (T5, GPT-2, LLaMA3) demonstrate that our approach significantly outperforms shadow model-based MIAs and individual-feature baselines. The results show its practicality for privacy attacks in LLM-driven recommender systems.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper8
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Membership Inference Attacks Against Machine Learning ModelsReza Shokri, Marco Stronati, Congzheng Song, Vitaly ShmatikovS&P 2017 · 被引用 5,137 次
- Extracting Training Data from Large Language ModelsNicholas Carlini, Florian Tramèr, Eric Wallace, Matthew Jagielski 等USENIX Security 2021 · 被引用 2,866 次
- Detecting Pretraining Data from Large Language ModelsWeijia Shi, Anirudh Ajith, Mengzhou Xia, Yangsibo Huang 等ICLR 2024 · 被引用 365 次
- Online Knowledge Distillation with Diverse PeersDefang Chen, Jian-Ping Mei, Can Wang, Yan Feng 等AAAI 2020 · 被引用 354 次
相关 Paper
- Decoding Web Memorization: A Semantic Membership Inference Attack on LLMsZhiyao Wu, Zi Liang, Haibo HuWWW 2026
- Debiasing Learning for Membership Inference Attacks Against Recommender SystemsZihan Wang, Na Huang, Fei Sun, Pengjie Ren 等KDD 2022 · 被引用 22 次
- DF-MIA: A Distribution-Free Membership Inference Attack on Fine-Tuned Large Language ModelsZhiheng Huang, Yannan Liu, Daojing He, Yu LiAAAI 2025 · 被引用 7 次
- Robust Membership Inference for Large Language Models under Adversarial Generative CorruptionYuanhong Huang, Huili Wang, Xueying Bai, Jinrui Wang 等ACL 2026
- Privacy Leaks by Adversaries: Adversarial Iterations for Membership Inference AttackJing Xue, Zhishen Sun, Haishan Ye, Luo Luo 等AAAI 2026
