Unlocking Memorization in Large Language Models with Dynamic Soft Prompting
Zhepeng Wang, Runxue Bao, Yawen Wu, Jackson Taylor, Cao Xiao, Feng Zheng, Weiwen Jiang, Shangqian Gao, Yanfu Zhang
摘要
Pretrained large language models (LLMs) have excelled in a variety of natural language processing (NLP) tasks, including summarization, question answering, and translation. However, LLMs pose significant security risks due to their tendency to memorize training data, leading to potential privacy breaches and copyright infringement. Therefore, accurate measurement of the memorization is essential to evaluate and mitigate these potential risks. However, previous attempts to characterize memorization are constrained by either using prefixes only or by prepending a constant soft prompt to the prefixes, which cannot react to changes in input. To address this challenge, we propose a novel method for estimating LLM memorization using dynamic, prefix-dependent soft prompts. Our approach involves training a transformer-based generator to produce soft prompts that adapt to changes in input, thereby enabling more accurate extraction of memorized data. Our method not only addresses the limitations of previous methods but also demonstrates superior performance in diverse experimental settings compared to state-of-theart techniques. In particular, our method can achieve the maximum relative improvement of 135.3% and 39.8% over the vanilla baseline on average in terms of discoverable memorization rate for the text generation task and code generation task, respectively. Our code is available at https://github.com/wangger/llmmemorization-dsp . Prefix Method Generations and the Stack is set to Main.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- AgentTTS: Large Language Model Agent for Test-time Compute-optimal Scaling Strategy in Complex TasksFali Wang, Hui Liu, Zhenwei Dai, Jingying Zeng 等NeurIPS 2025 · 被引用 20 次
- Retracing the Past: LLMs Emit Training Data When They Get LostMyeongseob Ko, Nikhil Reddy Billa, Adam Nguyen, Charles Fleming 等EMNLP 2025 · 被引用 1 次
- Catastrophic Failure of LLM Unlearning via QuantizationZhiwei Zhang, Fali Wang, Xiaomin Li, Zongyu Wu 等ICLR 2025
- Controllable Memorization in LLMs via Weight PruningChenjie Ni, Zhepeng Wang, Runxue Bao, Shangqian Gao 等EMNLP 2025
- A Text-Routed Sparse Mixture-of-Experts Model with Explanation and Temporal Alignment for Multi-Modal Sentiment AnalysisDongning Rao, Yunbiao Zeng, Zhihua Jiang, Jujian LvAAAI 2026
它引用的顶会 Paper23
- Deep Learning with Differential PrivacyMartín Abadi, Andy Chu, Ian J. Goodfellow, H. Brendan McMahan 等CCS 2016 · 被引用 7,620 次
- Extracting Training Data from Large Language ModelsNicholas Carlini, Florian Tramèr, Eric Wallace, Matthew Jagielski 等USENIX Security 2021 · 被引用 2,866 次
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley 等ICML 2023 · 被引用 1,822 次
- AdaptFormer: Adapting Vision Transformers for Scalable Visual RecognitionShoufa Chen, Chongjian Ge, Zhan Tong, Jiangliu Wang 等NeurIPS 2022 · 被引用 1,291 次
- Deduplicating Training Data Makes Language Models BetterKatherine Lee, Daphne Ippolito, Andrew Nystrom, Chiyuan Zhang 等ACL 2022 · 被引用 844 次
相关 Paper
- Decoding Secret Memorization in Code LLMs Through Token-Level CharacterizationYuqing Nie, Chong Wang, Kailong Wang, Guoai Xu 等ICSE 2025 · 被引用 10 次
- Quantifying Memorization Across Neural Language ModelsNicholas Carlini, Daphne Ippolito, Matthew Jagielski, Katherine Lee 等ICLR 2023 · 被引用 158 次
- Large Language Model Unlearning for Source CodeXue Jiang, Yihong Dong, Huangzhao Zhang, Tangxinyu Wang 等AAAI 2026
- Positional Fragility in LLMs: How Offset Effects Reshape Our Understanding of Memorization RisksYixuan Even Xu, Antoine Bosselut, Imanol SchlagNeurIPS 2025 · 被引用 2 次
- A Multi-Perspective Analysis of Memorization in Large Language ModelsBowen Chen, Namgi Han, Yusuke MiyaoEMNLP 2024 · 被引用 2 次
