Model Stealing for Any Low-Rank Language Model
Allen Liu, Ankur Moitra
Abstract
Model stealing, where a learner tries to recover an unknown model via carefully chosen queries, is a critical problem in machine learning, as it threatens the security of proprietary models and the privacy of data they are trained on. In recent years, there has been particular interest in stealing large language models (LLMs). In this paper, we aim to build a theoretical understanding of stealing language models by studying a simple and mathematically tractable setting. We study model stealing for Hidden Markov Models (HMMs), and more generally low-rank language models. We assume that the learner works in the conditional query model, introduced by Kakade, Krishnamurthy, Mahajan and Zhang [KKMZ24]. Our main result is an efficient algorithm in the conditional query model, for learning any low-rank distribution. In other words, our algorithm succeeds at stealing any language model whose output distribution is low-rank. This improves upon the result in [KKMZ24] which also requires the unknown distribution to have high "fidelity" -a property that holds only in restricted cases. There are two key insights behind our algorithm: First, we represent the conditional distributions at each timestep by constructing barycentric spanners among a collection of vectors of exponentially large dimension. Second, for sampling from our representation, we iteratively solve a sequence of convex optimization problems that involve projection in relative entropy to prevent compounding of errors over the length of the sequence. This is an interesting example where, at least theoretically, allowing a machine learning model to solve more complex problems at inference time can lead to drastic improvements in its performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b2648920-272e-428e-b25c-2f3508873fb7Cited by top-tier papers2
- Sequences of Logits Reveal the Low Rank Structure of Language ModelsNoah Golowich, Allen Liu, Abhishek ShettyICLR 2026 · 9 citations
- Language Identification in the Limit with Computational TraceBinghui Peng, Amin Saberi, Grigoris VelegkasICLR 2026
Builds on16
- Stealing Machine Learning Models via Prediction APIsFlorian Tramèr, Fan Zhang, Ari Juels, Michael K. Reiter et al.USENIX Security 2016 · 2,088 citations
- Stealing Hyperparameters in Machine LearningBinghui Wang, Neil Zhenqiang GongS&P 2018 · 504 citations
- Entangled Watermarks as a Defense against Model ExtractionHengrui Jia, Christopher A. Choquette-Choo, Varun Chandrasekaran, Nicolas PapernotUSENIX Security 2021 · 287 citations
- Stealing Links from Graph Neural NetworksXinlei He, Jinyuan Jia, Michael Backes, Neil Zhenqiang Gong et al.USENIX Security 2021 · 226 citations
- Prediction Poisoning: Towards Defenses Against DNN Model Stealing AttacksTribhuvanesh Orekondy, Bernt Schiele, Mario FritzICLR 2020 · 194 citations
Related papers
- LoRO: Real-Time on-Device Secure Inference for LLMs via TEE-Based Low Rank ObfuscationGaojian Xiong, Yu Sun, Jianhua Liu, Jian Cui et al.NeurIPS 2025 · 6 citations
- Medical Multimodal Model Stealing Attacks via Adversarial Domain AlignmentYaling Shen, Zhixiong Zhuang, Kun Yuan, Maria-Irina Nicolae et al.AAAI 2025 · 12 citations
- Towards Data-Free Model Stealing in a Hard Label SettingSunandini Sanyal, Sravanti Addepalli, R. Venkatesh BabuCVPR 2022 · 76 citations
- Mimicking the Familiar: Dynamic Command Generation for Information Theft Attacks in LLM Tool-Learning SystemZiyou Jiang, Mingyang Li, Guowei Yang, Junjie Wang et al.ACL 2025
- Marich: A Query-efficient Distributionally Equivalent Model Extraction AttackPratik Karmakar, Debabrota BasuNeurIPS 2023 · 11 citations
