No More, No Less: Least-Privilege Language Models
Paulius Rauba, Dominykas Seputis, Patrikas Vanagas, Mihaela van der Schaar
Abstract
Least privilege is a core security principle: grant each request only the minimum access needed to achieve its goal. Deployed language models almost never follow it, instead being exposed through a single API endpoint that serves all users and requests. This gap exists not because least privilege would be unhelpful—deployments would benefit greatly from reducing unnecessary capability exposure. The real obstacle is definitional and mechanistic: what does "access" mean inside a language model, and how can we enforce it without retraining or deploying multiple models? We take inspiration from least privilege in computer systems and define a class of models called least-privilege language models , where privilege is reachable internal computation during the forward pass. In this view, lowering privilege literally shrinks the model's accessible function class (as opposed to denying access via learned policies). We formalize deployment-time control as a monitor--allocator--enforcer stack, separating (i) request-time signals, (ii) a decision rule that allocates privilege, and (iii) an inference-time mechanism that selects privilege. We then propose Nested Least-Privilege Networks , a shape-preserving, rank-indexed intervention that provides a smooth, reversible control knob. We show that this knob yields policy-usable privilege--utility frontiers and enables selective suppression of targeted capabilities with limited collateral degradation across various policies. Most importantly, we see this as a defense of a completely new deployment paradigm which challenges the premise that we can only have output-level control of language models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 406095a9-e9b2-4b1a-bfdb-2adcf2f2f068Builds on26
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Locating and Editing Factual Associations in GPTKevin Meng, David Bau, Alex Andonian, Yonatan BelinkovNeurIPS 2022 · 3,415 citations
- Once-for-All: Train One Network and Specialize it for Efficient DeploymentHan Cai, Chuang Gan, Tianzhe Wang, Zhekai Zhang et al.ICLR 2020 · 1,522 citations
- Safe RLHF: Safe Reinforcement Learning from Human FeedbackJosef Dai, Xuehai Pan, Ruiyang Sun, Jiaming Ji et al.ICLR 2024 · 656 citations
Related papers
- SkillScope: Toward Fine-Grained Least-Privilege Enforcement for Agent SkillsJiangrong Wu, Yuhong Nan, Yixi Lin, Huaijin Wang et al.CCS 2026 · 5 citations
- ALPS: Automated Least-Privilege Enforcement for Securing Serverless FunctionsChanghee Shin, Bom Kim, Seungsoo LeeINFOCOM 2026 · 1 citation
- Automatically Reducing Privilege for Access Control PoliciesLoris D'Antoni, Shuo Ding, Amit Goel, Mathangi Ramesh et al.OOPSLA 2024 · 11 citations
- DOMBA: Double Model Balancing for Access-Controlled Language Models via Minimum-Bounded AggregationTom Segal, Asaf Shabtai, Yuval EloviciAAAI 2025 · 4 citations
- A Middle Path for On-Premises LLM Deployment: Preserving Privacy Without Sacrificing Model ConfidentialityHanbo Huang, Yihan Li, Bowen Jiang, Bo Jiang et al.EMNLP 2025 · 4 citations
