Prompt Risk Control: A Rigorous Framework for Responsible Deployment of Large Language Models
Thomas P. Zollo, Todd Morrill, Zhun Deng, Jake Snell, Toniann Pitassi, Richard S. Zemel
Abstract
The recent explosion in the capabilities of large language models has led to a wave of interest in how best to prompt a model to perform a given task. While it may be tempting to simply choose a prompt based on average performance on a validation set, this can lead to a deployment where unexpectedly poor responses are generated, especially for the worst-off users. To mitigate this prospect, we propose Prompt Risk Control, a lightweight framework for selecting a prompt based on rigorous upper bounds on families of informative risk measures. We offer methods for producing bounds on a diverse set of metrics, including quantities that measure worst-case responses and disparities in generation quality across the population of users. In addition, we extend the underlying statistical bounding techniques to accommodate the possibility of distribution shifts in deployment. Experiments on applications such as open-ended chat, medical question summarization, and code generation highlight how such a framework can foster responsible deployment by reducing the risk of the worst outcomes. Recent leaps in the capabilities of large language models (LLMs) such as GPT-4 (OpenAI, 2023), LLaMA (Touvron et al., 2023), and Claude have driven a wave of interest in constructing the best prompt for a given task, where a prompt generally refers to an input to the LLM. Various prompting strategies have been proposed, including but not limited to: in-context learning (Brown et al., 2020) , instruction following (Wei et al., 2022a), chain-of-thought prompting (Wei et al., 2022b), Published as a conference paper at ICLR 2024. * indicates equal contribution.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1f73afa8-42f1-48b7-91de-c2c93692846bCited by top-tier papers8
- Fast yet Safe: Early-Exiting with Risk ControlMetod Jazbec, Alexander Timans, Tin Hadzi Veljkovic, Kaspar Sakmann et al.NeurIPS 2024 · 35 citations
- SConU: Selective Conformal Uncertainty in Large Language ModelsZhiyuan Wang, Qingni Wang, Yue Zhang, Tianlong Chen et al.ACL 2025 · 18 citations
- Certifying Counterfactual Bias in LLMsIsha Chaudhary, Qian Hu, Manoj Kumar, Morteza Ziyadi et al.ICLR 2025 · 3 citations
- QuEst: Enhancing Estimates of Quantile-Based Distributional Measures Using Model PredictionsZhun Deng, Thomas P. Zollo, Benjamin Eyre, Amogh Inamdar et al.ICML 2025
- Conformal Tail Risk Control for Large Language Model AlignmentCatherine Yu-Chi Chen, Jingyan Shen, Zhun Deng, Lihua LeiICML 2025
Builds on13
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu et al.ICLR 2022 · 4,966 citations
- Adaptive Conformal Inference Under Distribution ShiftIsaac Gibbs, Emmanuel J. CandèsNeurIPS 2021 · 665 citations
- Confident Adaptive Language ModelingTal Schuster, Adam Fisch, Jai Gupta, Mostafa Dehghani et al.NeurIPS 2022 · 394 citations
Related papers
- Scam2Prompt: A Scalable Framework for Auditing Malicious Scam Endpoints in Production LLMsZhiyang Chen, Tara Saba, Xun Deng, Xujie Si et al.ICML 2026
- Ask Me Anything: A simple strategy for prompting language modelsSimran Arora, Avanika Narayan, Mayee F. Chen, Laurel J. Orr et al.ICLR 2023 · 74 citations
- Why Johnny Can't Prompt: How Non-AI Experts Try (and Fail) to Design LLM PromptsJ. D. Zamfirescu-Pereira, Richmond Y. Wong, Bjoern Hartmann, Qian YangCHI 2023 · 892 citations
- Pareto Optimal Learning for Estimating Large Language Model ErrorsTheodore Zhao, Mu Wei, Joseph Preston, Hoifung PoonACL 2024
- Analyzing and Modeling LLM Response Lengths with Extreme Value Theory: Anchoring Effects and Hybrid DistributionsLiuxuan Jiao, Chen Gao, Yiqian Yang, Chenliang Zhou et al.EMNLP 2025
