Conformal Tail Risk Control for Large Language Model Alignment
Catherine Yu-Chi Chen, Jingyan Shen, Zhun Deng, Lihua Lei
Abstract
Recent developments in large language models (LLMs) have led to their widespread usage for various tasks. The prevalence of LLMs in society implores the assurance on the reliability of their performance. In particular, risk-sensitive applications demand meticulous attention to unexpectedly poor outcomes, i.e., tail events, for instance, toxic answers, humiliating language, and offensive outputs. Due to the costly nature of acquiring human annotations, general-purpose scoring models have been created to automate the process of quantifying these tail events. This phenomenon introduces potential human-machine misalignment between the respective scoring mechanisms. In this work, we present a lightweight calibration framework for blackbox models that ensures the alignment of humans and machines with provable guarantees. Our framework provides a rigorous approach to controlling any distortion risk measure that is characterized by a weighted average of quantiles of the loss incurred by the LLM with high confidence. The theoretical foundation of our method relies on the connection between conformal risk control and a traditional family of statistics, i.e., L-statistics. To demonstrate the utility of our framework, we conduct comprehensive experiments that address the issue of human-machine misalignment.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8e1f30d2-25c9-428f-8aed-cb4998dedf9eCited by top-tier papers5
- Conformal Risk Training: End-to-End Optimization of Conformal Risk ControlChristopher Yeh, Nicolas Christianson, Adam Wierman, Yisong YueNeurIPS 2025 · 14 citations
- Conformal Arbitrage: Risk-Controlled Balancing of Competing Objectives in Language ModelsWilliam Overman, Mohsen BayatiNeurIPS 2025 · 12 citations
- COFT: Counterfactual–Conformal Decoding for Fair Chain‑of‑Thought Reasoning in Large Language ModelsArya Fayyazi, Mehdi Kamal, Massoud PedramICML 2026
- Adversarially Robust Control of Conditional Value-at-Risk via Rockafellar-Uryasev Conformal InferenceCatherine Chen, Jingyan Shen, Xinyu Yang, Lihua LeiICML 2026
- Calibrating Conservatism for Scalable OversightWilliam Overman, Mohsen BayatiICML 2026
Builds on6
- Conformal Risk ControlAnastasios Nikolas Angelopoulos, Stephen Bates, Adam Fisch, Lihua Lei et al.ICLR 2024 · 242 citations
- Conformal Language ModelingVictor Quach, Adam Fisch, Tal Schuster, Adam Yala et al.ICLR 2024 · 132 citations
- Large language model validity via enhanced conformal prediction methodsJohn J. Cherian, Isaac Gibbs, Emmanuel J. CandèsNeurIPS 2024 · 120 citations
- Language Models with Conformal Factuality GuaranteesChristopher Mohri, Tatsunori HashimotoICML 2024 · 107 citations
- Prompt Risk Control: A Rigorous Framework for Responsible Deployment of Large Language ModelsThomas P. Zollo, Todd Morrill, Zhun Deng, Jake Snell et al.ICLR 2024 · 14 citations
Related papers
- Beyond Expectations: Quantile-Guided Alignment for Risk-Calibrated Language ModelsXinran Wang, Jin Du, Azal Ahmad Khan, Qi Le et al.NeurIPS 2025 · 1 citation
- How to Correctly Report LLM-as-a-Judge EvaluationsChungpa Lee, Thomas Zeng, Jongwon Jeong, Jy-yong Sohn et al.ICML 2026 · 24 citations
- Safety Game: Inference-Time Alignment of Black-Box LLMs via Constrained OptimizationTuan Nguyen, Long Tran-ThanhICML 2026
- Noisy but Valid: Robust Statistical Evaluation of LLMs with Imperfect JudgesChen Feng, Minghe Shen, Ananth Balashankar, Carsten Gerner-Beuerle et al.ICLR 2026 · 24 citations
- Conformal Prediction Beyond the Seen: A Missing Mass Perspective for Uncertainty Quantification in Generative ModelsSima Noorani, Shayan Kiyani, George J. Pappas, Hamed HassaniNeurIPS 2025 · 10 citations
