Conformal Tail Risk Control for Large Language Model Alignment
Catherine Yu-Chi Chen, Jingyan Shen, Zhun Deng, Lihua Lei
摘要
Recent developments in large language models (LLMs) have led to their widespread usage for various tasks. The prevalence of LLMs in society implores the assurance on the reliability of their performance. In particular, risk-sensitive applications demand meticulous attention to unexpectedly poor outcomes, i.e., tail events, for instance, toxic answers, humiliating language, and offensive outputs. Due to the costly nature of acquiring human annotations, general-purpose scoring models have been created to automate the process of quantifying these tail events. This phenomenon introduces potential human-machine misalignment between the respective scoring mechanisms. In this work, we present a lightweight calibration framework for blackbox models that ensures the alignment of humans and machines with provable guarantees. Our framework provides a rigorous approach to controlling any distortion risk measure that is characterized by a weighted average of quantiles of the loss incurred by the LLM with high confidence. The theoretical foundation of our method relies on the connection between conformal risk control and a traditional family of statistics, i.e., L-statistics. To demonstrate the utility of our framework, we conduct comprehensive experiments that address the issue of human-machine misalignment.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Conformal Risk Training: End-to-End Optimization of Conformal Risk ControlChristopher Yeh, Nicolas Christianson, Adam Wierman, Yisong YueNeurIPS 2025 · 被引用 14 次
- Conformal Arbitrage: Risk-Controlled Balancing of Competing Objectives in Language ModelsWilliam Overman, Mohsen BayatiNeurIPS 2025 · 被引用 12 次
- COFT: Counterfactual–Conformal Decoding for Fair Chain‑of‑Thought Reasoning in Large Language ModelsArya Fayyazi, Mehdi Kamal, Massoud PedramICML 2026
- Adversarially Robust Control of Conditional Value-at-Risk via Rockafellar-Uryasev Conformal InferenceCatherine Chen, Jingyan Shen, Xinyu Yang, Lihua LeiICML 2026
- Calibrating Conservatism for Scalable OversightWilliam Overman, Mohsen BayatiICML 2026
它引用的顶会 Paper6
- Conformal Risk ControlAnastasios Nikolas Angelopoulos, Stephen Bates, Adam Fisch, Lihua Lei 等ICLR 2024 · 被引用 242 次
- Conformal Language ModelingVictor Quach, Adam Fisch, Tal Schuster, Adam Yala 等ICLR 2024 · 被引用 132 次
- Large language model validity via enhanced conformal prediction methodsJohn J. Cherian, Isaac Gibbs, Emmanuel J. CandèsNeurIPS 2024 · 被引用 120 次
- Language Models with Conformal Factuality GuaranteesChristopher Mohri, Tatsunori HashimotoICML 2024 · 被引用 107 次
- Prompt Risk Control: A Rigorous Framework for Responsible Deployment of Large Language ModelsThomas P. Zollo, Todd Morrill, Zhun Deng, Jake Snell 等ICLR 2024 · 被引用 14 次
相关 Paper
- Beyond Expectations: Quantile-Guided Alignment for Risk-Calibrated Language ModelsXinran Wang, Jin Du, Azal Ahmad Khan, Qi Le 等NeurIPS 2025 · 被引用 1 次
- How to Correctly Report LLM-as-a-Judge EvaluationsChungpa Lee, Thomas Zeng, Jongwon Jeong, Jy-yong Sohn 等ICML 2026 · 被引用 24 次
- Safety Game: Inference-Time Alignment of Black-Box LLMs via Constrained OptimizationTuan Nguyen, Long Tran-ThanhICML 2026
- Noisy but Valid: Robust Statistical Evaluation of LLMs with Imperfect JudgesChen Feng, Minghe Shen, Ananth Balashankar, Carsten Gerner-Beuerle 等ICLR 2026 · 被引用 24 次
- Conformal Prediction Beyond the Seen: A Missing Mass Perspective for Uncertainty Quantification in Generative ModelsSima Noorani, Shayan Kiyani, George J. Pappas, Hamed HassaniNeurIPS 2025 · 被引用 10 次
