Risk-Averse Fine-tuning of Large Language Models
Sapana Chaudhary, Ujwal Dinesha, Dileep Kalathil, Srinivas Shakkottai
Abstract
We consider the challenge of mitigating the generation of negative or toxic content by the Large Language Models (LLMs) in response to certain prompts. We propose integrating risk-averse principles into LLM fine-tuning to minimize the occurrence of harmful outputs, particularly rare but significant events. By optimizing the risk measure of Conditional Value at Risk (CVaR), our methodology trains LLMs to exhibit superior performance in avoiding toxic outputs while maintaining effectiveness in generative tasks. Empirical evaluations on sentiment modification and toxicity mitigation tasks demonstrate the efficacy of risk-averse reinforcement learning with human feedback (RLHF) in promoting a safer and more constructive online discourse environment.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 80b719d0-cbc2-42cb-a701-bcfe7470456dCited by top-tier papers6
- Uni-DPO: A Unified Paradigm for Dynamic Preference Optimization of LLMsShangpin Peng, Weinong Wang, Zhuotao Tian, Senqiao Yang et al.ICLR 2026 · 10 citations
- Efficient Tail-Aware Generative Optimization via Flow Model Fine-TuningZifan Wang, Riccardo De Santi, Xiaoyu Mo, Michael Zavlanos et al.ICML 2026 · 4 citations
- QStore: Quantization-Aware Compressed Model StorageRaunak Shah, Zhaoheng Li, Yongjoo ParkVLDB 2026 · 3 citations
- Risk-aware Direct Preference Optimization under Nested Risk MeasureLijun Zhang, Lin Li, Yajie Qi, Huizhong Song et al.NeurIPS 2025 · 3 citations
- Generating Informative Samples for Risk-Averse Fine-Tuning of Downstream TasksHeasung Kim, Taekyun Lee, Hyeji Kim, Gustavo de VecianaNeurIPS 2025 · 2 citations
Builds on13
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
Related papers
- Beyond Expectations: Quantile-Guided Alignment for Risk-Calibrated Language ModelsXinran Wang, Jin Du, Azal Ahmad Khan, Qi Le et al.NeurIPS 2025 · 1 citation
- Safe RLHF: Safe Reinforcement Learning from Human FeedbackJosef Dai, Xuehai Pan, Ruiyang Sun, Jiaming Ji et al.ICLR 2024 · 656 citations
- QUARK: Controllable Text Generation with Reinforced UnlearningXiming Lu, Sean Welleck, Jack Hessel, Liwei Jiang et al.NeurIPS 2022 · 290 citations
- Unveiling the Implicit Toxicity in Large Language ModelsJiaxin Wen, Pei Ke, Hao Sun, Zhexin Zhang et al.EMNLP 2023 · 21 citations
- Adversarially Robust Control of Conditional Value-at-Risk via Rockafellar-Uryasev Conformal InferenceCatherine Chen, Jingyan Shen, Xinyu Yang, Lihua LeiICML 2026
