PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference
Jiaming Ji, Donghai Hong, Borong Zhang, Boyuan Chen, Josef Dai, Boren Zheng, Tianyi Alex Qiu, Jiayi Zhou, Kaile Wang, Boxun Li, Sirui Han, Yike Guo, Yaodong Yang
摘要
In this work, we introduce the PKU-SAFERLHF dataset, designed to promote research on safety alignment in large language models (LLMs). As a sibling project to SafeRLHF and BEAVERTAILS, we separate annotations of helpfulness and harmlessness for question-answering pairs, providing distinct perspectives on these coupled attributes. Overall, we provide 44.6k refined prompts and 265k question-answer pairs with safety meta-labels for 19 harm categories and three severity levels ranging from minor to severe, with answers generated by Llama-family models. Based on this, we collected 166.8k preference data, including dual-preference (helpfulness and harmlessness decoupled) and singlepreference data (trade-off the helpfulness and harmlessness from scratch), respectively. Using the large-scale annotation data, we further train severity-sensitive moderation for the risk control of LLMs and safety-centric RLHF algorithms for the safety alignment of LLMs. We believe this dataset will be a valuable resource for the community, aiding in the safe deployment of LLMs. 1 Warning: this paper contains example data that may be offensive or harmful.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper64
- Aligner: Efficient Alignment by Learning to CorrectJiaming Ji, Boyuan Chen, Hantao Lou, Donghai Hong 等NeurIPS 2024 · 被引用 115 次
- AgenTracer: Who Is Inducing Failure in the LLM Agentic Systems?Guibin Zhang, Junhao Wang, Junjie Chen, Wangchunshu Zhou 等ICLR 2026 · 被引用 107 次
- Your Agent May Misevolve: Emergent Risks in Self-evolving LLM AgentsShuai Shao, Qihan Ren, Dongrui Liu, Chen Qian 等ICLR 2026 · 被引用 60 次
- Safe RLHF-V: Safe Reinforcement Learning from Multi-modal Human FeedbackJiaming Ji, Xinyu Chen, Rui Pan, Han Zhu 等NeurIPS 2025 · 被引用 28 次
- What's In My Human Feedback? Learning Interpretable Descriptions of Preference DataRajiv Movva, Smitha Milli, Sewon Min, Emma PiersonICLR 2026 · 被引用 27 次
它引用的顶会 Paper9
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!Xiangyu Qi, Yi Zeng, Tinghao Xie, Pin-Yu Chen 等ICLR 2024 · 被引用 1,104 次
- Towards Understanding Sycophancy in Language ModelsMrinank Sharma, Meg Tong, Tomasz Korbak, David Duvenaud 等ICLR 2024 · 被引用 762 次
相关 Paper
- SPA-VL: A Comprehensive Safety Preference Alignment Dataset for Vision Language ModelsYongting Zhang, Lu Chen, Guodong Zheng, Yifeng Gao 等CVPR 2025
- Safe RLHF: Safe Reinforcement Learning from Human FeedbackJosef Dai, Xuehai Pan, Ruiyang Sun, Jiaming Ji 等ICLR 2024 · 被引用 656 次
- Safety-Tuned LLaMAs: Lessons From Improving the Safety of Large Language Models that Follow InstructionsFederico Bianchi, Mirac Suzgun, Giuseppe Attanasio, Paul Röttger 等ICLR 2024 · 被引用 373 次
- MM-RLHF: The Next Step Forward in Multimodal LLM AlignmentYifan Zhang, Tao Yu, Haochen Tian, Chaoyou Fu 等ICML 2025
- PLLuM-Align: Polish Preference Dataset for Large Language Model AlignmentKarolina Seweryn, Anna Kolos, Agnieszka Karlinska, Katarzyna Lorenc 等EMNLP 2025
