Process for Adapting Language Models to Society (PALMS) with Values-Targeted Datasets
Irene Solaiman, Christy Dennison
Abstract
Language models can generate harmful and biased outputs and exhibit undesirable behavior according to a given cultural context. We propose a Process for Adapting Language Models to Society (PALMS) with Values-Targeted Datasets, an iterative process to significantly change model behavior by crafting and fine-tuning on a dataset that reflects a predetermined set of target values. We evaluate our process using three metrics: quantitative metrics with human evaluations that score output adherence to a target value, toxicity scoring on outputs; and qualitative metrics analyzing the most common word associated with a given social category. Through each iteration, we add additional training dataset examples based on observed shortcomings from evaluations. PALMS performs significantly better on all metrics compared to baseline and control models for a broad range of GPT-3 language model sizes without compromising capability integrity. We find that the effectiveness of PALMS increases with model size. We show that significantly adjusting language model behavior is feasible with a small, hand-curated dataset.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2ecdf10d-6881-4090-bfb1-377820dba92fCited by top-tier papers50
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu et al.ICLR 2022 · 4,966 citations
- TruthfulQA: Measuring How Models Mimic Human FalsehoodsStephanie Lin, Jacob Hilton, Owain EvansACL 2022 · 3,228 citations
- Jailbroken: How Does LLM Safety Training Fail?Alexander Wei, Nika Haghtalab, Jacob SteinhardtNeurIPS 2023 · 2,230 citations
- Principle-Driven Self-Alignment of Language Models from Scratch with Minimal Human SupervisionZhiqing Sun, Yikang Shen, Qinhong Zhou, Hongxin Zhang et al.NeurIPS 2023 · 463 citations
Builds on5
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes et al.ICLR 2020 · 4,112 citations
- Towards Debiasing Sentence RepresentationsPaul Pu Liang, Irene Mengze Li, Emily Zheng, Yao Chong Lim et al.ACL 2020 · 149 citations
- Learning from others' mistakes: Avoiding dataset biases without modeling themVictor Sanh, Thomas Wolf, Yonatan Belinkov, Alexander M. RushICLR 2021 · 123 citations
- Don't Stop Pretraining: Adapt Language Models to Domains and TasksSuchin Gururangan, Ana Marasovic, Swabha Swayamdipta, Kyle Lo et al.ACL 2020 · 93 citations
Related papers
- Exploring the Limits of Domain-Adaptive Training for Detoxifying Large-Scale Language ModelsBoxin Wang, Wei Ping, Chaowei Xiao, Peng Xu et al.NeurIPS 2022 · 89 citations
- Walking in Others' Shoes: How Perspective-Taking Guides Large Language Models in Reducing Toxicity and BiasRongwu Xu, Zi'an Zhou, Tianwei Zhang, Zehan Qi et al.EMNLP 2024 · 2 citations
- Pretraining Language Models with Human PreferencesTomasz Korbak, Kejian Shi, Angelica Chen, Rasika Vinayak Bhalerao et al.ICML 2023 · 287 citations
- SQuARe: A Large-Scale Dataset of Sensitive Questions and Acceptable Responses Created through Human-Machine CollaborationHwaran Lee, Seokhee Hong, Joonsuk Park, Takyoung Kim et al.ACL 2023 · 3 citations
- Code Red! On the Harmfulness of Applying Off-the-Shelf Large Language Models to Programming TasksAli Al-Kaswan, Sebastian Deatc, Begüm Koç, Arie van Deursen et al.FSE 2025 · 1 citation
