Training Socially Aligned Language Models on Simulated Social Interactions
Ruibo Liu, Ruixin Yang, Chenyan Jia, Ge Zhang, Diyi Yang, Soroush Vosoughi
Abstract
Social alignment in AI systems aims to ensure that these models behave according to established societal values. However, unlike humans, who derive consensus on value judgments through social interaction, current language models (LMs) are trained to rigidly replicate their training corpus in isolation, leading to subpar generalization in unfamiliar scenarios and vulnerability to adversarial attacks. This work presents a novel training paradigm that permits LMs to learn from simulated social interactions. In comparison to existing methodologies, our approach is considerably more scalable and efficient, demonstrating superior performance in alignment benchmarks and human evaluations. This paradigm shift in the training of LMs brings us a step closer to developing AI systems that can robustly and accurately reflect societal norms and values.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6d0187d3-9abd-4df4-b5d3-559377f8eb51Cited by top-tier papers18
- Self-Alignment of Large Language Models via Monopolylogue-based Social Scene SimulationXianghe Pang, Shuo Tang, Rui Ye, Yuxin Xiong et al.ICML 2024 · 50 citations
- Agents Under Siege: Breaking Pragmatic Multi-Agent LLM Systems with Optimized Prompt AttacksRana Muhammad Shahroz, Zhen Tan, Sukwon Yun, Charles Fleming et al.ACL 2025 · 18 citations
- Autonomous Agents for Collaborative Task under Information AsymmetryWei Liu, Chenxi Wang, Yifei Wang, Zihao Xie et al.NeurIPS 2024 · 17 citations
- INDICT: Code Generation with Internal Dialogues of Critiques for Both Security and HelpfulnessHung Le, Doyen Sahoo, Yingbo Zhou, Caiming Xiong et al.NeurIPS 2024 · 12 citations
- Finetuning LLMs for Human Behavior Prediction in Social Science ExperimentsAkaash Kolluri, Shengguang Wu, Joon Sung Park, Michael S. BernsteinEMNLP 2025 · 12 citations
Builds on17
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- TruthfulQA: Measuring How Models Mimic Human FalsehoodsStephanie Lin, Jacob Hilton, Owain EvansACL 2022 · 3,228 citations
- SimCSE: Simple Contrastive Learning of Sentence EmbeddingsTianyu Gao, Xingcheng Yao, Danqi ChenEMNLP 2021 · 2,496 citations
- Generative Agents: Interactive Simulacra of Human BehaviorJoon Sung Park, Joseph C. O'Brien, Carrie Jun Cai, Meredith Ringel Morris et al.UIST 2023 · 1,882 citations
Related papers
- Distributional LLM-as-a-JudgeLuyu Chen, Zeyu Zhang, Haoran Tan, Quanyu Dai et al.NeurIPS 2025 · 4 citations
- Cultural Learning-Based Culture Adaptation of Language ModelsChen Cecilia Liu, Anna Korhonen, Iryna GurevychACL 2025 · 14 citations
- Aligning Large Language Models through Synthetic FeedbackSungdong Kim, Sanghwan Bae, Jamin Shin, Soyoung Kang et al.EMNLP 2023 · 13 citations
- Second Thoughts are Best: Learning to Re-Align With Human Values from Text EditsRuibo Liu, Chenyan Jia, Ge Zhang, Ziyu Zhuang et al.NeurIPS 2022 · 46 citations
- ReMoDetect: Reward Models Recognize Aligned LLM's GenerationsHyunseok Lee, Jihoon Tack, Jinwoo ShinNeurIPS 2024 · 13 citations
