CoBRA: Programming Cognitive Bias in Social Agents Using Classic Social Science Experiments
Xuan Liu, HaoYang Shang, Haojian Jin
摘要
This paper introduces CoBRA, a novel toolkit for systematically specifying agent behavior in LLM-based social simulation. We found that conventional approaches that specify agent behavior through implicit natural-language descriptions often do not yield consistent behavior across models, and the resulting behavior does not capture the nuances of the descriptions. In contrast, CoBRA introduces a model-agnostic way to control agent behavior that lets researchers explicitly specify desired nuances and obtain consistent behavior across models. At the heart of CoBRA is a novel closed-loop system primitive with two components: (1) Cognitive Bias Index that measures the demonstrated cognitive bias of a social agent, by quantifying the agent’s reactions in a set of validated classic social science experiments; (2) Behavioral Regulation Engine that aligns the agent’s behavior to exhibit controlled cognitive bias. Through CoBRA, we show how to operationalize validated social-science knowledge (i.e., classical experiments) as reusable “gym” environments for AI—an approach that may generalize to richer social and affective simulations beyond bias alone.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- When Agents "Misremember" Collectively: Exploring the Mandela Effect in LLM-based Multi-Agent SystemsNaen Xu, Hengyu An, Shuo Shi, Jinghuai Zhang 等ICLR 2026 · 被引用 3 次
- Cross-Domain Molecular Relational Learning: Leveraging Chemical Structure-Activity AnalysisPeiliang Zhang, Jingling Yuan, Shiqing Wu, Mengqing Hu 等KDD 2026 · 被引用 1 次
- "I See What You Did There": Can Large Vision-Language Models Understand Multimodal Puns?Naen Xu, Jiayi Sheng, Changjiang Li, Chunyi Zhou 等ACL 2026 · 被引用 1 次
它引用的顶会 Paper38
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- CAMEL: Communicative Agents for "Mind" Exploration of Large Language Model SocietyGuohao Li, Hasan Hammoud, Hani Itani, Dmitrii Khizbullin 等NeurIPS 2023 · 被引用 1,975 次
相关 Paper
- Justice or Prejudice? Quantifying Biases in LLM-as-a-JudgeJiayi Ye, Yanbo Wang, Yue Huang, Dongping Chen 等ICLR 2025
- Capturing Failures of Large Language Models via Human Cognitive BiasesErik Jones, Jacob SteinhardtNeurIPS 2022 · 被引用 154 次
- Emulating Aggregate Human Choice Behavior and Biases with GPT Conversational AgentsStephen Pilli, Vivek NallurCHI 2026 · 被引用 2 次
- SOCIA-EVO: Automated Simulator Construction via Dual-Anchored Bi-Level OptimizationYuncheng Hua, Sion Weatherhead, Mehdi Jafari, Hao Xue 等ACL 2026
- BiasAsker: Measuring the Bias in Conversational AI SystemYuxuan Wan, Wenxuan Wang, Pinjia He, Jiazhen Gu 等FSE 2023 · 被引用 50 次
