Parametric Social Identity Injection and Diversification in Public Opinion Simulation
Hexi Wang, Yujia Zhou, Bangde Du, Qingyao Ai, Yiqun Liu
摘要
Large language models (LLMs) have recently been adopted as synthetic agents for public opinion simulation, offering a promising alternative to costly and slow human surveys. Despite their scalability, current LLM-based simulation methods fail to capture social diversity, producing flattened inter-group differences and overly homogeneous responses across demographic groups. We identify this limitation as a Diversity Collapse phenomenon in LLM hidden representations, where distinct social identities become increasingly indistinguishable across layers. Motivated by this observation, we propose Parametric Social Identity Injection (PSII), a general framework that injects explicit, parametric representations of demographic attributes and value orientations directly into intermediate hidden states of LLMs. Unlike prompt-based persona conditioning, PSII enables fine-grained and controllable identity modulation at the representation level. Extensive experiments on the World Values Survey using multiple open-source LLMs show that PSII significantly improves distributional fidelity and diversity, reducing KL divergence to real-world survey data while enhancing overall diversity. This work provides new insights into representation-level control of LLM agents and advances scalable, diversity-aware public opinion simulation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper11
- Whose Opinions Do Language Models Reflect?Shibani Santurkar, Esin Durmus, Faisal Ladhak, Cinoo Lee 等ICML 2023 · 被引用 764 次
- Increasing Diversity While Maintaining Accuracy: Text Data Generation with Large Language Models and Human InterventionsJohn Joon Young Chung, Ece Kamar, Saleema AmershiACL 2023 · 被引用 55 次
- Language Model Fine-Tuning on Scaled Survey Data for Predicting Distributions of Public OpinionsJoseph Suh, Erfan Jahanparast, Suhong Moon, Minwoo Kang 等ACL 2025 · 被引用 48 次
- Cultivating Pluralism In Algorithmic Monoculture: The Community Alignment DatasetLily H Zhang, Smitha Milli, Karen Long Jusko, Jonathan Smith 等ICLR 2026 · 被引用 41 次
- Parametric Retrieval Augmented GenerationWeihang Su, Yichen Tang, Qingyao Ai, Junxi Yan 等SIGIR 2025 · 被引用 25 次
相关 Paper
- Synthia: Scalable Grounded Persona Generation from Social Media DataVahid Rahimzadeh, Erfan Moosavi Monazzah, Mohammad Taher Pilehvar, Yadollah YaghoobzadehACL 2026 · 被引用 1 次
- Interview-Informed Generative Agents for Product Discovery: A Validation StudyZichao Wang, Alexa F. SiuCHI 2026 · 被引用 1 次
- From Values to Opinions: Predicting Human Behaviors and Stances Using Value-Injected Large Language ModelsDongjun Kang, Joonsuk Park, Yohan Jo, JinYeong BakEMNLP 2023 · 被引用 4 次
- Which Demographics do LLMs Default to During Annotation?Johannes Schäfer, Aidan Combs, Christopher Bagdon, Jiahui Li 等ACL 2025 · 被引用 11 次
- Survey Response Generation: Generating Closed-Ended Survey Responses In-Silico with Large Language ModelsGeorg Ahnert, Anna-Carolina Haensch, Barbara Plank, Markus StrohmaierACL 2026 · 被引用 4 次
