Parametric Social Identity Injection and Diversification in Public Opinion Simulation
Hexi Wang, Yujia Zhou, Bangde Du, Qingyao Ai, Yiqun Liu
Abstract
Large language models (LLMs) have recently been adopted as synthetic agents for public opinion simulation, offering a promising alternative to costly and slow human surveys. Despite their scalability, current LLM-based simulation methods fail to capture social diversity, producing flattened inter-group differences and overly homogeneous responses across demographic groups. We identify this limitation as a Diversity Collapse phenomenon in LLM hidden representations, where distinct social identities become increasingly indistinguishable across layers. Motivated by this observation, we propose Parametric Social Identity Injection (PSII), a general framework that injects explicit, parametric representations of demographic attributes and value orientations directly into intermediate hidden states of LLMs. Unlike prompt-based persona conditioning, PSII enables fine-grained and controllable identity modulation at the representation level. Extensive experiments on the World Values Survey using multiple open-source LLMs show that PSII significantly improves distributional fidelity and diversity, reducing KL divergence to real-world survey data while enhancing overall diversity. This work provides new insights into representation-level control of LLM agents and advances scalable, diversity-aware public opinion simulation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6257ee03-2a15-4209-84fc-514fa26b1fc0Builds on11
- Whose Opinions Do Language Models Reflect?Shibani Santurkar, Esin Durmus, Faisal Ladhak, Cinoo Lee et al.ICML 2023 · 764 citations
- Increasing Diversity While Maintaining Accuracy: Text Data Generation with Large Language Models and Human InterventionsJohn Joon Young Chung, Ece Kamar, Saleema AmershiACL 2023 · 55 citations
- Language Model Fine-Tuning on Scaled Survey Data for Predicting Distributions of Public OpinionsJoseph Suh, Erfan Jahanparast, Suhong Moon, Minwoo Kang et al.ACL 2025 · 48 citations
- Cultivating Pluralism In Algorithmic Monoculture: The Community Alignment DatasetLily H Zhang, Smitha Milli, Karen Long Jusko, Jonathan Smith et al.ICLR 2026 · 41 citations
- Parametric Retrieval Augmented GenerationWeihang Su, Yichen Tang, Qingyao Ai, Junxi Yan et al.SIGIR 2025 · 25 citations
Related papers
- Synthia: Scalable Grounded Persona Generation from Social Media DataVahid Rahimzadeh, Erfan Moosavi Monazzah, Mohammad Taher Pilehvar, Yadollah YaghoobzadehACL 2026 · 1 citation
- Interview-Informed Generative Agents for Product Discovery: A Validation StudyZichao Wang, Alexa F. SiuCHI 2026 · 1 citation
- From Values to Opinions: Predicting Human Behaviors and Stances Using Value-Injected Large Language ModelsDongjun Kang, Joonsuk Park, Yohan Jo, JinYeong BakEMNLP 2023 · 4 citations
- Which Demographics do LLMs Default to During Annotation?Johannes Schäfer, Aidan Combs, Christopher Bagdon, Jiahui Li et al.ACL 2025 · 11 citations
- Survey Response Generation: Generating Closed-Ended Survey Responses In-Silico with Large Language ModelsGeorg Ahnert, Anna-Carolina Haensch, Barbara Plank, Markus StrohmaierACL 2026 · 4 citations
