LiveCultureBench: a Multi-Agent, Multi-Cultural Benchmark for Large Language Models in Dynamic Social Simulations
Viet Thanh Pham, Lizhen Qu, Thuy-Trang Vu, Gholamreza Haffari, Dinh Q. Phung
Abstract
Large language models (LLMs) are increasingly deployed as autonomous agents, yet evaluations focus primarily on task success rather than cultural appropriateness or evaluator reliability. We introduce LiveCultureBench, a multi-cultural, dynamic benchmark that embeds LLMs as agents in a simulated town and evaluates them on both task completion and adherence to socio-cultural norms. The simulation models a small city as a location graph with synthetic residents having diverse demographic and cultural profiles. Each episode assigns one resident a daily goal while others provide social context. An LLM-based verifier generates structured judgments on norm violations and task progress, which we aggregate into metrics capturing task-norm trade-offs and verifier uncertainty. Using LiveCultureBench across models and cultural profiles, we study (i) cross-cultural robustness of LLM agents, (ii) how they balance effectiveness against norm sensitivity, and (iii) when LLM-as-a-judge evaluation is reliable for automated benchmarking versus when human oversight is needed.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 780e0c08-fb06-4314-8ac4-9d6ae32f1238Builds on4
- CAMEL: Communicative Agents for "Mind" Exploration of Large Language Model SocietyGuohao Li, Hasan Hammoud, Hani Itani, Dmitrii Khizbullin et al.NeurIPS 2023 · 1,975 citations
- Generative Agents: Interactive Simulacra of Human BehaviorJoon Sung Park, Joseph C. O'Brien, Carrie Jun Cai, Meredith Ringel Morris et al.UIST 2023 · 1,882 citations
- Conformal Language ModelingVictor Quach, Adam Fisch, Tal Schuster, Adam Yala et al.ICLR 2024 · 132 citations
- SurveyPilot: an Agentic Framework for Automated Human Opinion Collection from Social MediaViet Thanh Pham, Lizhen Qu, Zhuang Li, Suraj Sharma et al.ACL 2025 · 4 citations
Related papers
- SocialCC: Interactive Evaluation for Cultural Competence in Language AgentsJincenzi Wu, Jianxun Lian, Dingdong Wang, Helen M. MengACL 2025 · 8 citations
- GuideBench: Benchmarking Domain-Oriented Guideline Following for LLM AgentsLingxiao Diao, Xinyue Xu, Wanxuan Sun, Cheng Yang et al.ACL 2025
- CAR-bench: Evaluating the Consistency and Limit-Awareness of LLM Agents under Real-World UncertaintyJohannes Kirmayr, Lukas Stappen, Elisabeth AndréACL 2026 · 5 citations
- Copyright-Bench: Agentic Evaluation of Copyright Law ComplianceZheng Hui, Doni Bloomfield, Noam KoltICML 2026 · 1 citation
- LaoBench: A Large-Scale Multidimensional Lao Benchmark for Large Language ModelsJian Gao, Richeng Xuan, Zhaolu Kang, Dingshi Liao et al.ACL 2026 · 1 citation
