FAIRGAMER: Evaluating Social Biases in LLM-Based Video Game NPCs
Bingkang Shi, Jen-tse Huang, Luo Long, Tianyu Zong, Hongzhu Yi, Yuanxiang Wang, Songlin Hu, Xiaodan Zhang, Zhongjiang Yao
摘要
Large Language Models (LLMs) have increasingly enhanced or replaced traditional Non-Player Characters (NPCs) in video games. However, these LLM-based NPCs inherit underlying social biases (e.g., race or class), posing fairness risks during in-game interactions. To address the limited exploration of this issue, we introduce FairGamer, the first benchmark to evaluate social biases across three interaction patterns: transaction, cooperation, and competition. FairGamer assesses four bias types, including class, race, age, and nationality, across 12 distinct evaluation tasks using a novel metric, FairMCV. Our evaluation of seven frontier LLMs reveals that: (1) models exhibit biased decision-making, with Grok-4-Fast demonstrating the highest bias (average FairMCV = 76.9%); and (2) larger LLMs display more severe social biases, suggesting that increased model capacity inadvertently amplifies these biases. We release FairGamer at https://github.com/Anonymous999-xxx/FairGamer to facilitate future research on NPC fairness.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper9
- Is ChatGPT a General-Purpose Natural Language Processing Task Solver?Chengwei Qin, Aston Zhang, Zhuosheng Zhang, Jiaao Chen 等EMNLP 2023 · 被引用 449 次
- Can Large Language Models Serve as Rational Players in Game Theory? A Systematic AnalysisCaoyun Fan, Jindou Chen, Yaohui Jin, Hao HeAAAI 2024 · 被引用 123 次
- SmartPlay : A Benchmark for LLMs as Intelligent AgentsYue Wu, Xuan Tang, Tom M. Mitchell, Yuanzhi LiICLR 2024 · 被引用 121 次
- WinoQueer: A Community-in-the-Loop Benchmark for Anti-LGBTQ+ Bias in Large Language ModelsVirginia K. Felkner, Ho-Chun Herbert Chang, Eugene Jang, Jonathan MayACL 2023 · 被引用 46 次
- Systematic Biases in LLM Simulations of DebatesAmir Taubenfeld, Yaniv Dover, Roi Reichart, Ariel GoldsteinEMNLP 2024 · 被引用 35 次
相关 Paper
- FairMT-Bench: Benchmarking Fairness for Multi-turn Dialogue in Conversational LLMsZhiting Fan, Ruizhe Chen, Tianxiang Hu, Zuozhu LiuICLR 2025
- Fairness Testing of Large Language Models in Role-PlayingXinyue Li, Zhenpeng Chen, Jie M. Zhang, Ying Xiao 等FSE 2026 · 被引用 5 次
- The Impossibility of Fair LLMsJacy Reese Anthis, Kristian Lum, Michael D. Ekstrand, Avi Feller 等ACL 2025
- Orak: A Foundational Benchmark for Training and Evaluating LLM Agents on Diverse Video GamesDongmin Park, Minkyu Kim, Beongjun Choi, Junhyuck Kim 等ICLR 2026 · 被引用 30 次
- Competing Large Language Models in Multi-Agent Gaming EnvironmentsJen-tse Huang, Eric John Li, Man Ho Lam, Tian Liang 等ICLR 2025
