STEER-BENCH: A Benchmark for Evaluating the Steerability of Large Language Models
Kai Chen, Zihao He, Taiwei Shi, Kristina Lerman
摘要
Steerability, or the ability of large language models (LLMs) to adapt outputs to align with diverse community-specific norms, perspectives, and communication styles, is critical for real-world applications but remains underevaluated. We introduce STEER-BENCH, a benchmark for assessing population-specific steering using contrasting Reddit communities. Covering 30 contrasting subreddit pairs across 19 domains, STEER-BENCH includes over 10,000 instruction-response pairs and validated 5,500 multiple-choice questions with corresponding silver labels to test alignment with diverse community norms. It systematically assesses how effectively LLMs understand community-specific instructions, their resilience to adversarial steering attempts, and their ability to accurately represent diverse cultural and ideological perspectives. Our evaluation of 13 popular LLMs using STEER-BENCH reveals that while human experts achieve an accuracy of 81% with silver labels, the bestperforming models reach only around 65% accuracy depending on the domain and configuration. Some models lag behind humanlevel alignment by over 15 percentage points, highlighting significant gaps in communitysensitive steerability. 1 Steered LLM Vanilla LLM Steered LLM Instruction: Why might some users decide to switch to Linux? Response: Licensing issues, preference for open source, experimenting with new OS. Instruction: Why might some users decide to switch to Linux? Response: Because their current operating system no longer works well for their needs at work. What motivates some Windows users to try Linux? A. Curiosity about open-source software. B. Frustration with Windows updates. C. Influence from tech industry trends. D. Desire to explore new GUI options.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper13
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- WizardLM: Empowering Large Pre-Trained Language Models to Follow Complex InstructionsCan Xu, Qingfeng Sun, Kai Zheng, Xiubo Geng 等ICLR 2024 · 被引用 1,206 次
- Aligning AI With Shared Human ValuesDan Hendrycks, Collin Burns, Steven Basart, Andrew Critch 等ICLR 2021 · 被引用 878 次
- Whose Opinions Do Language Models Reflect?Shibani Santurkar, Esin Durmus, Faisal Ladhak, Cinoo Lee 等ICML 2023 · 被引用 764 次
- Self-Instruct: Aligning Language Models with Self-Generated InstructionsYizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu 等ACL 2023 · 被引用 540 次
相关 Paper
- SimBench: Benchmarking the Ability of Large Language Models to Simulate Human BehaviorsTiancheng Hu, Joachim Baumann, Lorenzo Lupo, Nigel Collier 等ICLR 2026 · 被引用 61 次
- How Controllable Are Large Language Models? A Unified Evaluation across Behavioral GranularitiesZiwen Xu, Kewei Xu, Haoming Xu, Haiwen Hong 等ACL 2026
- SteeringSafety: Benchmarking Representation Steering in LLMs Across Safety PerspectivesVincent Siu, Nicholas Crispino, David Park, Nathan Henry 等ICML 2026
- Model Editing as a Double-Edged Sword: Steering Agent Behavior Toward Beneficence or HarmBaixiang Huang, Zhen Tan, Haoran Wang, Zijie Liu 等AAAI 2026
- What's Producible May Not Be Reachable: Measuring the Steerability of Generative ModelsKeyon Vafa, Sarah Bentley, Jon M. Kleinberg, Sendhil MullainathanNeurIPS 2025 · 被引用 5 次
