How Controllable Are Large Language Models? A Unified Evaluation across Behavioral Granularities
Ziwen Xu, Kewei Xu, Haoming Xu, Haiwen Hong, Longtao Huang, Hui Xue, Ningyu Zhang, Yongliang Shen, Guozhou Zheng, Huajun Chen, Shumin Deng
摘要
Large Language Models (LLMs) are increasingly deployed in socially sensitive domains, yet their unpredictable behaviors, ranging from misaligned intent to inconsistent personality, pose significant risks. We introduce SteerEval, a hierarchical benchmark for evaluating LLM controllability across three domains: language features, sentiment, and personality. Each domain is structured into three specification levels: L1 (what to express), L2 (how to express), and L3 (how to instantiate), connecting highlevel behavioral intent to concrete textual output. Using SteerEval, we systematically evaluate contemporary steering methods, revealing that control often degrades at finer-grained levels. Our benchmark offers a principled and interpretable framework for safe and controllable LLM behavior, serving as a foundation for future research 1 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper19
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Inference-Time Intervention: Eliciting Truthful Answers from a Language ModelKenneth Li, Oam Patel, Fernanda B. Viégas, Hanspeter Pfister 等NeurIPS 2023 · 被引用 1,549 次
- Sparse Autoencoders Find Highly Interpretable Features in Language ModelsRobert Huben, Hoagy Cunningham, Logan Riggs Smith, Aidan Ewart 等ICLR 2024 · 被引用 1,072 次
- Rethinking the Role of Demonstrations: What Makes In-Context Learning Work?Sewon Min, Xinxi Lyu, Ari Holtzman, Mikel Artetxe 等EMNLP 2022 · 被引用 634 次
- ReFT: Representation Finetuning for Language ModelsZhengxuan Wu, Aryaman Arora, Zheng Wang, Atticus Geiger 等NeurIPS 2024 · 被引用 233 次
相关 Paper
- Psychological Steering in LLMs: An Evaluation of Effectiveness and TrustworthinessAmin Banayeeanzade, Ala N. Tak, Fatemeh Bahrani, Anahita Bolourani 等ACL 2026 · 被引用 4 次
- LexInstructEval: Lexical Instruction Following Evaluation for Large Language ModelsHuimin Ren, Yan Liang, Baiqiao Su, Chaobo Sun 等AAAI 2026
- STEER-BENCH: A Benchmark for Evaluating the Steerability of Large Language ModelsKai Chen, Zihao He, Taiwei Shi, Kristina LermanEMNLP 2025 · 被引用 1 次
- S-Eval: Towards Automated and Comprehensive Safety Evaluation for Large Language ModelsXiaohan Yuan, Jinfeng Li, Dongxia Wang, Yuefeng Chen 等ISSTA 2025 · 被引用 4 次
- SteeringSafety: Benchmarking Representation Steering in LLMs Across Safety PerspectivesVincent Siu, Nicholas Crispino, David Park, Nathan Henry 等ICML 2026
