Style Outweighs Substance: Failure Modes of LLM Judges in Alignment Benchmarking
Benjamin Feuer, Micah Goldblum, Teresa Datta, Sanjana Nambiar, Raz Besaleli, Samuel Dooley, Max Cembalest, John P. Dickerson
摘要
The release of ChatGPT in November 2022 sparked an explosion of interest in post-training and an avalanche of new preference optimization (PO) methods. These methods claim superior alignment by virtue of better correspondence with human pairwise preferences, often measured by LLM-judges. In this work, we attempt to answer the following question -- do LLM-judge preferences translate to progress on other, more concrete metrics for alignment, and if not, why not? We define a concrete metric for alignment, and introduce SOS-Bench (Substance Outweighs Style Benchmark), which is to the best of our knowledge the largest standardized, reproducible LLM meta-benchmark to date. We find that (1) LLM-judge preferences do not correlate with concrete measures of safety, world knowledge, and instruction following; (2) LLM-judges have powerful implicit biases, prioritizing style over factuality and safety; and (3) the supervised fine-tuning (SFT) stage of post-training, and not the PO stage, has the greatest impact on alignment, with data scaling and prompt diversity as the driving factors. Our codebase and complete results can be found at https://github.com/penfever/sos-bench.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- Flattery, Fluff, and Fog: Diagnosing and Mitigating Idiosyncratic Biases in Preference ModelsAnirudh Bharadwaj, Chaitanya Malaviya, Nitish Joshi, Mark YatskarICLR 2026 · 被引用 15 次
- A Coin Flip for Safety: LLM Judges Fail to Reliably Measure Adversarial RobustnessLeo Schwinn, Moritz Ladenburger, Tim Beyer, Mehrnaz Mofakhami 等ICML 2026 · 被引用 15 次
- J4R: Learning to Judge with Equivalent Initial State Group Relative Policy OptimizationAustin Xu, Yilun Zhou, Xuan-Phi Nguyen, Caiming Xiong 等ACL 2026 · 被引用 8 次
- Sparta Alignment: Collectively Aligning Multiple Language Models through CombatYuru Jiang, Wenxuan Ding, Shangbin Feng, Greg Durrett 等NeurIPS 2025 · 被引用 8 次
- Deep Value Benchmark: Measuring Whether Models Generalize Deep Values or Shallow PreferencesJoshua Ashkinaze, Hua Shen, Sai Avula, Eric Gilbert 等NeurIPS 2025 · 被引用 7 次
它引用的顶会 Paper19
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- WizardLM: Empowering Large Pre-Trained Language Models to Follow Complex InstructionsCan Xu, Qingfeng Sun, Kai Zheng, Xiubo Geng 等ICLR 2024 · 被引用 1,206 次
- SimPO: Simple Preference Optimization with a Reference-Free RewardYu Meng, Mengzhou Xia, Danqi ChenNeurIPS 2024 · 被引用 1,203 次
- LLM Evaluators Recognize and Favor Their Own GenerationsArjun Panickssery, Samuel R. Bowman, Shi FengNeurIPS 2024 · 被引用 865 次
相关 Paper
- SORRY-Bench: Systematically Evaluating Large Language Model Safety RefusalTinghao Xie, Xiangyu Qi, Yi Zeng, Yangsibo Huang 等ICLR 2025
- Select Before Use: On the Importance of Reference Model Selection in Preference AlignmentMuyang Li, Runze Wu, Xiangyu Zhao, Bo Han 等ACL 2026
- Dissecting Human and LLM PreferencesJunlong Li, Fan Zhou, Shichao Sun, Yikai Zhang 等ACL 2024 · 被引用 1 次
- Improve LLM-as-a-Judge Ability as a General AbilityJiachen Yu, Shaoning Sun, Xiaohui Hu, Jiaxu Yan 等EMNLP 2025 · 被引用 1 次
- Preference-Oriented Supervised Fine-Tuning: Favoring Target Model over Aligned Large Language ModelsYuchen Fan, Yuzhong Hong, Qiushi Wang, Junwei Bao 等AAAI 2025 · 被引用 7 次
