VibeCheck: Discover and Quantify Qualitative Differences in Large Language Models
Lisa Dunlap, Krishna Mandal, Trevor Darrell, Jacob Steinhardt, Joseph E. Gonzalez
摘要
Large language models (LLMs) often exhibit subtle yet distinctive characteristics in their outputs that users intuitively recognize, but struggle to quantify. These "vibes" -such as tone, formatting, or writing style -influence user preferences, yet traditional evaluations focus primarily on the singular vibe of correctness. We introduce VibeCheck, a system for automatically comparing a pair of LLMs by discovering identifying traits of a model ("vibes") that are well-defined, differentiating, and user-aligned. VibeCheck iteratively discovers vibes from model outputs and then utilizes a panel of LLM judges to quantitatively measure the utility of each vibe. We validate that the vibes generated by VibeCheck align with those found in human discovery and run VibeCheck on pairwise preference data from real-world user conversations with Llama-3-70b vs GPT-4. VibeCheck reveals that Llama has a friendly, funny, and somewhat controversial vibe. These vibes predict model identity with 80% accuracy and human preference with 61% accuracy. Lastly, we run VibeCheck on a variety of models and tasks including summarization, math, and captioning to provide insight into differences in model behavior. VibeCheck discovers vibes like Command X prefers to add concrete intros and conclusions when summarizing in comparison to TNGL, Llama-405b often overexplains its thought process on math problems compared to GPT-4o, and GPT-4 prefers to focus on the mood and emotions of the scene when captioning compared to Gemini-1.5-Flash. Code and vibe visualizer found at https://bench-mark.org/ INTRO vibe check : A process by which a group obtains a subjective assessment of another person, place, or thing. -Urban Dictionary How a large language model writes a story, explains a concept, or edits an essay can be evaluated along many different dimensions such as creativity, formatting, and writing style. However, most evaluations focus on one dimension: "correctness". State-of-the-art in evaluation methods remain largely focused on measuring accuracy for question answering and analytical reasoning tasks (Hendrycks et al., 2021a; Wang et al., 2019b;a; Hendrycks et al., 2021c), and methods which aim to provide a more holistic view of LLMs (Zhang et al., 2024; Padlewski et al., 2024; Mehri & Eskenazi, 2020b) rely on predefined concepts like conciseness, clarity, and trustworthiness to measure a model's performance. These evaluation approaches fail to capture the open-ended nature of LLM applications and the critical dependence on subjective user preferences and context of the task. For instance, tone and creativity might be crucial in creative writing, whereas efficiency and readability are crucial in coding tasks. To best inform users of which model would be best for their needs, we require flexible evaluation methods that can both discover and measure the relevant axes to evaluate for a given task. When interacting with a set of LLMs for an extended period, a user can often tell which model generated a particular response by looking at certain traits of the outputs. We define these identifying traits of models as "vibes". For instance, users have found Llama-3 outputs tend to be more friendly compared to outputs from GPT-4 and Claude which tend to be more formal (see Figure 1 ); in other words, Llama-3 ranks high on the friendliness vibe, defined by the axis formal → friendly. Using these insights, we might select Llama for customer service tasks and Claude for coding tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Can AI writing be salvaged? Mitigating Idiosyncrasies and Improving Human-AI Alignment in the Writing Process through EditsTuhin Chakrabarty, Philippe Laban, Chien-Sheng WuCHI 2025 · 被引用 14 次
- AutoMetrics: Approximate Human Judgments with Automatically Generated EvaluatorsMichael J Ryan, Yanzhe Zhang, Amol Salunkhe, Yi Chu 等ICLR 2026 · 被引用 3 次
- Evalet: Evaluating Large Language Models through Functional FragmentationTae Soo Kim, Heechan Lee, Yoonjoo Lee, Joseph Seering 等CHI 2026 · 被引用 1 次
- Adaptively profiling models with task elicitationDavis Brown, Prithvi Balehannina, Helen Jin, Shreya Havaldar 等EMNLP 2025 · 被引用 1 次
- HumT DumT: Measuring and controlling human-like language in LLMsMyra Cheng, Sunny Yu, Dan JurafskyACL 2025
它引用的顶会 Paper8
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- Chatbot Arena: An Open Platform for Evaluating LLMs by Human PreferenceWei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos 等ICML 2024 · 被引用 1,212 次
- AlpacaFarm: A Simulation Framework for Methods that Learn from Human FeedbackYann Dubois, Chen Xuechen Li, Rohan Taori, Tianyi Zhang 等NeurIPS 2023 · 被引用 948 次
- Goal Driven Discovery of Distributional Differences via Language DescriptionsRuiqi Zhong, Peter Zhang, Steve Li, Jinwoo Ahn 等NeurIPS 2023 · 被引用 81 次
- Towards Holistic and Automatic Evaluation of Open-Domain Dialogue GenerationBo Pang, Erik Nijkamp, Wenjuan Han, Linqi Zhou 等ACL 2020 · 被引用 69 次
相关 Paper
- Results-Actionability Gap: Understanding How Practitioners Evaluate LLM Products in the WildWillem van der Maden, Malak Sadek, Ziang Xiao, Aske Mottelson 等CHI 2026 · 被引用 2 次
- SWE-IF: Aligning Code Evaluation with Human PreferenceMing Zhong, Xiang Zhou, Ting-Yun Chang, Qingze Wang 等ICML 2026 · 被引用 3 次
- Building Software by Rolling the Dice: A Qualitative Study of Vibe CodingYi-Hung Chou, Boyuan Jiang, Yi Wen Chen, Mingyue Weng 等FSE 2026 · 被引用 1 次
- AesBiasBench: Evaluating Bias and Alignment in Multimodal Language Models for Personalized Image Aesthetic AssessmentKun Li, Lai Man Po, Hongzheng Yang, Xuyuan Xu 等EMNLP 2025
- Persona-Augmented Benchmarking: Evaluating LLMs Across Diverse Writing StylesKimberly Le Truong, Riccardo Fogliato, Hoda Heidari, Steven WuEMNLP 2025 · 被引用 2 次
