VibeCheck: Discover and Quantify Qualitative Differences in Large Language Models
Lisa Dunlap, Krishna Mandal, Trevor Darrell, Jacob Steinhardt, Joseph E. Gonzalez
Abstract
Large language models (LLMs) often exhibit subtle yet distinctive characteristics in their outputs that users intuitively recognize, but struggle to quantify. These "vibes" -such as tone, formatting, or writing style -influence user preferences, yet traditional evaluations focus primarily on the singular vibe of correctness. We introduce VibeCheck, a system for automatically comparing a pair of LLMs by discovering identifying traits of a model ("vibes") that are well-defined, differentiating, and user-aligned. VibeCheck iteratively discovers vibes from model outputs and then utilizes a panel of LLM judges to quantitatively measure the utility of each vibe. We validate that the vibes generated by VibeCheck align with those found in human discovery and run VibeCheck on pairwise preference data from real-world user conversations with Llama-3-70b vs GPT-4. VibeCheck reveals that Llama has a friendly, funny, and somewhat controversial vibe. These vibes predict model identity with 80% accuracy and human preference with 61% accuracy. Lastly, we run VibeCheck on a variety of models and tasks including summarization, math, and captioning to provide insight into differences in model behavior. VibeCheck discovers vibes like Command X prefers to add concrete intros and conclusions when summarizing in comparison to TNGL, Llama-405b often overexplains its thought process on math problems compared to GPT-4o, and GPT-4 prefers to focus on the mood and emotions of the scene when captioning compared to Gemini-1.5-Flash. Code and vibe visualizer found at https://bench-mark.org/ INTRO vibe check : A process by which a group obtains a subjective assessment of another person, place, or thing. -Urban Dictionary How a large language model writes a story, explains a concept, or edits an essay can be evaluated along many different dimensions such as creativity, formatting, and writing style. However, most evaluations focus on one dimension: "correctness". State-of-the-art in evaluation methods remain largely focused on measuring accuracy for question answering and analytical reasoning tasks (Hendrycks et al., 2021a; Wang et al., 2019b;a; Hendrycks et al., 2021c), and methods which aim to provide a more holistic view of LLMs (Zhang et al., 2024; Padlewski et al., 2024; Mehri & Eskenazi, 2020b) rely on predefined concepts like conciseness, clarity, and trustworthiness to measure a model's performance. These evaluation approaches fail to capture the open-ended nature of LLM applications and the critical dependence on subjective user preferences and context of the task. For instance, tone and creativity might be crucial in creative writing, whereas efficiency and readability are crucial in coding tasks. To best inform users of which model would be best for their needs, we require flexible evaluation methods that can both discover and measure the relevant axes to evaluate for a given task. When interacting with a set of LLMs for an extended period, a user can often tell which model generated a particular response by looking at certain traits of the outputs. We define these identifying traits of models as "vibes". For instance, users have found Llama-3 outputs tend to be more friendly compared to outputs from GPT-4 and Claude which tend to be more formal (see Figure 1 ); in other words, Llama-3 ranks high on the friendliness vibe, defined by the axis formal → friendly. Using these insights, we might select Llama for customer service tasks and Claude for coding tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7f10832a-9d8d-44bc-9c99-60dbf8a2b31aCited by top-tier papers7
- Can AI writing be salvaged? Mitigating Idiosyncrasies and Improving Human-AI Alignment in the Writing Process through EditsTuhin Chakrabarty, Philippe Laban, Chien-Sheng WuCHI 2025 · 14 citations
- AutoMetrics: Approximate Human Judgments with Automatically Generated EvaluatorsMichael J Ryan, Yanzhe Zhang, Amol Salunkhe, Yi Chu et al.ICLR 2026 · 3 citations
- Evalet: Evaluating Large Language Models through Functional FragmentationTae Soo Kim, Heechan Lee, Yoonjoo Lee, Joseph Seering et al.CHI 2026 · 1 citation
- Adaptively profiling models with task elicitationDavis Brown, Prithvi Balehannina, Helen Jin, Shreya Havaldar et al.EMNLP 2025 · 1 citation
- HumT DumT: Measuring and controlling human-like language in LLMsMyra Cheng, Sunny Yu, Dan JurafskyACL 2025
Builds on8
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Chatbot Arena: An Open Platform for Evaluating LLMs by Human PreferenceWei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos et al.ICML 2024 · 1,212 citations
- AlpacaFarm: A Simulation Framework for Methods that Learn from Human FeedbackYann Dubois, Chen Xuechen Li, Rohan Taori, Tianyi Zhang et al.NeurIPS 2023 · 948 citations
- Goal Driven Discovery of Distributional Differences via Language DescriptionsRuiqi Zhong, Peter Zhang, Steve Li, Jinwoo Ahn et al.NeurIPS 2023 · 81 citations
- Towards Holistic and Automatic Evaluation of Open-Domain Dialogue GenerationBo Pang, Erik Nijkamp, Wenjuan Han, Linqi Zhou et al.ACL 2020 · 69 citations
Related papers
- Results-Actionability Gap: Understanding How Practitioners Evaluate LLM Products in the WildWillem van der Maden, Malak Sadek, Ziang Xiao, Aske Mottelson et al.CHI 2026 · 2 citations
- SWE-IF: Aligning Code Evaluation with Human PreferenceMing Zhong, Xiang Zhou, Ting-Yun Chang, Qingze Wang et al.ICML 2026 · 3 citations
- Building Software by Rolling the Dice: A Qualitative Study of Vibe CodingYi-Hung Chou, Boyuan Jiang, Yi Wen Chen, Mingyue Weng et al.FSE 2026 · 1 citation
- AesBiasBench: Evaluating Bias and Alignment in Multimodal Language Models for Personalized Image Aesthetic AssessmentKun Li, Lai Man Po, Hongzheng Yang, Xuyuan Xu et al.EMNLP 2025
- Persona-Augmented Benchmarking: Evaluating LLMs Across Diverse Writing StylesKimberly Le Truong, Riccardo Fogliato, Hoda Heidari, Steven WuEMNLP 2025 · 2 citations
