LLM Comparator: Interactive Analysis of Side-by-Side Evaluation of Large Language Models
Minsuk Kahng, Ian Tenney, Mahima Pushkarna, Michael Xieyang Liu, James Wexler, Emily Reif, Krystal Kallarackal, Minsuk Chang, Michael Terry, Lucas Dixon
Abstract
Evaluating large language models (LLMs) presents unique challenges. While automatic side-by-side evaluation, also known as LLM-as-a-judge, has become a promising solution, model developers and researchers face difficulties with scalability and interpretability when analyzing these evaluation outcomes. To address these challenges, we introduce LLM Comparator, a new visual analytics tool designed for side-by-side evaluations of LLMs. This tool provides analytical workflows that help users understand when and why one LLM outperforms or underperforms another, and how their responses differ. Through close collaboration with practitioners developing LLMs at Google, we have iteratively designed, developed, and refined the tool. Qualitative feedback from these users highlights that the tool facilitates in-depth analysis of individual examples while enabling users to visually overview and flexibly slice data. This empowers users to identify undesirable patterns, formulate hypotheses about model behavior, and gain insights for model improvement. LLM Comparator has been integrated into Google's LLM evaluation platforms and open-sourced.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8f4fab04-0287-4dd6-a1aa-db695e081fbaCited by top-tier papers8
- Interactive Debugging and Steering of Multi-Agent AI SystemsWill Epperson, Gagan Bansal, Victor C. Dibia, Adam Fourney et al.CHI 2025 · 33 citations
- RAGTrace: Understanding and Refining Retrieval-Generation Dynamics in Retrieval-Augmented GenerationSizhe Cheng, Jiaping Li, Huanchen Wang, Yuxin MaUIST 2025 · 6 citations
- ConceptViz: A Visual Analytics Approach for Exploring Concepts in Large Language ModelsHaoxuan Li, Zhen Wen, Qiqi Jiang, Chenxiao Li et al.IEEE VIS 2025 · 3 citations
- DiLLS: Interactive Diagnosis of LLM-based Multi-agent Systems via Layered Summary of Agent BehaviorsRui Sheng, Yukun Yang, Chuhan Shi, Yanna Lin et al.CHI 2026 · 2 citations
- Evalet: Evaluating Large Language Models through Functional FragmentationTae Soo Kim, Heechan Lee, Yoonjoo Lee, Joseph Seering et al.CHI 2026 · 1 citation
Builds on15
- PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning OptimizationYidong Wang, Zhuohao Yu, Wenjin Yao, Zhengran Zeng et al.ICLR 2024 · 368 citations
- Interactive and Visual Prompt Engineering for Ad-hoc Task Adaptation with Large Language ModelsHendrik Strobelt, Albert Webson, Victor Sanh, Benjamin Hoover et al.IEEE VIS 2022 · 191 citations
- Generative Judge for Evaluating AlignmentJunlong Li, Shichao Sun, Weizhe Yuan, Run-Ze Fan et al.ICLR 2024 · 173 citations
- ChainForge: A Visual Toolkit for Prompt Engineering and LLM Hypothesis TestingIan Arawjo, Chelse Swoopes, Priyan Vaithilingam, Martin Wattenberg et al.CHI 2024 · 141 citations
- DECE: Decision Explorer with Counterfactual Explanations for Machine Learning ModelsFurui Cheng, Yao Ming, Huamin QuIEEE VIS 2020 · 118 citations
Related papers
- MLLM-as-a-Judge: Assessing Multimodal LLM-as-a-Judge with Vision-Language BenchmarkDongping Chen, Ruoxi Chen, Shilin Zhang, Yaochen Wang et al.ICML 2024 · 345 citations
- Lexara: A User-Centered Toolkit for Evaluating Large Language Models for Conversational Visual AnalyticsSrishti Palani, Vidya SetlurCHI 2026 · 1 citation
- WebDevJudge: Evaluating (M)LLMs as Critiques for Web Development QualityChunyang Li, Yilun Zheng, Xinting Huang, Tianqing Fang et al.ICLR 2026 · 14 citations
- Quantifying Biases in LLM-as-a-Judge EvaluationsMagda Dubois, Harry Coppock, Mario Giulianelli, Ole Jorgensen et al.ICML 2026
- Praetor: A Fine-Grained Generative LLM Evaluator with Instance-Level Customizable Evaluation CriteriaYongqi Leng, Renren Jin, Yue Chen, Zhuowen Han et al.ACL 2025
