A User-Centric Multi-Intent Benchmark for Evaluating Large Language Models
Jiayin Wang, Fengran Mo, Weizhi Ma, Peijie Sun, Min Zhang, Jian-Yun Nie
Abstract
Large language models (LLMs) are essential tools that users employ across various scenarios, so evaluating their performance and guiding users in selecting the suitable service is important. Although many benchmarks exist, they mainly focus on specific predefined model abilities, such as world knowledge, reasoning, etc. Based on these ability scores, it is hard for users to determine which LLM best suits their particular needs. To address these issues, we propose to evaluate LLMs from a user-centric perspective and design this benchmark to measure their efficacy in satisfying user needs under distinct intents. Firstly, we collect 1,846 real-world use cases from a user study with 712 participants from 23 countries. This first-hand data helps us understand actual user intents and needs in LLM interactions, forming the User Reported Scenarios (URS) dataset, which is categorized with six types of user intents. Secondly, based on this authentic dataset, we benchmark 10 LLM services with GPT-4-as-Judge. Thirdly, we show that benchmark scores align well with human preference in both real-world experience and pair-wise annotations, achieving Pearson correlations of 0.95 and 0.94, respectively. This alignment confirms that the URS dataset and our evaluation method establish an effective user-centric benchmark. The dataset and code are publicly available 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fabbdc37-dad8-4af9-9854-249dea11aab1Cited by top-tier papers6
- CORONA: A Coarse-to-Fine Framework for Graph-based Recommendation with Large Language ModelsJunze Chen, Xinjie Yang, Cheng Yang, Junfei Bao et al.SIGIR 2025 · 5 citations
- Simulating Dispute Mediation with LLM-Based Agents for Legal ResearchJunjie Chen, Haitao Li, Minghao Qin, Yujia Zhou et al.AAAI 2026 · 4 citations
- Boosting Data Utilization for Multilingual Dense RetrievalChao Huang, Fengran Mo, Yufeng Chen, Changhao Guan et al.EMNLP 2025 · 2 citations
- PICACO: Pluralistic In-Context Value Alignment via Total Correlation OptimizationHan Jiang, Dongyao Zhu, Xiaoyuan Yi, Ziang Xiao et al.ICML 2026 · 2 citations
- Expectation Alignment of Language Models for Real-World User ExpectationsMiaomiao Li, Yang Wang, Bin Liang, Shudong Liu et al.ICML 2026
Builds on9
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIsYujia Qin, Shihao Liang, Yining Ye, Kunlun Zhu et al.ICLR 2024 · 1,469 citations
- Chatbot Arena: An Open Platform for Evaluating LLMs by Human PreferenceWei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos et al.ICML 2024 · 1,212 citations
- A Non-Factoid Question-Answering TaxonomyValeria Bolotova, Vladislav Blinov, Falk Scholer, W. Bruce Croft et al.SIGIR 2022 · 35 citations
- MT-Bench-101: A Fine-Grained Benchmark for Evaluating Large Language Models in Multi-Turn DialoguesGe Bai, Jie Liu, Xingyuan Bu, Yancheng He et al.ACL 2024 · 35 citations
Related papers
- WebCoderBench: Benchmarking Web Application Generation with Comprehensive and Interpretable Evaluation MetricsChenxu Liu, Yingjie Fu, Wei Yang, Ying Zhang et al.ACL 2026 · 10 citations
- The Shifted and The Overlooked: A Task-oriented Investigation of User-GPT InteractionsSiru Ouyang, Shuohang Wang, Yang Liu, Ming Zhong et al.EMNLP 2023 · 10 citations
- MTU-Bench: A Multi-granularity Tool-Use Benchmark for Large Language ModelsPei Wang, Yanan Wu, Noah Wang, Jiaheng Liu et al.ICLR 2025
- LLMEval: A Preliminary Study on How to Evaluate Large Language ModelsYue Zhang, Ming Zhang, Haipeng Yuan, Shichun Liu et al.AAAI 2024 · 29 citations
- OmniBench: A Comprehensive Benchmark Integrating Real-World, Time-sensitive, and Multi-Hop Questions with a Multi-Dimensional Hybrid Evaluation FrameworkWenjie Wang, Yufeng Jiang, Ge Sun, Chenghang Dong et al.AAAI 2026
