MT-Eval: A Multi-Turn Capabilities Evaluation Benchmark for Large Language Models
Wai-Chung Kwan, Xingshan Zeng, Yuxin Jiang, Yufei Wang, Liangyou Li, Lifeng Shang, Xin Jiang, Qun Liu, Kam-Fai Wong
Abstract
Large language models (LLMs) are increasingly relied upon for complex multi-turn conversations across diverse real-world applications. However, existing benchmarks predominantly focus on single-turn evaluations, overlooking the models' capabilities in multi-turn interactions. To address this gap, we introduce MT-Eval, a comprehensive benchmark designed to evaluate multi-turn conversational abilities. By analyzing human-LLM conversations, we categorize interaction patterns into four types: recollection, expansion, refinement, and follow-up. We construct multi-turn queries for each category either by augmenting existing datasets or by creating new examples with GPT-4 to avoid data leakage. To study the factors impacting multi-turn abilities, we create singleturn versions of the 1170 multi-turn queries and compare performance. Our evaluation of 11 well-known LLMs shows that while closedsource models generally surpass open-source ones, certain open-source models exceed GPT-3.5-Turbo in specific tasks. We observe significant performance degradation in multi-turn settings compared to single-turn settings in most models, which is not correlated with the models' fundamental capabilities. Moreover, we identify the distance to relevant content and susceptibility to error propagation as the key factors influencing multi-turn performance. MT-Eval is released publicly to encourage future research towards more robust conversational models 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b8badbcb-3d6a-46a6-a80e-7466b4c7ed9dCited by top-tier papers30
- LLMs Get Lost In Multi-Turn ConversationPhilippe Laban, Hiroaki Hayashi, Yingbo Zhou, Jennifer NevilleICLR 2026 · 491 citations
- Know You First and Be You Better: Modeling Human-Like User Simulators via Implicit ProfilesKuang Wang, Xianfei Li, Shenghao Yang, Li Zhou et al.ACL 2025 · 24 citations
- RefineBench: Evaluating Refinement Capability of Language Models via ChecklistsYoung-Jun Lee, Seungone Kim, Byung-Kwan Lee, Minkyeong Moon et al.ICLR 2026 · 13 citations
- MDEval: Evaluating and Enhancing Markdown Awareness in Large Language ModelsZhongpu Chen, Yinfeng Liu, Long Shi, Zhi-Jie Wang et al.WWW 2025 · 12 citations
- FollowBench: A Multi-level Fine-grained Constraints Following Benchmark for Large Language ModelsYuxin Jiang, Yufei Wang, Xingshan Zeng, Wanjun Zhong et al.ACL 2024 · 10 citations
Builds on3
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- G-Eval: NLG Evaluation using Gpt-4 with Better Human AlignmentYang Liu, Dan Iter, Yichong Xu, Shuohang Wang et al.EMNLP 2023 · 549 citations
- GLM: General Language Model Pretraining with Autoregressive Blank InfillingZhengxiao Du, Yujie Qian, Xiao Liu, Ming Ding et al.ACL 2022
Related papers
- MT-Bench-101: A Fine-Grained Benchmark for Evaluating Large Language Models in Multi-Turn DialoguesGe Bai, Jie Liu, Xingyuan Bu, Yancheng He et al.ACL 2024 · 35 citations
- MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language FeedbackXingyao Wang, Zihan Wang, Jiateng Liu, Yangyi Chen et al.ICLR 2024 · 308 citations
- MTR-Bench: A Comprehensive Benchmark for Multi-Turn Reasoning EvaluationXiaoyuan Li, Keqin Bao, Yubo Ma, Moxin Li et al.ACL 2026 · 12 citations
- ConvCodeWorld: Benchmarking Conversational Code Generation in Reproducible Feedback EnvironmentsHojae Han, Seung-won Hwang, Rajhans Samdani, Yuxiong HeICLR 2025
- FairMT-Bench: Benchmarking Fairness for Multi-turn Dialogue in Conversational LLMsZhiting Fan, Ruizhe Chen, Tianxiang Hu, Zuozhu LiuICLR 2025
