Benchmarking and Learning Real-World Customer Service Dialogue
Tianhong Gao, Jundong Shen, Jiapeng Wang, Bei Shi, Ying Ju, Junfeng Yao, Huiyu Yu
Abstract
Existing benchmarks and training pipelines for industrial intelligent customer service (ICS) remain misaligned with real-world dialogue requirements, overemphasizing verifiable task success while under-measuring subjective service quality and realistic failure modes, leaving a gap between offline gains and deployable dialogue behavior. We close this gap with a benchmark-to-optimization loop: we first introduce OLABENCH, an ICS benchmark spanning retrieval-augmented generation, workflow-based systems, and agentic settings, which evaluates service capability, safety, and latency sensitivity; moreover, motivated by OLABENCH results showing state-of-the-art LLMs still fall short, we propose OLAMIND, which distills reusable reasoning patterns and service strategies from expert dialogues and applies staged exploration-exploitation reinforcement learning with instance-level rubric-aware guidance to improve model capability. OLA-MIND surpasses GPT-5.2 and Gemini 3 Pro on OLABENCH (83.64 vs. 70.58/70.84) and, in online A/B tests, delivers an average +23.67% issue resolution and -6.6% human transfer rate versus the baseline, bridging offline gains to deployment. Together, OLABENCH and OLA-MIND advance ICS systems toward more anthropomorphic, professional, and reliable deployment. The project page and evaluation are available at https://olamind-olabench. github.io .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on4
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Rubrics as Rewards: Reinforcement Learning Beyond Verifiable DomainsAnisha Gunjal, Anthony Wang, Elaine Lau, Vaskar Nath et al.ICLR 2026 · 340 citations
- An Efficient Task-Oriented Dialogue Policy: Evolutionary Reinforcement Learning Injected by Elite IndividualsYangyang Zhao, Ben Niu, Libo Qin, Shihan WangACL 2025 · 3 citations
- Bootstrapped Policy Learning for Task-oriented Dialogue through Goal ShapingYangyang Zhao, Ben Niu, Mehdi Dastani, Shihan WangEMNLP 2024
Related papers
- RefineBench: Evaluating Refinement Capability of Language Models via ChecklistsYoung-Jun Lee, Seungone Kim, Byung-Kwan Lee, Minkyeong Moon et al.ICLR 2026 · 13 citations
- LiveOIBench: Can Large Language Models Outperform Human Contestants in Informatics Olympiads?Kaijian Zou, Feiyang Xiong, Yunxiang Zhang, Xinliang Frederick Zhang et al.ICML 2026 · 3 citations
- τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World DomainsShunyu Yao, Noah Shinn, Pedram Razavi, Karthik R. NarasimhanICLR 2025
- CounselBench: A Large-Scale Expert Evaluation and Adversarial Benchmarking of Large Language Models in Mental Health Question AnsweringYahan Li, Jifan Yao, John Bosco S. Bunyi, Adam C. Frank et al.ICLR 2026 · 24 citations
- ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific DiscoveryZiru Chen, Shijie Chen, Yuting Ning, Qianheng Zhang et al.ICLR 2025 · 6 citations
