Dialogue Benchmark Generation from Knowledge Graphs with Cost-Effective Retrieval-Augmented LLMs
Reham Omar, Omij Mangukiya, Essam Mansour
Abstract
Dialogue benchmarks are crucial in training and evaluating chatbots engaging in domain-specific conversations. Knowledge graphs (KGs) represent semantically rich and well-organized data spanning various domains, such as DBLP, DBpedia, and YAGO. Traditionally, dialogue benchmarks have been manually created from documents, neglecting the potential of KGs in automating this process. Some question-answering benchmarks are automatically generated using extensive preprocessing from KGs, but they do not support dialogue generation. This paper introduces Chatty-Gen, a novel multi-stage retrieval-augmented generation platform for automatically generating high-quality dialogue benchmarks tailored to a specific domain using a KG. Chatty-Gen decomposes the generation process into manageable stages and uses assertion rules for automatic validation between stages. Our approach enables control over intermediate results to prevent time-consuming restarts due to hallucinations. It also reduces reliance on costly and more powerful commercial LLMs. Chatty-Gen eliminates upfront processing of the entire KG using efficient query-based retrieval to find representative subgraphs based on the dialogue context. Our experiments with several real and large KGs demonstrate that C hatty -G en significantly outperforms state-of-the-art systems and ensures consistent model and system performance across multiple LLMs of diverse capabilities, such as GPT-4o, Gemini 1.5, Llama 3, and Mistral.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fa9cc685-1fcf-4d7d-8aa9-221a157c0ffbCited by top-tier papers6
- Chatty-KG: A Multi-Agent AI System for On-Demand Conversational Question Answering over Knowledge GraphsReham Omar, Abdelghny Orogat, Ibrahim Abdelaziz, Omij Mangukiya et al.SIGMOD 2026 · 4 citations
- SPARTA: Scalable and Principled Benchmark of Tree-Structured Multi-hop QA over Text and TablesSungho Park, Jueun Kim, Wook-Shin HanICLR 2026 · 2 citations
- Accurate Table Question Answering with Accessible LLMsYangfan Jiang, Fei Wei, Ergute Bao, Yaliang Li et al.ICDE 2026 · 1 citation
- CRAFT: Corpus Relatedness Analysis Using Fourier TransformsKaiwen Chen, Nick KoudasVLDB 2026
- AGRAG: Advanced Graph-Based Retrieval-Augmented Generation for LLMsYubo Wang, Haoyang Li, Fei Teng, Lei ChenICDE 2026
Builds on18
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo et al.NeurIPS 2022 · 8,168 citations
- Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code GenerationJiawei Liu, Chunqiu Steven Xia, Yuyao Wang, Lingming ZhangNeurIPS 2023 · 2,317 citations
Related papers
- Can Knowledge Graphs Make Large Language Models More Trustworthy? An Empirical Study Over Open-ended Question AnsweringYuan Sui, Yufei He, Zifeng Ding, Bryan HooiACL 2025 · 29 citations
- Scaling Knowledge Graph Construction through Synthetic Data Generation and DistillationPrafulla Kumar Choubey, Xin Su, Man Luo, XIANGYU PENG et al.ICLR 2026 · 5 citations
- MultiHal: Multilingual Dataset for Knowledge-Graph Grounded Evaluation of LLM HallucinationsErnests Lavrinovics, Russa Biswas, Katja Hose, Johannes BjervaICML 2026 · 5 citations
- MP2D: An Automated Topic Shift Dialogue Generation Framework Leveraging Knowledge GraphsYerin Hwang, Yongil Kim, Yunah Jang, Jeesoo Bang et al.EMNLP 2024 · 2 citations
- Maestro: Automatic Generation of Comprehensive Benchmarks for Question Answering Over Knowledge GraphsAbdelghny Orogat, Ahmed El-RobySIGMOD 2023 · 10 citations
