ChronoPlay: A Framework for Modeling Dual Dynamics and Authenticity in Game RAG Benchmarks
Liyang He, Yuren Zhang, Ziwei Zhu, Zhenghui Li, Shiwei Tong
摘要
Retrieval Augmented Generation (RAG) systems are increasingly vital in dynamic domains like online gaming, yet the lack of a dedicated benchmark has impeded standardized evaluation in this area. The core difficulty lies in Dual Dynamics: the constant interplay between game content updates and the shifting focus of the player community. Furthermore, the necessity of automating such a benchmark introduces a critical requirement for player-centric authenticity to ensure generated questions are realistic. To address this integrated challenge, we introduce ChronoPlay, a novel framework for the automated and continuous generation of game RAG benchmarks. ChronoPlay utilizes a dual-dynamic update mechanism to track both forms of change, and a dual-source synthesis engine that draws from official sources and player community to ensure both factual correctness and authentic query patterns. We instantiate our framework on three distinct games to create the first dynamic RAG benchmark for the gaming domain, offering new insights into model performance under these complex and realistic conditions. Our code is available at: https://github.com/hly1998/ChronoPlay.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper9
- When Not to Trust Language Models: Investigating Effectiveness of Parametric and Non-Parametric MemoriesAlex Mallen, Akari Asai, Victor Zhong, Rajarshi Das 等ACL 2023 · 被引用 233 次
- Precise Zero-Shot Dense Retrieval without Relevance LabelsLuyu Gao, Xueguang Ma, Jimmy Lin, Jamie CallanACL 2023 · 被引用 211 次
- HoH: A Dynamic Benchmark for Evaluating the Impact of Outdated Information on Retrieval-Augmented GenerationJie Ouyang, Tingyue Pan, Mingyue Cheng, Ruiran Yan 等ACL 2025 · 被引用 14 次
- Towards Interest Drift-driven User Representation Learning in Sequential RecommendationXiaolin Lin, Weike Pan, Zhong MingSIGIR 2025 · 被引用 8 次
- Self-ICL: Zero-Shot In-Context Learning with Self-Generated DemonstrationsWei-Lin Chen, Cheng-Kuang Wu, Yun-Nung Chen, Hsin-Hsi ChenEMNLP 2023 · 被引用 8 次
相关 Paper
- REAL-MM-RAG: A Real-World Multi-Modal Retrieval BenchmarkNavve Wasserman, Roi Pony, Oshri Naparstek, Adi Raz Goldfarb 等ACL 2025 · 被引用 33 次
- MTR-Suite: A Framework for Evaluating and Synthesizing Conversational Retrieval BenchmarksJunhao Ruan, Abudukeyumu Abudula, Bei Li, Yongjing Yin 等ACL 2026
- DynamicER: Resolving Emerging Mentions to Dynamic Entities for RAGJinyoung Kim, Dayoon Ko, Gunhee KimEMNLP 2024 · 被引用 3 次
- When to use Graphs in RAG: A Comprehensive Analysis for Graph Retrieval-Augmented GenerationZhishang Xiang, Chuanjie Wu, Qinggang Zhang, Shengyuan Chen 等ICLR 2026 · 被引用 56 次
- Re³: Relevance & Recency Retrieval for Mitigating Temporal HallucinationJiawei Cao, Jie Ouyang, Mingyue Cheng, Zhaomeng Zhou 等ACL 2026
