Chinese SimpleQA: A Chinese Factuality Evaluation for Large Language Models
Yancheng He, Shilong Li, Jiaheng Liu, Yingshui Tan, Weixun Wang, Hui Huang, Xingyuan Bu, Hangyu Guo, Chengwei Hu, Boren Zheng, Zhuoran Lin, Dekai Sun
Abstract
New LLM evaluation benchmarks are important to align with the rapid development of Large Language Models (LLMs). In this work, we present Chinese Sim-pleQA, the first comprehensive Chinese benchmark to evaluate the factuality ability of language models to answer short questions, and Chinese SimpleQA mainly has five properties (i.e., Chinese, Diverse, High-quality, Static, Easy-to-evaluate). Specifically, first, we focus on the Chinese language over 6 major topics with 99 diverse subtopics. Second, we conduct a comprehensive quality control process to achieve high-quality questions and answers, where the reference answers are static and cannot be changed over time. Third, following SimpleQA, the questions and answers are very short, and the grading process is easy-to-evaluate based on OpenAI API. Based on Chinese SimpleQA, we perform a comprehensive evaluation on the factuality abilities of existing LLMs. Finally, we hope that Chinese SimpleQA could guide the developers to better understand the Chinese factuality abilities of their models and facilitate the growth of foundation models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3c2e7a50-b498-41db-8c41-39ec1fdd7bedCited by top-tier papers12
- When to use Graphs in RAG: A Comprehensive Analysis for Graph Retrieval-Augmented GenerationZhishang Xiang, Chuanjie Wu, Qinggang Zhang, Shengyuan Chen et al.ICLR 2026 · 56 citations
- DeepDiver: Adaptive Web-Search Intensity Scaling via Reinforcement LearningWenxuan Shi, Haochen Tan, Chuqiao Kuang, Xiaoguang Li et al.NeurIPS 2025 · 26 citations
- Lost in Benchmarks? Rethinking Large Language Model Benchmarking with Item Response TheoryHongli Zhou, Hui Huang, Ziqing Zhao, Lvyuan Han et al.AAAI 2026 · 15 citations
- KnowRL: Exploring Knowledgeable Reinforcement Learning for FactualityBaochang Ren, Shuofei Qiao, Ningyu Zhang, Da Zheng et al.ACL 2026 · 12 citations
- Chinese SafetyQA: A Safety Short-form Factuality Benchmark for Large Language ModelsYingshui Tan, Boren Zheng, Baihui Zheng, Kerui Cao et al.ACL 2025 · 7 citations
Builds on10
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Locating and Editing Factual Associations in GPTKevin Meng, David Bau, Alex Andonian, Yonatan BelinkovNeurIPS 2022 · 3,415 citations
- TruthfulQA: Measuring How Models Mimic Human FalsehoodsStephanie Lin, Jacob Hilton, Owain EvansACL 2022 · 3,228 citations
- CRITIC: Large Language Models Can Self-Correct with Tool-Interactive CritiquingZhibin Gou, Zhihong Shao, Yeyun Gong, Yelong Shen et al.ICLR 2024 · 699 citations
Related papers
- SimpleVQA: Multimodal Factuality Evaluation for Multimodal Large Language ModelsXianfu Cheng, Wei Zhang, Shiwei Zhang, Jian Yang et al.ICCV 2025 · 4 citations
- Video SimpleQA: Towards Factuality Evaluation in Large Video Language ModelsMeng Cao, Pengfei Hu, Yingyao Wang, Jihao Gu et al.AAAI 2026 · 1 citation
- SafetyBench: Evaluating the Safety of Large Language ModelsZhexin Zhang, Leqi Lei, Lindong Wu, Rui Sun et al.ACL 2024
- Harnessing Temporal Databases for Systematic Evaluation of Factual Time-Sensitive Question-Answering in LLMsSoyeon Kim, Jindong Wang, Xing Xie, Steven Euijong WhangICLR 2026
- SmartBench: Is Your LLM Truly a Good Chinese Smartphone Assistant?Xudong Lu, Haohao Gao, Renshou Wu, Shuai Ren et al.EMNLP 2025
