ProxyQA: An Alternative Framework for Evaluating Long-Form Text Generation with Large Language Models
Haochen Tan, Zhijiang Guo, Zhan Shi, Lu Xu, Zhili Liu, Yunlong Feng, Xiaoguang Li, Yasheng Wang, Lifeng Shang, Qun Liu, Linqi Song
摘要
Large Language Models (LLMs) have succeeded remarkably in understanding longform contents. However, exploring their capability for generating long-form contents, such as reports and articles, has been relatively unexplored and inadequately assessed by existing benchmarks. The prevalent evaluation methods, which predominantly rely on crowdsourcing, are recognized for their labor-intensive nature and lack of efficiency, whereas automated metrics, such as the ROUGE score, demonstrate discordance with human judgment criteria. In this paper, we propose PROXYQA, an innovative framework dedicated to assessing longtext generation. PROXYQA comprises in-depth human-curated meta-questions spanning various domains, each accompanied by specific proxy-questions with pre-annotated answers. LLMs are tasked to generate extensive content in response to these meta-questions, by engaging an evaluator and incorporating the generated texts as contextual background, PROX-YQA assesses the generated content's quality through the evaluator's accuracy in addressing the proxy-questions. We examine multiple LLMs, emphasizing PROXYQA's demanding nature as a high-quality assessment tool. Human evaluation demonstrates that the proxyquestion method is notably self-consistent and aligns closely with human evaluative standards. The dataset and leaderboard is available at https://proxy-qa.com .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper16
- LiveResearchBench: A Live Benchmark for User-Centric Deep Research in the WildJiayu Wang, Yifei Ming, Riya Dulepet, Qinglin Chen 等ICLR 2026 · 被引用 37 次
- DeepDiver: Adaptive Web-Search Intensity Scaling via Reinforcement LearningWenxuan Shi, Haochen Tan, Chuqiao Kuang, Xiaoguang Li 等NeurIPS 2025 · 被引用 26 次
- LIFBench: Evaluating the Instruction Following Performance and Stability of Large Language Models in Long-Context ScenariosXiaodong Wu, Minhao Wang, Yichen Liu, Xiaoming Shi 等ACL 2025 · 被引用 22 次
- DeepSolution: Boosting Complex Engineering Solution Design via Tree-based Exploration and Bi-point ThinkingZhuoqun Li, Haiyang Yu, Xuanang Chen, Hongyu Lin 等ACL 2025 · 被引用 13 次
- Improving Preference Extraction In LLMs By Identifying Latent Knowledge Through Classifying ProbesSharan Maiya, Yinhong Liu, Ramit Debnath, Anna KorhonenACL 2025 · 被引用 4 次
它引用的顶会 Paper25
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger 等ICLR 2020 · 被引用 8,443 次
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra 等NeurIPS 2022 · 被引用 5,493 次
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 被引用 2,600 次
- Train Short, Test Long: Attention with Linear Biases Enables Input Length ExtrapolationOfir Press, Noah A. Smith, Mike LewisICLR 2022 · 被引用 1,168 次
相关 Paper
- From General Reward to Targeted Reward: Improving Open-ended Long-context Generation ModelsZhihan Guo, Jiele Wu, Wenqian Cui, Yifei Zhang 等EMNLP 2025
- An Empirical Study of Evaluating Long-form Question AnsweringNing Xian, Yixing Fan, Ruqing Zhang, Maarten de Rijke 等SIGIR 2025 · 被引用 2 次
- LFQA-E: Carefully Benchmarking Long-form QA EvaluationYuchen Fan, Chen Ling, Xin Zhong, Shuo Zhang 等ICLR 2026 · 被引用 2 次
- Long-form factuality in large language modelsJerry Wei, Chengrun Yang, Xinying Song, Yifeng Lu 等NeurIPS 2024 · 被引用 182 次
- ExpertLongBench: Benchmarking Language Models on Expert-Level Long-Form Generation Tasks with Structured ChecklistsJie Ruan, Inderjeet Nair, Shuyang Cao, Amy Liu 等ICLR 2026 · 被引用 25 次
