SealQA: Raising the Bar for Reasoning in Search-Augmented Language Models
Thinh Pham, Nguyen Phan Nguyen, Pratibha Zunjare, Weiyuan Chen, Yu-Min Tseng, Tu Vu
摘要
We introduce SEALQA, a challenge benchmark for evaluating SEarch-Augmented Language models on fact-seeking questions where web search yields conflicting, noisy, or unhelpful results. SEALQA comes in three flavors: (1) SEAL-0 (main) and (2) SEAL-HARD, both of which assess factual accuracy and reasoning capabilities, where SEAL-0 targets the most challenging questions that frontier non-reasoning models (e.g., .1) answer with near-zero accuracy; and (3) LONGSEAL, which extends SEALQA to test long-context, multi-document reasoning in "needle-in-a-haystack" settings. Our evaluation reveals critical limitations in current models. Even frontier reasoning models face significant challenges across SEALQA flavors. On SEAL-0, GPT-5 with tools achieves only 43.2% accuracy at its best reasoning effort. We also find that even advanced reasoning models (e.g., DEEPSEEK-R1) can be vulnerable to noisy search results. Notably, increasing test-time compute does not yield reliable gains across GPT-5 and the O-series of models, with performance often plateauing or even declining early. Finally, while current models are less affected by the "lost-in-the-middle" issue, they still fail to reliably identify relevant documents in LONGSEAL when faced with numerous distractors. To facilitate future work, we release SEALQA at https://huggingface.co/datasets/vtllms/sealqa . 1 Each question required over an hour on average -roughly 45 minutes to draft, plus additional time for review and revision. Many initial ideas were discarded as they failed to meaningfully challenge frontier LLMS. 2 For example, the widely used GPQA-DIAMOND (Rein et al., 2024) , a compact set of 198 expert-vetted questions, demonstrates how a small, carefully curated dataset can effectively assess a model's reasoning ability. 3 Our questions often lead multiple models to fail across repeated attempts.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- A Survey of Large Language Model-Based Search AgentsYunjia Xi, Jianghao Lin, Yongzhao Xiao, Zheli Zhou 等ACL 2026 · 被引用 1,216 次
- Scaling Agents via Continual Pre-trainingLiangcai Su, Zhen Zhang, Guangyu Li, Zhuo Chen 等ICLR 2026 · 被引用 46 次
- IterResearch: Rethinking Long-Horizon Agents with Interaction ScalingGuoxin Chen, Zile Qiao, Xuanzhong Chen, Donglei Yu 等ICLR 2026 · 被引用 17 次
- Empowering Efficiency and Efficacy in WebAgent via Enabling Info-Rich SeekingZhengwei Tao, Haiyang SHEN, Baixuan Li, Wenbiao Yin 等ICLR 2026 · 被引用 14 次
- Evidence-Augmented Policy Optimization with Reward Co-Evolution for Long-Context ReasoningXin Guan, Zijian Li, Shen Huang, Pengjun Xie 等ACL 2026 · 被引用 6 次
它引用的顶会 Paper13
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- SWE-bench: Can Language Models Resolve Real-world Github Issues?Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao 等ICLR 2024 · 被引用 2,082 次
- Large Language Models Can Be Easily Distracted by Irrelevant ContextFreda Shi, Xinyun Chen, Kanishka Misra, Nathan Scales 等ICML 2023 · 被引用 970 次
- Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth ApproachJonas Geiping, Sean McLeish, Neel Jain, John Kirchenbauer 等NeurIPS 2025 · 被引用 431 次
- Adaptive Chameleon or Stubborn Sloth: Revealing the Behavior of Large Language Models in Knowledge ConflictsJian Xie, Kai Zhang, Jiangjie Chen, Renze Lou 等ICLR 2024 · 被引用 294 次
相关 Paper
- LakeQA: An Exploratory QA Benchmark over a Million-Scale Data LakeHaonan Wang, Jiaxiang Liu, Yurong Liu, Austin Wijaya 等ICML 2026
- One Thousand and One Pairs: A "novel" challenge for long-context language modelsMarzena Karpinska, Katherine Thai, Kyle Lo, Tanya Goyal 等EMNLP 2024 · 被引用 6 次
- MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Long-tail KnowledgeJie He, Nan Hu, Wanqiu Long, Jiaoyan Chen 等ACL 2026 · 被引用 1 次
- NoLiMa: Long-Context Evaluation Beyond Literal MatchingAli Modarressi, Hanieh Deilamsalehy, Franck Dernoncourt, Trung Bui 等ICML 2025
- RefineBench: Evaluating Refinement Capability of Language Models via ChecklistsYoung-Jun Lee, Seungone Kim, Byung-Kwan Lee, Minkyeong Moon 等ICLR 2026 · 被引用 13 次
