SealQA: Raising the Bar for Reasoning in Search-Augmented Language Models
Thinh Pham, Nguyen Phan Nguyen, Pratibha Zunjare, Weiyuan Chen, Yu-Min Tseng, Tu Vu
Abstract
We introduce SEALQA, a challenge benchmark for evaluating SEarch-Augmented Language models on fact-seeking questions where web search yields conflicting, noisy, or unhelpful results. SEALQA comes in three flavors: (1) SEAL-0 (main) and (2) SEAL-HARD, both of which assess factual accuracy and reasoning capabilities, where SEAL-0 targets the most challenging questions that frontier non-reasoning models (e.g., .1) answer with near-zero accuracy; and (3) LONGSEAL, which extends SEALQA to test long-context, multi-document reasoning in "needle-in-a-haystack" settings. Our evaluation reveals critical limitations in current models. Even frontier reasoning models face significant challenges across SEALQA flavors. On SEAL-0, GPT-5 with tools achieves only 43.2% accuracy at its best reasoning effort. We also find that even advanced reasoning models (e.g., DEEPSEEK-R1) can be vulnerable to noisy search results. Notably, increasing test-time compute does not yield reliable gains across GPT-5 and the O-series of models, with performance often plateauing or even declining early. Finally, while current models are less affected by the "lost-in-the-middle" issue, they still fail to reliably identify relevant documents in LONGSEAL when faced with numerous distractors. To facilitate future work, we release SEALQA at https://huggingface.co/datasets/vtllms/sealqa . 1 Each question required over an hour on average -roughly 45 minutes to draft, plus additional time for review and revision. Many initial ideas were discarded as they failed to meaningfully challenge frontier LLMS. 2 For example, the widely used GPQA-DIAMOND (Rein et al., 2024) , a compact set of 198 expert-vetted questions, demonstrates how a small, carefully curated dataset can effectively assess a model's reasoning ability. 3 Our questions often lead multiple models to fail across repeated attempts.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0ee981d2-f231-4d63-b96c-09025237e92bCited by top-tier papers8
- A Survey of Large Language Model-Based Search AgentsYunjia Xi, Jianghao Lin, Yongzhao Xiao, Zheli Zhou et al.ACL 2026 · 1,216 citations
- Scaling Agents via Continual Pre-trainingLiangcai Su, Zhen Zhang, Guangyu Li, Zhuo Chen et al.ICLR 2026 · 46 citations
- IterResearch: Rethinking Long-Horizon Agents with Interaction ScalingGuoxin Chen, Zile Qiao, Xuanzhong Chen, Donglei Yu et al.ICLR 2026 · 17 citations
- Empowering Efficiency and Efficacy in WebAgent via Enabling Info-Rich SeekingZhengwei Tao, Haiyang SHEN, Baixuan Li, Wenbiao Yin et al.ICLR 2026 · 14 citations
- Evidence-Augmented Policy Optimization with Reward Co-Evolution for Long-Context ReasoningXin Guan, Zijian Li, Shen Huang, Pengjun Xie et al.ACL 2026 · 6 citations
Builds on13
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- SWE-bench: Can Language Models Resolve Real-world Github Issues?Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao et al.ICLR 2024 · 2,082 citations
- Large Language Models Can Be Easily Distracted by Irrelevant ContextFreda Shi, Xinyun Chen, Kanishka Misra, Nathan Scales et al.ICML 2023 · 970 citations
- Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth ApproachJonas Geiping, Sean McLeish, Neel Jain, John Kirchenbauer et al.NeurIPS 2025 · 431 citations
- Adaptive Chameleon or Stubborn Sloth: Revealing the Behavior of Large Language Models in Knowledge ConflictsJian Xie, Kai Zhang, Jiangjie Chen, Renze Lou et al.ICLR 2024 · 294 citations
Related papers
- LakeQA: An Exploratory QA Benchmark over a Million-Scale Data LakeHaonan Wang, Jiaxiang Liu, Yurong Liu, Austin Wijaya et al.ICML 2026
- One Thousand and One Pairs: A "novel" challenge for long-context language modelsMarzena Karpinska, Katherine Thai, Kyle Lo, Tanya Goyal et al.EMNLP 2024 · 6 citations
- MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Long-tail KnowledgeJie He, Nan Hu, Wanqiu Long, Jiaoyan Chen et al.ACL 2026 · 1 citation
- NoLiMa: Long-Context Evaluation Beyond Literal MatchingAli Modarressi, Hanieh Deilamsalehy, Franck Dernoncourt, Trung Bui et al.ICML 2025
- RefineBench: Evaluating Refinement Capability of Language Models via ChecklistsYoung-Jun Lee, Seungone Kim, Byung-Kwan Lee, Minkyeong Moon et al.ICLR 2026 · 13 citations
