UIS-Digger: Towards Comprehensive Research Agent Systems for Real-world Unindexed Information Seeking
Chang Liu, Chuqiao Kuang, Tianyi Zhuang, Yuxin Cheng, Huichi Zhou, Xiaoguang Li, Lifeng Shang
Abstract
Recent advancements in LLM-based information-seeking agents have achieved record-breaking performance on established benchmarks. However, these agents remain heavily reliant on search-engine-indexed knowledge, leaving a critical blind spot: Unindexed Information Seeking (UIS). This paper identifies and explores the UIS problem, where vital information is not captured by search engine crawlers, such as overlooked content, dynamic webpages, and embedded files. Despite its significance, UIS remains an underexplored challenge. To address this gap, we introduce UIS-QA, the first dedicated UIS benchmark, comprising 110 expert-annotated QA pairs. Notably, even state-of-the-art agents experience a drastic performance drop on UIS-QA (e.g., from 70.90 on GAIA and 46.70 on BrowseComp-zh to 24.55 on UIS-QA), underscoring the severity of the problem. To mitigate this, we propose UIS-Digger, a novel multi-agent framework that incorporates dual-mode browsing and enables simultaneous webpage searching and file parsing. With a relatively small 30B-parameter backbone LLM optimized using SFT and RFT training strategies, UIS-Digger sets a strong baseline at 27.27%, outperforming systems integrating sophisticated LLMs such as O3 and GPT-4.1. This demonstrates the importance of proactive interaction with unindexed sources for effective and comprehensive information-seeking. Our work not only uncovers a fundamental limitation in current agent evaluation paradigms but also provides the first toolkit for advancing UIS research, defining a new and promising direction for robust information-seeking systems. The dataset has been released at: https://huggingface.co/datasets/UIS-Digger/UIS-QA.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on6
- GAIA: a benchmark for General AI AssistantsGrégoire Mialon, Clémentine Fourrier, Thomas Wolf, Yann LeCun et al.ICLR 2024 · 716 citations
- DeepResearch Bench: A Comprehensive Benchmark for Deep Research AgentsMingxuan Du, Benfeng Xu, Chiwei Zhu, Licheng Zhang et al.ICLR 2026 · 250 citations
- WideSearch: Benchmarking Agentic Broad Info-SeekingRyan Wong, Jiawei Wang, Junjie Zhao, Li Chen et al.ICLR 2026 · 66 citations
- DeepResearcher: Scaling Deep Research via Reinforcement Learning in Real-world EnvironmentsYuxiang Zheng, Dayuan Fu, Xiangkun Hu, Xiaojie Cai et al.EMNLP 2025 · 8 citations
- Browsing Lost Unformed Recollections: A Benchmark for Tip-of-the-Tongue Search and ReasoningSky CH-Wang, Darshan Girish Deshpande, Smaranda Muresan, Anand Kannappan et al.ACL 2025
Related papers
- Empowering Efficiency and Efficacy in WebAgent via Enabling Info-Rich SeekingZhengwei Tao, Haiyang SHEN, Baixuan Li, Wenbiao Yin et al.ICLR 2026 · 14 citations
- WebWalker: Benchmarking LLMs in Web TraversalJialong Wu, Wenbiao Yin, Yong Jiang, Zhenglin Wang et al.ACL 2025
- DeepDiver: Adaptive Web-Search Intensity Scaling via Reinforcement LearningWenxuan Shi, Haochen Tan, Chuqiao Kuang, Xiaoguang Li et al.NeurIPS 2025 · 26 citations
- InfoMosaic-Bench: Evaluating Multi-Source Information Seeking in Tool-Augmented AgentsYaxin Du, Yuanshuo Zhang, Xiyuan Yang, Yifan Zhou et al.ICLR 2026 · 3 citations
- BrowseComp-Plus: A Fair and Disentangled Evaluation Benchmark for Deep Search AgentsZijian Chen, Xueguang Ma, Shengyao Zhuang, Ping Nie et al.ACL 2026
