UIS-Digger: Towards Comprehensive Research Agent Systems for Real-world Unindexed Information Seeking
Chang Liu, Chuqiao Kuang, Tianyi Zhuang, Yuxin Cheng, Huichi Zhou, Xiaoguang Li, Lifeng Shang
摘要
Recent advancements in LLM-based information-seeking agents have achieved record-breaking performance on established benchmarks. However, these agents remain heavily reliant on search-engine-indexed knowledge, leaving a critical blind spot: Unindexed Information Seeking (UIS). This paper identifies and explores the UIS problem, where vital information is not captured by search engine crawlers, such as overlooked content, dynamic webpages, and embedded files. Despite its significance, UIS remains an underexplored challenge. To address this gap, we introduce UIS-QA, the first dedicated UIS benchmark, comprising 110 expert-annotated QA pairs. Notably, even state-of-the-art agents experience a drastic performance drop on UIS-QA (e.g., from 70.90 on GAIA and 46.70 on BrowseComp-zh to 24.55 on UIS-QA), underscoring the severity of the problem. To mitigate this, we propose UIS-Digger, a novel multi-agent framework that incorporates dual-mode browsing and enables simultaneous webpage searching and file parsing. With a relatively small 30B-parameter backbone LLM optimized using SFT and RFT training strategies, UIS-Digger sets a strong baseline at 27.27%, outperforming systems integrating sophisticated LLMs such as O3 and GPT-4.1. This demonstrates the importance of proactive interaction with unindexed sources for effective and comprehensive information-seeking. Our work not only uncovers a fundamental limitation in current agent evaluation paradigms but also provides the first toolkit for advancing UIS research, defining a new and promising direction for robust information-seeking systems. The dataset has been released at: https://huggingface.co/datasets/UIS-Digger/UIS-QA.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper6
- GAIA: a benchmark for General AI AssistantsGrégoire Mialon, Clémentine Fourrier, Thomas Wolf, Yann LeCun 等ICLR 2024 · 被引用 716 次
- DeepResearch Bench: A Comprehensive Benchmark for Deep Research AgentsMingxuan Du, Benfeng Xu, Chiwei Zhu, Licheng Zhang 等ICLR 2026 · 被引用 250 次
- WideSearch: Benchmarking Agentic Broad Info-SeekingRyan Wong, Jiawei Wang, Junjie Zhao, Li Chen 等ICLR 2026 · 被引用 66 次
- DeepResearcher: Scaling Deep Research via Reinforcement Learning in Real-world EnvironmentsYuxiang Zheng, Dayuan Fu, Xiangkun Hu, Xiaojie Cai 等EMNLP 2025 · 被引用 8 次
- Browsing Lost Unformed Recollections: A Benchmark for Tip-of-the-Tongue Search and ReasoningSky CH-Wang, Darshan Girish Deshpande, Smaranda Muresan, Anand Kannappan 等ACL 2025
相关 Paper
- Empowering Efficiency and Efficacy in WebAgent via Enabling Info-Rich SeekingZhengwei Tao, Haiyang SHEN, Baixuan Li, Wenbiao Yin 等ICLR 2026 · 被引用 14 次
- WebWalker: Benchmarking LLMs in Web TraversalJialong Wu, Wenbiao Yin, Yong Jiang, Zhenglin Wang 等ACL 2025
- DeepDiver: Adaptive Web-Search Intensity Scaling via Reinforcement LearningWenxuan Shi, Haochen Tan, Chuqiao Kuang, Xiaoguang Li 等NeurIPS 2025 · 被引用 26 次
- InfoMosaic-Bench: Evaluating Multi-Source Information Seeking in Tool-Augmented AgentsYaxin Du, Yuanshuo Zhang, Xiyuan Yang, Yifan Zhou 等ICLR 2026 · 被引用 3 次
- BrowseComp-Plus: A Fair and Disentangled Evaluation Benchmark for Deep Search AgentsZijian Chen, Xueguang Ma, Shengyao Zhuang, Ping Nie 等ACL 2026
