Browsing Lost Unformed Recollections: A Benchmark for Tip-of-the-Tongue Search and Reasoning
Sky CH-Wang, Darshan Girish Deshpande, Smaranda Muresan, Anand Kannappan, Rebecca Qian
Abstract
We introduce BROWSING LOST UNFORMED RECOLLECTIONS, a tip-of-the-tongue knownitem search and reasoning benchmark for general AI assistants. BLUR introduces a set of 573 real-world validated questions that demand searching and reasoning across multimodal and multilingual inputs, as well as proficient tool use, in order to excel on. Humans easily ace these questions (scoring on average 98%), while the best-performing system scores around 56%. To facilitate progress toward addressing this challenging and aspirational use case for general AI assistants, we release 350 questions through a public leaderboard, retain the answers to 250 of them, and have the rest as a private test set.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- InteractComp: Evaluating Search Agents With Ambiguous QueriesMingyi Deng, Lijun Huang, Yani Fan, Fanqi Kong et al.ICML 2026 · 11 citations
- UIS-Digger: Towards Comprehensive Research Agent Systems for Real-world Unindexed Information SeekingChang Liu, Chuqiao Kuang, Tianyi Zhuang, Yuxin Cheng et al.ICLR 2026 · 1 citation
Builds on8
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- SWE-bench: Can Language Models Resolve Real-world Github Issues?Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao et al.ICLR 2024 · 2,082 citations
- Gorilla: Large Language Model Connected with Massive APIsShishir G. Patil, Tianjun Zhang, Xin Wang, Joseph E. GonzalezNeurIPS 2024 · 1,715 citations
- GAIA: a benchmark for General AI AssistantsGrégoire Mialon, Clémentine Fourrier, Thomas Wolf, Yann LeCun et al.ICLR 2024 · 716 citations
- InfiAgent-DABench: Evaluating Agents on Data Analysis TasksXueyu Hu, Ziyu Zhao, Shuang Wei, Ziwei Chai et al.ICML 2024 · 110 citations
Related papers
- RBench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning EvaluationMeng-Hao Guo, Jiajun Xu, Yi Zhang, Jiaxi Song et al.ICML 2025
- SealQA: Raising the Bar for Reasoning in Search-Augmented Language ModelsThinh Pham, Nguyen Phan Nguyen, Pratibha Zunjare, Weiyuan Chen et al.ICLR 2026 · 69 citations
- BRIGHT: A Realistic and Challenging Benchmark for Reasoning-Intensive RetrievalHongjin Su, Howard Yen, Mengzhou Xia, Weijia Shi et al.ICLR 2025
- KnowMe-Bench: Benchmarking Person Understanding for Lifelong Digital CompanionsTingyu Wu, Zhisheng Chen, Ziyan Weng, Shuhe Wang et al.ACL 2026 · 6 citations
- MuMA-ToM: Multi-modal Multi-Agent Theory of MindHaojun Shi, Suyu Ye, Xinyu Fang, Chuanyang Jin et al.AAAI 2025 · 48 citations
