Towards Personalized Deep Research: Benchmarks and Evaluations
Yuan Liang, Jiaxian Li, Yuqing Wang, Piaohong Wang, Motong Tian, Pai Liu, Shuofei Qiao, Runnan Fang, He Zhu, Ge Zhang, Minghao Liu, Yuchen Eleanor Jiang
Abstract
Deep Research Agents (DRAs) can autonomously conduct complex investigations and generate comprehensive reports, demonstrating strong real-world potential. However, existing evaluations mostly rely on close-ended benchmarks, while open-ended deep research benchmarks remain scarce and typically neglect personalized scenarios. To bridge this gap, we introduce Personalized Deep Research Bench (PDR-Bench), the first benchmark for evaluating personalization in DRAs. It pairs 50 diverse research tasks across 10 domains with 25 authentic user profiles that combine structured persona attributes with dynamic real-world contexts, yielding 250 realistic user-task queries. To assess system performance, we propose the PQR Evaluation Framework, which jointly measures Personalization Alignment, Content Quality, and Factual Reliability. Our experiments on a range of systems highlight current capabilities and limitations in handling personalized deep research. This work establishes a rigorous foundation for developing and evaluating the next generation of truly personalized AI research assistants 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 824af6d5-1b77-4cdc-912c-892400e4d2b3Cited by top-tier papers4
- Language Models Don't Know What You Want: Evaluating Personalization in Deep Research Needs Real UsersNishant Balepur, Malachi Hamada, Varsha Kishore, Sergey Feldman et al.ACL 2026 · 1 citation
- EVOLVING ROLLOUTS: Harnessing Historical Experience for Web Agent Evolution in Reinforcement LearningSinuo Wang, WANG PIAOHONG, Tianrui Qin, Maojia Song et al.ICML 2026
- Towards Knowledgeable Deep Research: Framework and BenchmarkWenxuan Liu, Zixuan Li, Long Bai, Chunmao Zhang et al.SIGIR 2026
- IntentRL: Training Proactive User-intent Agents for Open-ended Deep Research via Reinforcement LearningHaohao Luo, Zexi Li, Yuexiang Xie, Wenhao Zhang et al.ICML 2026
Builds on12
- GAIA: a benchmark for General AI AssistantsGrégoire Mialon, Clémentine Fourrier, Thomas Wolf, Yann LeCun et al.ICLR 2024 · 716 citations
- DeepResearch Bench: A Comprehensive Benchmark for Deep Research AgentsMingxuan Du, Benfeng Xu, Chiwei Zhu, Licheng Zhang et al.ICLR 2026 · 250 citations
- FActScore: Fine-grained Atomic Evaluation of Factual Precision in Long Form Text GenerationSewon Min, Kalpesh Krishna, Xinxi Lyu, Mike Lewis et al.EMNLP 2023 · 225 citations
- Long-form factuality in large language modelsJerry Wei, Chengrun Yang, Xinying Song, Yifeng Lu et al.NeurIPS 2024 · 182 citations
- OWL: Optimized Workforce Learning for General Multi-Agent Assistance in Real-World Task AutomationMengkang Hu, Yuhang Zhou, Wendong Fan, Yuzhou Nie et al.NeurIPS 2025 · 158 citations
Related papers
- DRBench: A Realistic Benchmark for Enterprise Deep ResearchAmirhossein Abaskohi, Tianyi Chen, Miguel Muñoz-Mármol, Curtis Fox et al.ICLR 2026 · 18 citations
- Beyond Single-shot Writing: Deep Research Agents are Unreliable at Multi-turn Report RevisionBingsen Chen, Boyan Li, Ping Nie, Yuyu Zhang et al.ACL 2026 · 1 citation
- ResearchRubrics: A Benchmark of Prompts and Rubrics For Evaluating Deep Research AgentsManasi Sharma, Chen Bo Calvin Zhang, Chaithanya Bandi, Clinton Wang et al.ICLR 2026 · 83 citations
- LiveResearchBench: A Live Benchmark for User-Centric Deep Research in the WildJiayu Wang, Yifei Ming, Riya Dulepet, Qinglin Chen et al.ICLR 2026 · 37 citations
- DR-Arena: an Automated Evaluation Framework for Deep Research AgentsYiwen Gao, Ruochen Zhao, Yang Deng, Wenxuan ZhangACL 2026 · 2 citations
