Language Models Don't Know What You Want: Evaluating Personalization in Deep Research Needs Real Users
Nishant Balepur, Malachi Hamada, Varsha Kishore, Sergey Feldman, Amanpreet Singh, Pao Siangliulue, Joseph Chee Chang, Eunsol Choi, Jordan Lee Boyd-Graber, Aakanksha Naik
Abstract
Deep Research (DR) systems help researchers cope with ballooning publishing counts. Such tools synthesize scientific papers to answer research queries, but lack understanding of their users. We address this with MYSCHOLARQA (MYSQA), a personalized DR agent that: 1) infers a profile with a user's research interests; 2) proposes personalized actions for a user's input query; and 3) writes a multi-section report for the query that follows user-approved actions. We first test MYSQA with NLP's standard protocol: we build a benchmark with synthetic users and LLM judges, where MYSQA beats baselines in citation metrics and personalized action-following. However, we suspect this process does not cover all aspects of personalized DR users value, so we interview users in an online version of MYSQA to unmask them. We reveal nine nuanced errors of personalized DR undetectable by our LLM judges, and we study qualitative feedback to form lessons for future DR design. In all, we argue for a pillar of personalization that easy-to-use LLM judges can lead NLP to overlook: real progress in personalization is only possible with real users. 1 When Deep Research Gets to Know You Scholars increasingly turn to LLMs to support their scientific research (Liao et al., 2024) , such as to learn new concepts (August et al., 2023) or brainstorm ideas (Pu et al., 2025) . With publishing rates skyrocketing and literature becoming daunting to track (Parolo et al., 2015) , a new use case of LLMs emerges: Deep Research (DR) tools that answer researchers' queries by retrieving, organizing, and synthesizing papers into multi-section, attributed reports (Asai et al., 2024; Huang et al., 2025) . 1) Infer a Profile ( §2.1 ) 2) Propose Actions ( §2.2 ) 3) Synthesize a Report ( §2.3 ) Researcher Profile Knowledge: -You have expertise on training building scientific QA… [1, 3] Research Style: -You prefer to study models using black-box analysis… [2, 3] [1] [2] [3] Researcher-picked papers Query How can I build the first personalized Deep Research System? List of Actions Research Ideas: -Point out current QA limitations -Give design deployment steps Content: -Focus on scientific QA papers -Discuss deep learning findings
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a6ac8271-aa69-4ab6-bd33-2f7eee5e98a3Builds on26
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- Why Johnny Can't Prompt: How Non-AI Experts Try (and Fail) to Design LLM PromptsJ. D. Zamfirescu-Pereira, Richmond Y. Wong, Bjoern Hartmann, Qian YangCHI 2023 · 892 citations
- AutoSurvey: Large Language Models Can Automatically Write SurveysYidong Wang, Qi Guo, Wenjin Yao, Hongbo Zhang et al.NeurIPS 2024 · 151 citations
- Human Feedback is not Gold StandardTom Hosking, Phil Blunsom, Max BartoloICLR 2024 · 96 citations
- Evaluating the Factual Consistency of Abstractive Text SummarizationWojciech Kryscinski, Bryan McCann, Caiming Xiong, Richard SocherEMNLP 2020 · 67 citations
Related papers
- Towards Personalized Deep Research: Benchmarks and EvaluationsYuan Liang, Jiaxian Li, Yuqing Wang, Piaohong Wang et al.ICLR 2026 · 13 citations
- DeepResearch Bench: A Comprehensive Benchmark for Deep Research AgentsMingxuan Du, Benfeng Xu, Chiwei Zhu, Licheng Zhang et al.ICLR 2026 · 250 citations
- IntentRL: Training Proactive User-intent Agents for Open-ended Deep Research via Reinforcement LearningHaohao Luo, Zexi Li, Yuexiang Xie, Wenhao Zhang et al.ICML 2026
- Beyond Single-shot Writing: Deep Research Agents are Unreliable at Multi-turn Report RevisionBingsen Chen, Boyan Li, Ping Nie, Yuyu Zhang et al.ACL 2026 · 1 citation
- ResearchRubrics: A Benchmark of Prompts and Rubrics For Evaluating Deep Research AgentsManasi Sharma, Chen Bo Calvin Zhang, Chaithanya Bandi, Clinton Wang et al.ICLR 2026 · 83 citations
