Language Models Don't Know What You Want: Evaluating Personalization in Deep Research Needs Real Users
Nishant Balepur, Malachi Hamada, Varsha Kishore, Sergey Feldman, Amanpreet Singh, Pao Siangliulue, Joseph Chee Chang, Eunsol Choi, Jordan Lee Boyd-Graber, Aakanksha Naik
摘要
Deep Research (DR) systems help researchers cope with ballooning publishing counts. Such tools synthesize scientific papers to answer research queries, but lack understanding of their users. We address this with MYSCHOLARQA (MYSQA), a personalized DR agent that: 1) infers a profile with a user's research interests; 2) proposes personalized actions for a user's input query; and 3) writes a multi-section report for the query that follows user-approved actions. We first test MYSQA with NLP's standard protocol: we build a benchmark with synthetic users and LLM judges, where MYSQA beats baselines in citation metrics and personalized action-following. However, we suspect this process does not cover all aspects of personalized DR users value, so we interview users in an online version of MYSQA to unmask them. We reveal nine nuanced errors of personalized DR undetectable by our LLM judges, and we study qualitative feedback to form lessons for future DR design. In all, we argue for a pillar of personalization that easy-to-use LLM judges can lead NLP to overlook: real progress in personalization is only possible with real users. 1 When Deep Research Gets to Know You Scholars increasingly turn to LLMs to support their scientific research (Liao et al., 2024) , such as to learn new concepts (August et al., 2023) or brainstorm ideas (Pu et al., 2025) . With publishing rates skyrocketing and literature becoming daunting to track (Parolo et al., 2015) , a new use case of LLMs emerges: Deep Research (DR) tools that answer researchers' queries by retrieving, organizing, and synthesizing papers into multi-section, attributed reports (Asai et al., 2024; Huang et al., 2025) . 1) Infer a Profile ( §2.1 ) 2) Propose Actions ( §2.2 ) 3) Synthesize a Report ( §2.3 ) Researcher Profile Knowledge: -You have expertise on training building scientific QA… [1, 3] Research Style: -You prefer to study models using black-box analysis… [2, 3] [1] [2] [3] Researcher-picked papers Query How can I build the first personalized Deep Research System? List of Actions Research Ideas: -Point out current QA limitations -Give design deployment steps Content: -Focus on scientific QA papers -Discuss deep learning findings
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper26
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni 等NeurIPS 2020 · 被引用 19,162 次
- Why Johnny Can't Prompt: How Non-AI Experts Try (and Fail) to Design LLM PromptsJ. D. Zamfirescu-Pereira, Richmond Y. Wong, Bjoern Hartmann, Qian YangCHI 2023 · 被引用 892 次
- AutoSurvey: Large Language Models Can Automatically Write SurveysYidong Wang, Qi Guo, Wenjin Yao, Hongbo Zhang 等NeurIPS 2024 · 被引用 151 次
- Human Feedback is not Gold StandardTom Hosking, Phil Blunsom, Max BartoloICLR 2024 · 被引用 96 次
- Evaluating the Factual Consistency of Abstractive Text SummarizationWojciech Kryscinski, Bryan McCann, Caiming Xiong, Richard SocherEMNLP 2020 · 被引用 67 次
相关 Paper
- Towards Personalized Deep Research: Benchmarks and EvaluationsYuan Liang, Jiaxian Li, Yuqing Wang, Piaohong Wang 等ICLR 2026 · 被引用 13 次
- DeepResearch Bench: A Comprehensive Benchmark for Deep Research AgentsMingxuan Du, Benfeng Xu, Chiwei Zhu, Licheng Zhang 等ICLR 2026 · 被引用 250 次
- IntentRL: Training Proactive User-intent Agents for Open-ended Deep Research via Reinforcement LearningHaohao Luo, Zexi Li, Yuexiang Xie, Wenhao Zhang 等ICML 2026
- Beyond Single-shot Writing: Deep Research Agents are Unreliable at Multi-turn Report RevisionBingsen Chen, Boyan Li, Ping Nie, Yuyu Zhang 等ACL 2026 · 被引用 1 次
- ResearchRubrics: A Benchmark of Prompts and Rubrics For Evaluating Deep Research AgentsManasi Sharma, Chen Bo Calvin Zhang, Chaithanya Bandi, Clinton Wang 等ICLR 2026 · 被引用 83 次
