LaMP-QA: A Benchmark for Personalized Long-form Question Answering
Alireza Salemi, Hamed Zamani
摘要
Personalization is essential for question answering systems that are user-centric. Despite its importance, personalization in answer generation has been relatively underexplored. This is mainly due to lack of resources for training and evaluating personalized question answering systems. We address this gap by introducing LaMP-QA-a benchmark designed for evaluating personalized long-form answer generation. The benchmark covers questions from three major categories: (1) Arts & Entertainment, (2) Lifestyle & Personal Development, and (3) Society & Culture, encompassing over 45 subcategories in total. To assess the quality and potential impact of the LaMP-QA benchmark for personalized question answering, we conduct comprehensive human and automatic evaluations, to compare multiple evaluation strategies for evaluating generated personalized responses and measure their alignment with human preferences. Furthermore, we benchmark a number of non-personalized and personalized approaches based on open-source and proprietary large language models. Our results show that incorporating the personalized context provided leads to up to 39% performance improvements. The benchmark is publicly released to support future research in this area.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- Think-While-Generating: On-the-Fly Reasoning for Personalized Long-Form GenerationChengbing Wang, Yang Zhang, Wenjie Wang, Xiaoyan Zhao 等ICLR 2026 · 被引用 35 次
- P-GenRM: Personalized Generative Reward Model with Test-time User-based ScalingPinyi Zhang, Ting-En Lin, Yuchuan Wu, Jingyang Chen 等ICLR 2026 · 被引用 5 次
- Pathways of Thoughts: Multi-Directional Thinking for Long-form Personalized Question AnsweringAlireza Salemi, Cheng Li, Mingyang Zhang, Qiaozhu Mei 等WWW 2026 · 被引用 3 次
- PRISP: Privacy-Safe Few-Shot Personalization via Lightweight AdaptationJunho Park, Dohoon Kim, Taesup MoonACL 2026 · 被引用 1 次
- Language Models Don't Know What You Want: Evaluating Personalization in Deep Research Needs Real UsersNishant Balepur, Malachi Hamada, Varsha Kishore, Sergey Feldman 等ACL 2026 · 被引用 1 次
它引用的顶会 Paper9
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes 等ICLR 2020 · 被引用 4,112 次
- G-Eval: NLG Evaluation using Gpt-4 with Better Human AlignmentYang Liu, Dan Iter, Yichong Xu, Shuohang Wang 等EMNLP 2023 · 被引用 549 次
- FActScore: Fine-grained Atomic Evaluation of Factual Precision in Long Form Text GenerationSewon Min, Kalpesh Krishna, Xinxi Lyu, Mike Lewis 等EMNLP 2023 · 被引用 225 次
- Learning to Rewrite Prompts for Personalized Text GenerationCheng Li, Mingyang Zhang, Qiaozhu Mei, Weize Kong 等WWW 2024 · 被引用 54 次
相关 Paper
- LaMP: When Large Language Models Meet PersonalizationAlireza Salemi, Sheshera Mysore, Michael Bendersky, Hamed ZamaniACL 2024
- LLMs + Persona-Plug = Personalized LLMsJiongnan Liu, Yutao Zhu, Shuting Wang, Xiaochi Wei 等ACL 2025 · 被引用 19 次
- Optimization Methods for Personalizing Large Language Models through Retrieval AugmentationAlireza Salemi, Surya Kallumadi, Hamed ZamaniSIGIR 2024 · 被引用 52 次
- LFQA-E: Carefully Benchmarking Long-form QA EvaluationYuchen Fan, Chen Ling, Xin Zhong, Shuo Zhang 等ICLR 2026 · 被引用 2 次
- An Empirical Study of Evaluating Long-form Question AnsweringNing Xian, Yixing Fan, Ruqing Zhang, Maarten de Rijke 等SIGIR 2025 · 被引用 2 次
