IPQA: A Benchmark for Core Intent Identification in Personalized Question Answering
Jieyong Kim, Maryam Amirizaniani, Soojin Yoon, Dongha Lee
摘要
Intent identification serves as the foundation for generating appropriate responses in personalized question answering (PQA). However, existing benchmarks evaluate only response quality or retrieval performance without directly measuring intent identification capabilities. This gap is critical because without understanding which intents users prioritize, systems cannot generate responses satisfying individual information needs. To address this, we introduce the concept of core intents: intents users prioritize when selecting answers to satisfy their information needs. To evaluate these core intents, we propose IPQA, a benchmark for core Intent identification in Personalized Question Answering. Since users do not explicitly state their prioritized intents, we derive core intents from observable behavior patterns in answer selection, grounded in bounded rationality, where users satisfice by choosing answers meeting their acceptance thresholds. We construct a dataset with various domains through systematic filtering, LLM-based annotation, and rigorous quality control combining automated verification with human validation. Experimental evaluations across state-ofthe-art language models reveal that current systems struggle with core intent identification in personalized contexts. Models fail to identify core intents from user histories, with performance degrading as question complexity increases. [REPOSITORY]
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper10
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger 等ICLR 2020 · 被引用 8,443 次
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng 等SOSP 2023 · 被引用 1,016 次
- G-Eval: NLG Evaluation using Gpt-4 with Better Human AlignmentYang Liu, Dan Iter, Yichong Xu, Shuohang Wang 等EMNLP 2023 · 被引用 549 次
- Deep Open Intent Classification with Adaptive Decision BoundaryHanlei Zhang, Hua Xu, Ting-En LinAAAI 2021 · 被引用 127 次
- Review-driven Personalized Preference Reasoning with Large Language Models for RecommendationJieyong Kim, Hyunseo Kim, Hyunjin Cho, SeongKu Kang 等SIGIR 2025 · 被引用 13 次
相关 Paper
- LaMP-QA: A Benchmark for Personalized Long-form Question AnsweringAlireza Salemi, Hamed ZamaniEMNLP 2025 · 被引用 1 次
- Beyond Facts: Evaluating Intent Hallucination in Large Language ModelsYijie Hao, Haofei Yu, Jiaxuan YouACL 2025
- Pathways of Thoughts: Multi-Directional Thinking for Long-form Personalized Question AnsweringAlireza Salemi, Cheng Li, Mingyang Zhang, Qiaozhu Mei 等WWW 2026 · 被引用 3 次
- A User-Centric Multi-Intent Benchmark for Evaluating Large Language ModelsJiayin Wang, Fengran Mo, Weizhi Ma, Peijie Sun 等EMNLP 2024 · 被引用 10 次
- Persona2Web: Benchmarking Personalized Web Agents for Contextual Reasoning with User HistorySerin Kim, Sangam Lee, Dongha LeeICML 2026
