Exploring the Potential of LLMs as Personalized Assistants: Dataset, Evaluation, and Analysis
Jisoo Mok, Ik-hwan Kim, Sangkwon Park, Sungroh Yoon
摘要
Personalized AI assistants, a hallmark of the human-like capabilities of Large Language Models (LLMs), are a challenging application that intertwines multiple problems in LLM research. Despite the growing interest in the development of personalized assistants, the lack of an open-source conversational dataset tailored for personalization remains a significant obstacle for researchers in the field. To address this research gap, we introduce HiCU-PID, a new benchmark to probe and unleash the potential of LLMs to deliver personalized responses. Alongside a conversational dataset, HiCUPID provides a Llama-3.2-based automated evaluation model whose assessment closely mirrors human preferences. We release our dataset, evaluation model, and code at https://github.com/12kimih/HiCUPID .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Your Language Model Secretly Contains Personality SubnetworksRuimeng Ye, Zihan Wang, Zinan Ling, Yang Xiao 等ICLR 2026 · 被引用 4 次
- Latent Inter-User Difference Modeling for LLM PersonalizationYilun Qiu, Tianhao Shi, Xiaoyan Zhao, Fengbin Zhu 等EMNLP 2025 · 被引用 1 次
- Contextualized Visual Personalization in Vision-Language ModelsYeongtak Oh, Sangwon Yu, Junsung Park, Han Cheol Moon 等ICML 2026 · 被引用 1 次
- AgentCoMa: A Compositional Benchmark Mixing Commonsense and Mathematical Reasoning in Real-World ScenariosLisa Alazraki, Lihu Chen, Ana Brassard, Joe Stacey 等ACL 2026 · 被引用 1 次
- HingeMem: Boundary Guided Long-Term Memory with Query Adaptive Retrieval for Scalable DialoguesYijie Zhong, Yunfan Gao, Haofen WangWWW 2026
它引用的顶会 Paper10
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- MIND: A Large-scale Dataset for News RecommendationFangzhao Wu, Ying Qiao, Jiun-Hung Chen, Chuhan Wu 等ACL 2020 · 被引用 454 次
- DOC: Improving Long Story Coherence With Detailed Outline ControlKevin Yang, Dan Klein, Nanyun Peng, Yuandong TianACL 2023 · 被引用 23 次
相关 Paper
- Do LLMs Recognize Your Preferences? Evaluating Personalized Preference Following in LLMsSiyan Zhao, Mingyi Hong, Yang Liu, Devamanyu Hazarika 等ICLR 2025
- PersonalLLM: Tailoring LLMs to Individual PreferencesThomas P. Zollo, Andrew Wei Tung Siah, Naimeng Ye, Ang Li 等ICLR 2025
- FaST: Feature-aware Sampling and Tuning for Personalized Preference Alignment with Limited DataThibaut Thonet, Germán Kruszewski, Jos Rozen, Pierre Erbacher 等EMNLP 2025 · 被引用 1 次
- Learning Preference Model for LLMs via Automatic Preference Data GenerationShijia Huang, Jianqiao Zhao, Yanyang Li, Liwei WangEMNLP 2023 · 被引用 3 次
- Doing Personal LAPS: LLM-Augmented Dialogue Construction for Personalized Multi-Session Conversational SearchHideaki Joko, Shubham Chatterjee, Andrew Ramsay, Arjen P. de Vries 等SIGIR 2024 · 被引用 24 次
