Exploring the Potential of LLMs as Personalized Assistants: Dataset, Evaluation, and Analysis
Jisoo Mok, Ik-hwan Kim, Sangkwon Park, Sungroh Yoon
Abstract
Personalized AI assistants, a hallmark of the human-like capabilities of Large Language Models (LLMs), are a challenging application that intertwines multiple problems in LLM research. Despite the growing interest in the development of personalized assistants, the lack of an open-source conversational dataset tailored for personalization remains a significant obstacle for researchers in the field. To address this research gap, we introduce HiCU-PID, a new benchmark to probe and unleash the potential of LLMs to deliver personalized responses. Alongside a conversational dataset, HiCUPID provides a Llama-3.2-based automated evaluation model whose assessment closely mirrors human preferences. We release our dataset, evaluation model, and code at https://github.com/12kimih/HiCUPID .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c4c54b69-663e-4ffc-becd-0e10a21e8216Cited by top-tier papers5
- Your Language Model Secretly Contains Personality SubnetworksRuimeng Ye, Zihan Wang, Zinan Ling, Yang Xiao et al.ICLR 2026 · 4 citations
- Latent Inter-User Difference Modeling for LLM PersonalizationYilun Qiu, Tianhao Shi, Xiaoyan Zhao, Fengbin Zhu et al.EMNLP 2025 · 1 citation
- Contextualized Visual Personalization in Vision-Language ModelsYeongtak Oh, Sangwon Yu, Junsung Park, Han Cheol Moon et al.ICML 2026 · 1 citation
- AgentCoMa: A Compositional Benchmark Mixing Commonsense and Mathematical Reasoning in Real-World ScenariosLisa Alazraki, Lihu Chen, Ana Brassard, Joe Stacey et al.ACL 2026 · 1 citation
- HingeMem: Boundary Guided Long-Term Memory with Query Adaptive Retrieval for Scalable DialoguesYijie Zhong, Yunfan Gao, Haofen WangWWW 2026
Builds on10
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- MIND: A Large-scale Dataset for News RecommendationFangzhao Wu, Ying Qiao, Jiun-Hung Chen, Chuhan Wu et al.ACL 2020 · 454 citations
- DOC: Improving Long Story Coherence With Detailed Outline ControlKevin Yang, Dan Klein, Nanyun Peng, Yuandong TianACL 2023 · 23 citations
Related papers
- Do LLMs Recognize Your Preferences? Evaluating Personalized Preference Following in LLMsSiyan Zhao, Mingyi Hong, Yang Liu, Devamanyu Hazarika et al.ICLR 2025
- PersonalLLM: Tailoring LLMs to Individual PreferencesThomas P. Zollo, Andrew Wei Tung Siah, Naimeng Ye, Ang Li et al.ICLR 2025
- FaST: Feature-aware Sampling and Tuning for Personalized Preference Alignment with Limited DataThibaut Thonet, Germán Kruszewski, Jos Rozen, Pierre Erbacher et al.EMNLP 2025 · 1 citation
- Learning Preference Model for LLMs via Automatic Preference Data GenerationShijia Huang, Jianqiao Zhao, Yanyang Li, Liwei WangEMNLP 2023 · 3 citations
- Doing Personal LAPS: LLM-Augmented Dialogue Construction for Personalized Multi-Session Conversational SearchHideaki Joko, Shubham Chatterjee, Andrew Ramsay, Arjen P. de Vries et al.SIGIR 2024 · 24 citations
