Evaluating Personalized Tool-Augmented LLMs from the Perspectives of Personalization and Proactivity
Yupu Hao, Pengfei Cao, Zhuoran Jin, Huanxuan Liao, Yubo Chen, Kang Liu, Jun Zhao
摘要
Personalized tool utilization is essential for aligning large language models (LLMs) with user preference in interaction scenarios with various tools. However, most of the current benchmarks primarily focus on either personalization of text generation or direct toolutilizing, without considering both. In this work, we introduce a novel benchmark ETAPP for evaluating personalized tool invocation, establishing a sandbox environment, and a comprehensive dataset of 800 testing cases covering diverse user profiles. To improve the accuracy of our evaluation, we propose a keypoint-based LLM evaluation method, mitigating biases in the LLM-as-a-judge system by manually annotating key points for each test case and providing them to LLM as the reference. Additionally, we evaluate the excellent LLMs and provide an in-depth analysis. Furthermore, we investigate the impact of different tool-invoking strategies on LLMs' personalization performance and the effects of finetuning in our task. The effectiveness of our preference-setting and key-point-based evaluation method is also validated. Our findings offer insights into improving personalized LLM agents. Our code is available at https://github.com/hypasd-art/ETAPP .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Talk, Evaluate, Diagnose: User-aware Agent Evaluation with Automated Error AnalysisPenny Chong, Harshavardhan Abichandani, Jiyuan Shen, Atin Ghosh 等ICLR 2026 · 被引用 4 次
- AgentSelect: Benchmark for Narrative Query-to-Agent RecommendationYunxiao Shi, Wujiang Xu, Tingwei Chen, Haoning Shang 等ICML 2026 · 被引用 3 次
- Look Light, Think Heavy: What Multimodal Chain-of-Thought Reasoning Can and Cannot DoZhuoran Jin, Kejian Zhu, Hongbang Yuan, Yupu Hao 等ACL 2026
它引用的顶会 Paper11
- TravelPlanner: A Benchmark for Real-World Planning with Language AgentsJian Xie, Kai Zhang, Jiangjie Chen, Tinghui Zhu 等ICML 2024 · 被引用 376 次
- Aligning to Thousands of Preferences via System Message GeneralizationSeongyun Lee, Sue Hyun Park, Seungone Kim, Minjoon SeoNeurIPS 2024 · 被引用 102 次
- Character-LLM: A Trainable Agent for Role-PlayingYunfan Shao, Linyang Li, Junqi Dai, Xipeng QiuEMNLP 2023 · 被引用 97 次
- API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMsMinghao Li, Yingxiu Zhao, Bowen Yu, Feifan Song 等EMNLP 2023 · 被引用 72 次
- Large Language Models Empowered Personalized Web AgentsHongru Cai, Yongqi Li, Wenjie Wang, Fengbin Zhu 等WWW 2025 · 被引用 62 次
相关 Paper
- Do LLMs Recognize Your Preferences? Evaluating Personalized Preference Following in LLMsSiyan Zhao, Mingyi Hong, Yang Liu, Devamanyu Hazarika 等ICLR 2025
- Beyond Itinerary Planning - A Real-World Benchmark for Multi-Turn and Tool-Using Travel TasksXiang Cheng, Yulan Hu, Xiangwen Zhang, Lu Xu 等ACL 2026 · 被引用 4 次
- Learning to Ask: When LLM Agents Meet Unclear InstructionWenxuan Wang, Juluan Shi, Zixuan Ling, Yuk-Kit Chan 等EMNLP 2025 · 被引用 1 次
- MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to UseYue Huang, Jiawen Shi, Yuan Li, Chenrui Fan 等ICLR 2024 · 被引用 188 次
- LaMP: When Large Language Models Meet PersonalizationAlireza Salemi, Sheshera Mysore, Michael Bendersky, Hamed ZamaniACL 2024
