Learning Personalized Alignment for Evaluating Open-ended Text Generation
Danqing Wang, Kevin Yang, Hanlin Zhu, Xiaomeng Yang, Andrew Cohen, Lei Li, Yuandong Tian
摘要
Recent research has increasingly focused on evaluating large language models' (LLMs) alignment with diverse human values and preferences, particularly for open-ended tasks like story generation. Traditional evaluation metrics rely heavily on lexical similarity with humanwritten references, often showing poor correlation with human judgments and failing to account for alignment with the diversity of human preferences. To address these challenges, we introduce PERSE, an interpretable evaluation framework designed to assess alignment with specific human preferences. It is tuned to infer specific preferences from an in-context personal profile and evaluate the alignment between the generated content and personal preferences. PERSE enhances interpretability by providing detailed comments and fine-grained scoring, facilitating more personalized content generation. Our 13B LLaMA-2-based PERSE shows a 15.8% increase in Kendall correlation and a 13.7% rise in accuracy with zero-shot reviewers compared to GPT-4. It also outperforms GPT-4 by 46.01% in Kendall correlation on new domains, indicating its transferability 1 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Towards Personalized Deep Research: Benchmarks and EvaluationsYuan Liang, Jiaxian Li, Yuqing Wang, Piaohong Wang 等ICLR 2026 · 被引用 13 次
- Instant Personalized Large Language Model Adaptation via HypernetworkZhaoxuan Tan, Zixuan Zhang, Haoyang Wen, Zheng Li 等ACL 2026 · 被引用 7 次
- P-GenRM: Personalized Generative Reward Model with Test-time User-based ScalingPinyi Zhang, Ting-En Lin, Yuchuan Wu, Jingyang Chen 等ICLR 2026 · 被引用 5 次
- EvolvR: Self-Evolving Pairwise Reasoning for Story Evaluation to Enhance GenerationXinda Wang, Zhengxu Hou, Yangshijie Zhang, Bingren Yan 等ACL 2026 · 被引用 1 次
- Exploring Persona Sentiment Sensitivity in Personalized Dialogue GenerationYonghyun Jun, Hwanhee LeeACL 2025
它引用的顶会 Paper16
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger 等ICLR 2020 · 被引用 8,443 次
- Fantastically Ordered Prompts and Where to Find Them: Overcoming Few-Shot Prompt Order SensitivityYao Lu, Max Bartolo, Alastair Moore, Sebastian Riedel 等ACL 2022 · 被引用 1,494 次
- BARTScore: Evaluating Generated Text as Text GenerationWeizhe Yuan, Graham Neubig, Pengfei LiuNeurIPS 2021 · 被引用 1,143 次
- Re3: Generating Longer Stories With Recursive Reprompting and RevisionKevin Yang, Yuandong Tian, Nanyun Peng, Dan KleinEMNLP 2022 · 被引用 77 次
相关 Paper
- On Evaluating LLM Alignment by Evaluating LLMs as JudgesYixin Liu, Pengfei Liu, Arman CohanNeurIPS 2025 · 被引用 7 次
- Themis: A Reference-free NLG Evaluation Language Model with Flexibility and InterpretabilityXinyu Hu, Li Lin, Mingqi Gao, Xunjian Yin 等EMNLP 2024 · 被引用 2 次
- Exploring Precision and Recall to assess the quality and diversity of LLMsFlorian Le Bronnec, Alexandre Verine, Benjamin Négrevergne, Yann Chevaleyre 等ACL 2024 · 被引用 11 次
- Co-Eval: Augmenting LLM-based Evaluation with Machine MetricsLing-I Wu, Weijie Wu, Minyu Chen, Jianxin Xue 等EMNLP 2025
- FLASK: Fine-grained Language Model Evaluation based on Alignment Skill SetsSeonghyeon Ye, Doyoung Kim, Sungdong Kim, Hyeonbin Hwang 等ICLR 2024 · 被引用 176 次
