Learning Personalized Alignment for Evaluating Open-ended Text Generation
Danqing Wang, Kevin Yang, Hanlin Zhu, Xiaomeng Yang, Andrew Cohen, Lei Li, Yuandong Tian
Abstract
Recent research has increasingly focused on evaluating large language models' (LLMs) alignment with diverse human values and preferences, particularly for open-ended tasks like story generation. Traditional evaluation metrics rely heavily on lexical similarity with humanwritten references, often showing poor correlation with human judgments and failing to account for alignment with the diversity of human preferences. To address these challenges, we introduce PERSE, an interpretable evaluation framework designed to assess alignment with specific human preferences. It is tuned to infer specific preferences from an in-context personal profile and evaluate the alignment between the generated content and personal preferences. PERSE enhances interpretability by providing detailed comments and fine-grained scoring, facilitating more personalized content generation. Our 13B LLaMA-2-based PERSE shows a 15.8% increase in Kendall correlation and a 13.7% rise in accuracy with zero-shot reviewers compared to GPT-4. It also outperforms GPT-4 by 46.01% in Kendall correlation on new domains, indicating its transferability 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 017c5a5f-de1a-4b25-89f3-8303299bec35Cited by top-tier papers6
- Towards Personalized Deep Research: Benchmarks and EvaluationsYuan Liang, Jiaxian Li, Yuqing Wang, Piaohong Wang et al.ICLR 2026 · 13 citations
- Instant Personalized Large Language Model Adaptation via HypernetworkZhaoxuan Tan, Zixuan Zhang, Haoyang Wen, Zheng Li et al.ACL 2026 · 7 citations
- P-GenRM: Personalized Generative Reward Model with Test-time User-based ScalingPinyi Zhang, Ting-En Lin, Yuchuan Wu, Jingyang Chen et al.ICLR 2026 · 5 citations
- EvolvR: Self-Evolving Pairwise Reasoning for Story Evaluation to Enhance GenerationXinda Wang, Zhengxu Hou, Yangshijie Zhang, Bingren Yan et al.ACL 2026 · 1 citation
- Exploring Persona Sentiment Sensitivity in Personalized Dialogue GenerationYonghyun Jun, Hwanhee LeeACL 2025
Builds on16
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- Fantastically Ordered Prompts and Where to Find Them: Overcoming Few-Shot Prompt Order SensitivityYao Lu, Max Bartolo, Alastair Moore, Sebastian Riedel et al.ACL 2022 · 1,494 citations
- BARTScore: Evaluating Generated Text as Text GenerationWeizhe Yuan, Graham Neubig, Pengfei LiuNeurIPS 2021 · 1,143 citations
- Re3: Generating Longer Stories With Recursive Reprompting and RevisionKevin Yang, Yuandong Tian, Nanyun Peng, Dan KleinEMNLP 2022 · 77 citations
Related papers
- On Evaluating LLM Alignment by Evaluating LLMs as JudgesYixin Liu, Pengfei Liu, Arman CohanNeurIPS 2025 · 7 citations
- Themis: A Reference-free NLG Evaluation Language Model with Flexibility and InterpretabilityXinyu Hu, Li Lin, Mingqi Gao, Xunjian Yin et al.EMNLP 2024 · 2 citations
- Exploring Precision and Recall to assess the quality and diversity of LLMsFlorian Le Bronnec, Alexandre Verine, Benjamin Négrevergne, Yann Chevaleyre et al.ACL 2024 · 11 citations
- Co-Eval: Augmenting LLM-based Evaluation with Machine MetricsLing-I Wu, Weijie Wu, Minyu Chen, Jianxin Xue et al.EMNLP 2025
- FLASK: Fine-grained Language Model Evaluation based on Alignment Skill SetsSeonghyeon Ye, Doyoung Kim, Sungdong Kim, Hyeonbin Hwang et al.ICLR 2024 · 176 citations
