Iterative Dual-Model Alignment for Story Evaluation
Bruce Qin, Dan Goldwasser
摘要
Large language models (LLMs) can both evaluate and explain text quality; however, most existing evaluators operate as static classifiers and lack the ability to refine their reasoning through interaction. We propose an Iterative Alpha-Beta Learning framework that jointly trains two complementary 8B models: an Alpha (α) classifier that assesses pairwise story engagement, and a Beta (β) generator that produces structured, rubric-guided comparative explanations. The two models co-evolve within a closed feedback loop: α provides probabilistic preference signals to guide β's Direct Preference Optimization (DPO), while β's improved explanations are reintegrated to retrain α via a KL-based contrastive objective. This dual optimization enables mutual learning: α gains interpretability and robustness from β's textual rationales, while β acquires stronger alignment and discriminative precision from α's confidence deltas. Experiments on human-annotated storypair datasets (HANNA) show that the proposed system consistently outperforms strong singlemodel baselines in both accuracy and explanation quality across multiple iterative rounds 1 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper9
- G-Eval: NLG Evaluation using Gpt-4 with Better Human AlignmentYang Liu, Dan Iter, Yichong Xu, Shuohang Wang 等EMNLP 2023 · 被引用 549 次
- CoAuthor: Designing a Human-AI Collaborative Writing Dataset for Exploring Language Model CapabilitiesMina Lee, Percy Liang, Qian YangCHI 2022 · 被引用 340 次
- Social Dynamics of AI Support in Creative WritingKaty Ilonka Gero, Tao Long, Lydia B. ChiltonCHI 2023 · 被引用 125 次
- Shaping Human-AI Collaboration: Varied Scaffolding Levels in Co-writing with Language ModelsParamveer S. Dhillon, Somayeh Molaei, Jiaqi Li, Maximilian Golub 等CHI 2024 · 被引用 102 次
- UNION: An Unreferenced Metric for Evaluating Open-ended Story GenerationJian Guan, Minlie HuangEMNLP 2020 · 被引用 46 次
相关 Paper
- EvolvR: Self-Evolving Pairwise Reasoning for Story Evaluation to Enhance GenerationXinda Wang, Zhengxu Hou, Yangshijie Zhang, Bingren Yan 等ACL 2026 · 被引用 1 次
- iRULER: Intelligible Rubric-Based User-Defined LLM Evaluation for RevisionJingwen Bai, Wei Soon Cheong, Philippe Muller, Brian Y. LimCHI 2026 · 被引用 1 次
- HD-Eval: Aligning Large Language Model Evaluators Through Hierarchical Criteria DecompositionYuxuan Liu, Tianchi Yang, Shaohan Huang, Zihan Zhang 等ACL 2024 · 被引用 1 次
- Think Wise, Collaborate Effectively: A Rationale-Aware LLM-Based Recommender with Reinforcement Learning from Collaborative SignalsChung Park, Taesan Kim, Hyeongjun Yun, Dongjoon Hong 等AAAI 2026
- Co-Evolving LLMs and Embedding Models via Density-Guided Preference Optimization for Text ClusteringZetong Li, Qinliang Su, Minhua Huang, Yin YangEMNLP 2025
