Iterative Dual-Model Alignment for Story Evaluation
Bruce Qin, Dan Goldwasser
Abstract
Large language models (LLMs) can both evaluate and explain text quality; however, most existing evaluators operate as static classifiers and lack the ability to refine their reasoning through interaction. We propose an Iterative Alpha-Beta Learning framework that jointly trains two complementary 8B models: an Alpha (α) classifier that assesses pairwise story engagement, and a Beta (β) generator that produces structured, rubric-guided comparative explanations. The two models co-evolve within a closed feedback loop: α provides probabilistic preference signals to guide β's Direct Preference Optimization (DPO), while β's improved explanations are reintegrated to retrain α via a KL-based contrastive objective. This dual optimization enables mutual learning: α gains interpretability and robustness from β's textual rationales, while β acquires stronger alignment and discriminative precision from α's confidence deltas. Experiments on human-annotated storypair datasets (HANNA) show that the proposed system consistently outperforms strong singlemodel baselines in both accuracy and explanation quality across multiple iterative rounds 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on9
- G-Eval: NLG Evaluation using Gpt-4 with Better Human AlignmentYang Liu, Dan Iter, Yichong Xu, Shuohang Wang et al.EMNLP 2023 · 549 citations
- CoAuthor: Designing a Human-AI Collaborative Writing Dataset for Exploring Language Model CapabilitiesMina Lee, Percy Liang, Qian YangCHI 2022 · 340 citations
- Social Dynamics of AI Support in Creative WritingKaty Ilonka Gero, Tao Long, Lydia B. ChiltonCHI 2023 · 125 citations
- Shaping Human-AI Collaboration: Varied Scaffolding Levels in Co-writing with Language ModelsParamveer S. Dhillon, Somayeh Molaei, Jiaqi Li, Maximilian Golub et al.CHI 2024 · 102 citations
- UNION: An Unreferenced Metric for Evaluating Open-ended Story GenerationJian Guan, Minlie HuangEMNLP 2020 · 46 citations
Related papers
- EvolvR: Self-Evolving Pairwise Reasoning for Story Evaluation to Enhance GenerationXinda Wang, Zhengxu Hou, Yangshijie Zhang, Bingren Yan et al.ACL 2026 · 1 citation
- iRULER: Intelligible Rubric-Based User-Defined LLM Evaluation for RevisionJingwen Bai, Wei Soon Cheong, Philippe Muller, Brian Y. LimCHI 2026 · 1 citation
- HD-Eval: Aligning Large Language Model Evaluators Through Hierarchical Criteria DecompositionYuxuan Liu, Tianchi Yang, Shaohan Huang, Zihan Zhang et al.ACL 2024 · 1 citation
- Think Wise, Collaborate Effectively: A Rationale-Aware LLM-Based Recommender with Reinforcement Learning from Collaborative SignalsChung Park, Taesan Kim, Hyeongjun Yun, Dongjoon Hong et al.AAAI 2026
- Co-Evolving LLMs and Embedding Models via Density-Guided Preference Optimization for Text ClusteringZetong Li, Qinliang Su, Minhua Huang, Yin YangEMNLP 2025
