LCES: Zero-shot Automated Essay Scoring via Pairwise Comparisons Using Large Language Models
Takumi Shibata, Yuichi Miyamura
摘要
Recent advances in large language models (LLMs) have enabled zero-shot automated essay scoring (AES), providing a promising way to reduce the cost and effort of essay scoring in comparison with manual grading. However, most existing zero-shot approaches rely on LLMs to directly generate absolute scores, which often diverge from human evaluations owing to model biases and inconsistent scoring. To address these limitations, we propose LLMbased Comparative Essay Scoring (LCES), a method that formulates AES as a pairwise comparison task. Specifically, we instruct LLMs to judge which of two essays is better, collect many such comparisons, and convert them into continuous scores. Considering that the number of possible comparisons grows quadratically with the number of essays, we improve scalability by employing RankNet to efficiently transform LLM preferences into scalar scores. Experiments using AES benchmark datasets show that LCES outperforms conventional zero-shot methods in accuracy while maintaining computational efficiency. Moreover, LCES is robust across different LLM backbones, highlighting its applicability to real-world zero-shot AES.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper7
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo 等NeurIPS 2022 · 被引用 8,168 次
- Automated Cross-prompt Scoring of Essay TraitsRobert Ridley, Liang He, Xin-Yu Dai, Shujian Huang 等AAAI 2021 · 被引用 101 次
- From Generation to Judgment: Opportunities and Challenges of LLM-as-a-judgeDawei Li, Bohan Jiang, Liangjie Huang, Alimohammad Beigi 等EMNLP 2025 · 被引用 37 次
- PMAES: Prompt-mapping Contrastive Learning for Cross-prompt Automated Essay ScoringYuan Chen, Xia LiACL 2023 · 被引用 20 次
相关 Paper
- A Setwise Approach for Effective and Highly Efficient Zero-shot Ranking with Large Language ModelsShengyao Zhuang, Honglei Zhuang, Bevan Koopman, Guido ZucconSIGIR 2024 · 被引用 60 次
- Activations as Features: Probing LLMs for Generalizable Essay Scoring RepresentationsJinwei Chi, Ke Wang, Yu Chen, Xuanye Lin 等AAAI 2026 · 被引用 1 次
- KAES: Multi-aspect Shared Knowledge Finding and Aligning for Cross-prompt Automated Scoring of Essay TraitsXia Li, Wenjing PanAAAI 2025 · 被引用 5 次
- Aggregating Multiple Heuristic Signals as Supervision for Unsupervised Automated Essay ScoringCong Wang, Zhiwei Jiang, Yafeng Yin, Zifeng Cheng 等ACL 2023 · 被引用 4 次
- Precise Zero-Shot Pointwise Ranking with LLMs through Post-Aggregated Global Context InformationKehan Long, Shasha Li, Chen Xu, Jintao Tang 等SIGIR 2025
