Learning from Scoring Disagreements: Contrastive Error Mining for Efficient and Robust LLM-based Assessment
Lei Chen, Tengteng Cheng, Boyu Gao, Zitao Liu, Weiqi Luo
摘要
Automated grading of student responses still faces numerous challenges, particularly when dealing with complex and ambiguous answers. In particular, large models are prone to scoring bias when handling uncertain responses, and few-shot reasoning methods often lack stability, which limits their applicability in real educational scenarios. To tackle these challenges, we propose the Contrastive Error Mining and Fine-Tuning (CEM-FT) framework, which automatically identifies high-value hard samples by analyzing scoring disagreements between a full fine-tuned model and a few-shot model. A lightweight LoRA adapter is then trained on these samples to refine model performance with minimal computational overhead. Experiments on the SciEntsbank, Beetle, and Mohler datasets show that CEM-FT can improve QWK by up to 3.9% compared to the fine-tuned Qwen model on SciEntsbank datasets, which is a significant improvement over the few-shot baseline. The proposed framework substantially enhances both scoring accuracy and consistency, providing a practical, robust solution for reliable automated assessment with large language models.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper4
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- DeBERTaV3: Improving DeBERTa using ELECTRA-Style Pre-Training with Gradient-Disentangled Embedding SharingPengcheng He, Jianfeng Gao, Weizhu ChenICLR 2023 · 被引用 394 次
- Full Parameter Fine-tuning for Large Language Models with Limited ResourcesKai Lv, Yuqing Yang, Tengxiao Liu, Qipeng Guo 等ACL 2024 · 被引用 61 次
相关 Paper
- C-LoRA: Contextual Low-Rank Adaptation for Uncertainty Estimation in Large Language ModelsAmir Hossein Rahmati, Sanket R. Jantre, Weifeng Zhang, Yucheng Wang 等NeurIPS 2025 · 被引用 11 次
- Beyond Two-Stage Training: Cooperative SFT and RL for LLM ReasoningLiang Chen, Xueting Han, Li Shen, Jing Bai 等ICML 2026 · 被引用 24 次
- Automatically Generating Numerous Context-Driven SFT Data for LLMs Across Diverse GranularityShanghaoran QuanAAAI 2025 · 被引用 7 次
- LightReasoner: Can Small Language Models Teach Large Language Models Reasoning?Jingyuan Wang, Yankai Chen, Zhonghang Li, Chao HuangACL 2026
- Hard Sample Aware Prompt-TuningYuanjian Xu, Qi An, Jiahuan Zhang, Peng Li 等ACL 2023 · 被引用 4 次
