Are Large Language Models Good Classifiers? A Study on Edit Intent Classification in Scientific Document Revisions
Qian Ruan, Ilia Kuznetsov, Iryna Gurevych
摘要
Classification is a core NLP task architecture with many potential applications. While large language models (LLMs) have brought substantial advancements in text generation, their potential for enhancing classification tasks remains underexplored. To address this gap, we propose a framework for thoroughly investigating fine-tuning LLMs for classification, including both generation-and encoding-based approaches. We instantiate this framework in edit intent classification (EIC), a challenging and underexplored classification task. Our extensive experiments and systematic comparisons with various training approaches and a representative selection of LLMs yield new insights into their application for EIC. We investigate the generalizability of these findings on five further classification tasks. To demonstrate the proposed methods and address the data shortage for empirical edit analysis, we use our bestperforming EIC model to create Re3-Sci2.0, a new large-scale dataset of 1,780 scientific document revisions with over 94k labeled edits. The quality of the dataset is assessed through human evaluation. The new dataset enables an in-depth empirical study of human editing behavior in academic writing. We make our experimental framework 1 , models and data 2 publicly available. Revision Re3-Sci2.0 x1780 scientific papers
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper8
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 被引用 5,863 次
- TruthfulQA: Measuring How Models Mimic Human FalsehoodsStephanie Lin, Jacob Hilton, Owain EvansACL 2022 · 被引用 3,228 次
- Is ChatGPT a General-Purpose Natural Language Processing Task Solver?Chengwei Qin, Aston Zhang, Zhuosheng Zhang, Jiaao Chen 等EMNLP 2023 · 被引用 449 次
- Element-aware Summarization with Large Language Models: Expert-aligned Evaluation and Chain-of-Thought MethodYiming Wang, Zhuosheng Zhang, Rui WangACL 2023 · 被引用 39 次
- NLPeer: A Unified Resource for the Computational Study of Peer ReviewNils Dycke, Ilia Kuznetsov, Iryna GurevychACL 2023 · 被引用 17 次
相关 Paper
- Re3: A Holistic Framework and Dataset for Modeling Collaborative Document RevisionQian Ruan, Ilia Kuznetsov, Iryna GurevychACL 2024
- Intention is All you Need: Refining your Code from your IntentionQi Guo, Xiaofei Xie, Shangqing Liu, Ming Hu 等ICSE 2025 · 被引用 7 次
- Grace: Language Models Meet Code EditsPriyanshu Gupta, Avishree Khare, Yasharth Bajpai, Saikat Chakraborty 等FSE 2023 · 被引用 15 次
- No One-Size-Fits-All: Adaptive Code Editing with Feature-Based Strategy SelectionJun Wan, Zhongxin Liu, Dajun Chen, Wei Jiang 等ISSTA 2026
- XtraGPT: Context-Aware and Controllable Academic Paper Revision via Human-AI CollaborationNuo Chen, Andre Huikai Lin, Jiaying Wu, Junyi Hou 等ACL 2026 · 被引用 3 次
