arXivEdits: Understanding the Human Revision Process in Scientific Writing
Chao Jiang, Wei Xu, Samuel Stevens
Abstract
Scientific publications are the primary means to communicate research discoveries, where the writing quality is of crucial importance. However, prior work studying the human editing process in this domain mainly focused on the abstract or introduction sections, resulting in an incomplete picture. In this work, we provide a complete computational framework for studying text revision in scientific writing. We first introduce arXivEdits, a new annotated corpus of 751 full papers from arXiv with gold sentence alignment across their multiple versions of revision, as well as fine-grained span-level edits and their underlying intentions for 1,000 sentence pairs. It supports our data-driven analysis to unveil the common strategies practiced by researchers for revising their papers. To scale up the analysis, we also develop automatic methods to extract revision at document-, sentence-, and word-levels. A neural CRF sentence alignment model trained on our corpus achieves 93.8 F1, enabling the reliable matching of sentences between different versions. We formulate the edit extraction task as a span alignment problem, and our proposed method extracts more fine-grained and explainable edits, compared to the commonly used diff algorithm. An intention classifier trained on our dataset achieves 78.9 F1 on the fine-grained intent classification task. Our data and system are released at tiny.one/arxivedits.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 608605d3-7b2f-430f-8977-c9156e82fe1dCited by top-tier papers10
- Use of an AI-powered Rewriting Support Software in Context with Other Tools: A Study of Non-Native English SpeakersTakumi Ito, Naomi Yamashita, Tatsuki Kuribayashi, Masatoshi Hidaka et al.UIST 2023 · 19 citations
- Can AI writing be salvaged? Mitigating Idiosyncrasies and Improving Human-AI Alignment in the Writing Process through EditsTuhin Chakrabarty, Philippe Laban, Chien-Sheng WuCHI 2025 · 14 citations
- SWiPE: A Dataset for Document-Level Simplification of Wikipedia PagesPhilippe Laban, Jesse Vig, Wojciech Kryscinski, Shafiq Joty et al.ACL 2023 · 5 citations
- Synthia: Visually Interpreting and Synthesizing Feedback for Writing RevisionChao Zhang, Kexin Ju, Zhuolun Han, Yu-Chun Grace Yen et al.UIST 2025 · 4 citations
- ARIES: A Corpus of Scientific Paper Edits Made in Response to Peer ReviewsMike D'Arcy, Alexis Ross, Erin Bransom, Bailey Kuehl et al.ACL 2024 · 3 citations
Builds on6
- Neural CRF Model for Sentence Alignment in Text SimplificationChao Jiang, Mounica Maddela, Wuwei Lan, Yang Zhong et al.ACL 2020 · 103 citations
- A Supervised Word Alignment Method based on Cross-Language Span Prediction using Multilingual BERTMasaaki Nagata, Katsuki Chousa, Masaaki NishinoEMNLP 2020 · 38 citations
- Improving Large-scale Paraphrase Acquisition and GenerationYao Dou, Chao Jiang, Wei XuEMNLP 2022 · 11 citations
- Reformulating Unsupervised Style Transfer as Paraphrase GenerationKalpesh Krishna, John Wieting, Mohit IyyerEMNLP 2020 · 9 citations
- Neural semi-Markov CRF for Monolingual Word AlignmentWuwei Lan, Chao Jiang, Wei XuACL 2021
Related papers
- Understanding Iterative Revision from Human-Written TextWanyu Du, Vipul Raheja, Dhruv Kumar, Zae Myung Kim et al.ACL 2022
- Re3: A Holistic Framework and Dataset for Modeling Collaborative Document RevisionQian Ruan, Ilia Kuznetsov, Iryna GurevychACL 2024
- XtraGPT: Context-Aware and Controllable Academic Paper Revision via Human-AI CollaborationNuo Chen, Andre Huikai Lin, Jiaying Wu, Junyi Hou et al.ACL 2026 · 3 citations
- Improving Iterative Text Revision by Learning Where to Edit from Other Revision TasksZae Myung Kim, Wanyu Du, Vipul Raheja, Dhruv Kumar et al.EMNLP 2022 · 8 citations
- To Revise or Not to Revise: Learning to Detect Improvable Claims for Argumentative Writing SupportGabriella Skitalinskaya, Henning WachsmuthACL 2023 · 2 citations
