arXivEdits: Understanding the Human Revision Process in Scientific Writing
Chao Jiang, Wei Xu, Samuel Stevens
摘要
Scientific publications are the primary means to communicate research discoveries, where the writing quality is of crucial importance. However, prior work studying the human editing process in this domain mainly focused on the abstract or introduction sections, resulting in an incomplete picture. In this work, we provide a complete computational framework for studying text revision in scientific writing. We first introduce arXivEdits, a new annotated corpus of 751 full papers from arXiv with gold sentence alignment across their multiple versions of revision, as well as fine-grained span-level edits and their underlying intentions for 1,000 sentence pairs. It supports our data-driven analysis to unveil the common strategies practiced by researchers for revising their papers. To scale up the analysis, we also develop automatic methods to extract revision at document-, sentence-, and word-levels. A neural CRF sentence alignment model trained on our corpus achieves 93.8 F1, enabling the reliable matching of sentences between different versions. We formulate the edit extraction task as a span alignment problem, and our proposed method extracts more fine-grained and explainable edits, compared to the commonly used diff algorithm. An intention classifier trained on our dataset achieves 78.9 F1 on the fine-grained intent classification task. Our data and system are released at tiny.one/arxivedits.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Use of an AI-powered Rewriting Support Software in Context with Other Tools: A Study of Non-Native English SpeakersTakumi Ito, Naomi Yamashita, Tatsuki Kuribayashi, Masatoshi Hidaka 等UIST 2023 · 被引用 19 次
- Can AI writing be salvaged? Mitigating Idiosyncrasies and Improving Human-AI Alignment in the Writing Process through EditsTuhin Chakrabarty, Philippe Laban, Chien-Sheng WuCHI 2025 · 被引用 14 次
- SWiPE: A Dataset for Document-Level Simplification of Wikipedia PagesPhilippe Laban, Jesse Vig, Wojciech Kryscinski, Shafiq Joty 等ACL 2023 · 被引用 5 次
- Synthia: Visually Interpreting and Synthesizing Feedback for Writing RevisionChao Zhang, Kexin Ju, Zhuolun Han, Yu-Chun Grace Yen 等UIST 2025 · 被引用 4 次
- ARIES: A Corpus of Scientific Paper Edits Made in Response to Peer ReviewsMike D'Arcy, Alexis Ross, Erin Bransom, Bailey Kuehl 等ACL 2024 · 被引用 3 次
它引用的顶会 Paper6
- Neural CRF Model for Sentence Alignment in Text SimplificationChao Jiang, Mounica Maddela, Wuwei Lan, Yang Zhong 等ACL 2020 · 被引用 103 次
- A Supervised Word Alignment Method based on Cross-Language Span Prediction using Multilingual BERTMasaaki Nagata, Katsuki Chousa, Masaaki NishinoEMNLP 2020 · 被引用 38 次
- Improving Large-scale Paraphrase Acquisition and GenerationYao Dou, Chao Jiang, Wei XuEMNLP 2022 · 被引用 11 次
- Reformulating Unsupervised Style Transfer as Paraphrase GenerationKalpesh Krishna, John Wieting, Mohit IyyerEMNLP 2020 · 被引用 9 次
- Neural semi-Markov CRF for Monolingual Word AlignmentWuwei Lan, Chao Jiang, Wei XuACL 2021
相关 Paper
- Understanding Iterative Revision from Human-Written TextWanyu Du, Vipul Raheja, Dhruv Kumar, Zae Myung Kim 等ACL 2022
- Re3: A Holistic Framework and Dataset for Modeling Collaborative Document RevisionQian Ruan, Ilia Kuznetsov, Iryna GurevychACL 2024
- XtraGPT: Context-Aware and Controllable Academic Paper Revision via Human-AI CollaborationNuo Chen, Andre Huikai Lin, Jiaying Wu, Junyi Hou 等ACL 2026 · 被引用 3 次
- Improving Iterative Text Revision by Learning Where to Edit from Other Revision TasksZae Myung Kim, Wanyu Du, Vipul Raheja, Dhruv Kumar 等EMNLP 2022 · 被引用 8 次
- To Revise or Not to Revise: Learning to Detect Improvable Claims for Argumentative Writing SupportGabriella Skitalinskaya, Henning WachsmuthACL 2023 · 被引用 2 次
