ScholaWrite: A Dataset of End-to-End Scholarly Writing
Khanh Chi Le, Linghe Wang, Minhwa Lee, Ross Volkov, Luan Tuyen Chau, Dongyeop Kang
Abstract
Writing is a cognitively demanding activity that requires constant decision-making, heavy reliance on working memory, and frequent shifts between tasks of different goals. To build writing assistants that truly align with writers' cognition, it is necessary to capture and analyze the complete thought process behind how writers transform ideas into final texts. We present ScholaWrite, the first dataset of end-to-end scholarly writing, tracing the multi-month journey from initial drafts to final manuscripts. The dataset traces nearly 62K L A T E X-based edits from five computer science preprints over four months and is enriched with fine-grained annotations of cognitive writing intentions. We demonstrate the value of ScholaWrite through three complementary contributions: (1) analysis of real-world writing behavior reveals that scholarly writing is highly non-linear and multiintentional, blending rapid drafting bursts with cognitively sustained writing sessions; (2) evaluations of current large language models show that they struggle to provide meaningful support throughout the human writing process; and (3) models finetuned on SCHOLAWRITE demonstrate improved alignment with human writing workflows. SCHOLAWRITE underscores the value of capturing scientists' cognitive writing process and provides actionable insights and resources for the development of future writing assistants. All data and tools are available on our project page. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 23863a26-02fa-406e-89fd-d8ad4015ce8cCited by top-tier papers1
Ask how each one uses itBuilds on8
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 5,863 citations
- How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMsYi Zeng, Hongpeng Lin, Jingwen Zhang, Diyi Yang et al.ACL 2024 · 64 citations
- arXivEdits: Understanding the Human Revision Process in Scientific WritingChao Jiang, Wei Xu, Samuel StevensEMNLP 2022 · 10 citations
- Latxa: An Open Language Model and Evaluation Suite for BasqueJulen Etxaniz, Oscar Sainz, Naiara Miguel, Itziar Aldabe et al.ACL 2024 · 5 citations
Related papers
- Prototypical Human-AI Collaboration Behaviors from LLM-Assisted Writing in the WildSheshera Mysore, Debarati Das, Hancheng Cao, Bahareh SarrafzadehEMNLP 2025 · 3 citations
- XtraGPT: Context-Aware and Controllable Academic Paper Revision via Human-AI CollaborationNuo Chen, Andre Huikai Lin, Jiaying Wu, Junyi Hou et al.ACL 2026 · 3 citations
- Characterizing Stage-aware Writing Assistance for Collaborative Document AuthoringBahareh Sarrafzadeh, Sujay Kumar Jauhar, Michael Gamon, Edward Lank et al.CSCW 2020 · 12 citations
- CoAuthor: Designing a Human-AI Collaborative Writing Dataset for Exploring Language Model CapabilitiesMina Lee, Percy Liang, Qian YangCHI 2022 · 340 citations
- IGA: An Intent-Guided Authoring AssistantSimeng Sun, Wenlong Zhao, Varun Manjunatha, Rajiv Jain et al.EMNLP 2021
