Creative Writers' Attitudes on Writing as Training Data for Large Language Models
Katy Ilonka Gero, Meera A. Desai, Carly Schnitzler, Nayun Eom, Jack Cushman, Elena L. Glassman
Abstract
The use of creative writing as training data for large language models (LLMs) is highly contentious and many writers have expressed outrage at the use of their work without consent or compensation. In this paper, we seek to understand how creative writers reason about the real or hypothetical use of their writing as training data. We interviewed 33 writers with variation across genre, method of publishing, degree of professionalization, and attitudes toward and engagement with LLMs. We report on core principles that writers express (support of the creative chain, respect for writers and writing, and the human element of creativity) and how these principles can be at odds with their realistic expectations of the world (a lack of control, industry-scale impacts, and interpretation of scale). Collectively these findings demonstrate that writers have a nuanced understanding of LLMs and are more concerned with power imbalances than the technology itself.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8f7e83ad-eaee-42e8-b24d-82dfce168b5aCited by top-tier papers6
- Governance of AI-Generated Content: A Case Study on Social Media PlatformsLan Gao, Abani Ahmed, Oscar Chen, Margaux Reyl et al.CHI 2026 · 3 citations
- Tracing Everyday AI Literacy Discussions at Scale: How Online Creative Communities Make Sense of Generative AIHaidan Liu, Poorvi Bhatia, Nicholas Vincent, Parmit K. ChilanaCHI 2026 · 2 citations
- Can Good Writing Be Generative? Expert-Level AI Writing Emerges through Fine-Tuning on High Quality BooksTuhin Chakrabarty, Paramveer S. DhillonCHI 2026 · 2 citations
- A Design Space for Live Music AgentsYewon Kim, Stephen Brade, Alexander Wang, David Zhou et al.CHI 2026 · 2 citations
- Interpretive Cultures: Resonance, randomness, and negotiated meaning for AI-assisted tarot divinationMatthew Kieran Prock, Ziv Epstein, Hope Schroeder, Amy Smith et al.CHI 2026 · 1 citation
Builds on29
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Deduplicating Training Data Makes Language Models BetterKatherine Lee, Daphne Ippolito, Andrew Nystrom, Chiyuan Zhang et al.ACL 2022 · 844 citations
- "Everyone wants to do the model work, not the data work": Data Cascades in High-Stakes AINithya Sambasivan, Shivani Kapania, Hannah Highfill, Diana Akrong et al.CHI 2021 · 725 citations
- Deduplicating Training Data Mitigates Privacy Risks in Language ModelsNikhil Kandpal, Eric Wallace, Colin RaffelICML 2022 · 395 citations
- CoAuthor: Designing a Human-AI Collaborative Writing Dataset for Exploring Language Model CapabilitiesMina Lee, Percy Liang, Qian YangCHI 2022 · 340 citations
Related papers
- 'I'm Categorizing LLM as a Productivity Tool': Examining Ethics of LLM Use in HCI Research PracticesShivani Kapania, Ruiyi Wang, Toby Jia-Jun Li, Tianshi Li et al.CSCW 2025 · 31 citations
- The HaLLMark Effect: Supporting Provenance and Transparent Use of Large Language Models in Writing with Interactive VisualizationMd. Naimul Hoque, Tasfia Mashiat, Bhavya Ghai, Cecilia D. Shelton et al.CHI 2024 · 48 citations
- Social Dynamics of AI Support in Creative WritingKaty Ilonka Gero, Tao Long, Lydia B. ChiltonCHI 2023 · 125 citations
- Emerging Data Practices: Data Work in the Era of Large Language ModelsAdriana Alvarado Garcia, Heloisa Candello, Karla Badillo-Urquiola, Marisol Wong-VillacresCHI 2025 · 6 citations
- Governance of Generative AI in Creative Work: Consent, Credit, Compensation, and BeyondLin Kyi, Amruta Mahuli, Michael Six Silberman, Reuben Binns et al.CHI 2025 · 50 citations
