Creative Writers' Attitudes on Writing as Training Data for Large Language Models
Katy Ilonka Gero, Meera A. Desai, Carly Schnitzler, Nayun Eom, Jack Cushman, Elena L. Glassman
摘要
The use of creative writing as training data for large language models (LLMs) is highly contentious and many writers have expressed outrage at the use of their work without consent or compensation. In this paper, we seek to understand how creative writers reason about the real or hypothetical use of their writing as training data. We interviewed 33 writers with variation across genre, method of publishing, degree of professionalization, and attitudes toward and engagement with LLMs. We report on core principles that writers express (support of the creative chain, respect for writers and writing, and the human element of creativity) and how these principles can be at odds with their realistic expectations of the world (a lack of control, industry-scale impacts, and interpretation of scale). Collectively these findings demonstrate that writers have a nuanced understanding of LLMs and are more concerned with power imbalances than the technology itself.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Governance of AI-Generated Content: A Case Study on Social Media PlatformsLan Gao, Abani Ahmed, Oscar Chen, Margaux Reyl 等CHI 2026 · 被引用 3 次
- Tracing Everyday AI Literacy Discussions at Scale: How Online Creative Communities Make Sense of Generative AIHaidan Liu, Poorvi Bhatia, Nicholas Vincent, Parmit K. ChilanaCHI 2026 · 被引用 2 次
- Can Good Writing Be Generative? Expert-Level AI Writing Emerges through Fine-Tuning on High Quality BooksTuhin Chakrabarty, Paramveer S. DhillonCHI 2026 · 被引用 2 次
- A Design Space for Live Music AgentsYewon Kim, Stephen Brade, Alexander Wang, David Zhou 等CHI 2026 · 被引用 2 次
- Interpretive Cultures: Resonance, randomness, and negotiated meaning for AI-assisted tarot divinationMatthew Kieran Prock, Ziv Epstein, Hope Schroeder, Amy Smith 等CHI 2026 · 被引用 1 次
它引用的顶会 Paper29
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Deduplicating Training Data Makes Language Models BetterKatherine Lee, Daphne Ippolito, Andrew Nystrom, Chiyuan Zhang 等ACL 2022 · 被引用 844 次
- "Everyone wants to do the model work, not the data work": Data Cascades in High-Stakes AINithya Sambasivan, Shivani Kapania, Hannah Highfill, Diana Akrong 等CHI 2021 · 被引用 725 次
- Deduplicating Training Data Mitigates Privacy Risks in Language ModelsNikhil Kandpal, Eric Wallace, Colin RaffelICML 2022 · 被引用 395 次
- CoAuthor: Designing a Human-AI Collaborative Writing Dataset for Exploring Language Model CapabilitiesMina Lee, Percy Liang, Qian YangCHI 2022 · 被引用 340 次
相关 Paper
- 'I'm Categorizing LLM as a Productivity Tool': Examining Ethics of LLM Use in HCI Research PracticesShivani Kapania, Ruiyi Wang, Toby Jia-Jun Li, Tianshi Li 等CSCW 2025 · 被引用 31 次
- The HaLLMark Effect: Supporting Provenance and Transparent Use of Large Language Models in Writing with Interactive VisualizationMd. Naimul Hoque, Tasfia Mashiat, Bhavya Ghai, Cecilia D. Shelton 等CHI 2024 · 被引用 48 次
- Social Dynamics of AI Support in Creative WritingKaty Ilonka Gero, Tao Long, Lydia B. ChiltonCHI 2023 · 被引用 125 次
- Emerging Data Practices: Data Work in the Era of Large Language ModelsAdriana Alvarado Garcia, Heloisa Candello, Karla Badillo-Urquiola, Marisol Wong-VillacresCHI 2025 · 被引用 6 次
- Governance of Generative AI in Creative Work: Consent, Credit, Compensation, and BeyondLin Kyi, Amruta Mahuli, Michael Six Silberman, Reuben Binns 等CHI 2025 · 被引用 50 次
