ChatGPT to Replace Crowdsourcing of Paraphrases for Intent Classification: Higher Diversity and Comparable Model Robustness
Ján Cegin, Jakub Simko, Peter Brusilovsky
Abstract
<p>The emergence of generative large language models (LLMs) raises the question: what will be its impact on crowdsourcing? Traditionally, crowdsourcing has been used for acquiring solutions to a wide variety of human-intelligence tasks, including ones involving text generation, modification or evaluation. For some of these tasks, models like ChatGPT can potentially substitute human workers. In this study, we investigate whether this is the case for the task of paraphrase generation for intent classification. We apply data collection methodology of an existing crowdsourcing study (similar scale, prompts and seed data) using ChatGPT and Falcon-40B. We show that ChatGPT-created paraphrases are more diverse and lead to at least as robust models.</p>
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e2b9176a-aff3-4fb5-bd25-9a05b44677ffCited by top-tier papers10
- If in a Crowdsourced Data Annotation Pipeline, a GPT-4Zeyu He, Chieh-Yang Huang, Chien-Kuang Cornelia Ding, Shaurya Rohatgi et al.CHI 2024 · 31 citations
- Diversity-oriented Data Augmentation with Large Language ModelsZaitian Wang, Jinghan Zhang, Xinhao Zhang, Kunpeng Liu et al.ACL 2025 · 11 citations
- Effects of diversity incentives on sample diversity and downstream model performance in LLM-based text augmentationJán Cegin, Branislav Pecher, Jakub Simko, Ivan Srba et al.ACL 2024 · 5 citations
- Evaluating LLM-contaminated Crowdsourcing Data Without Ground TruthYichi Zhang, Jinlong Pang, Zhaowei Zhu, Yang LiuNeurIPS 2025 · 3 citations
- What Is Wrong with My Model? Identifying Systematic Problems with Semantic Data SlicingChenyang Yang, Yining Hong, Grace A. Lewis, Tongshuang Wu et al.ASE 2024 · 2 citations
Builds on7
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- Is ChatGPT a General-Purpose Natural Language Processing Task Solver?Chengwei Qin, Aston Zhang, Zhuosheng Zhang, Jiaao Chen et al.EMNLP 2023 · 449 citations
- Novelty Controlled Paraphrase Generation with Retrieval Augmented Conditional Prompt TuningJishnu Ray Chowdhury, Yong Zhuang, Shuyi WangAAAI 2022 · 39 citations
- Directed Diversity: Leveraging Language Embedding Distances for Collective Creativity in Crowd IdeationSamuel Rhys Cox, Yunlong Wang, Ashraf M. Abdul, Christian von der Weth et al.CHI 2021 · 24 citations
- Neural Syntactic Preordering for Controlled Paraphrase GenerationTanya Goyal, Greg DurrettACL 2020 · 9 citations
Related papers
- Evaluating Large Language Models in Generating Synthetic HCI Research Data: a Case StudyPerttu Hämäläinen, Mikke Tavast, Anton KunnariCHI 2023 · 244 citations
- Safeguarding Crowdsourcing Surveys from ChatGPT through Prompt InjectionChaofan Wang, Samuel Kernan Freire, Mo Zhang, Jing Wei et al.CSCW 2025 · 1 citation
- Current and Future Use of Large Language Models for Knowledge WorkMichelle Brachman, Amina H. El-Ashry, Casey Dugan, Werner GeyerCSCW 2025 · 10 citations
- Human-LLM Hybrid Text Answer Aggregation for Crowd AnnotationsJiyi LiEMNLP 2024 · 1 citation
- Monitoring AI-Modified Content at Scale: A Case Study on the Impact of ChatGPT on AI Conference Peer ReviewsWeixin Liang, Zachary Izzo, Yaohui Zhang, Haley Lepp et al.ICML 2024 · 213 citations
