On the Detectability of ChatGPT Content: Benchmarking, Methodology, and Evaluation through the Lens of Academic Writing
Zeyan Liu, Zijun Yao, Fengjun Li, Bo Luo
摘要
With ChatGPT under the spotlight, utilizing large language models (LLMs) to assist academic writing has drawn a significant amount of debate in the community. In this paper, we aim to present a comprehensive study of the detectability of ChatGPT-generated content within the academic literature, particularly focusing on the abstracts of scientific papers, to offer holistic support for the future development of LLM applications and policies in academia. Specifically, we first present GPABench2, a benchmarking dataset of over 2.8 million comparative samples of human-written, GPT-written, GPT-completed, and GPT-polished abstracts of scientific writing in computer science, physics, and humanities and social sciences. Second, we explore the methodology for detecting ChatGPT content. We start by examining the unsatisfactory performance of existing ChatGPT detecting tools and the challenges faced by human evaluators (including more than 240 researchers or students). We then test the hand-crafted linguistic features models as a baseline and develop a deep neural framework named CheckGPT to better capture the subtle and deep semantic and linguistic patterns in ChatGPT written literature. Last, we conduct comprehensive experiments to validate the proposed CheckGPT framework in each benchmarking task over different disciplines. To evaluate the detectability of ChatGPT content, we conduct extensive experiments on the transferability, prompt engineering, and robustness of CheckGPT.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- An Empirical Study to Understand How Students Use ChatGPT for Writing EssaysAndrew Jelson, Daniel Manesh, Alice Jang, Daniel Dunlap 等CHI 2026 · 被引用 3 次
- Multi-level Style Preference Optimization: An Adaptive Detection Framework for Human-Machine Hybrid TextZehao Wang, Lianwei Wu, Wenbo An, Hang Zhang 等AAAI 2026
- Discourse Realization of Generics in Human and LLM-generated TextsSøren Kirkegaard Fomsgaard, Martial Pastor, Gaël Dias, Nelleke OostdijkACL 2026
- OSTAR: Optimized Statistical Text-classifier with Adversarial ResistanceYuhan Yao, Feifei Kou, Lei Shi, Xiao Yang 等NeurIPS 2025
- Are We in the AI-Generated Text World Already? Quantifying and Monitoring AIGT on Social MediaZhen Sun, Zongmin Zhang, Xinyue Shen, Ziyi Zhang 等ACL 2025
它引用的顶会 Paper20
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo 等NeurIPS 2022 · 被引用 8,168 次
- Self-Refine: Iterative Refinement with Self-FeedbackAman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan 等NeurIPS 2023 · 被引用 4,972 次
- TruthfulQA: Measuring How Models Mimic Human FalsehoodsStephanie Lin, Jacob Hilton, Owain EvansACL 2022 · 被引用 3,228 次
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad 等ACL 2020 · 被引用 1,224 次
相关 Paper
- DEMASQ: Unmasking the ChatGPT WordsmithKavita Kumari, Alessandro Pegoraro, Hossein Fereidooni, Ahmad-Reza SadeghiNDSS 2024
- MGTBench: Benchmarking Machine-Generated Text DetectionXinlei He, Xinyue Shen, Zeyuan Chen, Michael Backes 等CCS 2024 · 被引用 30 次
- AI Wrote My Paper and All I Got was This False Negative:* Measuring the Efficacy of Commercial AI Text DetectorsSeth Layton, Bernardo B. P. Medeiros, Kevin R. B. Butler, Patrick TraynorS&P 2026 · 被引用 3 次
- SeqXGPT: Sentence-Level AI-Generated Text DetectionPengyu Wang, Linyang Li, Ke Ren, Botian Jiang 等EMNLP 2023 · 被引用 31 次
- An Empirical Study to Evaluate AIGC Detectors on Code ContentJian Wang, Shangqing Liu, Xiaofei Xie, Yi LiASE 2024 · 被引用 4 次
