Are We in the AI-Generated Text World Already? Quantifying and Monitoring AIGT on Social Media
Zhen Sun, Zongmin Zhang, Xinyue Shen, Ziyi Zhang, Yule Liu, Michael Backes, Yang Zhang, Xinlei He
摘要
Social media platforms are experiencing a growing presence of AI-Generated Texts (AIGTs). However, the misuse of AIGTs could have profound implications for public opinion, such as spreading misinformation and manipulating narratives. Despite its importance, it remains unclear how prevalent AIGTs are on social media. To address this gap, this paper aims to quantify and monitor the AIGTs on online social media platforms. We first collect a dataset (SM-D) with around 2.4M posts from 3 major social media platforms: Medium, Quora, and Reddit. Then, we construct a diverse dataset (AIGTBench) to train and evaluate AIGT detectors. AIGT-Bench combines popular open-source datasets and our AIGT datasets generated from social media texts by 12 LLMs, serving as a benchmark for evaluating mainstream detectors. With this setup, we identify the best-performing detector (OSM-Det). We then apply OSM-Det to SM-D to track AIGTs across social media platforms from January 2022 to October 2024, using the AI Attribution Rate (AAR) as the metric. Specifically, Medium and Quora exhibit marked increases in AAR, rising from 1.77% to 37.03% and 2.06% to 38.95%, respectively. In contrast, Reddit shows slower growth, with AAR increasing from 1.31% to 2.45% over the same period. Our further analysis indicates that AIGTs on social media differ from human-written texts across several dimensions, including linguistic patterns, topic distributions, engagement levels, and the follower distribution of authors. We envision our analysis and findings on AIGTs in social media can shed light on future research in this domain. Our code and dataset are publicly available. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- AI use in American newspapers is widespread, uneven, and rarely disclosedJenna Russell, Marzena Karpinska, Destiny Akinode, James Zhou 等ACL 2026 · 被引用 9 次
- Governance of AI-Generated Content: A Case Study on Social Media PlatformsLan Gao, Abani Ahmed, Oscar Chen, Margaux Reyl 等CHI 2026 · 被引用 3 次
- Mitigating GenAI-Powered Evidence Pollution for Out-Of-Context Misinformation DetectionZehong Yan, Peng Qi, Wynne Hsu, Mong-Li LeeICDE 2026 · 被引用 1 次
- UMPIRE: Unveiling LLM-generated Posts via Redundant ExpressionsXiaoquan Yi, Haixing Wu, Haozhao Wang, Yichen Li 等ACL 2026
- Characterizing an LLM-driven Social Network: The Case of Chirper.aiYiming Zhu, Yupeng He, Ehsan-Ul Haq, Gareth Tyson 等CSCW 2026
它引用的顶会 Paper14
- DetectGPT: Zero-Shot Machine-Generated Text Detection using Probability CurvatureEric Mitchell, Yoonho Lee, Alexander Khazatsky, Christopher D. Manning 等ICML 2023 · 被引用 988 次
- Synthetic Lies: Understanding AI-Generated Misinformation and Evaluating Algorithmic and Human SolutionsJiawei Zhou, Yixuan Zhang, Qianni Luo, Andrea G. Parker 等CHI 2023 · 被引用 283 次
- Document-Level Machine Translation with Large Language ModelsLongyue Wang, Chenyang Lyu, Tianbo Ji, Zhirui Zhang 等EMNLP 2023 · 被引用 129 次
- Evaluating Open-Domain Question Answering in the Era of Large Language ModelsEhsan Kamalloo, Nouha Dziri, Charles L. A. Clarke, Davood RafieiACL 2023 · 被引用 96 次
- Few-Shot Detection of Machine-Generated Text using Style RepresentationsRafael A. Rivera Soto, Kailin Koch, Aleem Khan, Barry Y. Chen 等ICLR 2024 · 被引用 49 次
相关 Paper
- MultiSocial: Multilingual Benchmark of Machine-Generated Text Detection of Social-Media TextsDominik Macko, Jakub Kopal, Róbert Móro, Ivan SrbaACL 2025 · 被引用 15 次
- An Empirical Study to Evaluate AIGC Detectors on Code ContentJian Wang, Shangqing Liu, Xiaofei Xie, Yi LiASE 2024 · 被引用 4 次
- Is Your Paper Being Reviewed by an LLM? Benchmarking AI Text Detection in Peer ReviewSungduk Yu, Man Luo, Avinash Madasu, Vasudev Lal 等ICLR 2026 · 被引用 24 次
- M4GT-Bench: Evaluation Benchmark for Black-Box Machine-Generated Text DetectionYuxia Wang, Jonibek Mansurov, Petar Ivanov, Jinyan Su 等ACL 2024
- OpenTuringBench: An Open-Model-based Benchmark and Framework for Machine-Generated Text Detection and AttributionLucio La Cava, Andrea TagarelliEMNLP 2025
