Can AI-Generated Persuasion Be Detected? Persuaficial Benchmark and AI vs. Human Linguistic Differences
Arkadiusz Modzelewski, Pawel Golik, Anna Kolos, Giovanni Da San Martino
Abstract
Large Language Models (LLMs) can generate highly persuasive text, raising concerns about their misuse for propaganda, manipulation, and other harmful purposes. This leads us to our central question: Is LLM-generated persuasion more difficult to automatically detect than human-written persuasion? To address this, we categorize controllable generation approaches for producing persuasive content with LLMs and introduce Persuaficial, a high-quality multilingual benchmark covering six languages: English, German, Polish, Italian, French and Russian. Using this benchmark, we conduct extensive empirical evaluations comparing human-authored and LLM-generated persuasive texts. We find that although overtly persuasive LLM-generated texts can be easier to detect than human-written ones, subtle LLM-generated persuasion consistently degrades automatic detection performance. Beyond detection performance, we provide the first comprehensive linguistic analysis contrasting human and LLM-generated persuasive texts, offering insights that may guide the development of more interpretable and robust detection tools.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b37c803c-ff3b-4e99-aaee-a6db08a94bb5Builds on4
- Towards Reliable Misinformation Mitigation: Generalization, Uncertainty, and GPT-4Kellin Pelrine, Anne Imouza, Camille Thibault, Meilina Reksoprodjo et al.EMNLP 2023 · 27 citations
- Multilingual Multifaceted Understanding of Online News in Terms of Genre, Framing, and Persuasion TechniquesJakub Piskorski, Nicolas Stefanovitch, Nikolaos Nikolaidis, Giovanni Da San Martino et al.ACL 2023 · 14 citations
- SocialEval: Evaluating Social Intelligence of Large Language ModelsJinfeng Zhou, Yuxuan Chen, Yihan Shi, Xuanming Zhang et al.ACL 2025 · 9 citations
- Better than Average: Paired Evaluation of NLP systemsMaxime Peyrard, Wei Zhao, Steffen Eger, Robert WestACL 2021
Related papers
- DetectRL-X: Towards Reliable Multilingual and Real-World LLM-Generated Text DetectionJunchao Wu, Yefeng Liu, Chenyu Zhu, Hao Zhang et al.ACL 2026
- Measuring And Improving Persuasiveness Of Large Language ModelsSomesh Kumar Singh, Yaman Kumar Singla, Harini S. I, Balaji KrishnamurthyICLR 2025
- M4GT-Bench: Evaluation Benchmark for Black-Box Machine-Generated Text DetectionYuxia Wang, Jonibek Mansurov, Petar Ivanov, Jinyan Su et al.ACL 2024
- AI 'News' Content Farms Are Easy to Make and Hard to Detect: A Case Study in ItalianGiovanni Puccetti, Anna Rogers, Chiara Alzetta, Felice Dell'Orletta et al.ACL 2024 · 1 citation
- MultiSocial: Multilingual Benchmark of Machine-Generated Text Detection of Social-Media TextsDominik Macko, Jakub Kopal, Róbert Móro, Ivan SrbaACL 2025 · 15 citations
