MultiSocial: Multilingual Benchmark of Machine-Generated Text Detection of Social-Media Texts
Dominik Macko, Jakub Kopal, Róbert Móro, Ivan Srba
Abstract
Recent LLMs are able to generate high-quality multilingual texts, indistinguishable for humans from authentic human-written ones. Research in machine-generated text detection is however mostly focused on the English language and longer texts, such as news articles, scientific papers or student essays. Social-media texts are usually much shorter and often feature informal language, grammatical errors, or distinct linguistic items (e.g., emoticons, hashtags). There is a gap in studying the ability of existing methods in detection of such texts, reflected also in the lack of existing multilingual benchmark datasets. To fill this gap we propose the first multilingual (22 languages) and multi-platform (5 social media platforms) dataset for benchmarking machine-generated text detection in the social-media domain, called MultiSocial. It contains 472,097 texts, of which about 58k are human-written and approximately the same amount is generated by each of 7 multilingual LLMs. We use this benchmark to compare existing detection methods in zero-shot as well as fine-tuned form. Our results indicate that the fine-tuned detectors have no problem to be trained on social-media texts and that the platform selection for training matters.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e731aeb1-a946-4117-a2f4-586002ff5d5cCited by top-tier papers5
- Authorship Attribution in Multilingual Machine-Generated TextsLucio La Cava, Dominik Macko, Róbert Móro, Ivan Srba et al.ACL 2026 · 7 citations
- OpenTuringBench: An Open-Model-based Benchmark and Framework for Machine-Generated Text Detection and AttributionLucio La Cava, Andrea TagarelliEMNLP 2025
- Learning From Dictionary: Enhancing Robustness of Machine-Generated Text Detection in Zero-Shot Language via Adversarial TrainingYuanfan Li, Qi Zhou, Zexuan XieICLR 2026
- Characterizing an LLM-driven Social Network: The Case of Chirper.aiYiming Zhu, Yupeng He, Ehsan-Ul Haq, Gareth Tyson et al.CSCW 2026
- Are We in the AI-Generated Text World Already? Quantifying and Monitoring AIGT on Social MediaZhen Sun, Zongmin Zhang, Xinyue Shen, Ziyi Zhang et al.ACL 2025
Builds on14
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 5,863 citations
- MAUVE: Measuring the Gap Between Neural Text and Human Text using Divergence FrontiersKrishna Pillutla, Swabha Swayamdipta, Rowan Zellers, John Thickstun et al.NeurIPS 2021 · 606 citations
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary et al.ACL 2020 · 539 citations
- DeBERTaV3: Improving DeBERTa using ELECTRA-Style Pre-Training with Gradient-Disentangled Embedding SharingPengcheng He, Jianfeng Gao, Weizhu ChenICLR 2023 · 394 citations
Related papers
- MULTITuDE: Large-Scale Multilingual Machine-Generated Text Detection BenchmarkDominik Macko, Róbert Móro, Adaku Uchendu, Jason Samuel Lucas et al.EMNLP 2023 · 25 citations
- TSM-Bench: Detecting LLM-Generated Text in Real-World Wikipedia Editing PracticesGerrit Quaremba, Elizabeth Black, Denny Vrandecic, Elena SimperlICLR 2026 · 2 citations
- M4GT-Bench: Evaluation Benchmark for Black-Box Machine-Generated Text DetectionYuxia Wang, Jonibek Mansurov, Petar Ivanov, Jinyan Su et al.ACL 2024
- MAGE: Machine-generated Text Detection in the WildYafu Li, Qintong Li, Leyang Cui, Wei Bi et al.ACL 2024 · 44 citations
- Can AI-Generated Persuasion Be Detected? Persuaficial Benchmark and AI vs. Human Linguistic DifferencesArkadiusz Modzelewski, Pawel Golik, Anna Kolos, Giovanni Da San MartinoACL 2026 · 1 citation
