On the Risk of Evidence Pollution for Malicious Social Text Detection in the Era of LLMs
Herun Wan, Minnan Luo, Zhixiong Su, Guang Dai, Xiang Zhao
Abstract
Evidence-enhanced detectors present remarkable abilities in identifying malicious social text. However, the rise of large language models (LLMs) brings potential risks of evidence pollution to confuse detectors. This paper explores potential manipulation scenarios including basic pollution, and rephrasing or generating evidence by LLMs. To mitigate the negative impact, we propose three defense strategies from the data and model sides, including machine-generated text detection, a mixture of experts, and parameter updating. Extensive experiments on four malicious social text detection tasks with ten datasets illustrate that evidence pollution significantly compromises detectors, where the generating strategy causes up to a 14.4% performance drop. Meanwhile, the defense strategies could mitigate evidence pollution, but they faced limitations for practical employment. Further analysis illustrates that polluted evidence (i) is of high quality, evaluated by metrics and humans; (ii) would compromise the model calibration, increasing expected calibration error up to 21.6%; and (iii) could be integrated to amplify the negative impact, especially for encoder-based LMs, where the accuracy drops by 21.8%.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bf0daf6f-e5bc-4634-ba81-08459f6fc350Cited by top-tier papers2
- Truth over Tricks: Measuring and Mitigating Shortcut Learning in Misinformation DetectionHerun Wan, Jiaying Wu, Minnan Luo, Zhi Zeng et al.NeurIPS 2025 · 14 citations
- How Do Social Bots Participate in Misinformation Spread? A Comprehensive Dataset and AnalysisHerun Wan, Minnan Luo, Zihan Ma, Guang Dai et al.EMNLP 2025 · 3 citations
Builds on31
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- DetectGPT: Zero-Shot Machine-Generated Text Detection using Probability CurvatureEric Mitchell, Yoonho Lee, Alexander Khazatsky, Christopher D. Manning et al.ICML 2023 · 988 citations
- Mining Dual Emotion for Fake News DetectionXueyao Zhang, Juan Cao, Xirong Li, Qiang Sheng et al.WWW 2021 · 332 citations
- Spotting LLMs With Binoculars: Zero-Shot Detection of Machine-Generated TextAbhimanyu Hans, Avi Schwarzschild, Valeriia Cherepanova, Hamid Kazemi et al.ICML 2024 · 262 citations
- When Not to Trust Language Models: Investigating Effectiveness of Parametric and Non-Parametric MemoriesAlex Mallen, Akari Asai, Victor Zhong, Rajarshi Das et al.ACL 2023 · 233 citations
Related papers
- Mitigating GenAI-Powered Evidence Pollution for Out-Of-Context Misinformation DetectionZehong Yan, Peng Qi, Wynne Hsu, Mong-Li LeeICDE 2026 · 1 citation
- Large Language Model (LLM)-driven Adversarial Social Influences in Online Information Spread: Risks and InterventionsZhuoran Lu, Gionnieve Lim, Ming YinCHI 2026
- Imperceptible Content Poisoning in LLM-Powered ApplicationsQuan Zhang, Chijin Zhou, Gwihwan Go, Binqi Zeng et al.ASE 2024 · 3 citations
- Enhancing LLM Text Detection with Retrieved Contexts and Logits Distribution ConsistencyZhaoheng Huang, Yutao Zhu, Ji-Rong Wen, Zhicheng DouEMNLP 2025
- What Does the Bot Say? Opportunities and Risks of Large Language Models in Social Media Bot DetectionShangbin Feng, Herun Wan, Ningnan Wang, Zhaoxuan Tan et al.ACL 2024 · 19 citations
