Who Writes What: Unveiling the Impact of Author Roles on AI-generated Text Detection
Jiatao Li, Xiaojun Wan
Abstract
The rise of Large Language Models (LLMs) necessitates accurate AI-generated text detection. However, current approaches largely overlook the influence of author characteristics. We investigate how sociolinguistic attributes-gender, CEFR proficiency, academic field, and language environment-impact stateof-the-art AI text detectors. Using the ICNALE corpus of human-authored texts and parallel AIgenerated texts from diverse LLMs, we conduct a rigorous evaluation employing multi-factor ANOVA and weighted least squares (WLS). Our results reveal significant biases: CEFR proficiency and language environment consistently affected detector accuracy, while gender and academic field showed detector-dependent effects. These findings highlight the crucial need for socially aware AI text detection to avoid unfairly penalizing specific demographic groups. We offer novel empirical evidence, a robust statistical framework, and actionable insights for developing more equitable and reliable detection systems in real-world, out-of-domain contexts. This work paves the way for future research on bias mitigation, inclusive evaluation benchmarks, and socially responsible LLM detectors. Detector CEFR Sex Academic Genre Language Env. binoculars Yes No No Yes chatgptroberta Yes No No No detectgpt Yes No Yes Yes fastdetectgpt Yes No No Yes fastdetectllm Yes No Yes Yes gltr Yes No No Yes gpt2-base Yes No Yes Yes gpt2-large Yes No Yes Yes llmdet Yes No Yes Yes radar Yes No No Yes
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 159e3531-6f79-40b8-a9b1-b5c950ea6c53Cited by top-tier papers1
Ask how each one uses itBuilds on8
- DetectGPT: Zero-Shot Machine-Generated Text Detection using Probability CurvatureEric Mitchell, Yoonho Lee, Alexander Khazatsky, Christopher D. Manning et al.ICML 2023 · 988 citations
- Paraphrasing evades detectors of AI-generated text, but retrieval is an effective defenseKalpesh Krishna, Yixiao Song, Marzena Karpinska, John Wieting et al.NeurIPS 2023 · 657 citations
- RADAR: Robust AI-Text Detection via Adversarial LearningXiaomeng Hu, Pin-Yu Chen, Tsung-Yi HoNeurIPS 2023 · 315 citations
- Spotting LLMs With Binoculars: Zero-Shot Detection of Machine-Generated TextAbhimanyu Hans, Avi Schwarzschild, Valeriia Cherepanova, Hamid Kazemi et al.ICML 2024 · 262 citations
- Few-Shot Detection of Machine-Generated Text using Style RepresentationsRafael A. Rivera Soto, Kailin Koch, Aleem Khan, Barry Y. Chen et al.ICLR 2024 · 49 citations
Related papers
- Identifying Bias in Machine-generated Text DetectionKevin Stowe, Svetlana Afanaseva, Rodolfo Raimundo, Yitao Sun et al.ACL 2026
- HLD: Approximate Hierarchical Linguistic Distribution Modeling for LLM-Generated Text DetectionRui Guo, Weibin Zeng, Fuzhang Wu, Yan Kong et al.ICLR 2026
- Profiler: Black-box AI-generated Text Origin Detection via Context-aware Inference Pattern AnalysisHanxi Guo, Siyuan Cheng, Xiaolong Jin, Zhuo Zhang et al.EMNLP 2025
- Enhancing LLM Text Detection with Retrieved Contexts and Logits Distribution ConsistencyZhaoheng Huang, Yutao Zhu, Ji-Rong Wen, Zhicheng DouEMNLP 2025
- AI Wrote My Paper and All I Got was This False Negative:* Measuring the Efficacy of Commercial AI Text DetectorsSeth Layton, Bernardo B. P. Medeiros, Kevin R. B. Butler, Patrick TraynorS&P 2026 · 3 citations
