It's Trying Too Hard To Look Real: Deepfake Moderation Mistakes and Identity-Based Bias
Jaron Mink, Miranda Wei, Collins W. Munyendo, Kurt Hugenberg, Tadayoshi Kohno, Elissa M. Redmiles, Gang Wang
Abstract
Online platforms employ manual human moderation to distinguish human-created social media profiles from deepfake-generated ones. Biased misclassification of real profiles as artificial can harm general users as well as specific identity groups; however, no work has yet systematically investigated such mistakes and biases. We conducted a user study (𝑛=695) that investigates how 1) the identity of the profile, 2) whether the moderator shares that identity, and 3) components of a profile shown affect the perceived artificiality of the profile. We find statistically significant biases in people's moderation of LinkedIn profiles based on all three factors. Further, upon examining how moderators make decisions, we find they rely on mental models of AI and attackers, as well as typicality expectations (how they think the world works). The latter includes reliance on race/gender stereotypes. Based on our findings, we synthesize recommendations for the design of moderation interfaces, moderation teams, and security training.
• Security and privacy → Social aspects of security and privacy; • Human-centered computing → Empirical studies in HCI .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 98809bf4-0167-47fb-b5fc-2bb71d6dca44Cited by top-tier papers8
- Characterizing Photorealism and Artifacts in Diffusion Model-Generated ImagesNegar Kamali, Karyn Nakamura, Aakriti Kumar, Angelos Chatzimparmpas et al.CHI 2025 · 23 citations
- Public Opinions About Copyright for AI-Generated Art: The Role of Egocentricity, Competition, and ExperienceGabriel Lima, Nina Grgic-Hlaca, Elissa M. RedmilesCHI 2025 · 18 citations
- 'There Has To Be a Lot That We're Missing': Moderating AI-Generated Content on RedditTravis Lloyd, Joseph Reagle, Mor NaamanCSCW 2025 · 6 citations
- Behind the Same Mask: Understanding the Practice of Spontaneous Collective Anonymity on Chinese Social PlatformsSuqi Lou, Weijun Li, Chao Zhang, Shi Chen et al.CSCW 2025 · 4 citations
- Governance of AI-Generated Content: A Case Study on Social Media PlatformsLan Gao, Abani Ahmed, Oscar Chen, Margaux Reyl et al.CHI 2026 · 3 citations
Builds on8
- Disproportionate Removals and Differing Content Moderation Experiences for Conservative, Transgender, and Black Social Media Users: Marginalization and Moderation Gray AreasOliver L. Haimson, Daniel Delmonaco, Peipei Nie, Andrea WegnerCSCW 2021 · 287 citations
- Self-supervised Learning of Adversarial Example: Towards Good Generalizations for Deepfake DetectionLiang Chen, Yong Zhang, Yibing Song, Lingqiao Liu et al.CVPR 2022 · 251 citations
- Conformity of Eating Disorders through Content ModerationJessica L. Feuston, Alex S. Taylor, Anne Marie PiperCSCW 2020 · 82 citations
- Seeing is Believing: Exploring Perceptual Differences in DeepFake VideosRashid Tahir, Brishna Batool, Hira Jamshed, Mahnoor Jameel et al.CHI 2021 · 78 citations
- How Experts Detect Phishing Scam EmailsRick WashCSCW 2020 · 76 citations
Related papers
- DeepPhish: Understanding User Trust Towards Artificially Generated Profiles in Online Social NetworksJaron Mink, Licheng Luo, Natã M. Barbosa, Olivia Figueira et al.USENIX Security 2022
- "Better Be Computer or I'm Dumb": A Large-Scale Evaluation of Humans as Audio Deepfake DetectorsKevin Warren, Tyler Tucker, Anna Crowder, Daniel Olszewski et al.CCS 2024 · 9 citations
- Effect of AI Performance, Risk Perception, and Trust on Human Dependence in Deepfake Detection AI SystemYingfan Zhou, Ester Chen, Manasa Pisipati, Aiping Xiong et al.CSCW 2025 · 2 citations
- "It Matches My Worldview": Examining Perceptions and Attitudes Around Fake VideosFarhana Shahid, Srujana Kamath, Annie Sidotam, Vivian Jiang et al.CHI 2022 · 38 citations
- Labeling Synthetic Content: User Perceptions of Label Designs for AI-Generated Content on Social MediaDilrukshi Gamage, Dilki Sewwandi, Min Zhang, Arosha K. BandaraCHI 2025 · 25 citations
