AAAI2020

Reasoning about Political Bias in Content Moderation

Shan Jiang, Ronald E. Robertson, Christo Wilson

被引用 39 次

摘要

Content moderation, the AI-human hybrid process of removing (toxic) content from social media to promote community health, has attracted increasing attention from lawmakers due to allegations of political bias. Hitherto, this allegation has been made based on anecdotes rather than logical reasoning and empirical evidence, which motivates us to audit its validity. In this paper, we first introduce two formal criteria to measure bias (i.e., independence and separation) and their contextual meanings in content moderation, and then use YouTube as a lens to investigate if the political leaning of a video plays a role in the moderation decision for its associated comments. Our results show that when justifiable target variables (e.g., hate speech and extremeness) are controlled with propensity scoring, the likelihood of comment moderation is equal across left-and right-leaning videos. Bad Content Moderation, Bad! Social media has long played host to problematic content such as partisan propaganda (Allcott and Gentzkow 2017), misinformation (Jiang and Wilson 2018), and violent hate speech (Olteanu et al. 2018) . In an attempt to police this content and improve the health of their user community, social media platforms publish sets of community guidelines that explain the types of content they prohibit, and remove or hide this content from their platforms. This practice is commonly referred to as content moderation. Content moderation is typically implemented as an AIhuman hybrid process. To scale with the large amount of toxic content generated online, an AI filtering layer first finds potential candidates for moderation (Gibbs 2017; Sloane 2018), and then sends them to human reviewers for a final determination (Levin 2017; Gershgorn and Murphy 2017). This content moderation process, however, has been criticized for potential bias: biased AI systems have been documented (Barocas, Hardt, and Narayanan 2019; Hutchinson and Mitchell 2019), and human moderators can bring their own biases into the moderation process (Diakopoulos and Naaman 2011). As a result, content moderation faces a backlash from ideological conservatives, who allege that social media platforms are biased against them and are censoring their views (Kamisar 2018; Usher 2018), e.g., Figure 1 . These