Validating Multimedia Content Moderation Software via Semantic Fusion
Wenxuan Wang, Jingyuan Huang, Chang Chen, Jiazhen Gu, Jianping Zhang, Weibin Wu, Pinjia He, Michael R. Lyu
Abstract
The exponential growth of social media platforms, such as Facebook, Instagram, Youtube, and TikTok, has revolutionized communication and content publication in human society. Users on these platforms can publish multimedia content that delivers information via the combination of text, audio, images, and video. Meanwhile, the multimedia content release facility has been increasingly exploited to propagate toxic content, such as hate speech, malicious advertisement, and pornography. To this end, content moderation software has been widely deployed on these platforms to detect and blocks toxic content. However, due to the complexity of content moderation models and the difficulty of understanding information across multiple modalities, existing content moderation software can fail to detect toxic content, which often leads to extremely negative impacts (e.g., harmful effects on teen mental health). We introduce Semantic Fusion, a general, effective methodology for validating multimedia content moderation software. Our key idea is to fuse two or more existing single-modal inputs (e.g., a textual sentence and an image) into a new input that combines the semantics of its ancestors in a novel manner and has toxic nature by construction. This fused input is then used for validating multimedia content moderation software. We realized Semantic Fusion as DUO, a practical content moderation software testing tool. In our evaluation, we employ DUO to test five commercial content moderation software and two state-of-the-art models against three kinds of toxic contents. The results show that DUO achieves up to 100% error finding rate (EFR) when testing moderation software and it obtains up to 94.1% EFR when testing the state-of-the-art models. In addition, we leverage the test cases generated by DUO to retrain the two models we explored, which largely improves model robustness (2.5%∼5.7% EFR) while maintaining the accuracy on the original test set.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext daaca6cd-80e4-4d36-af0f-f63d5506a178Cited by top-tier papers2
- Glitch Tokens in Large Language Models: Categorization Taxonomy and Effective DetectionYuxi Li, Yi Liu, Gelei Deng, Ying Zhang et al.FSE 2024 · 12 citations
- An Image is Worth a Thousand Toxic Words: A Metamorphic Testing Framework for Content Moderation SoftwareWenxuan Wang, Jingyuan Huang, Jen-tse Huang, Chang Chen et al.ASE 2023 · 7 citations
Builds on21
- The Hateful Memes Challenge: Detecting Hate Speech in Multimodal MemesDouwe Kiela, Hamed Firooz, Aravind Mohan, Vedanuj Goswami et al.NeurIPS 2020 · 1,022 citations
- Hidden Voice CommandsNicholas Carlini, Pratyush Mishra, Tavish Vaidya, Yuankai Zhang et al.USENIX Security 2016 · 672 citations
- BERT-ATTACK: Adversarial Attack Against BERT Using BERTLinyang Li, Ruotian Ma, Qipeng Guo, Xiangyang Xue et al.EMNLP 2020 · 529 citations
- Leveraging Automated Unit Tests for Unsupervised Code TranslationBaptiste Rozière, Jie Zhang, François Charton, Mark Harman et al.ICLR 2022 · 161 citations
- Improving Adversarial Transferability via Neuron Attribution-based AttacksJianping Zhang, Weibin Wu, Jen-tse Huang, Yizhan Huang et al.CVPR 2022 · 140 citations
Related papers
- MTTM: Metamorphic Testing for Textual Content Moderation SoftwareWenxuan Wang, Jen-tse Huang, Weibin Wu, Jianping Zhang et al.ICSE 2023 · 23 citations
- Metamorphic Testing for Audio Content Moderation SoftwareWenxuan Wang, Yongjiang Wu, Junyuan Zhang, Shuqing Li et al.ASE 2025
- Rethinking Implicit Hate Speech Detection: Focusing on Latent Hate Components via Dual-Process ArgumentationShiqi Sun, Du Su, Wei Chen, Xueqi ChengWWW 2026
- Text Detoxification: Data Efficiency, Semantic Preservation and Model GeneralizationJing Yu, Yibo Zhao, Jiapeng Zhu, Wenming Shao et al.EMNLP 2025
- DiffuFuse: Diffusion-Driven Dual-Stream Fusion Framework for Multimodal Sentiment AnalysisXiongjian Lv, Yimin Wen, Hang YuACM MM 2025
