INFOSHIELD: Generalizable Information-Theoretic Human-Trafficking Detection
Meng-Chieh Lee, Catalina Vajiac, Aayushi Kulshrestha, Sacha Levy, Namyong Park, Cara Jones, Reihaneh Rabbany, Christos Faloutsos
摘要
Given a million escort advertisements, how can we spot near-duplicates? Such micro-clusters of ads are usually signals of human trafficking. How can we summarize them, visually, to convince law enforcement to act? Can we build a general tool that works for different languages? Spotting micro-clusters of near-duplicate documents is useful in multiple, additional settings, including spam-bot detection in Twitter ads, plagiarism, and more.We present INFOSHIELD, which makes the following contributions: (a) Practical, being scalable and effective on real data, (b) Parameter-free and Principled, requiring no user-defined parameters, (c) Interpretable, finding a document to be the cluster representative, highlighting all the common phrases, and automatically detecting "slots", i.e. phrases that differ in every document; and (d) Generalizable, beating or matching domain-specific methods in Twitter bot detection and human trafficking detection respectively, as well as being language-independent finding clusters in Spanish, Italian, and Japanese. Interpretability is particularly important for the anti human-trafficking domain, where law enforcement must visually inspect ads.Our experiments on real data show that INFOSHIELD correctly identifies Twitter bots with an F1 score over 90% and detects human-trafficking ads with 84% precision. Moreover, it is scalable, requiring about 8 hours for 4 million documents on a stock laptop.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper2
- TrafficVis: Visualizing Organized Activity and Spatio-Temporal Patterns for Detecting and Labeling Human TraffickingCatalina Vajiac, Duen Horng Chau, Andreas M. Olligschlaeger, Rebecca Mackenzie 等IEEE VIS 2022 · 被引用 10 次
- IDTraffickers: An Authorship Attribution Dataset to link and connect Potential Human-Trafficking Operations on Text Escort AdvertisementsVageesh Saxena, Benjamin Bashpole, Gijs van Dijck, Gerasimos SpanakisEMNLP 2023 · 被引用 2 次
相关 Paper
- Scalable and Generalizable Social Bot Detection through Data SelectionKai-Cheng Yang, Onur Varol, Pik-Mai Hui, Filippo MenczerAAAI 2020 · 被引用 385 次
- #Twiti: Social Listening for Threat IntelligenceHyejin Shin, WooChul Shim, Saebom Kim, Sol Lee 等WWW 2021 · 被引用 32 次
- RETSim: Resilient and Efficient Text SimilarityMarina Zhang, Owen S. Vallis, Aysegul Bumin, Tanay Vakharia 等ICLR 2024 · 被引用 3 次
- Detecting and Understanding the Promotion of Illicit Goods and Services on TwitterHongyu Wang, Ying Li, Ronghong Huang, Xianghang MiWWW 2025 · 被引用 6 次
- MultiSocial: Multilingual Benchmark of Machine-Generated Text Detection of Social-Media TextsDominik Macko, Jakub Kopal, Róbert Móro, Ivan SrbaACL 2025 · 被引用 15 次
