IDTraffickers: An Authorship Attribution Dataset to link and connect Potential Human-Trafficking Operations on Text Escort Advertisements
Vageesh Saxena, Benjamin Bashpole, Gijs van Dijck, Gerasimos Spanakis
摘要
Human trafficking (HT) is a pervasive global issue affecting vulnerable individuals, violating their fundamental human rights. Investigations reveal that many HT cases are associated with online advertisements (ads), particularly in escort markets. Consequently, identifying and connecting HT vendors has become increasingly challenging for Law Enforcement Agencies (LEAs). To address this issue, we introduce IDTraffickers, an extensive dataset consisting of 87,595 text ads and 5,244 vendor labels to enable the verification and identification of potential HT vendors on online escort markets. To establish a benchmark for authorship identification, we train a DeCLUTR-small model, achieving a macro-F1 score of 0.8656 in a closed-set classification environment. Next, we leverage the style representations extracted from the trained classifier to conduct authorship verification, resulting in a mean r-precision score of 0.8852 in an open-set ranking environment. Finally, to encourage further research and ensure responsible data sharing, we plan to release IDTraffickers for the authorship attribution task to researchers under specific conditions, considering the sensitive nature of the data. We believe that the availability of our dataset and benchmarks will empower future researchers to utilize our findings, thereby facilitating the effective linkage of escort ads and the development of more robust approaches for identifying HT indicators (1).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper12
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel 等ICLR 2020 · 被引用 7,418 次
- SimCSE: Simple Contrastive Learning of Sentence EmbeddingsTianyu Gao, Xingcheng Yao, Danqi ChenEMNLP 2021 · 被引用 2,496 次
- Open-Set Recognition: A Good Closed-Set Classifier is All You NeedSagar Vaze, Kai Han, Andrea Vedaldi, Andrew ZissermanICLR 2022 · 被引用 594 次
- DeBERTaV3: Improving DeBERTa using ELECTRA-Style Pre-Training with Gradient-Disentangled Embedding SharingPengcheng He, Jianfeng Gao, Weizhu ChenICLR 2023 · 被引用 394 次
- Authorship Attribution for Neural Text GenerationAdaku Uchendu, Thai Le, Kai Shu, Dongwon LeeEMNLP 2020 · 被引用 110 次
相关 Paper
- TrafficVis: Visualizing Organized Activity and Spatio-Temporal Patterns for Detecting and Labeling Human TraffickingCatalina Vajiac, Duen Horng Chau, Andreas M. Olligschlaeger, Rebecca Mackenzie 等IEEE VIS 2022 · 被引用 10 次
- INFOSHIELD: Generalizable Information-Theoretic Human-Trafficking DetectionMeng-Chieh Lee, Catalina Vajiac, Aayushi Kulshrestha, Sacha Levy 等ICDE 2021 · 被引用 16 次
- VendorLink: An NLP approach for Identifying & Linking Vendor Migrants & Potential Aliases on Darknet MarketsVageesh Saxena, Nils Rethmeier, Gijs van Dijck, Gerasimos SpanakisACL 2023 · 被引用 6 次
- AI Wrote My Paper and All I Got was This False Negative:* Measuring the Efficacy of Commercial AI Text DetectorsSeth Layton, Bernardo B. P. Medeiros, Kevin R. B. Butler, Patrick TraynorS&P 2026 · 被引用 3 次
- Ethics and Efficacy of Unsolicited Anti-Trafficking SMS OutreachRasika Bhalerao, Nora McDonald, Hanna Barakat, Vaughn Hamilton 等CSCW 2022 · 被引用 7 次
