VendorLink: An NLP approach for Identifying & Linking Vendor Migrants & Potential Aliases on Darknet Markets
Vageesh Saxena, Nils Rethmeier, Gijs van Dijck, Gerasimos Spanakis
Abstract
The anonymity on the Darknet allows vendors to stay undetected by using multiple vendor aliases or frequently migrating between markets. Consequently, illegal markets and their connections are challenging to uncover on the Darknet. To identify relationships between illegal markets and their vendors, we propose VendorLink, an NLP-based approach that examines writing patterns to verify, identify, and link unique vendor accounts across text advertisements (ads) on seven public Darknet markets. In contrast to existing literature, Ven-dorLink utilizes the strength of supervised pretraining to perform closed-set vendor verification, open-set vendor identification, and lowresource market adaption tasks. Through Ven-dorLink, we uncover (i) 15 migrants and 71 potential aliases in the Alphabay-Dreams-Silk dataset, (ii) 17 migrants and 3 potential aliases in the Valhalla-Berlusconi dataset, and (iii) 75 migrants and 10 potential aliases in the Traderoute-Agora dataset. Altogether, our approach can help Law Enforcement Agencies (LEA) make more informed decisions by verifying and identifying migrating vendors and their potential aliases on existing and Low-Resource (LR) emerging Darknet markets. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cdedb6af-e5bd-4cdb-959b-e6c5e8263c4fCited by top-tier papers2
- IDTraffickers: An Authorship Attribution Dataset to link and connect Potential Human-Trafficking Operations on Text Escort AdvertisementsVageesh Saxena, Benjamin Bashpole, Gijs van Dijck, Gerasimos SpanakisEMNLP 2023 · 2 citations
- Covering Cracks in Content Moderation: Delexicalized Distant Supervision for Illicit Drug Jargon DetectionMinkyoo Song, Eugene Jang, Jaehan Kim, Seungwon ShinKDD 2025
Builds on10
- On the Sentence Embeddings from Pre-trained Language ModelsBohan Li, Hao Zhou, Junxian He, Mingxuan Wang et al.EMNLP 2020 · 538 citations
- PromptBERT: Improving BERT Sentence Embeddings with PromptsTing Jiang, Jian Jiao, Shaohan Huang, Zihan Zhang et al.EMNLP 2022 · 148 citations
- Improved Text Classification via Contrastive Adversarial TrainingLin Pan, Chung-Wei Hang, Avirup Sil, Saloni PotdarAAAI 2022 · 115 citations
- Authorship Attribution for Neural Text GenerationAdaku Uchendu, Thai Le, Kai Shu, Dongwon LeeEMNLP 2020 · 110 citations
- Plug and Prey? Measuring the Commoditization of Cybercrime via Online Anonymous MarketsRolf van Wegberg, Samaneh Tajalizadehkhoob, Kyle Soska, Ugur Akyazi et al.USENIX Security 2018 · 97 citations
Related papers
- eDarkFind: Unsupervised Multi-view Learning for Sybil Account DetectionRamnath Kumar, Shweta Yadav, Raminta Daniulaityte, Francois R. Lamy et al.WWW 2020 · 28 citations
- SYSML: StYlometry with Structure and Multitask Learning: Implications for Darknet Forum Migrant AnalysisPranav Maneriker, Yuntian He, Srinivasan ParthasarathyEMNLP 2021 · 6 citations
- Go See a Specialist? Predicting Cybercrime Sales on Online Anonymous Markets from Vendor and Product CharacteristicsRolf van Wegberg, Fieke Miedema, Ugur Akyazi, Arman Noroozian et al.WWW 2020 · 14 citations
- DarkBERT: A Language Model for the Dark Side of the InternetYoungjin Jin, Eugene Jang, Jian Cui, Jin-Woo Chung et al.ACL 2023 · 41 citations
- Uncovering and Mitigating the Hidden Chasm: A Study on the Text-Text Domain Gap in Euphemism IdentificationYuxue Hu, Junsong Li, Mingmin Wu, Zhongqiang Huang et al.AAAI 2024 · 1 citation
