DarkBERT: A Language Model for the Dark Side of the Internet
Youngjin Jin, Eugene Jang, Jian Cui, Jin-Woo Chung, Yongjae Lee, Seungwon Shin
摘要
Recent research has suggested that there are clear differences in the language used in the Dark Web compared to that of the Surface Web. As studies on the Dark Web commonly require textual analysis of the domain, language models specific to the Dark Web may provide valuable insights to researchers. In this work, we introduce DarkBERT, a language model pretrained on Dark Web data. We describe the steps taken to filter and compile the text data used to train DarkBERT to combat the extreme lexical and structural diversity of the Dark Web that may be detrimental to building a proper representation of the domain. We evaluate Dark-BERT and its vanilla counterpart along with other widely used language models to validate the benefits that a Dark Web domain specific model offers in various use cases. Our evaluations show that DarkBERT outperforms current language models and may serve as a valuable resource for future research on the Dark Web.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- KidLM: Advancing Language Models for Children - Early Insights and Future DirectionsMir Tafseer Nayeem, Davood RafieiEMNLP 2024 · 被引用 7 次
- Covering Cracks in Content Moderation: Delexicalized Distant Supervision for Illicit Drug Jargon DetectionMinkyoo Song, Eugene Jang, Jaehan Kim, Seungwon ShinKDD 2025
- SoK: Automated TTP Extraction from CTI Reports - Are We There Yet?Marvin Büchel, Tommaso Paladini, Stefano Longari, Michele Carminati 等USENIX Security 2025
它引用的顶会 Paper4
- Acing the IOC Game: Toward Automatic Discovery and Analysis of Open-Source Cyber Threat IntelligenceXiaojing Liao, Kan Yuan, XiaoFeng Wang, Zhou Li 等CCS 2016 · 被引用 308 次
- Cybercriminal Minds: An investigative study of cryptocurrency abuses in the Dark WebSeunghyeon Lee, Changhoon Yoon, Heedo Kang, Yeonkeun Kim 等NDSS 2019 · 被引用 100 次
- Reading Thieves' Cant: Automatically Identifying and Understanding Dark Jargons from Cybercrime MarketplacesKan Yuan, Haoran Lu, Xiaojing Liao, XiaoFeng WangUSENIX Security 2018 · 被引用 56 次
- Self-Supervised Euphemism Detection and Identification for Content ModerationWanzheng Zhu, Hongyu Gong, Rohan Bansal, Zachary Weinberg 等S&P 2021 · 被引用 56 次
相关 Paper
- PhishLang: A Real-Time, Fully Client-Side Phishing Detection Framework Using MobileBERTSayak Saha Roy, Shirin NilizadehNDSS 2026 · 被引用 9 次
- DrBERT: A Robust Pre-trained Model in French for Biomedical and Clinical domainsYanis Labrak, Adrien Bazoge, Richard Dufour, Mickael Rouvier 等ACL 2023 · 被引用 19 次
- Making Pre-trained Language Models Great on Tabular PredictionJiahuan Yan, Bo Zheng, Hongxia Xu, Yiheng Zhu 等ICLR 2024 · 被引用 72 次
- Table Search Using a Deep Contextualized Language ModelZhiyu Chen, Mohamed Trabelsi, Jeff Heflin, Yinan Xu 等SIGIR 2020 · 被引用 48 次
- Incorporating medical knowledge in BERT for clinical relation extractionArpita Roy, Shimei PanEMNLP 2021 · 被引用 56 次
