Uninvited Guests: Analyzing the Identity and Behavior of Certificate Transparency Bots
Brian Kondracki, Johnny So, Nick Nikiforakis
摘要
Since its creation, Certificate Transparency (CT) has served as a vital component of the secure web. However, with the increase in TLS adoption, CT has essentially become a defacto log for all newly-created websites, announcing to the public the existence of web endpoints, including those that could have otherwise remained hidden. As a result, web bots can use CT to probe websites in real time, as they are created. Little is known about these bots, their behaviors, and their intentions. In this paper we present CTPOT, a distributed honeypot system which creates new TLS certificates for the purpose of advertising previously non-existent domains, and records the activity generated towards them from a number of network vantage points. Using CTPOT, we create 4,657 TLS certificates over a period of ten weeks, attracting 1.5 million web requests from 31,898 unique IP addresses. We find that CT bots occupy a distinct subset of the overall web bot population, with less than 2% overlap between IP addresses of CT bots and traditional host-scanning web bots. By creating certificates with varying content types, we are able to further sub-divide the CT bot population into subsets of varying intentions, revealing a stark contrast in malicious behavior among these groups. Finally, we correlate observed bot IP addresses into campaigns using the file paths requested by each bot, and find 105 malicious campaigns targeting the domains we advertise. Our findings shed light onto the CT bot ecosystem, revealing that it is not only distinct to that of traditional IP-based bots, but is composed of numerous entities with varying targets and behaviors. categories, allowing us to observe the varying behaviors of bots with different goals. Using CTPOT, we create 4,657 TLS certificates across 12 measurement nodes for pseudo-random subdomains corresponding to: popular trademarks, common web endpoints of web application software, and dictionary words which act as a baseline through which we compare and interpret the results of our measurement groups. Over the course of ten weeks, CTPOT received a total of 1.5 million web requests from 31,898 unique IP addresses, as well as 22,839 requests to SSH, FTP, and Telnet honeypots. As each domain generated by CTPOT was previously unused and completely unguessable, our curated dataset consists entirely of bot traffic. By analyzing our curated dataset, we observe distinct behaviors from the bots targeting different types of subdomains, allowing us to learn about the intentions of these various bot populations. For instance, we observe bots targeting subdomains of common web application software send more than one request over 40% more often than bots targeting domains impersonating popular trademarks. Moreover, these bots exhibit more malicious behavior with over twice as many unique IP addresses attempting to authenticate with exposed network services such as SSH. Additionally, we correlate requests from seemingly isolated IP addresses into campaigns of related bots. Alarmingly, through this analysis we find 105 malicious campaigns attempting to perform malicious actions such as data exfiltration, fingerprinting, and vulnerability exploitation. Our main contributions are as follows: • We design and implement CTPOT, a honeypot-based system to create TLS certificates for pseudo-random subdomains, and analyze requests directed towards them. Using this system, we create 4,657 TLS certificates. • We curate the first public dataset of CT bots. Analysis of this dataset yields valuable insight into their populations and behaviors, including the varying behaviors of bots with distinct objectives and targets. • We correlate the behaviors of seemingly isolated bots into campaigns, finding 105 clusters of requests that are malicious in nature.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- The Poorest Man in Babylon: A Longitudinal Study of Cryptocurrency Investment ScamsMuhammad Muzammil, Abisheka Pitumpe, Xigao Li, Amir Rahmati 等WWW 2025 · 被引用 13 次
- C-Frame: Characterizing and measuring in-the-wild CAPTCHA attacksHoang Dai Nguyen, Karthika Subramani, Bhupendra Acharya, Roberto Perdisci 等S&P 2024 · 被引用 5 次
- SoK: Cryptographic Authenticated DictionariesHarjasleen Malvai, Francesca Falzon, Andrew Zitek-Estrada, Sarah Meiklejohn 等NDSS 2026 · 被引用 2 次
- LEAKYLINKS: Measuring the Security and Privacy Risks of URL Scanning ServicesAli Mustafa, Jannis Rautenstrauch, Florian Hantke, Shubham Agarwal 等S&P 2026
- The Power to Never Be Wrong: Evasions and Anachronistic Attacks Against Web ArchivesRobin Kirchner, Chris Tsoukaladelis, Martin Johns, Nick NikiforakisCCS 2025
它引用的顶会 Paper8
- Throwing Darts in the Dark? Detecting Bots with Limited Data using Neural Data AugmentationSteve T. K. Jan, Qingying Hao, Tianrui Hu, Jiameng Pu 等S&P 2020 · 被引用 88 次
- Resident Evil: Understanding Residential IP Proxy as a Dark ServiceXianghang Mi, Xuan Feng, Xiaojing Liao, Baojun Liu 等S&P 2019 · 被引用 80 次
- Good Bot, Bad Bot: Characterizing Automated Browsing ActivityXigao Li, Babak Amin Azad, Amir Rahmati, Nick NikiforakisS&P 2021 · 被引用 45 次
- Certificate Transparency in the Wild: Exploring the Reliability of MonitorsBingyu Li, Jingqiang Lin, Fengjun Li, Qiongxiao Wang 等CCS 2019 · 被引用 45 次
- Does Certificate Transparency Break the Web? Measuring Adoption and Error RateEmily Stark, Ryan Sleevi, Rijad Muminovic, Devon O'Brien 等S&P 2019 · 被引用 44 次
相关 Paper
- Domains Do Change Their Spots: Quantifying Potential Abuse of Residual TrustJohnny So, Najmeh Miramirkhani, Michael Ferdman, Nick NikiforakisS&P 2022 · 被引用 15 次
- Large-scale Evaluation of Malicious Tor Hidden Service Directory DiscoveryChunmian Wang, Zhen Ling, Wenjia Wu, Qi Chen 等INFOCOM 2022 · 被引用 10 次
- Catching Phishers By Their Bait: Investigating the Dutch Phishing Landscape through Phishing Kit DetectionHugo L. J. Bijmans, Tim M. Booij, Anneke Schwedersky, Aria Nedgabat 等USENIX Security 2021 · 被引用 61 次
- Certificate Transparency Revisited: The Public Inspections on Third-party MonitorsAozhuo Sun, Jingqiang Lin, Wei Wang, Zeyan Liu 等NDSS 2024
- Who's Calling? Characterizing Robocalls through Audio and Metadata AnalysisSathvik Prasad, Elijah Robert Bouma-Sims, Athishay Kiran Mylappan, Bradley ReavesUSENIX Security 2020
