USENIX Security2022Top-tier venue
Uninvited Guests: Analyzing the Identity and Behavior of Certificate Transparency Bots
Brian Kondracki, Johnny So, Nick Nikiforakis
Abstract
Since its creation, Certificate Transparency (CT) has served as a vital component of the secure web. However, with the increase in TLS adoption, CT has essentially become a defacto log for all newly-created websites, announcing to the public the existence of web endpoints, including those that could have otherwise remained hidden. As a result, web bots can use CT to probe websites in real time, as they are created. Little is known about these bots, their behaviors, and their intentions. In this paper we present CTPOT, a distributed honeypot system which creates new TLS certificates for the purpose of advertising previously non-existent domains, and records the activity generated towards them from a number of network vantage points. Using CTPOT, we create 4,657 TLS certificates over a period of ten weeks, attracting 1.5 million web requests from 31,898 unique IP addresses. We find that CT bots occupy a distinct subset of the overall web bot population, with less than 2% overlap between IP addresses of CT bots and traditional host-scanning web bots. By creating certificates with varying content types, we are able to further sub-divide the CT bot population into subsets of varying intentions, revealing a stark contrast in malicious behavior among these groups. Finally, we correlate observed bot IP addresses into campaigns using the file paths requested by each bot, and find 105 malicious campaigns targeting the domains we advertise. Our findings shed light onto the CT bot ecosystem, revealing that it is not only distinct to that of traditional IP-based bots, but is composed of numerous entities with varying targets and behaviors. categories, allowing us to observe the varying behaviors of bots with different goals. Using CTPOT, we create 4,657 TLS certificates across 12 measurement nodes for pseudo-random subdomains corresponding to: popular trademarks, common web endpoints of web application software, and dictionary words which act as a baseline through which we compare and interpret the results of our measurement groups. Over the course of ten weeks, CTPOT received a total of 1.5 million web requests from 31,898 unique IP addresses, as well as 22,839 requests to SSH, FTP, and Telnet honeypots. As each domain generated by CTPOT was previously unused and completely unguessable, our curated dataset consists entirely of bot traffic. By analyzing our curated dataset, we observe distinct behaviors from the bots targeting different types of subdomains, allowing us to learn about the intentions of these various bot populations. For instance, we observe bots targeting subdomains of common web application software send more than one request over 40% more often than bots targeting domains impersonating popular trademarks. Moreover, these bots exhibit more malicious behavior with over twice as many unique IP addresses attempting to authenticate with exposed network services such as SSH. Additionally, we correlate requests from seemingly isolated IP addresses into campaigns of related bots. Alarmingly, through this analysis we find 105 malicious campaigns attempting to perform malicious actions such as data exfiltration, fingerprinting, and vulnerability exploitation. Our main contributions are as follows: • We design and implement CTPOT, a honeypot-based system to create TLS certificates for pseudo-random subdomains, and analyze requests directed towards them. Using this system, we create 4,657 TLS certificates. • We curate the first public dataset of CT bots. Analysis of this dataset yields valuable insight into their populations and behaviors, including the varying behaviors of bots with distinct objectives and targets. • We correlate the behaviors of seemingly isolated bots into campaigns, finding 105 clusters of requests that are malicious in nature.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext da28c8dc-8304-43af-9bc9-89af85cb4dccCited by top-tier papers6
- The Poorest Man in Babylon: A Longitudinal Study of Cryptocurrency Investment ScamsMuhammad Muzammil, Abisheka Pitumpe, Xigao Li, Amir Rahmati et al.WWW 2025 · 13 citations
- C-Frame: Characterizing and measuring in-the-wild CAPTCHA attacksHoang Dai Nguyen, Karthika Subramani, Bhupendra Acharya, Roberto Perdisci et al.S&P 2024 · 5 citations
- SoK: Cryptographic Authenticated DictionariesHarjasleen Malvai, Francesca Falzon, Andrew Zitek-Estrada, Sarah Meiklejohn et al.NDSS 2026 · 2 citations
- LEAKYLINKS: Measuring the Security and Privacy Risks of URL Scanning ServicesAli Mustafa, Jannis Rautenstrauch, Florian Hantke, Shubham Agarwal et al.S&P 2026
- The Power to Never Be Wrong: Evasions and Anachronistic Attacks Against Web ArchivesRobin Kirchner, Chris Tsoukaladelis, Martin Johns, Nick NikiforakisCCS 2025
Builds on8
- Throwing Darts in the Dark? Detecting Bots with Limited Data using Neural Data AugmentationSteve T. K. Jan, Qingying Hao, Tianrui Hu, Jiameng Pu et al.S&P 2020 · 88 citations
- Resident Evil: Understanding Residential IP Proxy as a Dark ServiceXianghang Mi, Xuan Feng, Xiaojing Liao, Baojun Liu et al.S&P 2019 · 80 citations
- Good Bot, Bad Bot: Characterizing Automated Browsing ActivityXigao Li, Babak Amin Azad, Amir Rahmati, Nick NikiforakisS&P 2021 · 45 citations
- Certificate Transparency in the Wild: Exploring the Reliability of MonitorsBingyu Li, Jingqiang Lin, Fengjun Li, Qiongxiao Wang et al.CCS 2019 · 45 citations
- Does Certificate Transparency Break the Web? Measuring Adoption and Error RateEmily Stark, Ryan Sleevi, Rijad Muminovic, Devon O'Brien et al.S&P 2019 · 44 citations
Related papers
- Domains Do Change Their Spots: Quantifying Potential Abuse of Residual TrustJohnny So, Najmeh Miramirkhani, Michael Ferdman, Nick NikiforakisS&P 2022 · 15 citations
- Large-scale Evaluation of Malicious Tor Hidden Service Directory DiscoveryChunmian Wang, Zhen Ling, Wenjia Wu, Qi Chen et al.INFOCOM 2022 · 10 citations
- Catching Phishers By Their Bait: Investigating the Dutch Phishing Landscape through Phishing Kit DetectionHugo L. J. Bijmans, Tim M. Booij, Anneke Schwedersky, Aria Nedgabat et al.USENIX Security 2021 · 61 citations
- Certificate Transparency Revisited: The Public Inspections on Third-party MonitorsAozhuo Sun, Jingqiang Lin, Wei Wang, Zeyan Liu et al.NDSS 2024
- Who's Calling? Characterizing Robocalls through Audio and Metadata AnalysisSathvik Prasad, Elijah Robert Bouma-Sims, Athishay Kiran Mylappan, Bradley ReavesUSENIX Security 2020
