CTBench: A Library and Benchmark for Certified Training
Yuhao Mao, Stefan Balauca, Martin T. Vechev
Abstract
Training certifiably robust neural networks is an important but challenging task. While many algorithms for (deterministic) certified training have been proposed, they are often evaluated on different training schedules, certification methods, and systematically under-tuned hyperparameters, making it difficult to compare their performance. To address this challenge, we introduce CTBENCH, a unified library and a high-quality benchmark for certified training that evaluates all algorithms under fair settings and systematically tuned hyperparameters. We show that (1) almost all algorithms in CTBENCH surpass the corresponding reported performance in literature in the magnitude of algorithmic improvements, thus establishing new state-of-the-art, and (2) the claimed advantage of recent algorithms drops significantly when we enhance the outdated baselines with a fair training schedule, a fair certification method and well-tuned hyperparameters. Based on CTBENCH, we provide new insights into the current state of certified training, including (1) certified models have less fragmented loss surface, (2) certified models share many mistakes, (3) certified models have more sparse activations, (4) reducing regularization cleverly is crucial for certified training especially for large radii and ( 5 ) certified training has the potential to improve outof-distribution generalization. We are confident that CTBENCH will serve as a benchmark and testbed for future research in certified training.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e93cc246-8747-4fc0-ba86-fb559ede4314Cited by top-tier papers4
- Expressiveness of Multi-Neuron Convex Relaxations in Neural Network CertificationYuhao Mao, Yani Zhang, Martin T. VechevICLR 2026 · 4 citations
- Dual Randomized Smoothing: Beyond Global Noise VarianceChenhao Sun, Yuhao Mao, Martin VechevICLR 2026 · 1 citation
- MIBP-Cert: Certified Training against Data Perturbations with Mixed-Integer Bilinear ProgramsTobias Lorenz, Marta Kwiatkowska, Mario FritzNeurIPS 2025 · 1 citation
- Average Certified Radius is a Poor Metric for Randomized SmoothingChenhao Sun, Yuhao Mao, Mark Niklas Müller, Martin T. VechevICML 2025
Builds on18
- Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacksFrancesco Croce, Matthias HeinICML 2020 · 2,337 citations
- On Adaptive Attacks to Adversarial Example DefensesFlorian Tramèr, Nicholas Carlini, Wieland Brendel, Aleksander MadryNeurIPS 2020 · 1,026 citations
- AI2: Safety and Robustness Certification of Neural Networks with Abstract InterpretationTimon Gehr, Matthew Mirman, Dana Drachsler-Cohen, Petar Tsankov et al.S&P 2018 · 987 citations
- Automatic Perturbation Analysis for Scalable Certified Robustness and BeyondKaidi Xu, Zhouxing Shi, Huan Zhang, Yihan Wang et al.NeurIPS 2020 · 415 citations
- Towards Stable and Efficient Training of Verifiably Robust Neural NetworksHuan Zhang, Hongge Chen, Chaowei Xiao, Sven Gowal et al.ICLR 2020 · 384 citations
Related papers
- Rethinking Evaluation Paradigms in IBP-based Certified TrainingKonstantin Kaulen, Hadar Shavit, Holger HoosICML 2026
- SoK: Certified Robustness for Deep Neural NetworksLinyi Li, Tao Xie, Bo LiS&P 2023
- Boosting Verified Training for Robust Image Classifications via AbstractionZhaodi Zhang, Zhiyi Xue, Yang Chen, Si Liu et al.CVPR 2023
- Provably Cost-Sensitive Adversarial Defense via Randomized SmoothingYuan Xin, Dingfan Chen, Michael Backes, Xiao ZhangICML 2025
- Towards Better Understanding of Training Certifiably Robust Models against Adversarial ExamplesSungyoon Lee, Woojin Lee, Jinseong Park, Jaewook LeeNeurIPS 2021 · 27 citations
