OVLR: Efficient, Scalable, and Robust Training via Output-Level Variance-Reduced Likelihood Ratio
Minhao Zou, Tao Ren, Jinyang Jiang, Rui Tao, Zehao Li, Jiale Fu, Hui Shao, Xianhua Liu, Yijie Peng
Abstract
Gradient-based optimization via backpropagation (BP) is inherently limited by the requirement of differentiability, rendering it inapplicable for piecewise-constant objectives with vanishing gradients (e.g., the hard 0-1 loss) or black-box feedback. While likelihood ratio (LR) methods offer a theoretical alternative, their high variance in high-dimensional spaces undermines training stability and scalability. We propose OVLR, a framework that makes direct optimization of gradient-agnostic objectives practical for modern deep networks by performing perturbations and antithetic sampling in the low-dimensional output space. OVLR achieves dramatic variance reduction while requiring only a single deterministic forward pass, with additional costs restricted to evaluating the loss function across multiple samples. On problems where BP provides gradients, OVLR remains competitive; on problems where BP fails to provide reliable learning signals, OVLR enables the direct optimization of objectives such as the 0-1 loss for noise-tolerant classification and truncated losses for outlier-resistant regression. Extensive empirical results across classification, generative modeling, language modeling, robot imitation learning, and black-box optimization confirm that OVLR is an effective tool for settings where standard gradient-based optimization is inapplicable. Code is available at https://github.com/MinhZou/OVLR.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on4
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- One Forward is Enough for Neural Network Training via Likelihood Ratio MethodJinyang Jiang, Zeliang Zhang, Chenliang Xu, Zhaofei Yu et al.ICLR 2024 · 14 citations
- Half-order Fine-Tuning for Diffusion Model: A Recursive Likelihood Ratio OptimizerTao Ren, Zishi Zhang, Jinyang Jiang, Zehao Li et al.ICLR 2026 · 5 citations
- Unbalanced Optimal Total Variation Transport: A Theoretical Approach to Spatial Resource Allocation ProblemsNhan-Phu Chung, Jinhui Han, Bohan Li, Zehao LiNeurIPS 2025 · 1 citation
Related papers
- Natural Perturbations for Black-box Training of Neural Networks by Zeroth-Order OptimizationHiroshi Sawada, Kazuo Aoyama, Yuya HikimaICML 2025
- Learning to Learn by Zeroth-Order OracleYangjun Ruan, Yuanhao Xiong, Sashank J. Reddi, Sanjiv Kumar et al.ICLR 2020 · 21 citations
- On the Design of Black-Box Adversarial Examples by Leveraging Gradient-Free Optimization and Operator Splitting MethodPu Zhao, Sijia Liu, Pin-Yu Chen, Nghia Hoang et al.ICCV 2019 · 61 citations
- Towards Constituting Mathematical Structures for Learning to OptimizeJialin Liu, Xiaohan Chen, Zhangyang Wang, Wotao Yin et al.ICML 2023 · 18 citations
- Path Sample-Analytic Gradient Estimators for Stochastic Binary NetworksAlexander Shekhovtsov, Viktor Yanush, Boris FlachNeurIPS 2020 · 14 citations
