Zeroth-Order Optimization with Trajectory-Informed Derivative Estimation
Yao Shu, Zhongxiang Dai, Weicong Sng, Arun Verma, Patrick Jaillet, Bryan Kian Hsiang Low
摘要
Zeroth-order (ZO) optimization, in which the derivative is unavailable, has recently succeeded in many important machine learning applications. Existing algorithms rely on finite difference (FD) methods for derivative estimation and gradient descent (GD)-based approaches for optimization. However, these algorithms suffer from query inefficiency because many additional function queries are required for derivative estimation in their every GD update, which typically hinders their deployment in real-world applications where every function query is expensive. To this end, we propose a trajectory-informed derivative estimation method which only employs the optimization trajectory (i.e., the history of function queries during optimization) and hence can eliminate the need for additional function queries to estimate a derivative. Moreover, based on our derivative estimation, we propose the technique of dynamic virtual updates, which allows us to reliably perform multiple steps of GD updates without reapplying derivative estimation. Based on these two contributions, we introduce the zeroth-order optimization with trajectory-informed derivative estimation (ZORD) algorithm for query-efficient ZO optimization. We theoretically demonstrate that our trajectory-informed derivative estimation and our ZORD algorithm improve over existing approaches, which is then supported by our real-world experiments such as black-box adversarial attack, non-differentiable metric optimization, and derivative-free reinforcement learning. * Equal contribution.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- Revisiting Zeroth-Order Optimization for Memory-Efficient LLM Fine-Tuning: A BenchmarkYihua Zhang, Pingzhi Li, Junyuan Hong, Jiaxiang Li 等ICML 2024 · 被引用 134 次
- DeepZero: Scaling Up Zeroth-Order Optimization for Deep Model TrainingAochuan Chen, Yimeng Zhang, Jinghan Jia, James Diffenderfer 等ICLR 2024 · 被引用 88 次
- Sparse MeZO: Less Parameters for Better Performance in Zeroth-Order LLM Fine-TuningYong Liu, Zirui Zhu, Chaoyu Gong, Minhao Cheng 等NeurIPS 2025 · 被引用 66 次
- Localized Zeroth-Order Prompt OptimizationWenyang Hu, Yao Shu, Zongmin Yu, Zhaoxuan Wu 等NeurIPS 2024 · 被引用 30 次
- The Behavior and Convergence of Local Bayesian OptimizationKaiwen Wu, Kyurae Kim, Roman Garnett, Jacob R. GardnerNeurIPS 2023 · 被引用 27 次
它引用的顶会 Paper7
- Re-Examining Linear Embeddings for High-Dimensional Bayesian OptimizationBenjamin Letham, Roberto Calandra, Akshara Rai, Eytan BakshyNeurIPS 2020 · 被引用 152 次
- Federated Bayesian Optimization via Thompson SamplingZhongxiang Dai, Bryan Kian Hsiang Low, Patrick JailletNeurIPS 2020 · 被引用 144 次
- Gradientless Descent: High-Dimensional Zeroth-Order OptimizationDaniel Golovin, John Karro, Greg Kochanski, Chansoo Lee 等ICLR 2020 · 被引用 85 次
- BayesOpt Adversarial AttackBinxin Ru, Adam D. Cobb, Arno Blaas, Yarin GalICLR 2020 · 被引用 85 次
- On the Convergence of Prior-Guided Zeroth-Order Optimization AlgorithmsShuyu Cheng, Guoqiang Wu, Jun ZhuNeurIPS 2021 · 被引用 27 次
相关 Paper
- Learning to Learn by Zeroth-Order OracleYangjun Ruan, Yuanhao Xiong, Sashank J. Reddi, Sanjiv Kumar 等ICLR 2020 · 被引用 21 次
- Zeroth-Order Optimization Finds Flat MinimaLiang Zhang, Bingcong Li, Kiran Koshy Thekumparampil, Sewoong Oh 等NeurIPS 2025 · 被引用 8 次
- ReLIZO: Sample Reusable Linear Interpolation-based Zeroth-order OptimizationXiaoxing Wang, Xiaohan Qin, Xiaokang Yang, Junchi YanNeurIPS 2024 · 被引用 10 次
- Turning Stale Gradients into Stable Gradients: Coherent Coordinate Descent with Implicit Landscape Smoothing for Lightweight Zeroth-Order OptimizationChen Liang, Xiatao Sun, Qian Wang, Daniel RakitaICML 2026
- Towards Query-Efficient Black-Box Adversary with Zeroth-Order Natural Gradient DescentPu Zhao, Pin-Yu Chen, Siyue Wang, Xue LinAAAI 2020 · 被引用 42 次
