Zeroth-Order Optimization Meets Human Feedback: Provable Learning via Ranking Oracles
Zhiwei Tang, Dmitry Rybin, Tsung-Hui Chang
Abstract
In this study, we delve into an emerging optimization challenge involving a blackbox objective function that can only be gauged via a ranking oracle-a situation frequently encountered in real-world scenarios, especially when the function is evaluated by human judges. Such challenge is inspired from Reinforcement Learning with Human Feedback (RLHF), an approach recently employed to enhance the performance of Large Language Models (LLMs) using human guidance (Ouyang et al., 2022; Liu et al., 2023; OpenAI, 2022; Bai et al., 2022) . We introduce ZO-RankSGD, an innovative zeroth-order optimization algorithm designed to tackle this optimization problem, accompanied by theoretical assurances. Our algorithm utilizes a novel rank-based random estimator to determine the descent direction and guarantees convergence to a stationary point. Moreover, ZO-RankSGD is readily applicable to policy optimization problems in Reinforcement Learning (RL), particularly when only ranking oracles for the episode reward are available. Last but not least, we demonstrate the effectiveness of ZO-RankSGD in a novel application: improving the quality of images generated by a diffusion generative model with human ranking feedback. Throughout experiments, we found that ZO-RankSGD can significantly enhance the detail of generated images with only a few rounds of human feedback. Overall, our work advances the field of zeroth-order optimization by addressing the problem of optimizing functions with only ranking feedback, and offers a new and effective approach for aligning Artificial Intelligence (AI) with human intentions.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers18
- Fine-Tuning Language Models with Just Forward PassesSadhika Malladi, Tianyu Gao, Eshaan Nichani, Alex Damian et al.NeurIPS 2023 · 495 citations
- Accelerating Parallel Sampling of Diffusion ModelsZhiwei Tang, Jiasheng Tang, Hao Luo, Fan Wang et al.ICML 2024 · 30 citations
- Multimodal Large Language Models Make Text-to-Image Generative Models Align BetterXun Wu, Shaohan Huang, Guolong Wang, Jing Xiong et al.NeurIPS 2024 · 26 citations
- Boosting Text-to-Video Generative Model with MLLMs FeedbackXun Wu, Shaohan Huang, Guolong Wang, Jing Xiong et al.NeurIPS 2024 · 23 citations
- ComPO: Preference Alignment via Comparison OraclesPeter Chen, Xi Chen, Wotao Yin, Tianyi LinNeurIPS 2025 · 20 citations
Builds on17
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
Related papers
- Modality-Agnostic Zeroth-Order LoRA Fine-Tuning for Black-Box Prompt OptimizationXingchen Li, Jia Zhang, Tianxing Man, Wenkang Wang et al.KDD 2026
- Score as Action: Fine Tuning Diffusion Generative Models by Continuous-time Reinforcement LearningHanyang Zhao, Haoxian Chen, Ji Zhang, David D. Yao et al.ICML 2025
- Zeroth-Order Policy Gradient for Reinforcement Learning from Human Feedback without Reward InferenceQining Zhang, Lei YingICLR 2025
- Reinforcement Learning for Fine-tuning Text-to-Image Diffusion ModelsYing Fan, Olivia Watkins, Yuqing Du, Hao Liu et al.NeurIPS 2023 · 372 citations
- Training Diffusion Models with Reinforcement LearningKevin Black, Michael Janner, Yilun Du, Ilya Kostrikov et al.ICLR 2024 · 816 citations
