Gradient-Guided Multi-Judge Prompt Optimization
Chenzhuo Zhao, Xinda Wang, Pu Zhao, Yue Huang, Junting Lu, Ziqian Liu, Qingwei Lin, Saravan Rajmohan, Dongmei Zhang
摘要
Automatic prompt optimization is a practical alternative to fine-tuning for adapting large language models (LLMs), yet existing approaches often trade off signal quality against computational cost. Methods that rely on generative feedback can be informative but expensive to scale, while sampling-based optimization typically requires many evaluations and exhibits high variance. Even loss-driven prompt optimization remains limited by costly segment attribution that scales with prompt length and by overfitting to a single evaluator, which weakens transfer across model families and domains. We propose Gradient-guided Multi-judge Prompt Optimization (GMPO), a scalable framework that improves both efficiency and robustness. GMPO uses a firstorder gradient approximation to score segment importance in a continuous masking direction, requiring only one forward and one backward pass. GMPO further employs a generate multi-judge design in which candidate prompt edits are proposed by a generator and selected using cross-entropy losses aggregated from multiple lightweight judge models, reducing evaluator bias and improving generalization. Experiments across math, reasoning, instruction-following evaluation, and safety robustness benchmarks demonstrate consistent gains with substantially lower optimization overhead. Our code is available at https://github.com/cyzcz/GMPO . * Corresponding author. * "X days have passed since [date]" today = [date] plus X days → -Always explicitly state what "today" is before proceeding → 3. PARSE TEMPORAL RELATIONSHIPS -Understand directional phrases: * "before" means subtract days * "after" means add days * "since" typically means counting forward from a past date → -Be precise about whether events are described relative to "today" or to other dates → 4. PERFORM DATE CALCULATIONS -Use correct month lengths:
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper7
- Large Language Models as OptimizersChengrun Yang, Xuezhi Wang, Yifeng Lu, Hanxiao Liu 等ICLR 2024 · 被引用 817 次
- Take a Step Back: Evoking Reasoning via Abstraction in Large Language ModelsHuaixiu Steven Zheng, Swaroop Mishra, Xinyun Chen, Heng-Tze Cheng 等ICLR 2024 · 被引用 216 次
- RLPrompt: Optimizing Discrete Text Prompts with Reinforcement LearningMingkai Deng, Jianyu Wang, Cheng-Ping Hsieh, Yihan Wang 等EMNLP 2022 · 被引用 141 次
- On the Worst Prompt Performance of Large Language ModelsBowen Cao, Deng Cai, Zhisong Zhang, Yuexian Zou 等NeurIPS 2024 · 被引用 65 次
- Prompt Optimization with EASE? Efficient Ordering-aware Automated Selection of ExemplarsZhaoxuan Wu, Xiaoqiang Lin, Zhongxiang Dai, Wenyang Hu 等NeurIPS 2024 · 被引用 44 次
相关 Paper
- Unleashing the Potential of Large Language Models as Prompt Optimizers: Analogical Analysis with Gradient-based Model OptimizersXinyu Tang, Xiaolei Wang, Wayne Xin Zhao, Siyuan Lu 等AAAI 2025 · 被引用 36 次
- GReaTer: Gradients Over Reasoning Makes Smaller Language Models Strong Prompt OptimizersSarkar Snigdha Sarathi Das, Ryo Kamoi, Bo Pang, Yusen Zhang 等ICLR 2025
- ZERA: Zero-init Instruction Evolving Refinement Agent - From Zero Instructions to Structured Prompts via Principle-based OptimizationSeungyoun Yi, Minsoo Khang, Sungrae ParkEMNLP 2025
- PRompt Optimization in Multi-Step Tasks (PROMST): Integrating Human Feedback and Heuristic-based SamplingYongchao Chen, Jacob Arkin, Yilun Hao, Yang Zhang 等EMNLP 2024 · 被引用 6 次
- MASPO: Joint Prompt Optimization for LLM-based Multi-Agent SystemsZhexuan Wang, Xuebo Liu, Li Wang, Zifei Shan 等ICML 2026 · 被引用 3 次
