Dynamic Rewarding with Prompt Optimization Enables Tuning-free Self-Alignment of Language Models
Somanshu Singla, Zhen Wang, Tianyang Liu, Abdullah Ashfaq, Zhiting Hu, Eric P. Xing
Abstract
Aligning Large Language Models (LLMs) traditionally relies on costly training and human preference annotations.Self-alignment aims to reduce these expenses by aligning models by themselves.To further minimize the cost and enable LLM alignment without any expensive tuning and annotations, we introduce a new tuning-free approach for self-alignment, called Dynamic Rewarding with Prompt Optimization (DRPO).Our approach leverages a search-based optimization framework that allows LLMs to iteratively self-improve and design the best alignment instructions without the need for additional training or human intervention.The core of DRPO is a dynamic rewarding mechanism, which identifies and rectifies model-specific alignment weaknesses, allowing LLMs to adapt efficiently to diverse alignment challenges.Empirical evaluations on eight recent LLMs, both open-and closed-source, reveal that DRPO significantly enhances alignment performance, with base models outperforming their SFT/RLHF-tuned counterparts.Moreover, DRPO's automatically optimized prompts surpass those curated by human experts, further validating the effectiveness of our approach.Our findings highlight the great potential of current LLMs to be adaptively selfaligned through inference-time optimization, complementing existing tuning-based alignment research.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bd532e51-d88c-4524-a5e7-4cd57c827d75Cited by top-tier papers5
- What Makes a Good Natural Language Prompt?Do Xuan Long, Duy Dinh, Ngoc-Hai Nguyen, Kenji Kawaguchi et al.ACL 2025 · 13 citations
- AdAEM: An Adaptively and Automated Extensible Measurement of LLMs' Value DifferenceJing Yao, Shitong Duan, Xiaoyuan Yi, Dongkuan Xu et al.ICLR 2026 · 4 citations
- See Further, Think Deeper: Advancing VLM's Reasoning Ability with Low-level Visual Cues and ReflectionZhiheng Wu, Tong Wang, Shuning Wang, Naiming Liu et al.CVPR 2026 · 3 citations
- LLMdoctor: Token-Level Flow-Guided Preference Optimization for Efficient Test-Time Alignment of Large Language ModelsTiesunlong Shen, Rui Mao, Jin Wang, Heming Sun et al.AAAI 2026 · 2 citations
- CritiQ: Mining Data Quality Criteria from Human PreferencesHonglin Guo, Kai Lv, Qipeng Guo, Tianyi Liang et al.ACL 2025
Builds on19
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Self-Refine: Iterative Refinement with Self-FeedbackAman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan et al.NeurIPS 2023 · 4,972 citations
- Inference-Time Intervention: Eliciting Truthful Answers from a Language ModelKenneth Li, Oam Patel, Fernanda B. Viégas, Hanspeter Pfister et al.NeurIPS 2023 · 1,549 citations
- Large Language Models as OptimizersChengrun Yang, Xuezhi Wang, Yifeng Lu, Hanxiao Liu et al.ICLR 2024 · 817 citations
Related papers
- Direct Large Language Model Alignment Through Self-Rewarding Contrastive Prompt DistillationAiwei Liu, Haoping Bai, Zhiyun Lu, Xiang Kong et al.ACL 2024 · 4 citations
- LLM Alignment as Retriever Optimization: An Information Retrieval PerspectiveBowen Jin, Jinsung Yoon, Zhen Qin, Ziqi Wang et al.ICML 2025
- Provably Mitigating Overoptimization in RLHF: Your SFT Loss is Implicitly an Adversarial RegularizerZhihan Liu, Miao Lu, Shenao Zhang, Boyi Liu et al.NeurIPS 2024 · 119 citations
- Inverse Reinforcement Learning with Dynamic Reward Scaling for LLM AlignmentRuoxi Cheng, Haoxuan Ma, Weixin Wang, Ranjie Duan et al.ICLR 2026 · 23 citations
- Aligning Large Language Models via Fully Self-Synthetic DataShangjian Yin, Zhepei Wei, Xinyu Zhu, Wei-Lin Chen et al.ACL 2026 · 2 citations
