DARA: Few-shot Budget Allocation in Online Advertising via In-Context Decision Making with RL-Finetuned LLMs
Mingxuan Song, Yusen Huo, Bohan Zhou, Shenglin Yin, Zhen Xiao, Jieyi Long, Zhilin Zhang, Chuan Yu
Abstract
Optimizing the advertiser's cumulative value of winning impressions under budget constraints poses a complex challenge in online advertising, under the paradigm of AI-Generated Bidding (AIGB). Advertisers often have personalized objectives but limited historical interaction data, resulting in few-shot scenarios where traditional reinforcement learning (RL) methods struggle to perform effectively. Large Language Models (LLMs) offer a promising alternative for AIGB by leveraging their in-context learning capabilities to generalize from limited data. However, they lack the numerical precision required for fine-grained optimization. To address this limitation, we introduce GRPO-Adaptive, an efficient LLM post-training strategy that enhances both reasoning and numerical precision by dynamically updating the reference policy during training. Built upon this foundation, we further propose DARA, a novel dual-phase framework that decomposes the decision-making process into two stages: a few-shot reasoner that generates initial plans via in-context prompting, and a fine-grained optimizer that refines these plans using feedback-driven reasoning. This separation allows DARA to combine LLMs' in-context learning strengths with precise adaptability required by AIGB tasks. Extensive experiments on both real-world and synthetic data environments demonstrate that our approach consistently outperforms existing baselines in terms of cumulative advertiser value under budget constraints.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1bbd283a-ca00-440a-afd9-1fa91b63337bBuilds on7
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- LayoutPrompter: Awaken the Design Ability of Large Language ModelsJiawei Lin, Jiaqi Guo, Shizhao Sun, Zijiang Yang et al.NeurIPS 2023 · 71 citations
- SPRING: Improving the Throughput of Sharding Blockchain via Deep Reinforcement Learning Based State PlacementPengze Li, Mingxuan Song, Mingzhe Xing, Zhen Xiao et al.WWW 2024 · 32 citations
- MAP the Blockchain World: A Trustless and Scalable Blockchain Interoperability Protocol for Cross-chain ApplicationsYinfeng Cao, Jiannong Cao, Dongbin Bai, Long Wen et al.WWW 2025 · 18 citations
Related papers
- LBM: Hierarchical Large Auto-Bidding Model via Reasoning and ActingYewen Li, Zhiyi Lyu, Peng Jiang, Qingpeng Cai et al.WWW 2026
- Enhancing Generative Auto-bidding with Offline Reward Evaluation and Policy SearchZhiyu Mou, Yiqin Lv, Miao Xu, Qi Wang et al.ICLR 2026 · 4 citations
- Autobidding Auctions with LLM-Powered CreativesBingzhe Wang, Bowei Zhang, Changyuan Yu, Qi QiICML 2026
- Reason for Future, Act for Now: A Principled Architecture for Autonomous LLM AgentsZhihan Liu, Hao Hu, Shenao Zhang, Hongyi Guo et al.ICML 2024 · 17 citations
- Group Distributionally Robust Optimization-Driven RL for LLM ReasoningKishan Panaganti, Zhenwen Liang, Wenhao Yu, Haitao Mi et al.ICML 2026 · 5 citations
