Demonstration Selection for In-Context Learning via Reinforcement Learning
Xubin Wang, Jianfei Wu, Yichen Yuan, Deyu Cai, Mingzhe Li, Weijia Jia
摘要
Diversity in demonstration selection is critical for enhancing model generalization by enabling broader coverage of structures and concepts. Constructing appropriate demonstration sets remains a key research challenge. This paper introduces the Relevance-Diversity Enhanced Selection (RDES), an innovative approach that leverages reinforcement learning (RL) frameworks to optimize the selection of diverse reference demonstrations for tasks amenable to in-context learning (ICL), particularly text classification and reasoning, in fewshot prompting scenarios. RDES employs frameworks like Q-learning and a PPO-based variant to dynamically identify demonstrations that maximize both diversity (quantified by label distribution) and relevance to the task objective. This strategy ensures a balanced representation of reference data, leading to improved accuracy and generalization. Through extensive experiments on multiple benchmark datasets, including diverse reasoning tasks, and involving 14 closedsource and open-source LLMs, we demonstrate that RDES significantly enhances performance compared to ten established baselines. Our evaluation includes analysis of performance across varying numbers of demonstrations on selected datasets. Furthermore, we investigate incorporating Chain-of-Thought (CoT) reasoning, which further boosts predictive performance. The results highlight the potential of RL for adaptive demonstration selection and addressing challenges in ICL.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- DICE: Dynamic In-Context Example Selection in LLM Agents via Efficient Knowledge TransferRuoyu Wang, Junda Wu, Yu Xia, Tong Yu 等KDD 2026 · 被引用 6 次
- SAFE-SQL: Self-Augmented In-Context Learning with Fine-grained Example Selection for Text-to-SQLJimin Lee, Ingeol Baek, Byeongjeong Kim, Hyunkyung Bae 等EMNLP 2025 · 被引用 1 次
- Auto-regressive In-context Demonstration SelectionYunzhe Qi, Sirui Chen, Jiaru Zou, Jingrui HeICML 2026
- ContextIF: Enhancing Instruction-Following through Context RewardYule Zhong, Jiacheng Yao, Guoxiu HeICLR 2026
- Unsupervised Process-Aware Coreset Selection for In-Context LearningWei Zheng, Zijie Wang, Xin Li, Bin Gong 等ICML 2026
它引用的顶会 Paper18
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Self-Refine: Iterative Refinement with Self-FeedbackAman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan 等NeurIPS 2023 · 被引用 4,972 次
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu 等ICLR 2022 · 被引用 4,966 次
- Calibrate Before Use: Improving Few-shot Performance of Language ModelsZihao Zhao, Eric Wallace, Shi Feng, Dan Klein 等ICML 2021 · 被引用 1,843 次
相关 Paper
- Provoking Multi-modal Few-Shot LVLM via Exploration-Exploitation In-Context LearningCheng Chen, Yunpeng Zhai, Yifan Zhao, Jinyang Gao 等CVPR 2025
- Representative Demonstration Selection for In-Context Learning with Two-Stage Determinantal Point ProcessZhao Yang, Yuanzhe Zhang, Dianbo Sui, Cao Liu 等EMNLP 2023 · 被引用 3 次
- Many-Shot CoT-ICL: Making In-Context Learning Truly LearnTsz Ting Chung, Lemao Liu, Mo Yu, Dit-Yan YeungICML 2026 · 被引用 3 次
- Universal Self-Adaptive PromptingXingchen Wan, Ruoxi Sun, Hootan Nakhost, Hanjun Dai 等EMNLP 2023 · 被引用 4 次
- Generating Diverse Training Samples for Relation Extraction with Large Language ModelsZexuan Li, Hongliang Dai, Piji LiACL 2025
