Unsupervised Process-Aware Coreset Selection for In-Context Learning
Wei Zheng, Zijie Wang, Xin Li, Bin Gong, Yuqing Sun
摘要
We address the challenge of unsupervised coreset selection for few-shot in-context learning (ICL). The goal is to select a small subset of examples under a fixed annotation budget to yield effective prompts for large language models. Existing geometry-based methods often yield coresets that suffer from a skewed distribution, due to the oversampling of peripheral examples and high local redundancy. To address these issues, we propose a process-aware framework for coreset selection. It jointly optimizes the diversity and representativeness of selected samples via an adaptive submodular objective. It ensures representativeness by selecting samples based on local density awareness, while promoting diversity by imposing a redundancy penalty relative to the evolving selected set. Thus, it performs progress-aware balancing of representativeness and diversity based on the selection context. Extensive experiments on 7 NLP datasets demonstrate that our method consistently outperforms state-of-the-art coreset selection methods in downstream ICL performance. Further analysis validates that our approach better balances diversity and representativeness in the selection process, while retaining the theoretical guarantees of adaptive submodular optimization.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper14
- Deep Learning on a Data Diet: Finding Important Examples Early in TrainingMansheej Paul, Surya Ganguli, Gintare Karolina DziugaiteNeurIPS 2021 · 被引用 806 次
- Compositional Exemplars for In-context LearningJiacheng Ye, Zhiyong Wu, Jiangtao Feng, Tao Yu 等ICML 2023 · 被引用 188 次
- Batch Calibration: Rethinking Calibration for In-Context Learning and Prompt EngineeringHan Zhou, Xingchen Wan, Lev Proleev, Diana Mincu 等ICLR 2024 · 被引用 90 次
- Selective Annotation Makes Language Models Better Few-Shot LearnersHongjin Su, Jungo Kasai, Chen Henry Wu, Weijia Shi 等ICLR 2023 · 被引用 63 次
- Towards Sustainable Learning: Coresets for Data-efficient Deep LearningYu Yang, Hao Kang, Baharan MirzasoleimanICML 2023 · 被引用 58 次
相关 Paper
- Context Tuning for In-Context OptimizationJack Lu, Ryan Teehan, Zhenbang Yang, Mengye RenICML 2026
- Meta-Adaptive Prompt Distillation for Few-Shot Visual Question AnsweringAkash Gupta, Amos Storkey, Mirella LapataICLR 2026
- Optimization Inspired Few-Shot Adaptation for Large Language ModelsBoyan Gao, Xin Wang, Yibo Yang, David A. CliftonNeurIPS 2025 · 被引用 3 次
- Effective Demonstration Annotation for In-Context Learning via Language Model-Based Determinantal Point ProcessPeng Wang, Xiaobin Wang, Chao Lou, Shengyu Mao 等EMNLP 2024
- In-Context Learning as Rate–Distortion OptimizationJiayu Zhang, Changbang Li, Canran XiaoICML 2026
