Effective Demonstration Annotation for In-Context Learning via Language Model-Based Determinantal Point Process
Peng Wang, Xiaobin Wang, Chao Lou, Shengyu Mao, Pengjun Xie, Yong Jiang
Abstract
In-context learning (ICL) is a few-shot learning paradigm that involves learning mappings through input-output pairs and appropriately applying them to new instances. Despite the remarkable ICL capabilities demonstrated by Large Language Models (LLMs), existing works are highly dependent on large-scale labeled support sets, not always feasible in practical scenarios. To refine this approach, we focus primarily on an innovative selective annotation mechanism, which precedes the standard demonstration retrieval. We introduce the Language Model-based Determinant Point Process (LM-DPP) that simultaneously considers the uncertainty and diversity of unlabeled instances for optimal selection. Consequently, this yields a subset for annotation that strikes a trade-off between the two factors. We apply LM-DPP to various language models, including GPT-J, LlaMA, and GPT-3. Experimental results on 9 NLU and 2 Generation datasets demonstrate that LM-DPP can effectively select canonical examples. Further analysis reveals that LLMs benefit most significantly from subsets that are both low uncertainty and high diversity.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7f840919-cc25-4144-99ee-24f66901b77cCited by top-tier papers3
- What Makes a Good Natural Language Prompt?Do Xuan Long, Duy Dinh, Ngoc-Hai Nguyen, Kenji Kawaguchi et al.ACL 2025 · 13 citations
- Principled Content Selection to Generate Diverse and Personalized Multi-Document SummariesVishakh Padmakumar, Zichao Wang, David Arbour, Jennifer HealeyACL 2025 · 1 citation
- Demonstration Selection for In-Context Learning via Reinforcement LearningXubin Wang, Jianfei Wu, Yichen Yuan, Deyu Cai et al.ICML 2025
Builds on11
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel et al.ICLR 2020 · 7,418 citations
- Fantastically Ordered Prompts and Where to Find Them: Overcoming Few-Shot Prompt Order SensitivityYao Lu, Max Bartolo, Alastair Moore, Sebastian Riedel et al.ACL 2022 · 1,494 citations
- Rethinking the Role of Demonstrations: What Makes In-Context Learning Work?Sewon Min, Xinxi Lyu, Ari Holtzman, Mikel Artetxe et al.EMNLP 2022 · 634 citations
- Language Modeling Is CompressionGrégoire Delétang, Anian Ruoss, Paul-Ambroise Duquenne, Elliot Catt et al.ICLR 2024 · 243 citations
- Compositional Exemplars for In-context LearningJiacheng Ye, Zhiyong Wu, Jiangtao Feng, Tao Yu et al.ICML 2023 · 188 citations
Related papers
- Revisiting In-context Learning Inference Circuit in Large Language ModelsHakaze Cho, Mariko Kato, Yoshihiro Sakai, Naoya InoueICLR 2025
- Representative Demonstration Selection for In-Context Learning with Two-Stage Determinantal Point ProcessZhao Yang, Yuanzhe Zhang, Dianbo Sui, Cao Liu et al.EMNLP 2023 · 3 citations
- Correlation-Aware Example Selection for In-Context Learning with Nonsymmetric Determinantal Point ProcessesQiunan Du, Zhiliang Tian, Zhen Huang, Kailun Bian et al.EMNLP 2025
- Large Language Models are Demonstration Pre-Selectors for ThemselvesJiarui Jin, Yuwei Wu, Haoxuan Li, Xiaoting He et al.ICML 2025
- CoverICL: Selective Annotation for In-Context Learning via Active Graph CoverageCostas Mavromatis, Balasubramaniam Srinivasan, Zhengyuan Shen, Jiani Zhang et al.EMNLP 2024
