Few-Shot Open-Set Classification via Reasoning-Aware Decomposition
Avyav Kumar Singh, Helen Yannakoudakis
Abstract
Large language models (LLMs) excel at fewshot learning, but their ability to reject out-ofdistribution examples remains under-explored. We study this challenge under the setting of few-shot open-set classification, where a model must not only classify examples from a small set of seen classes but also reject unseen ones at inference time. This setting is more realistic and challenging than traditional closedset supervised learning, requiring both finegrained classification and robust rejection. We show that, for small LLMs, neither chain-ofthought (CoT) prompting nor supervised finetuning (SFT) alone are sufficient to generalise reliably, particularly when class semantics are anonymised. We introduce Wasserstein GFN (W-GFN), a novel amortised Generative Flow Network framework that uses latent trajectories to approximate the Bayesian posterior. With as few as 4 examples per class, W-GFN substantially improves performance, enabling Llama 3.2 3B to achieve up to ≥ 80% of the performance of Llama 3.3 70B in complex datasets, despite being ∼ 23 times smaller, which highlights the importance of reasoning-aware approaches for robust open-set few-shot learning. 1 This is a lenient evaluation setting: any incorrect prediction to an out-of-set label is accepted as 'None of these', artificially inflating unseen F1. We also prompt the model that some examples may be out-of-set (Appendix A.1). Thus, these results represent an upper bound on unseen performance. 2 Anonymised labels are a common practice in prior work (e.g., encoder-based classification and meta-learning; Liu et al. (2020); Snell et al. (2017); Bansal et al. (2020a)). overfitting, especially when Y lacks semantic cues, by promoting abstraction (i.e., intermediate concepts) over memorisation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c562e06f-be52-4fa2-86e7-dce0203d04ffBuilds on27
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo et al.NeurIPS 2022 · 8,168 citations
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 2,600 citations
Related papers
- LaFTer: Label-Free Tuning of Zero-shot Classifier using Language and Unlabeled Image CollectionsMuhammad Jehanzeb Mirza, Leonid Karlinsky, Wei Lin, Horst Possegger et al.NeurIPS 2023 · 63 citations
- Coarse-to-Fine Open-Set Graph Node Classification with Large Language ModelsXueqi Ma, Xingjun Ma, Sarah Monazam Erfani, Danilo P. Mandic et al.AAAI 2026
- DSV-LFS: Unifying LLM-Driven Semantic Cues with Visual Features for Robust Few-Shot SegmentationAmin Karimi, Charalambos PoullisCVPR 2025
- Tuning Language Models as Training Data Generators for Augmentation-Enhanced Few-Shot LearningYu Meng, Martin Michalski, Jiaxin Huang, Yu Zhang et al.ICML 2023 · 64 citations
- Making Large Language Models Better Reasoners with Orchestrated Streaming ExperiencesXiangyang Liu, Junliang He, Xipeng QiuEMNLP 2024
