Active Learning with LLMs for Partially Observed and Cost-Aware Scenarios
Nicolás Astorga, Tennison Liu, Nabeel Seedat, Mihaela van der Schaar
Abstract
Conducting experiments and collecting data for machine learning models is a complex and expensive endeavor, particularly when confronted with limited information. Typically, extensive experiments to obtain features and labels come with a significant acquisition cost, making it impractical to carry out all of them. Therefore, it becomes crucial to strategically determine what to acquire to maximize the predictive performance while minimizing costs. To perform this task, existing data acquisition methods assume the availability of an initial dataset that is both fully-observed and labeled, crucially overlooking the partial observability of features characteristic of many real-world scenarios. In response to this challenge, we present Partially Observable Cost-Aware Active-Learning (POCA), a new learning approach aimed at improving model generalization in data-scarce and data-costly scenarios through label and/or feature acquisition. Introducing µ POCA as an instantiation, we maximize the uncertainty reduction in the predictive model when obtaining labels and features, considering associated costs. µ POCA enhance traditional Active Learning metrics based solely on the observed features by generating the unobserved features through Generative Surrogate Models, particularly Large Language Models (LLMs). We empirically validate µ POCA across diverse tabular datasets, varying data availability, acquisition costs, and LLMs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f579f2b0-a51c-47eb-8a75-e2282be4592eCited by top-tier papers6
- ProbeLLM: Automating Principled Diagnosis of LLM FailuresYue Huang, Zhengzhe Jiang, Yuchen Ma, Yu Jiang et al.ICML 2026 · 4 citations
- Active Task Disambiguation with LLMsKasia Kobalczyk, Nicolás Astorga, Tennison Liu, Mihaela van der SchaarICLR 2025
- Continuously Updating Digital Twins using Large Language ModelsHarry Amad, Nicolás Astorga, Mihaela van der SchaarICML 2025
- Autoformulation of Mathematical Optimization Models Using LLMsNicolás Astorga, Tennison Liu, Yuanzhang Xiao, Mihaela van der SchaarICML 2025
- Sample Lottery: Unsupervised Discovery of Critical Instances for LLM ReasoningZhiping Xiao, Yusheng Zhao, Qixin Zhang, Jiaye Xie et al.ICLR 2026
Builds on25
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Language Models as Zero-Shot Planners: Extracting Actionable Knowledge for Embodied AgentsWenlong Huang, Pieter Abbeel, Deepak Pathak, Igor MordatchICML 2022 · 1,539 citations
- Large Language Models as OptimizersChengrun Yang, Xuezhi Wang, Yifeng Lu, Hanxiao Liu et al.ICLR 2024 · 817 citations
- Large Language Models as Commonsense Knowledge for Large-Scale Task PlanningZirui Zhao, Wee Sun Lee, David HsuNeurIPS 2023 · 423 citations
Related papers
- Acquisition Conditioned Oracle for Nongreedy Active Feature AcquisitionMichael Valancius, Max Lennon, Junier OlivaICML 2024 · 7 citations
- Active Feature Acquisition with Generative Surrogate ModelsYang Li, Junier OlivaICML 2021 · 52 citations
- Learning-To-Measure: In-Context Active Feature AcquisitionYuta Kobayashi, Zilin Jing, Jiayu Yao, Hongseok Namkoong et al.ICML 2026 · 2 citations
- Scaling Up Active Testing to Large Language ModelsGabrielle Berrada, Jannik Kossen, Freddie Bickford Smith, Muhammed Razzak et al.NeurIPS 2025 · 11 citations
- Progressive Generalization Risk Reduction for Data-Efficient Causal Effect EstimationHechuan Wen, Tong Chen, Guanhua Ye, Li Kheng Chai et al.KDD 2025 · 1 citation
