Active Learning with LLMs for Partially Observed and Cost-Aware Scenarios
Nicolás Astorga, Tennison Liu, Nabeel Seedat, Mihaela van der Schaar
摘要
Conducting experiments and collecting data for machine learning models is a complex and expensive endeavor, particularly when confronted with limited information. Typically, extensive experiments to obtain features and labels come with a significant acquisition cost, making it impractical to carry out all of them. Therefore, it becomes crucial to strategically determine what to acquire to maximize the predictive performance while minimizing costs. To perform this task, existing data acquisition methods assume the availability of an initial dataset that is both fully-observed and labeled, crucially overlooking the partial observability of features characteristic of many real-world scenarios. In response to this challenge, we present Partially Observable Cost-Aware Active-Learning (POCA), a new learning approach aimed at improving model generalization in data-scarce and data-costly scenarios through label and/or feature acquisition. Introducing µ POCA as an instantiation, we maximize the uncertainty reduction in the predictive model when obtaining labels and features, considering associated costs. µ POCA enhance traditional Active Learning metrics based solely on the observed features by generating the unobserved features through Generative Surrogate Models, particularly Large Language Models (LLMs). We empirically validate µ POCA across diverse tabular datasets, varying data availability, acquisition costs, and LLMs.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- ProbeLLM: Automating Principled Diagnosis of LLM FailuresYue Huang, Zhengzhe Jiang, Yuchen Ma, Yu Jiang 等ICML 2026 · 被引用 4 次
- Active Task Disambiguation with LLMsKasia Kobalczyk, Nicolás Astorga, Tennison Liu, Mihaela van der SchaarICLR 2025
- Continuously Updating Digital Twins using Large Language ModelsHarry Amad, Nicolás Astorga, Mihaela van der SchaarICML 2025
- Autoformulation of Mathematical Optimization Models Using LLMsNicolás Astorga, Tennison Liu, Yuanzhang Xiao, Mihaela van der SchaarICML 2025
- Sample Lottery: Unsupervised Discovery of Critical Instances for LLM ReasoningZhiping Xiao, Yusheng Zhao, Qixin Zhang, Jiaye Xie 等ICLR 2026
它引用的顶会 Paper25
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Language Models as Zero-Shot Planners: Extracting Actionable Knowledge for Embodied AgentsWenlong Huang, Pieter Abbeel, Deepak Pathak, Igor MordatchICML 2022 · 被引用 1,539 次
- Large Language Models as OptimizersChengrun Yang, Xuezhi Wang, Yifeng Lu, Hanxiao Liu 等ICLR 2024 · 被引用 817 次
- Large Language Models as Commonsense Knowledge for Large-Scale Task PlanningZirui Zhao, Wee Sun Lee, David HsuNeurIPS 2023 · 被引用 423 次
相关 Paper
- Acquisition Conditioned Oracle for Nongreedy Active Feature AcquisitionMichael Valancius, Max Lennon, Junier OlivaICML 2024 · 被引用 7 次
- Active Feature Acquisition with Generative Surrogate ModelsYang Li, Junier OlivaICML 2021 · 被引用 52 次
- Learning-To-Measure: In-Context Active Feature AcquisitionYuta Kobayashi, Zilin Jing, Jiayu Yao, Hongseok Namkoong 等ICML 2026 · 被引用 2 次
- Scaling Up Active Testing to Large Language ModelsGabrielle Berrada, Jannik Kossen, Freddie Bickford Smith, Muhammed Razzak 等NeurIPS 2025 · 被引用 11 次
- Progressive Generalization Risk Reduction for Data-Efficient Causal Effect EstimationHechuan Wen, Tong Chen, Guanhua Ye, Li Kheng Chai 等KDD 2025 · 被引用 1 次
