CS4ML: A general framework for active learning with arbitrary data based on Christoffel functions
Juan M. Cardenas, Ben Adcock, Nick C. Dexter
摘要
We introduce a general framework for active learning in regression problems. Our framework extends the standard setup by allowing for general types of data, rather than merely pointwise samples of the target function. This generalization covers many cases of practical interest, such as data acquired in transform domains (e.g., Fourier data), vector-valued data (e.g., gradient-augmented data), data acquired along continuous curves, and, multimodal data (i.e., combinations of different types of measurements). Our framework considers random sampling according to a finite number of sampling measures and arbitrary nonlinear approximation spaces (model classes). We introduce the concept of generalized Christoffel functions and show how these can be used to optimize the sampling measures. We prove that this leads to near-optimal sample complexity in various important cases. This paper focuses on applications in scientific computing, where active learning is often desirable, since it is usually expensive to generate data. We demonstrate the efficacy of our framework for gradient-augmented learning with polynomials, Magnetic Resonance Imaging (MRI) using generative models and adaptive sampling for solving PDEs using Physics-Informed Neural Networks (PINNs). Example 2.3 (Active learning in standard regression) The above framework extends the standard active learning problem in regression. In the classic regression problem, D ⊆ R d is a domain and X = L 2 ρ (D) is the space of square-integrable functions f * : D → R with respect to a measure ρ. Note that ρ is considered fixed -it is the measure with respect to which we measure the error. To embed this problem into the above framework, we let X 0 = C(D) be the space of continuous functions on D, C = 1, the measurement domain D c = D be equal to the domain of the function, ρ 1 = ρ and the measurement space Y 1 = R (with the Euclidean inner product). We then define the sampling operator L 1 (θ)(f * ) = f * (θ) as the pointwise evaluation operator. In particular, for a measure µ = µ 1 satisfying Assumption 2.2, the training data (2.1) is Hence, the aim is to choose the measure µ (or equivalently, its Radon-Nikodym derivative ν) to ensure as good generalization as possible. Next, we let F ⊆ X 0 be a subset within which we seek to learn f * . We term F the approximation space. Note this could be a linear space such as a space of algebraic or trigonometric polynomials, or a nonlinear space such as the space of sparse Fourier functions, a space of functions with sparse
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- A Unified Framework for Learning with Nonlinear Model Classes from Arbitrary Linear SamplesBen Adcock, Juan M. Cardenas, Nick C. DexterICML 2024 · 被引用 6 次
- How many measurements are enough? Bayesian recovery in inverse problems with general distributionsBen Adcock, Zi Yuan (Nick) HuangNeurIPS 2025 · 被引用 2 次
- Provably Accurate Shapley Value Estimation via Leverage Score SamplingChristopher Musco, R. Teal WitterICLR 2025
- TESSAR: Geometry-Aware Active Regression via Dynamic Voronoi TessellationSeong Jin Cho, Gwangsu Kim, Junghyun Lee, Hee Suk Yoon 等ICLR 2026
它引用的顶会 Paper6
- Robust Compressed Sensing MRI with Deep Generative PriorsAjil Jalal, Marius Arvinte, Giannis Daras, Eric Price 等NeurIPS 2021 · 被引用 483 次
- Generic bounds on the approximation error for physics-informed (and) operator learningTim De Ryck, Siddhartha MishraNeurIPS 2022 · 被引用 93 次
- Instance-Optimal Compressed Sensing via Posterior SamplingAjil Jalal, Sushrut Karmalkar, Alex Dimakis, Eric PriceICML 2021 · 被引用 62 次
- Fourier Sparse Leverage Scores and Approximate Kernel LearningTamás Erdélyi, Cameron Musco, Christopher MuscoNeurIPS 2020 · 被引用 28 次
- Non-Iterative Recovery from Nonlinear Observations using Generative ModelsJiulong Liu, Zhaoqiang LiuCVPR 2022 · 被引用 8 次
相关 Paper
- Accelerated Training of Physics-Informed Neural Networks (PINNs) using Meshless DiscretizationsRamansh Sharma, Varun ShankarNeurIPS 2022 · 被引用 81 次
- Sobolev Acceleration and Statistical Optimality for Learning Elliptic Equations via Gradient DescentYiping Lu, José H. Blanchet, Lexing YingNeurIPS 2022 · 被引用 15 次
- Physics-informed Neural Networks for Functional Differential Equations: Cylindrical Approximation and Its Convergence GuaranteesTaiki Miyagawa, Takeru YokotaNeurIPS 2024 · 被引用 8 次
- RoPINN: Region Optimized Physics-Informed Neural NetworksHaixu Wu, Huakun Luo, Yuezhou Ma, Jianmin Wang 等NeurIPS 2024 · 被引用 50 次
- Machine Learning For Elliptic PDEs: Fast Rate Generalization Bound, Neural Scaling Law and Minimax OptimalityYiping Lu, Haoxuan Chen, Jianfeng Lu, Lexing Ying 等ICLR 2022 · 被引用 54 次
