PSD Representations for Effective Probability Models
Alessandro Rudi, Carlo Ciliberto
Abstract
Finding a good way to model probability densities is key to probabilistic inference. An ideal model should be able to concisely approximate any probability while being also compatible with two main operations: multiplications of two models (product rule) and marginalization with respect to a subset of the random variables (sum rule). In this work, we show that a recently proposed class of positive semi-definite (PSD) models for non-negative functions is particularly suited to this end. In particular, we characterize both approximation and generalization capabilities of PSD models, showing that they enjoy strong theoretical guarantees. Moreover, we show that we can perform efficiently both sum and product rule in closed form via matrix operations, enjoying the same versatility of mixture models. Our results open the way to applications of PSD models to density estimation, decision theory and inference. base point matrices X ∈ R n×d and X ′ ∈ R m×d , we denote by K X,X ′ ,η ∈ R n×m the kernel matrix with entries (K X,X ′ ,η ) ij = k η (x i , x ′ j ) where x i , x ′ j are the i-th and j-th rows of X, X ′ respectively. When clear from context, in the following we will refer to Gaussian PSD models as PSD models. Remark 1 (PSD models generalize Mixture models). Mixture models (a mixture of Gaussian distributions) are a special case of PSD models. Let A = diag(a) be a diagonal matrix of n positive weights a ∈ R n ++ . We have f ( Remark 2 (PSD models allow negative weights). From (2), we immediately see that PSD models generalize mixture models by allowing also for negative weights: e.g., f ( x-1) 2 , i.e. a mixture of Gaussians with also negative weights. Operations with PSD models In Sec. 3 we will show that PSD models can approximate a wide class of probability densities, significantly outperforming mixture models. Here we show that this improvement does not come at the expenses of computations. In particular, we show that PSD models enjoy the same flexibility of mixture models: i) they are closed with respect to key operations such as marginalization and multiplication and ii) these operations can be performed efficiently in terms of matrix sums/products. The derivation of the results reported in the following is provided in Appendix F. They follow from well-known properties of the Gaussian function. Evaluation. Evaluating a PSD model in a point x 0 ∈ X corresponds to f (x = x 0 ; A, X, η) = K ⊤ X,x 0 ,η AK X,x 0 ,η . Moreover, partially evaluating a PSD in one variable yields Note that f (x ; B, X, η 1 ) is still a PSD model since B is positive semidefinite. Sum Rule (Marginalization and Integration ). The integral of a PSD model can be computed as where c η = π d/2 det(diag(η)) -1/2 . This is particularly useful to model probabiliy densities with PSD models. Let Z = f (x ; A, X, η)dx, then the function f (x ; A/Z, X, η) = 1 Z f (x ; A, X, η) is a probability density. Integrating only one variable of a PSD model we obtain the sum rule. Then, the following integral is a PSD model f (x, y ; A, [X, Y ], (η, η ′ )) dx = f (y ; B, Y, η ′ ), with B = c 2η A • K X,X, η 2 , (5) The result above shows that we can efficiently marginalize a PSD model with respect to one variable by means of an entry-wise multiplication between two matrices.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 20379eef-b0a8-475b-86aa-2be8f066a07dCited by top-tier papers11
- Subtractive Mixture Models via Squaring: Representation and LearningLorenzo Loconte, Aleksanteri M. Sladek, Stefan Mengel, Martin Trapp et al.ICLR 2024 · 42 citations
- On the Relationship Between Monotone and Squared Probabilistic CircuitsBenjie Wang, Guy Van den BroeckAAAI 2025 · 16 citations
- Squared Neural Families: A New Class of Tractable Density ModelsRussell Tsuchida, Cheng Soon Ong, Dino SejdinovicNeurIPS 2023 · 15 citations
- Exact, Fast and Expressive Poisson Point Processes via Squared Neural FamiliesRussell Tsuchida, Cheng Soon Ong, Dino SejdinovicAAAI 2024 · 7 citations
- How to Square Tensor Networks and Circuits Without Squaring ThemLorenzo Loconte, Adrián Javaloy, Antonio VergariICLR 2026 · 6 citations
Builds on1
Related papers
- Faster Kernel Matrix Algebra via Density EstimationArturs Backurs, Piotr Indyk, Cameron Musco, Tal WagnerICML 2021 · 10 citations
- Sum of Squares CircuitsLorenzo Loconte, Stefan Mengel, Antonio VergariAAAI 2025 · 20 citations
- Dimensionality Reduction for General KDE Mode FindingXinyu Luo, Christopher Musco, Cas WiddershovenICML 2023 · 2 citations
- Inverse M-Kernels for Linear Universal Approximators of Non-Negative FunctionsHideaki KimNeurIPS 2024 · 2 citations
- A Non-commutative Extension of Lee-Seung's Algorithm for Positive Semidefinite FactorizationsYong Sheng Soh, Antonios VarvitsiotisNeurIPS 2021 · 2 citations
