Score-Based Density Estimation from Pairwise Comparisons
Petrus Mikkola, Luigi Acerbi, Arto Klami
Abstract
We study density estimation from pairwise comparisons, motivated by expert knowledge elicitation and learning from human feedback. We relate the unobserved target density to a tempered winner density (marginal density of preferred choices), learning the winner's score via score-matching. This allows estimating the target by `de-tempering' the estimated winner density's score. We prove that the score vectors of the belief and the winner density are collinear, linked by a position-dependent tempering field. We give analytical formulas for this field and propose an estimator for it under the Bradley–Terry model. Using a diffusion model trained on tempered samples generated via score-scaled annealed Langevin dynamics, we can learn complex multivariate belief densities of simulated experts, from only hundreds to thousands of pairwise comparisons.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on19
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Elucidating the Design Space of Diffusion-Based Generative ModelsTero Karras, Miika Aittala, Timo Aila, Samuli LaineNeurIPS 2022 · 3,959 citations
Related papers
- Preferential Normalizing FlowsPetrus Mikkola, Luigi Acerbi, Arto KlamiNeurIPS 2024 · 3 citations
- Reward Modeling with Ordinal Feedback: Wisdom of the CrowdShang Liu, Yu Pan, Guanting Chen, Xiaocheng LiICML 2025
- Online Compatible Reward Identification from Preference FeedbackSimone Drago, Marco Mussi, Alberto Maria MetelliICML 2026
- Minimax Rate for Learning From Pairwise Comparisons in the BTL ModelJulien M. Hendrickx, Alex Olshevsky, Venkatesh SaligramaICML 2020 · 16 citations
- Towards Cognitively-Faithful Decision-Making Models to Improve AI AlignmentCyrus Cousins, Vijay Keswani, Vincent Conitzer, Hoda Heidari et al.ICLR 2026 · 2 citations
