Lune

ICML2026顶会

Graph-Preference Learning: Debiasing Network-Sampled Human Feedback for Target Welfare Estimation

Guangrui Fan, DanDan Liu, AZNUL SABRI, Pan Lihu

出版方
2026年份

摘要

Preference-based reward modeling is a core component of RLHF and DPO pipelines. In practice, the humans providing preference feedback are rarely an i.i.d. sample: recruitment and exposure often follow social, institutional, or spatial structure, inducing non-uniform inclusion probabilities that correlate with graph centrality. We formalize preference learning with network-sampled annotators and show that identity-agnostic scalar reward modeling implicitly represents an inclusion-weighted welfare, over-representing structurally central communities when the inclusion distribution qq differs from a designer-chosen target weighting π\pi. We propose Graph-Preference Learning, which combines (i) a graph-personalized reward model that shares statistical strength across neighboring annotators and (ii) graph-balanced aggregation that computes stabilized importance weights to target π\pi. Our analysis characterizes the induced welfare represented by the learned aggregate reward and bounds its deviation from the target in terms of weight mismatch, reward-model approximation, and finite-sample effects. Experiments on synthetic graphs and a semi-synthetic case study on the LMArena preference dataset, where biased inclusion is induced via graph-based sampling, demonstrate up to 62% reduction in target-welfare recovery error and 17% reduction in cross-language performance gaps under biased inclusion.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

lune papers fulltext fdaaa827-bf16-4fbf-bc99-7b0d65656cc9

它引用的顶会 Paper14

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖