Enhancing Preference-based Linear Bandits via Human Response Time
Shen Li, Yuyang Zhang, Zhaolin Ren, Claire Liang, Na Li, Julie A. Shah
摘要
Interactive preference learning systems infer human preferences by presenting queries as pairs of options and collecting binary choices. Although binary choices are simple and widely used, they provide limited information about preference strength. To address this, we leverage human response times, which are inversely related to preference strength, as an additional signal. We propose a computationally efficient method that combines choices and response times to estimate human utility functions, grounded in the EZ diffusion model from psychology. Theoretical and empirical analyses show that for queries with strong preferences, response times complement choices by providing extra information about preference strength, leading to significantly improved utility estimation. We incorporate this estimator into preference-based linear bandits for fixed-budget best-arm identification. Simulations on three real-world datasets demonstrate that using response times significantly accelerates preference learning compared to choice-only approaches. Additional materials, such as code, slides, and talk video, are available at https://shenlirobot.github.io/pages/NeurIPS24.html .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Understanding LoRA as Knowledge Memory: An Empirical AnalysisSeungju Back, Dongwoo Lee, Naun Kang, Taehee Lee 等ICML 2026 · 被引用 10 次
- ResponseRank: Data-Efficient Reward Modeling through Preference Strength LearningTimo Kaufmann, Yannick Metz, Daniel A. Keim, Eyke HüllermeierNeurIPS 2025 · 被引用 2 次
- Comparing Comparisons: Informative and Easy Human Feedback with Distinguishability QueriesXuening Feng, Zhaohui Jiang, Timo Kaufmann, Eyke Hüllermeier 等ICML 2025
它引用的顶会 Paper9
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Learning to summarize with human feedbackNisan Stiennon, Long Ouyang, Jeffrey Wu, Daniel M. Ziegler 等NeurIPS 2020 · 被引用 124 次
- Gamification of Pure Exploration for Linear BanditsRémy Degenne, Pierre Ménard, Xuedong Shang, Michal ValkoICML 2020 · 被引用 86 次
- High-dimensional Experimental Design and Kernel BanditsRomain Camilleri, Kevin Jamieson, Julian Katz-SamuelsICML 2021 · 被引用 63 次
- Minimax Optimal Fixed-Budget Best Arm Identification in Linear BanditsJunwen Yang, Vincent Y. F. TanNeurIPS 2022 · 被引用 38 次
相关 Paper
- Preference Learning with Response Time: Robust Losses and GuaranteesAyush Sawarni, Sahasrajit Sarmasarkar, Vasilis SyrgkanisNeurIPS 2025
- Preference Is More than Comparisons: Rethinking Dueling Bandits with Augmented Human FeedbackShengbo Wang, Hong Sun, Ke LiAAAI 2026
- Pragmatic Feature Preferences: Learning Reward-Relevant Preferences from Human InputAndi Peng, Yuying Sun, Tianmin Shu, David AbelICML 2024 · 被引用 7 次
- Neural Dueling Bandits: Preference-Based Optimization with Human FeedbackArun Verma, Zhongxiang Dai, Xiaoqiang Lin, Patrick Jaillet 等ICLR 2025
- Interactive Classification by Asking Informative QuestionsLili Yu, Howard Chen, Sida I. Wang, Tao Lei 等ACL 2020 · 被引用 13 次
