Online Decision Mediation
Daniel Jarrett, Alihan Hüyük, Mihaela van der Schaar
摘要
Consider learning a decision support assistant to serve as an intermediary between (oracle) expert behavior and (imperfect) human behavior: At each time, the algorithm observes an action chosen by a fallible agent, and decides whether to accept that agent's decision, intervene with an alternative, or request the expert's opinion. For instance, in clinical diagnosis, fully-autonomous machine behavior is often beyond ethical affordances, thus real-world decision support is often limited to monitoring and forecasting. Instead, such an intermediary would strike a prudent balance between the former (purely prescriptive) and latter (purely descriptive) approaches, while providing an efficient interface between human mistakes and expert feedback. In this work, we first formalize the sequential problem of online decision mediation -- that is, of simultaneously learning and evaluating mediator policies from scratch with abstentive feedback: In each round, deferring to the oracle obviates the risk of error, but incurs an upfront penalty, and reveals the otherwise hidden expert action as a new training data point. Second, we motivate and propose a solution that seeks to trade off (immediate) loss terms against (future) improvements in generalization error; in doing so, we identify why conventional bandit algorithms may fail. Finally, through experiments and sensitivities on a variety of datasets, we illustrate consistent gains over applicable benchmarks on performance measures with respect to the mediator policy, the learned model, and the decision-making system as a whole.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Cascaded Language Models for Cost-Effective Human-AI Decision-MakingClaudio Fanconi, Mihaela van der SchaarNeurIPS 2025 · 被引用 12 次
- Revisiting Task-Oriented Dataset Search in the Era of Large Language Models: Challenges, Benchmark, and SolutionZixin Wei, Yucan Guo, Jinyang Li, Xiaolin Han 等VLDB 2026 · 被引用 1 次
- Bayesian Inference for Correlated Human Experts and ClassifiersMarkelle Kelly, Alex James Boyd, Samuel Showalter, Mark Steyvers 等ICML 2025
它引用的顶会 Paper13
- Consistent Estimators for Learning to Defer to an ExpertHussein Mozannar, David A. SontagICML 2020 · 被引用 267 次
- Is the Most Accurate AI the Best Teammate? Optimizing AI for TeamworkGagan Bansal, Besmira Nushi, Ece Kamar, Eric Horvitz 等AAAI 2021 · 被引用 185 次
- Classification with Rejection Based on Cost-sensitive ClassificationNontawat Charoenphakdee, Zhenghang Cui, Yivan Zhang, Masashi SugiyamaICML 2021 · 被引用 78 次
- Interactive Label Cleaning with Example-based ExplanationsStefano Teso, Andrea Bontempelli, Fausto Giunchiglia, Andrea PasseriniNeurIPS 2021 · 被引用 59 次
- Teaching Humans When to Defer to a Classifier via ExemplarsHussein Mozannar, Arvind Satyanarayan, David A. SontagAAAI 2022 · 被引用 49 次
相关 Paper
- Designing Decision Support Systems using Counterfactual Prediction SetsEleni Straitouri, Manuel Gomez RodriguezICML 2024 · 被引用 24 次
- Learning Personalized Decision Support PoliciesUmang Bhatt, Valerie Chen, Katherine M. Collins, Parameswaran Kamalaruban 等AAAI 2025 · 被引用 14 次
- Algorithmic Recourse for Long-Term ImprovementKentaro Kanamori, Ken Kobayashi, Satoshi Hara, Takuya TakagiICML 2025
- Contextual Bandits and Imitation Learning with Preference-Based Active QueriesAyush Sekhari, Karthik Sridharan, Wen Sun, Runzhe WuNeurIPS 2023 · 被引用 18 次
- Inverse Contextual Bandits: Learning How Behavior Evolves over TimeAlihan Hüyük, Daniel Jarrett, Mihaela van der SchaarICML 2022 · 被引用 14 次
