Online Decision Mediation
Daniel Jarrett, Alihan Hüyük, Mihaela van der Schaar
Abstract
Consider learning a decision support assistant to serve as an intermediary between (oracle) expert behavior and (imperfect) human behavior: At each time, the algorithm observes an action chosen by a fallible agent, and decides whether to accept that agent's decision, intervene with an alternative, or request the expert's opinion. For instance, in clinical diagnosis, fully-autonomous machine behavior is often beyond ethical affordances, thus real-world decision support is often limited to monitoring and forecasting. Instead, such an intermediary would strike a prudent balance between the former (purely prescriptive) and latter (purely descriptive) approaches, while providing an efficient interface between human mistakes and expert feedback. In this work, we first formalize the sequential problem of online decision mediation -- that is, of simultaneously learning and evaluating mediator policies from scratch with abstentive feedback: In each round, deferring to the oracle obviates the risk of error, but incurs an upfront penalty, and reveals the otherwise hidden expert action as a new training data point. Second, we motivate and propose a solution that seeks to trade off (immediate) loss terms against (future) improvements in generalization error; in doing so, we identify why conventional bandit algorithms may fail. Finally, through experiments and sensitivities on a variety of datasets, we illustrate consistent gains over applicable benchmarks on performance measures with respect to the mediator policy, the learned model, and the decision-making system as a whole.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b51d8241-61e2-4f4b-a5e5-25c62aedb4d3Cited by top-tier papers3
- Cascaded Language Models for Cost-Effective Human-AI Decision-MakingClaudio Fanconi, Mihaela van der SchaarNeurIPS 2025 · 12 citations
- Revisiting Task-Oriented Dataset Search in the Era of Large Language Models: Challenges, Benchmark, and SolutionZixin Wei, Yucan Guo, Jinyang Li, Xiaolin Han et al.VLDB 2026 · 1 citation
- Bayesian Inference for Correlated Human Experts and ClassifiersMarkelle Kelly, Alex James Boyd, Samuel Showalter, Mark Steyvers et al.ICML 2025
Builds on13
- Consistent Estimators for Learning to Defer to an ExpertHussein Mozannar, David A. SontagICML 2020 · 267 citations
- Is the Most Accurate AI the Best Teammate? Optimizing AI for TeamworkGagan Bansal, Besmira Nushi, Ece Kamar, Eric Horvitz et al.AAAI 2021 · 185 citations
- Classification with Rejection Based on Cost-sensitive ClassificationNontawat Charoenphakdee, Zhenghang Cui, Yivan Zhang, Masashi SugiyamaICML 2021 · 78 citations
- Interactive Label Cleaning with Example-based ExplanationsStefano Teso, Andrea Bontempelli, Fausto Giunchiglia, Andrea PasseriniNeurIPS 2021 · 59 citations
- Teaching Humans When to Defer to a Classifier via ExemplarsHussein Mozannar, Arvind Satyanarayan, David A. SontagAAAI 2022 · 49 citations
Related papers
- Designing Decision Support Systems using Counterfactual Prediction SetsEleni Straitouri, Manuel Gomez RodriguezICML 2024 · 24 citations
- Learning Personalized Decision Support PoliciesUmang Bhatt, Valerie Chen, Katherine M. Collins, Parameswaran Kamalaruban et al.AAAI 2025 · 14 citations
- Algorithmic Recourse for Long-Term ImprovementKentaro Kanamori, Ken Kobayashi, Satoshi Hara, Takuya TakagiICML 2025
- Contextual Bandits and Imitation Learning with Preference-Based Active QueriesAyush Sekhari, Karthik Sridharan, Wen Sun, Runzhe WuNeurIPS 2023 · 18 citations
- Inverse Contextual Bandits: Learning How Behavior Evolves over TimeAlihan Hüyük, Daniel Jarrett, Mihaela van der SchaarICML 2022 · 14 citations
