When to Show a Suggestion? Integrating Human Feedback in AI-Assisted Programming
Hussein Mozannar, Gagan Bansal, Adam Fourney, Eric Horvitz
摘要
AI powered code-recommendation systems, such as Copilot and CodeWhisperer, provide code suggestions inside a programmer's environment (e.g., an IDE) with the aim of improving productivity. We pursue mechanisms for leveraging signals about programmers' acceptance and rejection of code suggestions to guide recommendations. We harness data drawn from interactions with GitHub Copilot, a system used by millions of programmers, to develop interventions that can save time for programmers. We introduce a utility-theoretic framework to drive decisions about suggestions to display versus withhold. The approach, conditional suggestion display from human feedback (CDHF), relies on a cascade of models that provide the likelihood that recommended code will be accepted. These likelihoods are used to selectively hide suggestions, reducing both latency and programmer verification time. Using data from 535 programmers, we perform a retrospective evaluation of CDHF and show that we can avoid displaying a significant fraction of suggestions that would have been rejected. We further demonstrate the importance of incorporating the programmer's latent unobserved state in decisions about when to display suggestions through an ablation study. Finally, we showcase how using suggestion acceptance as a reward signal for guiding the display of suggestions can lead to suggestions of reduced quality, indicating an unexpected pitfall.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- Need Help? Designing Proactive AI Assistants for ProgrammingValerie Chen, Alan Zhu, Sebastian Zhao, Hussein Mozannar 等CHI 2025 · 被引用 23 次
- Navigating Rifts in Human-LLM Grounding: Study and BenchmarkOmar Shaikh, Hussein Mozannar, Gagan Bansal, Adam Fourney 等ACL 2025 · 被引用 21 次
- Utilizing Human Behavior Modeling to Manipulate Explanations in AI-Assisted Decision Making: The Good, the Bad, and the ScaryZhuoyan Li, Ming YinNeurIPS 2024 · 被引用 15 次
- When Neural Code Completion Models Size up the Situation: Attaining Cheaper and Faster Completion through Dynamic Model InferenceZhensu Sun, Xiaoning Du, Fu Song, Shangwen Wang 等ICSE 2024 · 被引用 13 次
- "My productivity is boosted, but ..." Demystifying Users' Perception on AI Coding AssistantsYunbo Lyu, Zhou Yang, Jieke Shi, Jianming Chang 等ASE 2025 · 被引用 8 次
它引用的顶会 Paper8
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Grounded Copilot: How Programmers Interact with Code-Generating ModelsShraddha Barke, Michael B. James, Nadia PolikarpovaOOPSLA 2023 · 被引用 408 次
- Consistent Estimators for Learning to Defer to an ExpertHussein Mozannar, David A. SontagICML 2020 · 被引用 267 次
- Differentiable Learning Under TriageNastaran Okati, Abir De, Manuel Gomez-RodriguezNeurIPS 2021 · 被引用 99 次
相关 Paper
- Reading Between the Lines: Modeling User Behavior and Costs in AI-Assisted ProgrammingHussein Mozannar, Gagan Bansal, Adam Fourney, Eric HorvitzCHI 2024 · 被引用 88 次
- A Large-Scale Survey on the Usability of AI Programming Assistants: Successes and ChallengesJenny T. Liang, Chenyang Yang, Brad A. MyersICSE 2024 · 被引用 126 次
- An Eye for AI: Eye-Tracking the Micro-Interruptions of GenAI Code SuggestionsTarek Alakmeh, Sarah D’Angelo, Thomas FritzICSE 2026 · 被引用 1 次
- Validating AI-Generated Code with Live ProgrammingKasra Ferdowsi, Ruanqianqian (Lisa) Huang, Michael B. James, Nadia Polikarpova 等CHI 2024 · 被引用 27 次
- An Empirical Study of Knowledge Transfer in AI Pair ProgrammingAlisa Welter, Niklas Schneider, Tobias Dick, Kallistos Weis 等ASE 2025
