Play to Grade: Testing Coding Games as Classifying Markov Decision Process
Allen Nie, Emma Brunskill, Chris Piech
摘要
Contemporary coding education often presents students with the task of developing programs that have user interaction and complex dynamic systems, such as mouse based games. While pedagogically compelling, there are no contemporary autonomous methods for providing feedback. Notably, interactive programs are impossible to grade by traditional unit tests. In this paper we formalize the challenge of providing feedback to interactive programs as a task of classifying Markov Decision Processes (MDPs). Each student's program fully specifies an MDP where the agent needs to operate and decide, under reasonable generalization, if the dynamics and reward model of the input MDP should be categorized as correct or broken. We demonstrate that by designing a cooperative objective between an agent and an autoregressive model, we can use the agent to sample differential trajectories from the input MDP that allows a classifier to determine membership: Play to Grade. Our method enables an automatic feedback system for interactive code assignments. We release a dataset of 711,274 anonymized student submissions to a single assignment with hand-coded bug labels to support future research.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Robust -Divergence MDPsChin Pang Ho, Marek Petrik, Wolfram WiesemannNeurIPS 2022 · 被引用 13 次
- Giving Feedback on Interactive Student Programs with Meta-ExplorationEvan Zheran Liu, Moritz Stephan, Allen Nie, Chris Piech 等NeurIPS 2022 · 被引用 8 次
- Get a Head Start: On-Demand Pedagogical Policy Selection in Intelligent TutoringGe Gao, Xi Yang, Min ChiAAAI 2024 · 被引用 1 次
它引用的顶会 Paper2
相关 Paper
- ErrorCLR: Semantic Error Classification, Localization and Repair for Introductory Programming AssignmentsSiqi Han, Yu Wang, Xuesong LuSIGIR 2023 · 被引用 9 次
- Multi-Turn Code Generation Through Single-Step RewardsArnav Kumar Jain, Gonzalo Gonzalez-Pumariega, Wayne Chen, Alexander M. Rush 等ICML 2025
- PlayCoder: Making LLM-Generated GUI Code PlayableZhiyuan Peng, Wei Tao, Xin Yin, Chenhao Ying 等FSE 2026
- VizProg: Identifying Misunderstandings By Visualizing Students' Coding ProgressAshley Ge Zhang, Yan Chen, Steve OneyCHI 2023 · 被引用 31 次
- PuzzleMe: Leveraging Peer Assessment for In-Class Programming ExercisesApril Yi Wang, Yan Chen, John Joon Young Chung, Christopher Brooks 等CSCW 2021 · 被引用 22 次
