Play to Grade: Testing Coding Games as Classifying Markov Decision Process
Allen Nie, Emma Brunskill, Chris Piech
Abstract
Contemporary coding education often presents students with the task of developing programs that have user interaction and complex dynamic systems, such as mouse based games. While pedagogically compelling, there are no contemporary autonomous methods for providing feedback. Notably, interactive programs are impossible to grade by traditional unit tests. In this paper we formalize the challenge of providing feedback to interactive programs as a task of classifying Markov Decision Processes (MDPs). Each student's program fully specifies an MDP where the agent needs to operate and decide, under reasonable generalization, if the dynamics and reward model of the input MDP should be categorized as correct or broken. We demonstrate that by designing a cooperative objective between an agent and an autoregressive model, we can use the agent to sample differential trajectories from the input MDP that allows a classifier to determine membership: Play to Grade. Our method enables an automatic feedback system for interactive code assignments. We release a dataset of 711,274 anonymized student submissions to a single assignment with hand-coded bug labels to support future research.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9df54540-09c8-4e4d-bf4c-534938a4a750Cited by top-tier papers3
- Robust -Divergence MDPsChin Pang Ho, Marek Petrik, Wolfram WiesemannNeurIPS 2022 · 13 citations
- Giving Feedback on Interactive Student Programs with Meta-ExplorationEvan Zheran Liu, Moritz Stephan, Allen Nie, Chris Piech et al.NeurIPS 2022 · 8 citations
- Get a Head Start: On-Demand Pedagogical Policy Selection in Intelligent TutoringGe Gao, Xi Yang, Min ChiAAAI 2024 · 1 citation
Builds on2
Related papers
- ErrorCLR: Semantic Error Classification, Localization and Repair for Introductory Programming AssignmentsSiqi Han, Yu Wang, Xuesong LuSIGIR 2023 · 9 citations
- Multi-Turn Code Generation Through Single-Step RewardsArnav Kumar Jain, Gonzalo Gonzalez-Pumariega, Wayne Chen, Alexander M. Rush et al.ICML 2025
- PlayCoder: Making LLM-Generated GUI Code PlayableZhiyuan Peng, Wei Tao, Xin Yin, Chenhao Ying et al.FSE 2026
- VizProg: Identifying Misunderstandings By Visualizing Students' Coding ProgressAshley Ge Zhang, Yan Chen, Steve OneyCHI 2023 · 31 citations
- PuzzleMe: Leveraging Peer Assessment for In-Class Programming ExercisesApril Yi Wang, Yan Chen, John Joon Young Chung, Christopher Brooks et al.CSCW 2021 · 22 citations
