Learning Task Decomposition to Assist Humans in Competitive Programming
Jiaxin Wen, Ruiqi Zhong, Pei Ke, Zhihong Shao, Hongning Wang, Minlie Huang
Abstract
When using language models (LMs) to solve complex problems, humans might struggle to understand the LM-generated solutions and repair the flawed ones. To better assist humans in repairing them, we propose to automatically decompose complex solutions into simpler pieces corresponding to specific subtasks. We introduce a novel objective for learning task decomposition, termed assistive value (AssistV), which measures the feasibility and speed for humans to repair the decomposed solution. We collect a dataset of human repair experiences on different decomposed solutions. Utilizing the collected data as in-context examples, we then learn to critique, refine, and rank decomposed solutions to improve AssistV. We validate our method under competitive programming problems: under 177 hours of human study, our method enables non-experts to solve 33.3% more problems, speeds them up by 3.3x, and empowers them to match unassisted experts.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- From Code to Correctness: Closing the Last Mile of Code Generation with Hierarchical DebuggingYuling Shi, Songsong Wang, Chengcheng Wan, Min Wang et al.ICSE 2026 · 4 citations
- Language Models Learn to Mislead Humans via RLHFJiaxin Wen, Ruiqi Zhong, Akbir Khan, Ethan Perez et al.ICLR 2025
- A Good Plan is Hard to Find: Aligning Models with Preferences is Misaligned with What Helps UsersNishant Balepur, Matthew Shu, Yoo Yeon Sung, Seraphina Goldfarb-Tarrant et al.EMNLP 2025
Builds on8
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Self-Refine: Iterative Refinement with Self-FeedbackAman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan et al.NeurIPS 2023 · 4,972 citations
- Teaching Large Language Models to Self-DebugXinyun Chen, Maxwell Lin, Nathanael Schärli, Denny ZhouICLR 2024 · 1,085 citations
- Large Language Models Cannot Self-Correct Reasoning YetJie Huang, Xinyun Chen, Swaroop Mishra, Huaixiu Steven Zheng et al.ICLR 2024 · 858 citations
- Principle-Driven Self-Alignment of Language Models from Scratch with Minimal Human SupervisionZhiqing Sun, Yikang Shen, Qinhong Zhou, Hongxin Zhang et al.NeurIPS 2023 · 463 citations
Related papers
- Beyond Problem Solving: UOJ-Bench for Evaluating Code Generation, Hacking, and Repair in Competitive ProgrammingTingqiang Xu, Hangrui Zhou, Tianle Cai, Alex Gu et al.ICML 2026 · 1 citation
- AutoCode: LLMs as Problem Setters for Competitive ProgrammingShang Zhou, Zihan Zheng, Kaiyuan Liu, Zeyu Shen et al.ICLR 2026 · 12 citations
- Automated Repair of Programs from Large Language ModelsZhiyu Fan, Xiang Gao, Martin Mirchev, Abhik Roychoudhury et al.ICSE 2023 · 213 citations
- Ivie: Lightweight Anchored Explanations of Just-Generated CodeLitao Yan, Alyssa Hwang, Zhiyuan Wu, Andrew HeadCHI 2024 · 40 citations
- TaskLAMA: Probing the Complex Task Understanding of Language ModelsQuan Yuan, Mehran Kazemi, Xin Xu, Isaac Noble et al.AAAI 2024 · 22 citations
