Random Selection Reveals Implicit Knowledge Consensus in Code Generation
Ren-Biao Liu, Li Xin-Ye, Hui Sun, Yali Du, Jiang-Tian Xue, Ming Li
Abstract
Training large language models for code generation often involves selecting data from verifiable multi-solution pools, where each problem admits multiple correct implementations. Conventional studies on data selection suggest that complex selection strategies, such as diversity maximization or difficulty ranking, should outperform naive random sampling. In this work, we systematically evaluate within-problem solution selection strategies across different representation spaces, including continuous embeddings, discrete tokens, and syntactic structures, using various base language models. Instead, simple random sampling achieves consistently competitive performance across all models, exhibiting greater cross-model stability than complex methods. We interpret these results through the lens of implicit knowledge consensus : verified solution pools may contain representative algorithmic patterns that random sampling can preserve. Our findings suggest that practitioners should treat random sampling as a low-cost default for verifiable code-generation fine-tuning and move to complex selectors when hard-tail or constraint-focused coverage is the target.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on11
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Deduplicating Training Data Makes Language Models BetterKatherine Lee, Daphne Ippolito, Andrew Nystrom, Chiyuan Zhang et al.ACL 2022 · 844 citations
- Coresets for Data-efficient Training of Machine Learning ModelsBaharan Mirzasoleiman, Jeff A. Bilmes, Jure LeskovecICML 2020 · 494 citations
- Selection via Proxy: Efficient Data Selection for Deep LearningCody Coleman, Christopher Yeh, Stephen Mussmann, Baharan Mirzasoleiman et al.ICLR 2020 · 462 citations
Related papers
- Dynamic-Static Synergistic Selection Method for Candidate Code Solutions with Generated Test CasesRenbiao Liu, Jiang-Tian Xue, Chao-Zeng Ma, Hui Sun et al.AAAI 2026 · 2 citations
- On LLMs’ Internal Representation of Code CorrectnessFrancisco Ribeiro, Claudio Spiess, Premkumar Devanbu, Sarah NadiICSE 2026
- Oracle-Guided Program Selection from Large Language ModelsZhiyu Fan, Haifeng Ruan, Sergey Mechtaev, Abhik RoychoudhuryISSTA 2024 · 4 citations
- Code Repair with LLMs gives an Exploration-Exploitation TradeoffHao Tang, Keya Hu, Jin Zhou, Sicheng Zhong et al.NeurIPS 2024 · 85 citations
- TreeCoder: Systematic Exploration and Optimisation of Decoding and Constraints for LLM Code GenerationHenrijs Princis, Arindam Sharma, Cristina DavidPLDI 2026
