OpenCoder: The Open Cookbook for Top-Tier Code Large Language Models
Siming Huang, Tianhao Cheng, Jason Klein Liu, Weidi Xu, Jiaran Hao, Liuyihan Song, Yang Xu, Jian Yang, Jiaheng Liu, Chenchen Zhang, Linzheng Chai, Ruifeng Yuan
摘要
Code LLMs have been widely used in various domains, including code generation, logical reasoning, and agent systems. However, openaccess code LLMs mostly only release weights, lacking key features such as reproducible data pipelines and transparent training protocols, which are crucial for advancing deeper and more reliable investigations. To address the gap, we introduce OpenCoder, a top-tier code LLM that not only achieves performance comparable to industrial leading models but also serves as an "open cookbook" for the research community. Unlike most prior efforts, we release not only model weights and inference code, but also the reproducible training data, complete data processing pipeline, rigorous experimental ablation results, and detailed training protocols for open scientific research. Our work identifies the key ingredients for building a top-tier code LLM are: language-specific filtering rules, file-level deduplication , highquality synthetic data and two-stage supervised fine-tuning strategy. By offering high level of openness, we aim to broaden access to all aspects of a top-tier code LLM, with Open-Coder serving as both a powerful model and an open foundation to accelerate research, enabling reproducible advancements in code intelligence. The released resource is available at https://opencoder-llm.github.io.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper51
- DiffuCoder: Understanding and Improving Masked Diffusion Models for Code GenerationShansan Gong, Ruixiang Zhang, Huangjie Zheng, Jiatao Gu 等ICLR 2026 · 被引用 198 次
- OmniSQL: Synthesizing High-quality Text-to-SQL Data at ScaleHaoyang Li, Shang Wu, Xiaokang Zhang, Xinmei Huang 等VLDB 2025 · 被引用 90 次
- Adaptive Code Watermarking Through Reinforcement LearningZhimeng Guo, Huaisheng Zhu, Siyuan Xu, Hangfan Zhang 等ICML 2026 · 被引用 73 次
- ACECODER: Acing Coder RL via Automated Test-Case SynthesisHuaye Zeng, Dongfu Jiang, Haozhe Wang, Ping Nie 等ACL 2025 · 被引用 72 次
- Any-Order Flexible Length Masked DiffusionJaeyeon Kim, Cheuk Lee Kit, Carles Domingo-Enrich, Yilun Du 等ICLR 2026 · 被引用 51 次
它引用的顶会 Paper19
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code GenerationJiawei Liu, Chunqiu Steven Xia, Yuyao Wang, Lingming ZhangNeurIPS 2023 · 被引用 2,317 次
- WizardCoder: Empowering Code Large Language Models with Evol-InstructZiyang Luo, Can Xu, Pu Zhao, Qingfeng Sun 等ICLR 2024 · 被引用 945 次
- Deduplicating Training Data Makes Language Models BetterKatherine Lee, Daphne Ippolito, Andrew Nystrom, Chiyuan Zhang 等ACL 2022 · 被引用 844 次
- WildChat: 1M ChatGPT Interaction Logs in the WildWenting Zhao, Xiang Ren, Jack Hessel, Claire Cardie 等ICLR 2024 · 被引用 504 次
相关 Paper
- LMR-BENCH: Evaluating LLM Agent's Ability on Reproducing Language Modeling ResearchShuo Yan, Ruochen Li, Ziming Luo, Zimu Wang 等EMNLP 2025
- AlchemistCoder: Harmonizing and Eliciting Code Capability by Hindsight Tuning on Multi-source DataZifan Song, Yudong Wang, Wenwei Zhang, Kuikun Liu 等NeurIPS 2024 · 被引用 8 次
- Paper2Code: Automating Code Generation from Scientific Papers in Machine LearningMinju Seo, Jinheon Baek, Seongyun Lee, Sung Ju HwangICLR 2026 · 被引用 86 次
- Magicoder: Empowering Code Generation with OSS-InstructYuxiang Wei, Zhe Wang, Jiawei Liu, Yifeng Ding 等ICML 2024 · 被引用 246 次
- CodeSense: a Real-World Benchmark and Dataset for Code Semantic ReasoningMonoshi Kumar Roy, Simin Chen, Benjamin Steenhoek, Jinjun Peng 等ICLR 2026 · 被引用 18 次
