BashCoder-R1: Towards Robust and Explainable Bash Script Generation with Robustness-Aware Group Relative Policy Optimization
Lei Yu, Peng Wang, Jia Xu, Jingyuan Zhang, Xin Wang, Jiajia Ma, Li Yang, Changzhi Deng, Zenghua Wang, Fengjun Zhang
摘要
Bash scripts underpin system administration, DevOps automation, and Continuous Integration/Continuous Deployment (CI/CD), where code quality directly determines system stability and security. Yet when Large Language Models (LLMs) are tasked with generating such scripts, two problems compound each other: models produce code without any accompanying justification for their design choices, and this absence of scrutiny correlates with scripts that silently harbor robustness flaws, including mishandled edge cases, unchecked failure modes, and fragile assumptions about the execution environment. We present BashCoder-R1, a framework that tackles both problems jointly, treating explainability as a design goal rather than a byproduct of correctness. The training pipeline proceeds in three stages. Continual Pre-training (CPT) first adapts the base model to the syntactic conventions and idioms specific to Bash. We then curate thousands of expert-validated reasoning-and-code samples and use them for Long Chain-of-Thought Supervised Fine-Tuning (L-CoT SFT), training the model to reproduce the deliberative, risk-averse thinking pattern of an experienced system administrator before it writes any code. Finally, Robustness-Aware Group Relative Policy Optimization (R-GRPO) directly refines the generation policy against a weighted reward combining syntax correctness, robustness as verified by the static analyzer shellcheck , and adherence to the required reasoning format. We evaluate BashCoder-R1 on BashBench , a benchmark we construct from 952 real-world automation tasks (773 single-line commands and 179 multi-line scripts), against a wide range of state-of-the-art baselines. BashCoder-R1 attains SyntaxPass of 100.00% / 94.97%, RobustWarnRate of 4.01% / 16.47%, RobustPass of 95.99% / 79.33%, FuncRate of 93.01% / 93.85%, and FullRate of 90.04% / 73.18% on single-line / multi-line tasks respectively, a relative FullRate improvement of 37.82% and 20.18% over the strongest baseline, DeepSeek-V3.2 (Reasoning). Human evaluation across Functionality, Robustness, and Clarity further confirms that BashCoder-R1's reasoning chains are rated highest in quality among all compared systems.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- SmartCoder-R1: Towards Secure and Explainable Smart Contract Generation with Security-Aware Group Relative Policy OptimizationLei Yu, Jingyuan Zhang, Xin Wang, Li Yang 等FSE 2026 · 被引用 3 次
- Bash-Commenter: Leveraging Syntax-Aware Preference Optimization to Reinforce Large Language Model for Bash Code Comment GenerationLei Yu, Jingyuan Zhang, Xin Wang, Li Yang 等FSE 2026 · 被引用 1 次
- AutoBaxBuilder: Bootstrapping Code Security BenchmarkingTobias von Arx, Niels Mündler, Mark Vero, Maximilian Baader 等ICML 2026 · 被引用 1 次
- BaxBench: Can LLMs Generate Correct and Secure Backends?Mark Vero, Niels Mündler, Victor Chibotaru, Veselin Raychev 等ICML 2025
- SEER: Enhancing Chain-of-Thought Code Generation through Self-Exploring Deep ReasoningShuzheng Gao, Chaozheng Wang, Cuiyun Gao, Michael R. LyuICSE 2026 · 被引用 1 次
