Knowledge-Graph-Driven Data Synthesis for Low-Resource Software Development: A HarmonyOS Case Study
Mingwei Liu, Zheng Pei, Yanlin Wang, Zihao Wang, Zikang Li, Enci Lin, Xin Peng, Zibin Zheng
摘要
In low-resource framework development (e.g., HarmonyOS), large language models (LLMs) often lack sufficient pre-training exposure, resulting in poor code generation performance. Although they generally preserve programming logic across languages, they frequently fail on framework-specific APIs and syntax, revealing a gap between learned algorithmic knowledge and unfamiliar framework conventions. Consequently, even advanced models such as GPT-4o struggle to produce correct code without prior exposure.
Inspired by these challenges, we propose APIKG4Syn, a framework that leverages API knowledge graphs to synthesize API-oriented question-code pairs without requiring executable environments. It incorporates both single-API and multi-API information, with the latter guided by uncertainty estimation (UE) and Monte Carlo Tree Search (MCTS), to construct high-quality fine-tuning data. For evaluation, we select HarmonyOS as a case study due to its accessible documentation and growing ecosystem, and build the first benchmark for its code generation. Experimental results show that fine-tuning Qwen2.5-Coder-7B with APIKG4Syn achieves a pass@1 of 25.00%, outperforming untuned GPT-4o (17.59%). We further observe that larger volumes of data generated by APIKG4Syn consistently lead to better fine-tuning performance, and that the optimal Single-API to Multi-API ratio is 8:2. Ablation studies also confirm the necessity and effectiveness of each component in our framework. These findings highlight the effectiveness of API-oriented data in enhancing LLM performance for low-resource software development scenarios.
CCS Concepts: • Software and its engineering → Automatic programming.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper17
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Magicoder: Empowering Code Generation with OSS-InstructYuxiang Wei, Zhe Wang, Jiawei Liu, Yifeng Ding 等ICML 2024 · 被引用 246 次
- LongCoder: A Long-Range Pre-trained Language Model for Code CompletionDaya Guo, Canwen Xu, Nan Duan, Jian Yin 等ICML 2023 · 被引用 150 次
- Large Language Models for Test-Free Fault LocalizationAidan Z. H. Yang, Claire Le Goues, Ruben Martins, Vincent J. HellendoornICSE 2024 · 被引用 98 次
- Lost in Translation: A Study of Bugs Introduced by Large Language Models while Translating CodeRangeet Pan, Ali Reza Ibrahimzada, Rahul Krishna, Divya Sankar 等ICSE 2024 · 被引用 96 次
相关 Paper
- ExploraCoder: Advancing Code Generation for Multiple Unseen APIs via Planning and Chained ExplorationYunkun Wang, Yue Zhang, Zhen Qin, Chen Zhi 等ACL 2025 · 被引用 13 次
- API-Guided Dataset Synthesis to Finetune Large Code ModelsZongjie Li, Daoyuan Wu, Shuai Wang, Zhendong SuOOPSLA 2025 · 被引用 6 次
- Knowledge Graph Finetuning Enhances Knowledge Manipulation in Large Language ModelsHanzhu Chen, Xu Shen, Jie Wang, Zehao Wang 等ICLR 2025
- MoT: Modularization-of-Thought Prompting for Effective Code GenerationRuwei Pan, Hongyu ZhangISSTA 2026
- HELO-APR: Enhancing Low-Resource Program Repair through Cross-Lingual Knowledge TransferZhipeng Wang, Boyang Yang, Yidong Wan, Liuye Guo 等ISSTA 2026
