Adaptive Self-improvement LLM Agentic System for ML Library Development
Genghan Zhang, Weixin Liang, Olivia Hsu, Kunle Olukotun
摘要
ML libraries, often written in architecture-specific programming languages (ASPLs) that target domain-specific architectures, are key to efficient ML systems. However, writing these highperformance ML libraries is challenging because it requires expert knowledge of both ML algorithms and the ASPL. Large language models (LLMs), on the other hand, have shown general coding capabilities. However, challenges remain when using LLMs for generating ML libraries using ASPLs because 1) this task is complicated even for human experts and 2) there are limited code examples due to the esoteric and evolving nature of ASPLs. We present an adaptive selfimprovement agentic system that enables LLMs to perform such complex reasoning under limited data by iteratively improving their capability through self-generated experience. In order to evaluate the effectiveness of our system, we construct a benchmark of a typical ML library and generate ASPL code with both open and closedsource LLMs on this benchmark. Our results show improvements of up to 3.9× over a baseline single LLM 1 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Agentic Context Engineering: Evolving Contexts for Self-Improving Language ModelsQizheng Zhang, Changran Hu, Shubhangi Upasani, Boyuan Ma 等ICLR 2026 · 被引用 374 次
- KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent WorkflowsZaifeng Pan, Ajjkumar Patel, Yipeng Shen, Zhengding Hu 等NeurIPS 2025 · 被引用 77 次
- Agentic Plan Caching: Test-Time Memory for Fast and Cost-Efficient LLM AgentsQizheng Zhang, Michael Wornow, Kunle OlukotunNeurIPS 2025 · 被引用 27 次
- Can Dependencies Induced by LLM-Agent Workflows Be Trusted?Yu Yao, Yiliao Song, Yian Xie, Mengdan Fan 等NeurIPS 2025 · 被引用 4 次
- CubeBench: Diagnosing Interactive, Long-Horizon Physical Intelligence under Partial ObservationsHuan-ang Gao, Zikang Zhang, Tianwei Luo, Kaisen Yang 等ICLR 2026 · 被引用 2 次
它引用的顶会 Paper28
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Toolformer: Language Models Can Teach Themselves to Use ToolsTimo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu 等NeurIPS 2023 · 被引用 5,989 次
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan 等NeurIPS 2023 · 被引用 5,828 次
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra 等NeurIPS 2022 · 被引用 5,493 次
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran 等NeurIPS 2023 · 被引用 5,068 次
相关 Paper
- ARBench: Algorithmic Reasoner or API Alchemist? Evaluating LLMs Beyond API CallsRenbiao Liu, Chao-Zeng Ma, Anqi Li, Hui Sun 等AAAI 2026
- STARK: Strategic Team of Agents for Refining KernelsJuncheng Dong, Yang Yang, Tao Liu, Yang Wang 等ICLR 2026 · 被引用 26 次
- MLE-STAR: Machine Learning Engineering Agent via Search and Targeted RefinementJaehyun Nam, Jinsung Yoon, Jiefeng Chen, Jinwoo Shin 等NeurIPS 2025 · 被引用 58 次
- Enhancing Open-Domain Task-Solving Capability of LLMs via Autonomous Tool Integration from GitHubBohan Lyu, Xin Cong, Heyang Yu, Pan Yang 等ACL 2025
- An Agentic Framework with LLMs for Solving Complex Vehicle Routing ProblemsNi Zhang, Zhiguang Cao, Jianan Zhou, Cong Zhang 等ICLR 2026 · 被引用 8 次
