SC2025Top-tier venue
STELLAR: Storage Tuning Engine Leveraging LLM Autonomous Reasoning for High Performance Parallel File Systems
Chris Egersdoerfer, Philip H. Carns, Shane Snyder, Robert B. Ross, Dong Dai
Abstract
I/O performance is crucial to efficiency in data-intensive scientific computing; but tuning large-scale storage systems is complex, costly, and notoriously manpower-intensive, making it inaccessible for most domain scientists. To address this problem, we propose STELLAR, an autonomous tuner for high-performance parallel file systems. Our evaluations show that STELLAR almost always selects near-optimal configurations for the parallel file systems within the first five attempts, even for previously unseen applications.
STELLAR's human-like efficiency is fundamentally different from existing autotuning methods, which often require hundreds of thousands of iterations to converge. STELLAR achieves this through autonomous end-to-end agentic tuning. Powered by large language models (LLMs), STELLAR is capable of (1) accurately extracting tunable parameters from software manuals, (2) analyzing I/O trace logs generated by applications, (3) selecting initial tuning strategies, (4) rerunning applications on real systems and collecting I/O performance feedback, (5) adjusting tuning strategies and repeating the tuning cycle, and (6) reflecting on and summarizing tuning experiences into reusable knowledge for future optimizations.
STELLAR integrates retrieval-augmented generation (RAG), external tool execution, LLM-based reasoning, and a multiagent design to stabilize reasoning and combat hallucinations. We evaluate how each of these components impacts optimization outcomes, thus providing insight into the design of similar systems for other optimization problems. STELLAR's architecture and empirical validation open new avenues for tackling complex system optimization challenges, especially those characterized by vast search spaces and high exploration costs. Its highly efficient autonomous tuning will broaden access to I/O performance optimizations for domain scientists with minimal additional resource investment.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on7
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan et al.NeurIPS 2023 · 5,828 citations
- In-Context Impersonation Reveals Large Language Models' Strengths and BiasesLeonard Salewski, Stephan Alaniz, Isabel Rio-Torto, Eric Schulz et al.NeurIPS 2023 · 259 citations
- GPTuner: A Manual-Reading Database Tuning System via GPT-Guided Bayesian OptimizationJiale Lao, Yibo Wang, Yufei Li, Jianping Wang et al.VLDB 2024 · 76 citations
Related papers
- GLANCED-IO: Taming I/O Optimization for Deep Learning at ScaleRay A. O. Sinurat, William Nixon, Philip H. Carns, Huihuo Zheng et al.HPDC 2026
- Carver: Finding Important Parameters for Storage System TuningZhen Cao, Geoff Kuenning, Erez ZadokFAST 2020 · 52 citations
- CARBS: Compiler Autotuning via Randomized Biased SearchWei Li, Bin Gao, Weng-Fai WongHPDC 2026
- TuneAgent: Agentic Operating System Kernel Tuning with Reinforcement LearningHongyu Lin, Yuchen Li, Haoran Luo, Zhenghong Lin et al.KDD 2026 · 3 citations
- Improving Parallel Program Performance with LLM Optimizers via Agent-System InterfacesAnjiang Wei, Allen Nie, Thiago S. F. X. Teixeira, Rohan Yadav et al.ICML 2025
