Cognify: Supercharging Gen-AI Workflows With Hierarchical Autotuning
Zijian He, Reyna Abhyankar, Vikranth Srivatsa, Yiying Zhang
Abstract
Today's gen-AI workflows that involve multiple ML model calls, tool/API calls, data retrieval, or generic code execution are often tuned manually in an ad-hoc way that is both time-consuming and error-prone. In this paper, we propose a systematic approach for automatically tuning gen-AI workflows. Our key insight is that gen-AI workflows can benefit from structure, operator, and prompt changes, but unique properties of gen-AI workflows require new optimization techniques. We propose AdaSeek, an adaptive hierarchical search algorithm for autotuning gen-AI workflows. AdaSeek organizes workflow tuning methods into different layers based on the user-specified total search budget and distributes the budget across different layers based on the complexity of each layer. During its hierarchical search, AdaSeek redistributes the search budget from less useful to more promising tuning configurations based on workflow-level evaluation results. We implement AdaSeek in a workflow autotuning framework called Cognify and evaluate Cognify using six types of workflows such as RAG-based QA and text-to-SQL transformation. Overall, Cognify improves these workflows' generation quality by up to 2.8×, reduces execution monetary cost by up to 10×, and reduces end-to-end latency by 2.7×.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a7a49fb8-6922-4890-86f2-70133ce12f81Cited by top-tier papers3
- KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent WorkflowsZaifeng Pan, Ajjkumar Patel, Yipeng Shen, Zhengding Hu et al.NeurIPS 2025 · 77 citations
- Cache-to-Cache: Direct Semantic Communication Between Large Language ModelsTianyu Fu, Zihan Min, Hanling Zhang, Jichao Yan et al.ICLR 2026 · 59 citations
- SCOPE: Cost-Efficient Model Selection for Compound AI Systems under Quality ConstraintsYiqian Huang, Shiqi Zhang, Tianyuan Jin, Xiaokui XiaoKDD 2026
Builds on12
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran et al.NeurIPS 2023 · 5,068 citations
- Language Models as Zero-Shot Planners: Extracting Actionable Knowledge for Embodied AgentsWenlong Huang, Pieter Abbeel, Deepak Pathak, Igor MordatchICML 2022 · 1,539 citations
- Large Language Models as OptimizersChengrun Yang, Xuezhi Wang, Yifeng Lu, Hanxiao Liu et al.ICLR 2024 · 817 citations
Related papers
- Bootstrapping in-situ workflow auto-tuning via combining performance models of component applicationsTong Shu, Yanfei Guo, Justin M. Wozniak, Xiaoning Ding et al.SC 2021 · 9 citations
- HEXGEN-FLOW: Optimizing LLM Inference Request Scheduling for Agentic Text-to-SQLYou Peng, Youhe Jiang, Wenqi Jiang, Chen Wang et al.ICDE 2026
- Swift: Fast Performance Tuning with GAN-Generated ConfigurationsChao Chen, Shixin Huang, Xuehai Qian, Zhibin YuUSENIX ATC 2025 · 1 citation
- FusionFlow: Enabling Deep Structural Exploration for Automated Agentic Workflow GenerationXiang Wang, Zongtao Yang, Zhuojian Hong, Shuhao Zhang et al.ACL 2026
- METIS: Fast Quality-Aware RAG Systems with Configuration AdaptationSiddhant Ray, Rui Pan, Zhuohan Gu, Kuntai Du et al.SOSP 2025 · 3 citations
