ComfyBench: Benchmarking LLM-based Agents in ComfyUI for Autonomously Designing Collaborative AI Systems
Xiangyuan Xue, Zeyu Lu, Di Huang, Zidong Wang, Wanli Ouyang, Lei Bai
Abstract
Much previous AI research has focused on developing monolithic models to maximize their intelligence, with the primary goal of enhancing performance on specific tasks. In contrast, this work attempts to study using LLM-based agents to design collaborative AI systems autonomously. To explore this problem, we first introduce ComfyBench to evaluate agents's ability to design collaborative AI systems in ComfyUI. ComfyBench is a comprehensive benchmark comprising 200 diverse tasks covering various instructionfollowing generation challenges, along with detailed annotations for 3,205 nodes and 20 workflows. Based on Comfy-Bench, we further develop ComfyAgent, a novel framework that empowers LLM-based agents to autonomously design collaborative AI systems by generating workflows. Com-fyAgent is based on two core concepts. First, it represents workflows with code, which can be reversibly converted into workflows and executed as collaborative systems by the interpreter. Second, it constructs a multi-agent system that cooperates to learn from existing workflows and generate new workflows for a given task. While experimental results demonstrate that ComfyAgent achieves a comparable resolve rate to o1-preview and significantly surpasses other agents on ComfyBench, ComfyAgent has resolved only 15% of creative tasks. LLM-based agents still have a long way to go in autonomously designing collaborative AI systems. Progress with ComfyBench is paving the way for more intelligent and autonomous collaborative AI systems. Our code is available at: https://github.com/xxyQwQ/ComfyBench .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bb7181e9-fb77-4b4c-8504-633afbfad22eCited by top-tier papers7
- ComfyMind: Toward General-Purpose Generation via Tree-Based Planning and Reactive FeedbackLitao Guo, Xinli Xu, Luozhou Wang, Jiantao Lin et al.NeurIPS 2025 · 18 citations
- AgentBreeder: Mitigating the AI Safety Risks of Multi-Agent Scaffolds via Self-ImprovementJ. Rosser, Jakob N. FoersterNeurIPS 2025 · 12 citations
- ReSo: A Reward-driven Self-organizing LLM-based Multi-Agent System for Reasoning TasksHeng Zhou, Hejia Geng, Xiangyuan Xue, Li Kang et al.EMNLP 2025 · 4 citations
- SoPE: Spherical Coordinate-Based Positional Embedding for Enhancing Spatial Perception of 3D LVLMsKoonting Yip, Qiyan Zhao, Wenhao Yu, Liangyu Yuan et al.CVPR 2026 · 3 citations
- Policy Optimized Text-to-Image Pipeline DesignUri Gadot, Rinon Gal, Yftah Ziser, Gal Chechik et al.NeurIPS 2025 · 1 citation
Builds on23
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
Related papers
- CO-Bench: Benchmarking Language Model Agents in Algorithm Search for Combinatorial OptimizationWeiwei Sun, Shengyu Feng, Shanda Li, Yiming YangAAAI 2026 · 20 citations
- CollabBench: Benchmarking and Unleashing Collaborative Ability of LLMs with Diverse Players via Proactive EngagementHong Qian, Yuanhao Liu, Zihan Zhou, Zongbao Zhang et al.ICML 2026
- MultiAgentBench : Evaluating the Collaboration and Competition of LLM agentsKunlun Zhu, Hongyi Du, Zhaochen Hong, Xiaocheng Yang et al.ACL 2025 · 97 citations
- CompileAgent: Automated Real-World Repo-Level Compilation with Tool-Integrated LLM-based Agent SystemLi Hu, Guoqiang Chen, Xiuwei Shang, Shaoyin Cheng et al.ACL 2025
- AgencyBench: Benchmarking the Frontiers of Autonomous Agents in 1M-Token Real-World ContextsKeyu Li, Junhao Shi, Yang Xiao, Mohan Jiang et al.ACL 2026 · 14 citations
