DiffAgent: Fast and Accurate Text-to-Image API Selection with Large Language Model
Lirui Zhao, Yue Yang, Kaipeng Zhang, Wenqi Shao, Yuxin Zhang, Yu Qiao, Ping Luo, Rongrong Ji
摘要
Text-to-image (T2I) generative models have attracted significant attention and found extensive applications within and beyond academic research. For example, the Civitai community, a platform for T2I innovation, currently hosts an impressive array of 74,492 distinct models. However, this diversity presents a formidable challenge in selecting the most appropriate model and parameters, a process that typically requires numerous trials. Drawing inspiration from the tool usage research of large language models (LLMs), we introduce DiffAgent, an LLM agent designed to screen the accurate selection in seconds via API calls. DiffAgent leverages a novel two-stage training framework, SFTA, enabling it to accurately align T2I API responses with user input in accordance with human preferences. To train and evaluate DiffAgent's capabilities, we present DABench, a comprehensive dataset encompassing an extensive range of T2I APIs from the community. Our evaluations reveal that DiffAgent not only excels in identifying the appropriate T2I API but also underscores the effectiveness of the SFTA training framework. Codes are available at https://github.com/OpenGVLab/DiffAgent.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- DiffGraph: An Automated Agent-driven Model Merging Framework for In-the-Wild Text-to-Image GenerationZhuoling Li, Hossein Rahmani, Jiarui Zhang, Yu Xue 等CVPR 2026 · 被引用 5 次
- OctoT2I: A Self-Evolving Agentic Text-to-Image RouterXu Jiang, Bin Chen, Gehui Li, Yule Duan 等CVPR 2026 · 被引用 4 次
- MuseScorer: Idea Originality Scoring At ScaleAli Sarosh Bangash, Krish Veera, Ishfat Abrar Islam, Raiyan Abdul BatenEMNLP 2025 · 被引用 2 次
- Information Theoretic Text-to-Image AlignmentChao Wang, Giulio Franzese, Alessandro Finamore, Massimo Gallo 等ICLR 2025
- JarvisIR: Elevating Autonomous Driving Perception with Intelligent Image RestorationYunlong Lin, Zixu Lin, Haoyu Chen, Panwang Pan 等CVPR 2025
它引用的顶会 Paper20
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
相关 Paper
- DiffBench Meets DiffAgent: End-to-End LLM-Driven Diffusion Acceleration Code GenerationJiajun Jiao, Haowei Zhu, Puyuan Yang, Jianghui Wang 等AAAI 2026 · 被引用 1 次
- T2I-Copilot: A Training-Free Multi-Agent Text-to-Image System for Enhanced Prompt Interpretation and Interactive GenerationChieh-Yun Chen, Min Shi, Gong Zhang, Humphrey ShiICCV 2025 · 被引用 1 次
- AGFSync: Leveraging AI-Generated Feedback for Preference Optimization in Text-to-Image GenerationJingkun An, Yinghao Zhu, Zongjian Li, Enshen Zhou 等AAAI 2025 · 被引用 10 次
- CodeAgent: Enhancing Code Generation with Tool-Integrated Agent Systems for Real-World Repo-level Coding ChallengesKechi Zhang, Jia Li, Ge Li, Xianjie Shi 等ACL 2024
- ChatGen: Automatic Text-to-Image Generation From FreeStyle ChattingChengyou Jia, Changliang Xia, Zhuohang Dang, Weijia Wu 等CVPR 2025
