Evaluation Agent: Efficient and Promptable Evaluation Framework for Visual Generative Models
Fan Zhang, Shulin Tian, Ziqi Huang, Yu Qiao, Ziwei Liu
摘要
Recent advancements in visual generative models have enabled high-quality image and video generation, opening diverse applications. However, evaluating these models often demands sampling hundreds or thousands of images or videos, making the process computationally expensive, especially for diffusion-based models with inherently slow sampling. Moreover, existing evaluation methods rely on rigid pipelines that overlook specific user needs and provide numerical results without clear explanations. In contrast, humans can quickly form impressions of a model's capabilities by observing only a few samples. To mimic this, we propose the Evaluation Agent framework, which employs human-like strategies for efficient, dynamic, multi-round evaluations using only a few samples per round, while offering detailed, user-tailored analyses. It offers four key advantages: 1) efficiency, 2) promptable evaluation tailored to diverse user needs, 3) explainability beyond single numerical scores, and 4) scalability across various models and tools. Experiments show that Evaluation Agent reduces evaluation time to 10% of traditional methods while delivering comparable results. The Evaluation Agent framework is fully open-sourced to advance research in visual generative models and their efficient evaluation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper16
- SANA-Video: Efficient Video Generation with Block Linear Diffusion TransformerJunsong Chen, Yuyang Zhao, Jincheng Yu, Ruihang Chu 等ICLR 2026 · 被引用 96 次
- SLA: Beyond Sparsity in Diffusion Transformers via Fine-Tunable Sparse–Linear AttentionJintao Zhang, Haoxu Wang, Kai Jiang, Shuo Yang 等ICLR 2026 · 被引用 57 次
- EditVerse: Unifying Image and Video Editing and Generation with In-Context LearningXuan Ju, Tianyu Wang, Yuqian Zhou, He Zhang 等ICLR 2026 · 被引用 56 次
- InfinityStar: Unified Spacetime AutoRegressive Modeling for Visual GenerationJinlai Liu, Jian Han, Bin Yan, Hui Wu 等NeurIPS 2025 · 被引用 45 次
- VISTA: A Test-Time Self-Improving Video Generation AgentDo Xuan Long, Xingchen Wan, Hootan Nakhost, Chen-Yu Lee 等CVPR 2026 · 被引用 30 次
它引用的顶会 Paper23
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo 等NeurIPS 2022 · 被引用 8,168 次
- Toolformer: Language Models Can Teach Themselves to Use ToolsTimo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu 等NeurIPS 2023 · 被引用 5,989 次
相关 Paper
- ManipEvalAgent: Promptable and Efficient Evaluation Framework for Robotic Manipulation PoliciesYiteng Chen, Huiping Zhuang, Wenbo Li, Shiyi Wang 等ICLR 2026
- DiffBench Meets DiffAgent: End-to-End LLM-Driven Diffusion Acceleration Code GenerationJiajun Jiao, Haowei Zhu, Puyuan Yang, Jianghui Wang 等AAAI 2026 · 被引用 1 次
- Fast Sampling of Diffusion Models with Exponential IntegratorQinsheng Zhang, Yongxin ChenICLR 2023 · 被引用 58 次
- A Simple Early Exiting Framework for Accelerated Sampling in Diffusion ModelsTae Hong Moon, Moonseok Choi, EungGu Yun, Jongmin Yoon 等ICML 2024 · 被引用 10 次
- EdiVal-Agent: An Object-Centric Framework for Automated, Fine-Grained Evaluation of Multi-Turn EditingTianyu Chen, Yasi Zhang, Zhi Zhang, Peiyu Yu 等ICLR 2026 · 被引用 11 次
