Benchmarking and Learning Multi-Dimensional Quality Evaluator for Text-To-3D Generation
Yujie Zhang, Bingyang Cui, Qi Yang, Zhu Li, Yiling Xu
Abstract
Text-to-3D generation has achieved remarkable progress in recent years, yet evaluating these methods remains challenging for two reasons: i) Existing benchmarks lack fine-grained evaluation on different prompt categories and evaluation dimensions. ii) Previous evaluation metrics only focus on a single aspect (e.g., text-3D alignment) and fail to perform multi-dimensional quality assessment. To address these problems, we first propose a comprehensive benchmark named MATE-3D. The benchmark contains eight well-designed prompt categories that cover single and multiple object generation, resulting in 1,280 generated textured meshes. We have conducted a large-scale subjective experiment from four different evaluation dimensions and collected 107,520 annotations, followed by detailed analyses of the results. Based on MATE-3D, we propose a novel quality evaluator named HyperScore. Utilizing hypernetwork to generate specified mapping functions for each evaluation dimension, our metric can effectively perform multi-dimensional quality assessment. HyperScore presents superior performance over existing metrics on MATE-3D, making it a promising metric for assessing and improving text-to-3D generation. The project is available at https://mate-3d.github.io/.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Refining Few-Step Text-to-Multiview Diffusion via Reinforcement LearningZiyi Zhang, Li Shen, Deheng Ye, Yong Luo et al.CVPR 2026 · 2 citations
- Multimodal Semantic Bias Mitigation for Diverse Text-To-3D GenerationYukuan Min, Muli Yang, Jinhao Zhang, Yuxuan Wang et al.CVPR 2026
Builds on41
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 5,687 citations
Related papers
- Learning Multi-Dimensional Human Preference for Text-to-Image GenerationSixian Zhang, Bohan Wang, Junqiang Wu, Yan Li et al.CVPR 2024 · 15 citations
- Gen3DEval: Using vLLMs for Automatic Evaluation of Generated 3D ObjectsShalini Maiti, Lourdes Agapito, Filippos KokkinosCVPR 2025
- GPT-4V(ision) is a Human-Aligned Evaluator for Text-to-3D GenerationTong Wu, Guandao Yang, Zhibing Li, Kai Zhang et al.CVPR 2024 · 42 citations
- VBench: Comprehensive Benchmark Suite for Video Generative ModelsZiqi Huang, Yinan He, Jiashuo Yu, Fan Zhang et al.CVPR 2024
- WISE: World Knowledge-Informed Semantic Evaluation for Text-to-Image GenerationYuwei Niu, Munan Ning, Mengren Zheng, Weiyang Jin et al.ICML 2026 · 195 citations
