Gen3DEval: Using vLLMs for Automatic Evaluation of Generated 3D Objects
Shalini Maiti, Lourdes Agapito, Filippos Kokkinos
2025Year
1Top-tier citations
Abstract
Object 1 Object 2 Object 1 Object 2 Gen3DEval Object 1 Object 2 Object 1 Object 2 Object 1 Object 2
Figure 1. Gen3DEval: A holistic ranking metric to assess the quality of generated 3D objects on appearance, surface quality and text fidelity using a vision large language model (vLLM) which is trained to choose the better out of two objects on the three evaluation dimensions (appearance, text fidelity or surface quality).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on33
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
Related papers
- Benchmarking and Learning Multi-Dimensional Quality Evaluator for Text-To-3D GenerationYujie Zhang, Bingyang Cui, Qi Yang, Zhu Li et al.ICCV 2025 · 3 citations
- Multi-Dimensional Text-to-Face Image Quality Assessment Using LLM: Database and MethodYixuan Gao, Xiongkuo Min, Jinliang Han, Yuqin Cao et al.ACM MM 2025 · 2 citations
- SaMer: A Scenario-aware Multi-dimensional Evaluator for Large Language ModelsKehua Feng, Keyan Ding, Jing Yu, Yiwen Qu et al.ICLR 2025
- CG-MLLM: Captioning and Generating 3D content via Multi-modal Large Language ModelsJunming Huang, Chi Wang, Letian Li, Guangkai Xu et al.ICML 2026 · 2 citations
- LLMScore: Unveiling the Power of Large Language Models in Text-to-Image Synthesis EvaluationYujie Lu, Xianjun Yang, Xiujun Li, Xin Eric Wang et al.NeurIPS 2023 · 119 citations
