Armor: A Benchmark for Meta-evaluation of Artificial Music
Songhe Wang, Zheng Bao, Jingtong E
Abstract
Objective evaluation (OE) is essential to artificial music, but it's often very hard to determine the quality of OEs. Hitherto, subjective evaluation (SE) remains reliable and prevailing but suffers inevitable disadvantages that OEs may overcome. Therefore, a metaevaluation system is necessary for designers to test the effectiveness of OEs. In this paper, we present Armor, a complex and crossdomain benchmark dataset that serves for this purpose. Since OEs should correlate with human judgment, we provide music as test cases for OEs and human judgment scores as touchstones. We also provide two meta-evaluation scenarios and their corresponding testing methods to assess the effectiveness of OEs. To the best of our knowledge, Armor is the first comprehensive and rigorous framework that future works could follow, take example by, and improve upon for the task of evaluating computer-generated music and the field of computational music as a whole. By analyzing different OE methods on our dataset, we observe that there is still a huge gap between SE and OE, meaning that hard-coded algorithms are far from catching human's judgment to the music.
• Applied computing → Sound and music computing.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on1
Related papers
- An Order-Complexity Aesthetic Assessment Model for Aesthetic-aware Music RecommendationXin Jin, Wu Zhou, Jinyu Wang, Duo Xu et al.ACM MM 2023 · 4 citations
- Video Background Music Generation: Dataset, Method and EvaluationLe Zhuo, Zhaokai Wang, Baisen Wang, Yue Liao et al.ICCV 2023 · 51 citations
- CMI-RewardBench: Evaluating Music Reward Models with Compositional Multimodal InstructionYinghao Ma, Haiwen Xia, Hewei Gao, Weixiong Chen et al.ICML 2026 · 4 citations
- MMAU-Pro: A Challenging and Comprehensive Benchmark for Holistic Evaluation of Audio General IntelligenceSonal Kumar, Simon Sedlácek, Vaibhavi Lokegaonkar, Fernando López et al.AAAI 2026 · 1 citation
- Understanding the Potentials and Limitations of Prompt-based Music Generative AIYoujin Choi, JaeYoung Moon, Jinyoung Yoo, Jin-Hyuk HongCHI 2025 · 7 citations
