ARECHO: Autoregressive Evaluation via Chain-Based Hypothesis Optimization for Speech Multi-Metric Estimation
Jiatong Shi, Yifan Cheng, Bo-Hao Su, Hye-jin Shim, Jinchuan Tian, Samuele Cornell, Yiwen Zhao, Siddhant Arora, Shinji Watanabe
摘要
Speech signal analysis poses significant challenges, particularly in tasks such as speech quality evaluation and profiling, where the goal is to predict multiple perceptual and objective metrics. For instance, metrics like PESQ (Perceptual Evaluation of Speech Quality), STOI (Short-Time Objective Intelligibility), and MOS (Mean Opinion Score) each capture different aspects of speech quality. However, these metrics often have different scales, assumptions, and dependencies, making joint estimation non-trivial. To address these issues, we introduce ARECHO (Autoregressive Evaluation via Chain-based Hypothesis Optimization), a chain-based, versatile evaluation system for speech assessment grounded in autoregressive dependency modeling. ARECHO is distinguished by three key innovations: (1) a comprehensive speech information tokenization pipeline; (2) a dynamic classifier chain that explicitly captures inter-metric dependencies; and (3) a two-step confidence-oriented decoding algorithm that enhances inference reliability. Experiments demonstrate that ARECHO significantly outperforms the baseline framework across diverse evaluation scenarios, including enhanced speech analysis, speech generation evaluation, and, noisy speech evaluation. Furthermore, its dynamic dependency modeling improves interpretability by capturing inter-metric relationships. Across tasks, ARECHO offers reference-free evaluation using its dynamic classifier chain to support subset queries (single or multiple metrics) and reduces error propagation via confidence-oriented decoding.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper11
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman 等ICML 2023 · 被引用 6,966 次
- SCOREQ: Speech Quality Assessment with Contrastive RegressionAlessandro Ragano, Jan Skoglund, Andrew HinesNeurIPS 2024 · 被引用 90 次
- NORESQA: A Framework for Speech Quality Assessment using Non-Matching ReferencesPranay Manocha, Buye Xu, Anurag KumarNeurIPS 2021 · 被引用 67 次
- Mellow: a small audio language model for reasoningSoham Deshmukh, Satvik Dixit, Rita Singh, Bhiksha RajNeurIPS 2025 · 被引用 39 次
- GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning AbilitiesSreyan Ghosh, Sonal Kumar, Ashish Seth, Chandra Kiran Reddy Evuru 等EMNLP 2024 · 被引用 34 次
相关 Paper
- SpeechIQ: Speech-Agentic Intelligence Quotient Across Cognitive Levels in Voice Understanding by Large Language ModelsZhen Wan, Chao-Han Huck Yang, Yahan Yu, Jinchuan Tian 等ACL 2025 · 被引用 2 次
- TTSDS2: Resources and Benchmark for Evaluating Human-Quality Text to Speech SystemsChristoph Minixhofer, Ondrej Klejch, Peter BellICLR 2026 · 被引用 16 次
- Speech Recognition Model Improves Text-to-Speech Synthesis Using Fine-Grained RewardGuansu Wang, Peijie SunAAAI 2026
- UniSRM: A Unified Speech Reward Model for Reasoning-Based Fine-grained AssessmentYuanyuan Wang, Dongchao Yang, Yayue Deng, Zhiyong Wu 等ACL 2026
- SpeechLLM-as-Judges: Towards General and Interpretable Speech Quality EvaluationHui Wang, Jinghua Zhao, Yifan Yang, Shujie Liu 等ACL 2026 · 被引用 21 次
