Cued-Agent: A Collaborative Multi-Agent System for Automatic Cued Speech Recognition
Guanjie Huang, Danny H. K. Tsang, Shan Yang, Guangzhi Lei, Li Liu
摘要
Cued Speech (CS) is a visual communication system that combines lip-reading with hand coding to facilitate communication for individuals with hearing impairments. Automatic CS Recognition (ACSR) aims to convert CS hand gestures and lip movements into text via AI-driven methods. Traditionally, the temporal asynchrony between hand and lip movements requires the design of complex modules to facilitate effective multimodal fusion. However, constrained by limited data availability, current methods demonstrate insufficient capacity for adequately training these fusion mechanisms, resulting in suboptimal performance. Recently, multi-agent systems have shown promising capabilities in handling complex tasks with limited data availability. To this end, we propose the first collaborative multi-agent system for ACSR, named Cued-Agent. It integrates four specialized sub-agents: a Multimodal Large Language Model-based Hand Recognition agent that employs keyframe screening and CS expert prompt strategies to decode hand movements, a pretrained Transformer-based Lip Recognition agent that extracts lip features from the input video, a Hand Prompt Decoding agent that dynamically integrates hand prompts with lip features during inference in a training-free manner, and a Self-Correction Phoneme-to-Word agent that enables post-processing and end-to-end conversion from phoneme sequences to natural language sentences for the first time through semantic refinement. To support this study, we expand the existing Mandarin CS dataset by collecting data from eight hearing-impaired cuers, establishing a mixed dataset of fourteen subjects. Extensive experiments demonstrate that our Cued-Agent performs superbly in both normal and hearing-impaired scenarios compared with state-of-the-art methods. The implementation is available at https://github.com/DennisHgj/Cued-Agent.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper2
- GenArtist: Multimodal LLM as an Agent for Unified Image Generation and EditingZhenyu Wang, Aoxue Li, Zhenguo Li, Xihui LiuNeurIPS 2024 · 被引用 162 次
- MCCD: Multi-Agent Collaboration-based Compositional Diffusion for Complex Text-to-Image GenerationMingcheng Li, Xiaolu Hou, Ziyang Liu, Dingkang Yang 等CVPR 2025
相关 Paper
- Cuing Without Sharing: A Federated Cued Speech Recognition Framework via Mutual Knowledge DistillationYuxuan Zhang, Lei Liu, Li LiuACM MM 2023 · 被引用 8 次
- Cueing Without Gapping: Cuer-Independent Cued Speech Recognition Powered by Cross-Cuer Invariant ModelingFengji Ma, Chenxing Li, Li LiuAAAI 2026
- UniCUE: Unified Recognition and Generation Framework for Chinese Cued Speech Video-to-Speech GenerationJinting Wang, Shan Yang, Chenxing Li, Dong Yu 等AAAI 2026
- VALLR: Visual ASR Language Model for Lip ReadingMarshall Thomas, Edward Fish, Richard BowdenICCV 2025 · 被引用 6 次
- Talk With Human-like Agents: Empathetic Dialogue Through Perceptible Acoustic Reception and ReactionHaoqiu Yan, Yongxin Zhu, Kai Zheng, Bing Liu 等ACL 2024
