Towards AI-driven Sign Language Generation with Non-manual Markers
Han Zhang, Rotem Shalev-Arkushin, Vasileios Baltatzis, Connor Gillis, Gierad Laput, Raja S. Kushalnagar, Lorna C. Quandt, Leah Findlater, Abdelkareem Bedri, Colin Lea
摘要
Sign languages are essential for the Deaf and Hard-of-Hearing (DHH) community. Sign language generation systems have the potential to support communication by translating from written languages, such as English, into signed videos. However, current systems often fail to meet user needs due to poor translation of grammatical structures, the absence of facial cues and body language, and insufficient visual and motion fidelity. We address these challenges by building on recent advances in LLMs and video generation models to translate English sentences into natural-looking AI ASL signers. The text component of our model extracts information for manual and non-manual components of ASL, which are used to synthesize skeletal pose sequences and corresponding video frames. Our findings from a user study with 30 DHH participants and thorough technical evaluations demonstrate significant progress and identify critical areas necessary to meet user needs.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- ImageRAG: Dynamic Image Retrieval for Reference-Guided Image GenerationRotem Shalev-Arkushin, Rinon Gal, Amit Bermano, Ohad FriedICLR 2026 · 被引用 25 次
- ImageSentinel: Protecting Visual Datasets from Unauthorized Retrieval-Augmented Image GenerationZiyuan Luo, Yangyi Zhao, Ka Chun Cheung, Simon See 等NeurIPS 2025 · 被引用 5 次
- Stable Signer: Hierarchical Sign Language Generative ModelSen Fang, Yalin Feng, Hongbin Zhong, Yanxin Zhang 等ACL 2026 · 被引用 3 次
- Reimagining Sign Language Technologies: Analyzing Translation Work of Chinese Deaf Online Content CreatorsXinru Tang, Anne Marie PiperCHI 2026 · 被引用 2 次
- ASL Educators' Perspectives on AI for Enhancing Student Learning in American Sign Language EducationSaad Hassan, Laleh Nourian, Caluã de Lacerda Pataca, Michelle M. Olson 等CHI 2026 · 被引用 1 次
它引用的顶会 Paper26
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni 等NeurIPS 2020 · 被引用 19,162 次
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 被引用 6,759 次
- T2I-Adapter: Learning Adapters to Dig Out More Controllable Ability for Text-to-Image Diffusion ModelsChong Mou, Xintao Wang, Liangbin Xie, Yanze Wu 等AAAI 2024 · 被引用 1,641 次
相关 Paper
- Social App Accessibility for Deaf SignersKelly Mack, Danielle Bragg, Meredith Ringel Morris, Maarten W. Bos 等CSCW 2020 · 被引用 52 次
- Signs as Tokens: A Retrieval-Enhanced Multilingual Sign Language GeneratorRonglai Zuo, Rolandos Alexandros Potamias, Evangelos Ververas, Jiankang Deng 等ICCV 2025 · 被引用 9 次
- Customizing Generated Signs and Voices of AI Avatars: Deaf-Centric Mixed-Reality Design for Deaf-Hearing CommunicationSi Chen, Haocong Cheng, Suzy Su, Stephanie Patterson 等CSCW 2025 · 被引用 9 次
- Exploring the Impact of Emotional Voice Integration in Sign-to-Speech Translators for Deaf-to-Hearing CommunicationHyunchul Lim, Minghan Gao, Franklin Mingzhe Li, Nam Anh Dang 等CSCW 2025 · 被引用 2 次
- Perceptions and Preferences: Deaf ASL-Signing Users' Insights on Video Elements, Styles and LayoutsKhulood Alkhudaidi, Tish Burke, Rachel Boll, Shruti Mahajan 等CHI 2025 · 被引用 1 次
