LLMs are Good Sign Language Translators
Jia Gong, Lin Geng Foo, Yixuan He, Hossein Rahmani, Jun Liu
摘要
Sign Language Translation (SLT) is a challenging task that aims to translate sign videos into spoken language. Inspired by the strong translation capabilities of large language models (LLMs) that are trained on extensive multilingual text corpora, we aim to harness off-the-shelf LLMs to handle SLT. In this paper, we regularize the sign videos to embody linguistic characteristics of spoken language, and propose a novel SignLLM framework to transform sign videos into a language-like representation for improved readability by off-the-shelf LLMs. SignLLM comprises two key modules: (1) The Vector-Quantized Visual Sign module converts sign videos into a sequence of discrete characterlevel sign tokens, and (2) the Codebook Reconstruction and Alignment module converts these character-level tokens into word-level sign representations using an optimal transport formulation. A sign-text alignment loss further bridges the gap between sign and text tokens, enhancing semantic compatibility. We achieve state-of-the-art gloss-free results on two widely-used SLT benchmarks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper24
- LLMs are Good Action RecognizersHaoxuan Qu, Yujun Cai, Jun LiuCVPR 2024 · 被引用 37 次
- Bridging Sign and Spoken Languages: Pseudo Gloss Generation for Sign Language TranslationJianyuan Guo, Peike Li, Trevor CohnNeurIPS 2025 · 被引用 17 次
- Geo-Sign: Hyperbolic Contrastive Regularisation for Geometrically Aware Sign Language TranslationEdward Fish, Richard BowdenNeurIPS 2025 · 被引用 15 次
- MixSignGraph: A Sign Sequence is Worth Mixed Graphs of NodesShiwei Gan, Yafeng Yin, Zhiwei Jiang, Lei Xie 等NeurIPS 2025 · 被引用 11 次
- Leveraging the Power of MLLMs for Gloss-Free Sign Language TranslationJungeun Kim, Hyeongwoo Jeon, Jongseong Bae, Ha Young KimICCV 2025 · 被引用 10 次
它引用的顶会 Paper24
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 被引用 9,451 次
- data2vec: A General Framework for Self-supervised Learning in Speech, Vision and LanguageAlexei Baevski, Wei-Ning Hsu, Qiantong Xu, Arun Babu 等ICML 2022 · 被引用 1,123 次
- LLM-Planner: Few-Shot Grounded Planning for Embodied Agents with Large Language ModelsChan Hee Song, Brian M. Sadler, Jiaman Wu, Wei-Lun Chao 等ICCV 2023 · 被引用 685 次
相关 Paper
- Lost in Translation, Found in Context: Sign Language Translation with Contextual CuesYoungjoon Jang, Haran Raajesh, Liliane Momeni, Gül Varol 等CVPR 2025
- Gloss Attention for Gloss-free Sign Language TranslationAoxiong Yin, Tianyun Zhong, Li Tang, Weike Jin 等CVPR 2023
- Conditional Variational Autoencoder for Sign Language Translation with Cross-Modal AlignmentRui Zhao, Liang Zhang, Biao Fu, Cong Hu 等AAAI 2024 · 被引用 36 次
- Sign2GPT: Leveraging Large Language Models for Gloss-Free Sign Language TranslationRyan Wong, Necati Cihan Camgöz, Richard BowdenICLR 2024 · 被引用 58 次
- Scaling Sign Language TranslationBiao Zhang, Garrett Tanzer, Orhan FiratNeurIPS 2024 · 被引用 21 次
