Sign Language Video Retrieval with Free-Form Textual Queries
Amanda Cardoso Duarte, Samuel Albanie, Xavier Giró-i-Nieto, Gül Varol
摘要
Systems that can efficiently search collections of sign language videos have been highlighted as a useful application of sign language technology. However, the problem of searching videos beyond individual keywords has received limited attention in the literature. To address this gap, in this work we introduce the task of sign language retrieval with free-form <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sup> <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sup> The terminology “natural language query” is commonly used to describe unconstrained textual queries in spoken languages. However, since sign languages are also natural languages, we adopt for the term “free-form textual query” instead. textual queries: given a written query (e.g. a sentence) and a large collection of sign language videos, the objective is to find the signing video that best matches the written query. We propose to tackle this task by learning cross-modal embeddings on the recently introduced large-scale How2Sign dataset of American Sign Language (ASL). We identify that a key bottleneck in the performance of the system is the quality of the sign video embedding which suffers from a scarcity of labelled training data. We, therefore, propose SPOT-ALIGN, a framework for interleaving iterative rounds of sign spotting and feature alignment to expand the scope and scale of available training data. We validate the effectiveness of SPOT-ALIGN for learning a robust sign video embedding through improvements in both sign recognition and the proposed video retrieval task.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Understanding Co-Speech Gestures in-the-WildSindhu B. Hegde, K. R. Prajwal, Taein Kwon, Andrew ZissermanICCV 2025 · 被引用 4 次
- SEDS: Semantically Enhanced Dual-Stream Encoder for Sign Language RetrievalLongtao Jiang, Min Wang, Zecheng Li, Yao Fang 等ACM MM 2024 · 被引用 2 次
- CiCo: Domain-Aware Sign Language Retrieval via Cross-Lingual Contrastive LearningYiting Cheng, Fangyun Wei, Jianmin Bao, Dong Chen 等CVPR 2023
- Natural Language-Assisted Sign Language RecognitionRonglai Zuo, Fangyun Wei, Brian MakCVPR 2023
- Towards Privacy-Aware Sign Language Translation at ScalePhillip Rust, Bowen Shi, Skyler Wang, Necati Cihan Camgöz 等ACL 2024
它引用的顶会 Paper14
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel 等ICLR 2020 · 被引用 7,418 次
- On the Variance of the Adaptive Learning Rate and BeyondLiyuan Liu, Haoming Jiang, Pengcheng He, Weizhu Chen 等ICLR 2020 · 被引用 2,210 次
- Frozen in Time: A Joint Video and Image Encoder for End-to-End RetrievalMax Bain, Arsha Nagrani, Gül Varol, Andrew ZissermanICCV 2021 · 被引用 1,550 次
- HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video ClipsAntoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi 等ICCV 2019 · 被引用 1,437 次
- Support-set bottlenecks for video-text representation learningMandela Patrick, Po-Yao Huang, Yuki Markus Asano, Florian Metze 等ICLR 2021 · 被引用 269 次
相关 Paper
- Open-Domain Sign Language Translation Learned from Online VideoBowen Shi, Diane Brentari, Gregory Shakhnarovich, Karen LivescuEMNLP 2022 · 被引用 39 次
- Searching for fingerspelled content in American Sign LanguageBowen Shi, Diane Brentari, Greg Shakhnarovich, Karen LivescuACL 2022 · 被引用 8 次
- SignCLIP: Connecting Text and Sign Language by Contrastive LearningZifan Jiang, Gerard Sant, Amit Moryossef, Mathias Müller 等EMNLP 2024 · 被引用 4 次
- Sentence-level Segmentation for Long Sign Language Videos with CaptionsBowen Guo, Shiwei Gan, Yafeng Yin, Xiao Liu 等ACM MM 2025
- Segment, Embed, and Align: A Universal Recipe for Aligning Subtitles to SigningZifan Jiang, Youngjoon Jang, Liliane Momeni, Gül Varol 等ACL 2026
