Sentence-level Segmentation for Long Sign Language Videos with Captions
Bowen Guo, Shiwei Gan, Yafeng Yin, Xiao Liu, Zhiwei Jiang, Shunmei Meng
摘要
In existing Sign Language (SL) research, most datasets and backbone models focus on sentence-level samples. However, the annotated sentence-level SL datasets are rather limited, and it is in great need to expand sentence-level SL datasets. When considering the large-scale long SL videos with captions, we propose a new task, i.e., Sentence-level Sign Language Segmentation (SSLS), which splits the long videos into consecutive sentence-level videos. SSLS is an important and meaningful task, which can greatly reduce the labor costs in data annotation for sentence-level SL datasets. However, SSLS is a very challenging task, since it is rather dicult to accurately nd the boundary of each sentence in a long video. To address this issue, we formalize, learn, and optimize the boundaries of sentences step by step. First, to distinguish the boundary and the inside of a sentence, we formalize SSLS as a frame-level classication task and design a boundary annotation scheme. Second, to learn the boundary of each sentence from the long video, we design a multimodal framework, SignBD, which correlates the local features and global features through dual dilated attention, while aligning visual and textual (i.e., sentences) modalities through gated cross-attention. Third, to alleviate the widely existed over-segmentation and under-segmentation problems in segmentation tasks, we propose a boundary optimization strategy, which utilizes the number of sentences provided by captions to optimize (i.e., insert or delete) boundaries based on information uncertainty. Extensive experimental results demonstrate the superiority of our solution. Codes are publicly available at: https://github.com/newbg/Sign-Language-Segmentation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper18
- Visual Alignment Constraint for Continuous Sign Language RecognitionYuecong Min, Aiming Hao, Xiujuan Chai, Xilin ChenICCV 2021 · 被引用 211 次
- Gloss-free Sign Language Translation: Improving from Visual-Language PretrainingBenjia Zhou, Zhigang Chen, Albert Clapés, Jun Wan 等ICCV 2023 · 被引用 123 次
- Weakly Supervised Energy-Based Learning for Action SegmentationJun Li, Peng Lei, Sinisa TodorovicICCV 2019 · 被引用 109 次
- SKiT: a Fast Key Information Video Transformer for Online Surgical Phase RecognitionYang Liu, Jiayu Huo, Jingjing Peng, Rachel Sparks 等ICCV 2023 · 被引用 57 次
- Improving Gloss-free Sign Language Translation by Reducing Representation DensityJinhui Ye, Xing Wang, Wenxiang Jiao, Junwei Liang 等NeurIPS 2024 · 被引用 49 次
相关 Paper
- Gloss Attention for Gloss-free Sign Language TranslationAoxiong Yin, Tianyun Zhong, Li Tang, Weike Jin 等CVPR 2023
- Conditional Variational Autoencoder for Sign Language Translation with Cross-Modal AlignmentRui Zhao, Liang Zhang, Biao Fu, Cong Hu 等AAAI 2024 · 被引用 36 次
- TSPNet: Hierarchical Feature Learning via Temporal Semantic Pyramid for Sign Language TranslationDongxu Li, Chenchen Xu, Xin Yu, Kaihao Zhang 等NeurIPS 2020 · 被引用 171 次
- Sign Language Video Retrieval with Free-Form Textual QueriesAmanda Cardoso Duarte, Samuel Albanie, Xavier Giró-i-Nieto, Gül VarolCVPR 2022 · 被引用 27 次
- BoostSLT: Boosting Sign Language Translation via a Plug-and-Play Diffusion-Based Semantic EnhancerChangzhou Han, Wanlun Ma, Xi Tang, Kun Hu 等CVPR 2026
