Towards Noise-Tolerant Speech-Referring Video Object Segmentation: Bridging Speech and Text
Xiang Li, Jinglu Wang, Xiaohao Xu, Muqiao Yang, Fan Yang, Yizhou Zhao, Rita Singh, Bhiksha Raj
摘要
Linguistic communication is prevalent in Human-Computer Interaction (HCI). Speech (spoken language) serves as a convenient yet potentially ambiguous form due to noise and accents, exposing a gap compared to text. In this study, we investigate the prominent HCI task, Referring Video Object Segmentation (R-VOS), which aims to segment and track objects using linguistic references. While text input is well-investigated, speech input is under-explored. Our objective is to bridge the gap between speech and text, enabling the adaptation of existing text-input R-VOS models to accommodate noisy speech input effectively. Specifically, we propose a method to align the semantic spaces between speech and text by incorporating two key modules: 1) Noise-Aware Semantic Adjustment (NSA) for clear semantics extraction from noisy speech; and 2) Semantic Jitter Suppression (SJS) enabling R-VOS models to tolerate noisy queries. Comprehensive experiments conducted on the challenging AVOS benchmarks reveal that our proposed method outperforms state-of-the-art approaches.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- SPARTAN: Data-Adaptive Symbolic Time-Series ApproximationFan Yang, John PaparrizosSIGMOD 2025 · 被引用 11 次
- Completing Visual Objects via Bridging Generation and SegmentationXiang Li, Yinpeng Chen, Chung-Ching Lin, Hao Chen 等ICML 2024 · 被引用 3 次
- Matting Anything 2: Towards Video Matting for AnythingChenyi Zhang, Yiheng Lin, Yunchao Wei, Hongsong Wang 等ICLR 2026
- QDFormer: Towards Robust Audiovisual Segmentation in Complex Environments with Quantization-based Semantic DecompositionXiang Li, Jinglu Wang, Xiaohao Xu, Xiulian Peng 等CVPR 2024
它引用的顶会 Paper26
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 被引用 9,451 次
- YOLACT: Real-Time Instance SegmentationDaniel Bolya, Chong Zhou, Fanyi Xiao, Yong Jae LeeICCV 2019 · 被引用 2,075 次
- Video Object Segmentation Using Space-Time Memory NetworksSeoung Wug Oh, Joon-Young Lee, Ning Xu, Seon Joo KimICCV 2019 · 被引用 845 次
- Dice Loss for Data-imbalanced NLP TasksXiaoya Li, Xiaofei Sun, Yuxian Meng, Junjun Liang 等ACL 2020 · 被引用 575 次
相关 Paper
- Robust Referring Video Object Segmentation with Cyclic Structural ConsensusXiang Li, Jinglu Wang, Xiaohao Xu, Xiao Li 等ICCV 2023 · 被引用 65 次
- Tracking-forced Referring Video Object SegmentationRuxue Yan, Wenya Guo, Xubo Liu, Xumeng Liu 等ACM MM 2024 · 被引用 3 次
- SOC: Semantic-Assisted Object Cluster for Referring Video Object SegmentationZhuoyan Luo, Yicheng Xiao, Yong Liu, Shuyan Li 等NeurIPS 2023 · 被引用 89 次
- HTML: Hybrid Temporal-scale Multimodal Learning Framework for Referring Video Object SegmentationMingfei Han, Yali Wang, Zhihui Li, Lina Yao 等ICCV 2023 · 被引用 42 次
- Language as Queries for Referring Video Object SegmentationJiannan Wu, Yi Jiang, Peize Sun, Zehuan Yuan 等CVPR 2022 · 被引用 143 次
