Towards Noise-Tolerant Speech-Referring Video Object Segmentation: Bridging Speech and Text
Xiang Li, Jinglu Wang, Xiaohao Xu, Muqiao Yang, Fan Yang, Yizhou Zhao, Rita Singh, Bhiksha Raj
Abstract
Linguistic communication is prevalent in Human-Computer Interaction (HCI). Speech (spoken language) serves as a convenient yet potentially ambiguous form due to noise and accents, exposing a gap compared to text. In this study, we investigate the prominent HCI task, Referring Video Object Segmentation (R-VOS), which aims to segment and track objects using linguistic references. While text input is well-investigated, speech input is under-explored. Our objective is to bridge the gap between speech and text, enabling the adaptation of existing text-input R-VOS models to accommodate noisy speech input effectively. Specifically, we propose a method to align the semantic spaces between speech and text by incorporating two key modules: 1) Noise-Aware Semantic Adjustment (NSA) for clear semantics extraction from noisy speech; and 2) Semantic Jitter Suppression (SJS) enabling R-VOS models to tolerate noisy queries. Comprehensive experiments conducted on the challenging AVOS benchmarks reveal that our proposed method outperforms state-of-the-art approaches.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 35b807c4-b7b5-4d52-bdf4-7be0c6e9323dCited by top-tier papers4
- SPARTAN: Data-Adaptive Symbolic Time-Series ApproximationFan Yang, John PaparrizosSIGMOD 2025 · 11 citations
- Completing Visual Objects via Bridging Generation and SegmentationXiang Li, Yinpeng Chen, Chung-Ching Lin, Hao Chen et al.ICML 2024 · 3 citations
- Matting Anything 2: Towards Video Matting for AnythingChenyi Zhang, Yiheng Lin, Yunchao Wei, Hongsong Wang et al.ICLR 2026
- QDFormer: Towards Robust Audiovisual Segmentation in Complex Environments with Quantization-based Semantic DecompositionXiang Li, Jinglu Wang, Xiaohao Xu, Xiulian Peng et al.CVPR 2024
Builds on26
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 9,451 citations
- YOLACT: Real-Time Instance SegmentationDaniel Bolya, Chong Zhou, Fanyi Xiao, Yong Jae LeeICCV 2019 · 2,075 citations
- Video Object Segmentation Using Space-Time Memory NetworksSeoung Wug Oh, Joon-Young Lee, Ning Xu, Seon Joo KimICCV 2019 · 845 citations
- Dice Loss for Data-imbalanced NLP TasksXiaoya Li, Xiaofei Sun, Yuxian Meng, Junjun Liang et al.ACL 2020 · 575 citations
Related papers
- Robust Referring Video Object Segmentation with Cyclic Structural ConsensusXiang Li, Jinglu Wang, Xiaohao Xu, Xiao Li et al.ICCV 2023 · 65 citations
- Tracking-forced Referring Video Object SegmentationRuxue Yan, Wenya Guo, Xubo Liu, Xumeng Liu et al.ACM MM 2024 · 3 citations
- SOC: Semantic-Assisted Object Cluster for Referring Video Object SegmentationZhuoyan Luo, Yicheng Xiao, Yong Liu, Shuyan Li et al.NeurIPS 2023 · 89 citations
- HTML: Hybrid Temporal-scale Multimodal Learning Framework for Referring Video Object SegmentationMingfei Han, Yali Wang, Zhihui Li, Lina Yao et al.ICCV 2023 · 42 citations
- Language as Queries for Referring Video Object SegmentationJiannan Wu, Yi Jiang, Peize Sun, Zehuan Yuan et al.CVPR 2022 · 143 citations
