Segment beyond View: Handling Partially Missing Modality for Audio-Visual Semantic Segmentation
Renjie Wu, Hu Wang, Feras Dayoub, Hsiang-Ting Chen
Abstract
Augmented Reality (AR) devices, emerging as prominent mobile interaction platforms, face challenges in user safety, particularly concerning oncoming vehicles. While some solutions leverage onboard camera arrays, these cameras often have limited field-of-view (FoV) with front or downward perspectives. Addressing this, we propose a new out-of-view semantic segmentation task and Segment Beyond View (SBV), a novel audio-visual semantic segmentation method. SBV supplements the visual modality, which miss the information beyond FoV, with the auditory information using a teacher-student distillation model (Omni2Ego). The model consists of a vision teacher utilising panoramic information, an auditory teacher with 8-channel audio, and an audio-visual student that takes views with limited FoV and binaural audio as input and produce semantic segmentation for objects outside FoV. SBV outperforms existing models in comparative evaluations and shows a consistent performance across varying FoV ranges and in monaural audio settings.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on15
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar et al.NeurIPS 2021 · 9,661 citations
- MDETR - Modulated Detection for End-to-End Multi-Modal UnderstandingAishwarya Kamath, Mannat Singh, Yann LeCun, Gabriel Synnaeve et al.ICCV 2021 · 1,114 citations
- SMIL: Multimodal Learning with Severely Missing ModalityMengmeng Ma, Jian Ren, Long Zhao, Sergey Tulyakov et al.AAAI 2021 · 393 citations
- Dual Attention Matching for Audio-Visual Event LocalizationYu Wu, Linchao Zhu, Yan Yan, Yi YangICCV 2019 · 233 citations
- Self-Supervised Moving Vehicle Tracking With Stereo SoundChuang Gan, Hang Zhao, Peihao Chen, David D. Cox et al.ICCV 2019 · 157 citations
Related papers
- GeoDistill: Geometry-Guided Self-Distillation for Weakly Supervised Cross-View LocalizationShaowen Tong, Zimin Xia, Alexandre Alahi, Xuming He et al.ICCV 2025 · 3 citations
- Dense 2D-3D Indoor Prediction with Sound via Aligned Cross-Modal DistillationHeeseung Yun, Joonil Na, Gunhee KimICCV 2023 · 8 citations
- OmniSAM: Omnidirectional Segment Anything Model for UDA in Panoramic Semantic SegmentationDing Zhong, Xu Zheng, Chenfei Liao, Yuanhuiyi Lyu et al.ICCV 2025 · 4 citations
- Entropy-Monitored Kernelized Token Distillation for Audio-Visual CompressionHyoungseob Park, Lipeng Ke, Pritish Mohapatra, Huajun Ying et al.ICLR 2026
- Open-Vocabulary Audio-Visual Semantic SegmentationRuohao Guo, Liao Qu, Dantong Niu, Yanyu Qi et al.ACM MM 2024 · 4 citations
