Dense 2D-3D Indoor Prediction with Sound via Aligned Cross-Modal Distillation
Heeseung Yun, Joonil Na, Gunhee Kim
摘要
Sound can convey significant information for spatial reasoning in our daily lives. To endow deep networks with such ability, we address the challenge of dense indoor prediction with sound in both 2D and 3D via cross-modal knowledge distillation. In this work, we propose a Spatial Alignment via Matching (SAM) distillation framework that elicits local correspondence between the two modalities in vision-to-audio knowledge transfer. SAM integrates audio features with visually coherent learnable spatial embeddings to resolve inconsistencies in multiple layers of a student model. Our approach does not rely on a specific input representation, allowing for flexibility in the input shapes or dimensions without performance degradation. With a newly curated benchmark named Dense Auditory Prediction of Surroundings (DAPS), we are the first to tackle dense indoor prediction of omnidirectional surroundings in both 2D and 3D with audio observations. Specifically, for audio-based depth estimation, semantic segmentation, and challenging 3D scene reconstruction, the proposed distillation framework consistently achieves state-of-the-art performance across various metrics and backbone architectures.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- FIND: Few-Shot Anomaly Inspection with Normal-Only Multi-Modal DataYiting Li, Fayao Liu, Jingyi Liao, Sichao Tian 等ICCV 2025 · 被引用 5 次
- Cross-Modal Knowledge Distillation without Paired Data: Theoretical Foundation and AlgorithmT. K Tran, Duc Chu Anh, Quang Hung Pham, Phi Le Nguyen 等ICML 2026
- C2KD: Bridging the Modality Gap for Cross-Modal Knowledge DistillationFushuo Huo, Wenchao Xu, Jingcai Guo, Haozhao Wang 等CVPR 2024
它引用的顶会 Paper11
- Vision Transformers for Dense PredictionRené Ranftl, Alexey Bochkovskiy, Vladlen KoltunICCV 2021 · 被引用 2,647 次
- Habitat: A Platform for Embodied AI ResearchManolis Savva, Jitendra Malik, Devi Parikh, Dhruv Batra 等ICCV 2019 · 被引用 1,863 次
- Self-Supervised Moving Vehicle Tracking With Stereo SoundChuang Gan, Hang Zhao, Peihao Chen, David D. Cox 等ICCV 2019 · 被引用 157 次
- Hearing Lips: Improving Lip Reading by Distilling Speech RecognizersYa Zhao, Rui Xu, Xinchao Wang, Peng Hou 等AAAI 2020 · 被引用 106 次
- Image2Reverb: Cross-Modal Reverb Impulse Response SynthesisNikhil Singh, Jeff Mentch, Jerry Ng, Matthew Beveridge 等ICCV 2021 · 被引用 61 次
相关 Paper
- Sound Source Localization is All about Cross-Modal AlignmentArda Senocak, Hyeonggon Ryu, Junsik Kim, Tae-Hyun Oh 等ICCV 2023 · 被引用 39 次
- SGPFeat: Semantic and Geometric Priors for Multi-modal Image MatchingYuxin Deng, Botian Wang, Kaining Zhang, Hao Zhang 等AAAI 2026
- VLScene: Vision-Language Guidance Distillation for Camera-Based 3D Semantic Scene CompletionMeng Wang, Huilong Pi, Ruihui Li, Yunchuan Qin 等AAAI 2025 · 被引用 11 次
- Entropy-Monitored Kernelized Token Distillation for Audio-Visual CompressionHyoungseob Park, Lipeng Ke, Pritish Mohapatra, Huajun Ying 等ICLR 2026
- Omnidirectional Information Gathering for Knowledge Transfer-based Audio-Visual NavigationJinyu Chen, Wenguan Wang, Si Liu, Hongsheng Li 等ICCV 2023 · 被引用 21 次
