Lune

CVPR2026Top-tier venue

SOUPLE: Enhancing Audio-Visual Localization and Segmentation with Learnable Prompt Contexts

Khanh Binh Nguyen, Chae Jung Park

2026Year

Abstract

Large-scale pre-trained image-text models exhibit robust multimodal representation, yet applying contrastive language-image pretraining (CLIP) to audio-visual localization remains challenging. Replacing the classification token ([CLS][CLS]) with an audio-embedded token ([VA][V_A])struggles to capture semantic cues, and the prompt “a photo of a [VA][V_A]” fails to establish meaningful connections between audio embeddings and context tokens. To address these issues, we propose sound-aware prompt learning (SouPLe), which replaces fixed prompts with learnable context tokens. These tokens incorporate visual features to generate conditional context for a mask decoder, effectively bridging semantic correspondence between audio and visual inputs. Experiments on VGGSound, SoundNet, and AVSBench confirm that SouPLe significantly improves localization and segmentation performance.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext ccc222d2-ec30-4a0a-a2ea-ff9a05c91e13

Builds on18

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines