MARS-Sep: Multimodal-Aligned Reinforced Sound Separation
Zihan Zhang, Xize Cheng, Zhennan Jiang, Dongjie Fu, Jingyuan Chen, Zhou Zhao, Tao Jin
摘要
Universal sound separation faces a fundamental misalignment: models optimized for low-level signal metrics often produce semantically contaminated outputs, failing to suppress perceptually salient interference from acoustically similar sources. We introduce a preference alignment perspective, analogous to aligning LLMs with human intent. To address this, we introduce MARS-Sep, a reinforcement learning framework that reformulates separation as decision making. Instead of simply regressing ground-truth masks, MARS-Sep learns a factorized Beta mask policy that is steered by a preference reward model and optimized by a stable, clipped trust-region surrogate. The reward, derived from a progressivelyaligned audio-text-vision encoder, directly incentivizes semantic consistency with query prompts. Extensive experiments on multiple benchmarks demonstrate consistent gains in Text-, Audio-, and Image-Queried separation, with notable improvements in signal metrics and semantic quality. Our code is available at https://github.com/mars-sep/MARS-Sep . Sound separation samples are available at https://mars-sep.github.io/ .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper18
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language ModelsWenxuan Huang, Bohan Jia, Shaosheng Cao, Zheyu Ye 等ICLR 2026 · 被引用 670 次
- R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement LearningYifan Zhang, Xingyu Lu, Xiao Hu, Chaoyou Fu 等ICLR 2026 · 被引用 65 次
- Zero-Shot Audio Source Separation through Query-Based Learning from Weakly-Labeled DataKe Chen, Xingjian Du, Bilei Zhu, Zejun Ma 等AAAI 2022 · 被引用 58 次
相关 Paper
- Closing the Modality Reasoning Gap for Speech Large Language ModelsChaoren Wang, Heng Lu, Xueyao Zhang, Shujie Liu 等ACL 2026 · 被引用 10 次
- OpenSep: Leveraging Large Language Models with Textual Inversion for Open World Audio SeparationTanvir Mahmud, Diana MarculescuEMNLP 2024 · 被引用 1 次
- Optimal Transport for LLM Reward Modeling from Noisy FeedbackLicheng Pan, Haocheng Yang, Haoxuan Li, Yunsheng Lu 等ICML 2026
- Align-SLM: Textless Spoken Language Models with Reinforcement Learning from AI FeedbackGuan-Ting Lin, Prashanth Gurunath Shivakumar, Aditya Gourav, Yile Gu 等ACL 2025 · 被引用 33 次
- Move2Hear: Active Audio-Visual Source SeparationSagnik Majumder, Ziad Al-Halah, Kristen GraumanICCV 2021 · 被引用 48 次
