MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks
Sanjoy Chowdhury, Mohamed Elmoghany, Yohan Abeysinghe, Junjie Fei, Sayan Nag, Salman H. Khan, Mohamed Elhoseiny, Dinesh Manocha
摘要
Large multimodal models (LMMs) have shown remarkable progress in audio-visual understanding, yet they struggle with real-world scenarios that require complex reasoning across extensive video collections. Existing benchmarks for video question answering remain limited in scope, typically involving one clip per query, which falls short of representing the challenges of large-scale, audio-visual retrieval and reasoning encountered in practical applications. To bridge this gap, we introduce a novel task named AV-HaystacksQA, where the goal is to identify salient segments across different videos in response to a query and link them together to generate the most informative answer. To this end, we present AVHaystacks, an audio-visual benchmark comprising 3100 annotated QA pairs designed to assess the capabilities of LMMs in multi-video retrieval and temporal grounding task. Additionally, we propose a model-agnostic, multi-agent framework MAGNET to address this challenge, achieving up to 89% and 65% relative improvements over baseline methods on BLEU@4 and GPT evaluation scores in QA task on our proposed AVHaystacks. To enable robust evaluation of multi-video retrieval and temporal grounding for optimal response generation, we introduce two new metrics, STEM, which captures alignment errors between a ground truth and a predicted step sequence and MTGS, to facilitate balanced and interpretable evaluation of segment-level grounding performance. Project: https://schowdhury671.github.io/magnet_project/
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- EgoSound: Benchmarking Sound Understanding in Egocentric VideosBingwen Zhu, Yuqian Fu, Qiaole Dong, Guolei Sun 等CVPR 2026 · 被引用 7 次
- AMusE: Audio-Visual Benchmark and Alignment Framework for Agentic Multi-Speaker UnderstandingSanjoy Chowdhury, Karren Dai Yang, Xudong Liu, Fartash Faghri 等CVPR 2026 · 被引用 5 次
- Agentic Design Review SystemSayan Nag, K. J. Joseph, Koustava Goswami, Vlad I. Morariu 等AAAI 2026 · 被引用 1 次
- EgoAdapt: Adaptive Multisensory Distillation and Policy Learning for Efficient Egocentric PerceptionSanjoy Chowdhury, Subrata Biswas, Sayan Nag, Tushar Nagarajan 等ICCV 2025
- CFVBench: A Comprehensive Video Benchmark for Fine-grained Multimodal Retrieval-Augmented GenerationKaiwen Wei, Xiao Liu, Jie Zhang, Zijian Wang 等WWW 2026
它引用的顶会 Paper55
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni 等NeurIPS 2020 · 被引用 19,162 次
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li 等ICLR 2024 · 被引用 3,079 次
- Self-RAG: Learning to Retrieve, Generate, and Critique through Self-ReflectionAkari Asai, Zeqiu Wu, Yizhong Wang, Avirup Sil 等ICLR 2024 · 被引用 1,798 次
- MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual ContextsPan Lu, Hritik Bansal, Tony Xia, Jiacheng Liu 等ICLR 2024 · 被引用 1,472 次
- High-Fidelity Audio Compression with Improved RVQGANRithesh Kumar, Prem Seetharaman, Alejandro Luebs, Ishaan Kumar 等NeurIPS 2023 · 被引用 910 次
相关 Paper
- Visual Haystacks: A Vision-Centric Needle-In-A-Haystack BenchmarkTsung-Han Wu, Giscard Biamby, Jerome Quenum, Ritwik Gupta 等ICLR 2025
- Grounded Multi-Hop VideoQA in Long-Form Egocentric VideosQirui Chen, Shangzhe Di, Weidi XieAAAI 2025 · 被引用 35 次
- Needle In A Video Haystack: A Scalable Synthetic Evaluator for Video MLLMsZijia Zhao, Haoyu Lu, Yuqi Huo, Yifan Du 等ICLR 2025
- Video SimpleQA: Towards Factuality Evaluation in Large Video Language ModelsMeng Cao, Pengfei Hu, Yingyao Wang, Jihao Gu 等AAAI 2026 · 被引用 1 次
- Learning to Answer Questions in Dynamic Audio-Visual ScenariosGuangyao Li, Yake Wei, Yapeng Tian, Chenliang Xu 等CVPR 2022 · 被引用 101 次
