Lune

CVPR2026顶会

AVATAR: Reinforcement Learning to See, Hear, and Reason Over Video

Yogesh Kulkarni, Pooyan Fazli

2026年份
15被引次数
2顶会引用

摘要

Multimodal reasoning over long-horizon video is challenging due to the need for precise spatiotemporal fusion and alignment across modalities. While recent methods such as Group Relative Policy Optimization (GRPO) have shown promise in this domain, they suffer from three key limitations: (1) data inefficiency from their on-policy design, (2) a vanishing advantage problem, where identical or near-identical rewards within a group eliminate the learning signal by producing zero-valued advantages, and (3) uniform credit assignment that fails to emphasize critical reasoning steps.We introduce AVATAR\textbf{AVATAR} (A\textbf{A}udio-V\textbf{V}ideo A\textbf{A}gent\textbf{t} for A\textbf{A}lignment and R\textbf{R}easoning), a framework that addresses these limitations through two core components: (1) an off-policy training architecture that improves sample efficiency and resolves vanishing advantages by reusing past experiences with greater reward diversity, and (2) Temporal Advantage Shaping (TAS), a novel credit assignment strategy that upweights key reasoning phases during learning.AVATAR\textbf{AVATAR} achieves strong performance across various benchmarks, outperforming the Qwen2.5-Omni baseline by +5.4\textbf{+5.4} on MMVU, +4.9\textbf{+4.9} on OmniBench, and +4.5\textbf{+4.5} on Video-Holmes, while demonstrating 5 sample efficiency\textbf{5$$ sample efficiency}, requiring 8080% fewer generated completions to reach target performance.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper2

问问它们各自怎么用它

它引用的顶会 Paper39

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖