Lune

NeurIPS2025顶会

Aha! - Predicting What Matters Next: Online Highlight Detection Without Looking Ahead

Aiden Chang, Celso de Melo, Stephanie M. Lukin

2025年份
3被引次数

摘要

Real-time understanding of continuous video streams is essential for intelligent agents operating in high-stakes environments, including autonomous vehicles, surveillance drones, and disaster response robots. Yet, most existing video understanding and highlight detection methods assume access to the entire video during inference, making them unsuitable for online or streaming scenarios. In particular, current models optimize for offline summarization, failing to support step-by-step reasoning needed for real-time decision-making. We introduce AHA, an autoregressive highlight detection framework that predicts the relevance of each video frame against a task described in natural language. Without accessing future video frames, AHA utilizes a multimodal vision-language model and lightweight, decoupled heads trained on a large, curated dataset of human-centric video labels. To enable scalability, we introduce the Dynamic SinkCache mechanism that achieves constant memory usage across infinite-length streams without degrading performance on standard benchmarks. This encourages the hidden representation to capture high-level task objectives, enabling effective frame-level rankings for informativeness, relevance, and uncertainty with respect to the natural language task. AHA achieves state-of-the-art (SOTA) performance on highlight detection benchmarks, surpassing even prior offline, full-context approaches and video-language models by +5.9% on TVSum and +8.3% on Mr.Hisum in mAP (mean Average Precision). We explore AHA's potential for real-world robotics applications given a task-oriented natural language input and a continuous, robot-centric video. Both experiments demonstrate AHA's potential effectiveness as a real-time reasoning module for downstream planning and long-horizon understanding. * Conducted research as a fellow at the DEVCOM Army Research Laboratory. 1 github.com/aiden200/Aha-39th Conference on Neural Information Processing Systems (NeurIPS 2025).

natural language queries operate offline assume the entire video is available during inference [9,10]. This fundamental reliance on full-context renders existing HD methods unsuitable for streaming applications requiring step-by-step reasoning and immediate action based on unfolding events. Concurrently, a separate area of research has explored streaming video analysis, often leveraging Large Language Models (Video-LLMs) for tasks like dense video captioning or generating dialogue responses about ongoing events [11,12]. While some of these models have explored HD as an auxiliary capability, their application to OHD faces significant limitations. These Video-LLMs often necessitate modifications to standard HD benchmarks, employ post-hoc smoothing techniques that violate strict online constraints by implicitly using future information, and ultimately yield suboptimal HD performance [13]. This leaves a critical gap: a robust method designed specifically for accurate, online, task-conditioned highlight detection on standard benchmarks.

We address this gap by introducing a novel framework built for OHD. We define OHD as the method of analyzing a streaming video by observing frames strictly one at a time and, for each current frame, predicting its highlight score using only past and present information, without accessing any future frames. This sequential, causal processing is fundamental for enabling real-time decision-making in dynamic environments. Given a natural language task description, our model, AHA, performs OHD by employing a lightweight, autoregressive scoring mechanism focused directly on highlight detection. This allows AHA to operate effectively on traditional HD benchmarks in a truly online fashion, without requiring benchmark modifications or non-causal smoothing, and achieving SOTA performance even in zero-shot settings. Our main contributions are: AHA Framework for Efficient OHD: We propose AHA, an autoregressive framework featuring lightweight prediction heads (scoring relevance, informativeness, uncertainty) and our novel Dynamic SinkCache memory for efficient, constant-cost, OHD under natural language conditioning, and a video quality dropout mechanism to enhance robustness against real-world noise.

A Large-Scale Dataset for OHD 2 : We construct and release the Human Intuition Highlight Dataset (HIHD), a novel dataset of 23k videos incorporating user engagement signals and task-driven captions, specifically designed to train and benchmark task-conditioned OHD models.

SOTA OHD: AHA surpasses prior methods, including offline approaches, on the HD benchmarks TVSum [14] (+5.9% mAP) and Mr.Hisum [15] (+8.3% mAP). We validate AHA's robustness and real-world applicability through comprehensive experiments, ablations, and on a challenging longhorizon, noisy robotics video from SCOUT [16], demonstrating task-relevant understanding where offline processing is infeasible.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

它引用的顶会 Paper22

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖