Lune

NeurIPS2025Top-tier venue

Aha! - Predicting What Matters Next: Online Highlight Detection Without Looking Ahead

Aiden Chang, Celso de Melo, Stephanie M. Lukin

2025Year
3Citations

Abstract

Real-time understanding of continuous video streams is essential for intelligent agents operating in high-stakes environments, including autonomous vehicles, surveillance drones, and disaster response robots. Yet, most existing video understanding and highlight detection methods assume access to the entire video during inference, making them unsuitable for online or streaming scenarios. In particular, current models optimize for offline summarization, failing to support step-by-step reasoning needed for real-time decision-making. We introduce AHA, an autoregressive highlight detection framework that predicts the relevance of each video frame against a task described in natural language. Without accessing future video frames, AHA utilizes a multimodal vision-language model and lightweight, decoupled heads trained on a large, curated dataset of human-centric video labels. To enable scalability, we introduce the Dynamic SinkCache mechanism that achieves constant memory usage across infinite-length streams without degrading performance on standard benchmarks. This encourages the hidden representation to capture high-level task objectives, enabling effective frame-level rankings for informativeness, relevance, and uncertainty with respect to the natural language task. AHA achieves state-of-the-art (SOTA) performance on highlight detection benchmarks, surpassing even prior offline, full-context approaches and video-language models by +5.9% on TVSum and +8.3% on Mr.Hisum in mAP (mean Average Precision). We explore AHA's potential for real-world robotics applications given a task-oriented natural language input and a continuous, robot-centric video. Both experiments demonstrate AHA's potential effectiveness as a real-time reasoning module for downstream planning and long-horizon understanding. * Conducted research as a fellow at the DEVCOM Army Research Laboratory. 1 github.com/aiden200/Aha-39th Conference on Neural Information Processing Systems (NeurIPS 2025).

natural language queries operate offline assume the entire video is available during inference [9,10]. This fundamental reliance on full-context renders existing HD methods unsuitable for streaming applications requiring step-by-step reasoning and immediate action based on unfolding events. Concurrently, a separate area of research has explored streaming video analysis, often leveraging Large Language Models (Video-LLMs) for tasks like dense video captioning or generating dialogue responses about ongoing events [11,12]. While some of these models have explored HD as an auxiliary capability, their application to OHD faces significant limitations. These Video-LLMs often necessitate modifications to standard HD benchmarks, employ post-hoc smoothing techniques that violate strict online constraints by implicitly using future information, and ultimately yield suboptimal HD performance [13]. This leaves a critical gap: a robust method designed specifically for accurate, online, task-conditioned highlight detection on standard benchmarks.

We address this gap by introducing a novel framework built for OHD. We define OHD as the method of analyzing a streaming video by observing frames strictly one at a time and, for each current frame, predicting its highlight score using only past and present information, without accessing any future frames. This sequential, causal processing is fundamental for enabling real-time decision-making in dynamic environments. Given a natural language task description, our model, AHA, performs OHD by employing a lightweight, autoregressive scoring mechanism focused directly on highlight detection. This allows AHA to operate effectively on traditional HD benchmarks in a truly online fashion, without requiring benchmark modifications or non-causal smoothing, and achieving SOTA performance even in zero-shot settings. Our main contributions are: AHA Framework for Efficient OHD: We propose AHA, an autoregressive framework featuring lightweight prediction heads (scoring relevance, informativeness, uncertainty) and our novel Dynamic SinkCache memory for efficient, constant-cost, OHD under natural language conditioning, and a video quality dropout mechanism to enhance robustness against real-world noise.

A Large-Scale Dataset for OHD 2 : We construct and release the Human Intuition Highlight Dataset (HIHD), a novel dataset of 23k videos incorporating user engagement signals and task-driven captions, specifically designed to train and benchmark task-conditioned OHD models.

SOTA OHD: AHA surpasses prior methods, including offline approaches, on the HD benchmarks TVSum [14] (+5.9% mAP) and Mr.Hisum [15] (+8.3% mAP). We validate AHA's robustness and real-world applicability through comprehensive experiments, ablations, and on a challenging longhorizon, noisy robotics video from SCOUT [16], demonstrating task-relevant understanding where offline processing is infeasible.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 684fc903-d30d-4afb-9cd2-5cca01ff5706

Builds on22

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines