SEAL: Semantic Attention Learning for Long Video Representation
Lan Wang, Yujia Chen, Du Tran, Vishnu Naresh Boddeti, Wen-Sheng Chu
摘要
Long video understanding presents challenges due to the inherent high computational complexity and redundant temporal information. An effective representation for long videos must efficiently process such redundancy while preserving essential contents for downstream tasks. This paper introduces SEmantic Attention Learning (SEAL), a novel unified representation for long videos. To reduce computational complexity, long videos are decomposed into three distinct types of semantic entities: scenes, objects, and actions, allowing models to operate on a compact set of entities rather than a large number of frames or pixels. To further address redundancy, we propose an attention learning module that balances token relevance with diversity, formulated as a subset selection optimization problem. Our representation is versatile and applicable across various long video understanding tasks. Extensive experiments demonstrate that SEAL significantly outperforms state-of-the-art methods in video question answering and temporal grounding tasks across diverse benchmarks, including LVBench, MovieChat-1K, and Ego4D.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Failures to Surface Harmful Contents in Video Large Language ModelsYuxin Cao, Wei Song, Derui Wang, Jingling Xue 等AAAI 2026 · 被引用 3 次
- SAVA-X: Ego-to-Exo Imitation Error Detection via Scene-Adaptive View Alignment and Bidirectional Cross View FusionXiang Li, Heqian Qiu, Lanxiao Wang, Benliu Qiu 等CVPR 2026 · 被引用 1 次
- Craw: A Unified and Efficient Querying Framework for Large-Scale Video DatasetsZiqi Zhou, Hanjian Jiang, Zihao Zeng, Xupuzhe Shao 等VLDB 2026
它引用的顶会 Paper17
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- Egocentric Video-Language PretrainingKevin Qinghong Lin, Jinpeng Wang, Mattia Soldan, Michael Wray 等NeurIPS 2022 · 被引用 306 次
- Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language ModelsMuhammad Maaz, Hanoona Abdul Rasheed, Salman Khan, Fahad KhanACL 2024 · 被引用 279 次
- Streaming Long Video Understanding with Large Language ModelsRui Qian, Xiaoyi Dong, Pan Zhang, Yuhang Zang 等NeurIPS 2024 · 被引用 216 次
相关 Paper
- Flow4Agent: Long-form Video Understanding via Motion Prior from Optical FlowRuyang Liu, Shangkun Sun, Haoran Tang, Wei Gao 等ICCV 2025 · 被引用 15 次
- From Content to Knowledge: Lightning Fast Long-Video Understanding with Neural Knowledge RepresentationsYuchen Guan, Xiao Li, Zongyu Guo, Xiaoyi Zhang 等ICML 2026
- Compositional Video Understanding with Spatiotemporal Structure-based TransformersHoyeoung Yun, Jinwoo Ahn, Minseo Kim, Eun-Sol KimCVPR 2024 · 被引用 4 次
- Keyframe-Oriented Vision Token Pruning: Enhancing Efficiency of Large Vision Language Models on Long-form Video ProcessingYudong Liu, Jingwei Sun, Yueqian Lin, Jianyi Zhang 等ICCV 2025 · 被引用 22 次
- Seeing the Scene Matters: Revealing Forgetting in Video Understanding Models with a Scene-Aware Long-Video BenchmarkSeng Nam Chen, Hao Chen, Chenglam Ho, Xinyu Mao 等CVPR 2026
