Hawkeye: Discovering and Grounding Implicit Anomalous Sentiment in Recon-videos via Scene-enhanced Video Large Language Model
Jianing Zhao, Jingjing Wang, Yujie Jin, Jiamin Luo, Guodong Zhou
摘要
In real-world recon-videos such as surveillance and drone reconnaissance videos, commonly used explicit language, acoustic and facial expressions information is often missing. However, these videos are always rich in anomalous sentiments (e.g., criminal tendencies), which urgently requires the implicit scene information (e.g., actions and object relations) to fast and precisely identify these anomalous sentiments. Motivated by this, this paper proposes a new chat-paradigm Implicit anomalous sentiment Discovering and grounding (IasDig) task, aiming to interactively, fast discovering and grounding anomalous sentiments in recon-videos via leveraging the implicit scene information (i.e., actions and object relations). Furthermore, this paper believes that this IasDig task faces two key challenges, i.e., scene modeling and scene balancing. To this end, this paper proposes a new Scene-enhanced Video Large Language Model named Hawkeye, i.e., acting like a raptor (e.g., a Hawk) to discover and locate prey, for the IasDig task. Specifically, this approach designs a graph-structured scene modeling module and a balanced heterogeneous MoE module to address the above two challenges, respectively. Extensive experimental results on our constructed scene-sparsity and scene-density IasDig datasets demonstrate the great advantage of Hawkeye to IasDig over the advanced Video-LLM baselines, especially on the metric of false negative rates. This justifies the importance of the scene information for identifying implicit anomalous sentiments and the impressive practicality of Hawkeye for real-world applications.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper3
- Rethinking Occlusion in FER: A Semantic-Aware Perspective and Go BeyondHuiyu Zhai, Xingxing Yang, Yalan Ye, Chenyang Li 等ACM MM 2025 · 被引用 5 次
- Omni-SILA: Towards Omni-scene Driven Visual Sentiment Identifying, Locating and Attributing in VideosJiamin Luo, Jingjing Wang, Junxiao Ma, Yujie Jin 等WWW 2025 · 被引用 2 次
- DeepSVU: Towards In-depth Security-oriented Video Understanding via Unified Physical-world Regularized MoEYujie Jin, Wenxin Zhang, Jingjing Wang, Guodong ZhouWWW 2026
相关 Paper
- Sherlock: Towards Multi-scene Video Abnormal Event Extraction and Localization via a Global-local Spatial-sensitive LLMJunxiao Ma, Jingjing Wang, Jiamin Luo, Peiying Yu 等WWW 2025 · 被引用 11 次
- HAWK: Learning to Understand Open-World Video AnomaliesJiaqi Tang, Hao Lu, Ruizheng Wu, Xiaogang Xu 等NeurIPS 2024 · 被引用 71 次
- HoloTrace: LLM-based Bidirectional Causal Knowledge Graph for Edge-Cloud Video Anomaly DetectionHanling Wang, Qing Li, Li Chen, Haidong Kang 等ACM MM 2025 · 被引用 2 次
- Skynet-V1: Towards Early Warning of Video Abnormal Events via A Spatial-temporal Causal-enhanced MoE FrameworkJunxiao Ma, Jingjing Wang, Min Zhang, Guodong ZhouACM MM 2025
- Do LVLMs Truly Understand Video Anomalies? Revealing Hallucination via Co-Occurrence PatternsMenghao Zhang, Huazheng Wang, Pengfei Ren, Kangheng Lin 等NeurIPS 2025 · 被引用 4 次
