Learnable Expansion of Graph Operators for Multi-Modal Feature Fusion
Dexuan Ding, Lei Wang, Liyun Zhu, Tom Gedeon, Piotr Koniusz
摘要
In computer vision tasks, features often come from diverse representations, domains (e.g., indoor and outdoor), and modalities (e.g., text, images, and videos). Effectively fusing these features is essential for robust performance, especially with the availability of powerful pre-trained models like vision-language models. However, common fusion methods, such as concatenation, element-wise operations, and non-linear techniques, often fail to capture structural relationships, deep feature interactions, and suffer from inefficiency or misalignment of features across domains or modalities. In this paper, we shift from high-dimensional feature space to a lower-dimensional, interpretable graph space by constructing relationship graphs that encode feature relationships at different levels, e.g., clip, frame, patch, token, etc. To capture deeper interactions, we expand graphs through iterative graph relationship updates and introduce a learnable graph fusion operator to integrate these expanded relationships for more effective fusion. Our approach is relationship-centric, operates in a homogeneous space, and is mathematically principled, resembling element-wise relationship score aggregation via multilinear polynomials. We demonstrate the effectiveness of our graph-based fusion method on video anomaly detection, showing strong performance across multi-representational, multi-modal, and multi-domain feature fusion tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- A Unified Reasoning Framework for Holistic Zero-Shot Video Anomaly AnalysisDongheng Lin, Mengxue Qu, Kunyang Han, Jianbo Jiao 等NeurIPS 2025 · 被引用 14 次
- Graph Self-Supervised Learning with Learnable Structural and Positional EncodingsAsiri Wijesinghe, Hao Zhu, Piotr KoniuszWWW 2025 · 被引用 3 次
- Learning Time in Static ClassifiersXi Ding, Lei Wang, Piotr Koniusz, Yongsheng GaoAAAI 2026 · 被引用 2 次
- Subspace Kernel Learning on Tensor SequencesLei Wang, Xi Ding, Yongsheng Gao, Piotr KoniuszICLR 2026 · 被引用 1 次
- Privacy-Aware Video Anomaly Detection: Guided Orthogonal Projection and a Comprehensive Evaluation FrameworkWenxiang Diao, Lei Wang, Andrew Busch, Jun Zhou 等ICML 2026
它引用的顶会 Paper11
- SimCSE: Simple Contrastive Learning of Sentence EmbeddingsTianyu Gao, Xingcheng Yao, Danqi ChenEMNLP 2021 · 被引用 2,496 次
- Video Swin TransformerZe Liu, Jia Ning, Yue Cao, Yixuan Wei 等CVPR 2022 · 被引用 1,847 次
- VaTeX: A Large-Scale, High-Quality Multilingual Dataset for Video-and-Language ResearchXin Wang, Jiawei Wu, Jun-Kun Chen, Lei Li 等ICCV 2019 · 被引用 688 次
- Weakly-supervised Video Anomaly Detection with Robust Temporal Feature Magnitude LearningYu Tian, Guansong Pang, Yuanhong Chen, Rajvinder Singh 等ICCV 2021 · 被引用 495 次
- SwinBERT: End-to-End Transformers with Sparse Attention for Video CaptioningKevin Lin, Linjie Li, Chung-Ching Lin, Faisal Ahmed 等CVPR 2022 · 被引用 263 次
相关 Paper
- Aligning Effective Tokens with Video Anomaly in Large Language ModelsYingxian Chen, Jiahui Liu, Ruidi Fan, Yanwei Li 等ICCV 2025 · 被引用 2 次
- Learning to Represent Image and Text with Denotation GraphBowen Zhang, Hexiang Hu, Vihan Jain, Eugene Ie 等EMNLP 2020 · 被引用 22 次
- Graph-Based Video-Language Learning with Multi-Grained Audio-Visual AlignmentChenyang Lyu, Wenxi Li, Tianbo Ji, Longyue Wang 等ACM MM 2023 · 被引用 6 次
- Harnessing Large Language Models for Training-Free Video Anomaly DetectionLuca Zanella, Willi Menapace, Massimiliano Mancini, Yiming Wang 等CVPR 2024 · 被引用 57 次
- Alert-CLIP: Abnormality-aware Latent-Enhanced Representation Tuning of CLIP for Video Anomaly DetectionYiyan Zhu, Menghao Zhang, Haifeng Sun, Pengfei Ren 等CVPR 2026
