Learnable Expansion of Graph Operators for Multi-Modal Feature Fusion
Dexuan Ding, Lei Wang, Liyun Zhu, Tom Gedeon, Piotr Koniusz
Abstract
In computer vision tasks, features often come from diverse representations, domains (e.g., indoor and outdoor), and modalities (e.g., text, images, and videos). Effectively fusing these features is essential for robust performance, especially with the availability of powerful pre-trained models like vision-language models. However, common fusion methods, such as concatenation, element-wise operations, and non-linear techniques, often fail to capture structural relationships, deep feature interactions, and suffer from inefficiency or misalignment of features across domains or modalities. In this paper, we shift from high-dimensional feature space to a lower-dimensional, interpretable graph space by constructing relationship graphs that encode feature relationships at different levels, e.g., clip, frame, patch, token, etc. To capture deeper interactions, we expand graphs through iterative graph relationship updates and introduce a learnable graph fusion operator to integrate these expanded relationships for more effective fusion. Our approach is relationship-centric, operates in a homogeneous space, and is mathematically principled, resembling element-wise relationship score aggregation via multilinear polynomials. We demonstrate the effectiveness of our graph-based fusion method on video anomaly detection, showing strong performance across multi-representational, multi-modal, and multi-domain feature fusion tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- A Unified Reasoning Framework for Holistic Zero-Shot Video Anomaly AnalysisDongheng Lin, Mengxue Qu, Kunyang Han, Jianbo Jiao et al.NeurIPS 2025 · 14 citations
- Graph Self-Supervised Learning with Learnable Structural and Positional EncodingsAsiri Wijesinghe, Hao Zhu, Piotr KoniuszWWW 2025 · 3 citations
- Learning Time in Static ClassifiersXi Ding, Lei Wang, Piotr Koniusz, Yongsheng GaoAAAI 2026 · 2 citations
- Subspace Kernel Learning on Tensor SequencesLei Wang, Xi Ding, Yongsheng Gao, Piotr KoniuszICLR 2026 · 1 citation
- Privacy-Aware Video Anomaly Detection: Guided Orthogonal Projection and a Comprehensive Evaluation FrameworkWenxiang Diao, Lei Wang, Andrew Busch, Jun Zhou et al.ICML 2026
Builds on11
- SimCSE: Simple Contrastive Learning of Sentence EmbeddingsTianyu Gao, Xingcheng Yao, Danqi ChenEMNLP 2021 · 2,496 citations
- Video Swin TransformerZe Liu, Jia Ning, Yue Cao, Yixuan Wei et al.CVPR 2022 · 1,847 citations
- VaTeX: A Large-Scale, High-Quality Multilingual Dataset for Video-and-Language ResearchXin Wang, Jiawei Wu, Jun-Kun Chen, Lei Li et al.ICCV 2019 · 688 citations
- Weakly-supervised Video Anomaly Detection with Robust Temporal Feature Magnitude LearningYu Tian, Guansong Pang, Yuanhong Chen, Rajvinder Singh et al.ICCV 2021 · 495 citations
- SwinBERT: End-to-End Transformers with Sparse Attention for Video CaptioningKevin Lin, Linjie Li, Chung-Ching Lin, Faisal Ahmed et al.CVPR 2022 · 263 citations
Related papers
- Aligning Effective Tokens with Video Anomaly in Large Language ModelsYingxian Chen, Jiahui Liu, Ruidi Fan, Yanwei Li et al.ICCV 2025 · 2 citations
- Learning to Represent Image and Text with Denotation GraphBowen Zhang, Hexiang Hu, Vihan Jain, Eugene Ie et al.EMNLP 2020 · 22 citations
- Graph-Based Video-Language Learning with Multi-Grained Audio-Visual AlignmentChenyang Lyu, Wenxi Li, Tianbo Ji, Longyue Wang et al.ACM MM 2023 · 6 citations
- Harnessing Large Language Models for Training-Free Video Anomaly DetectionLuca Zanella, Willi Menapace, Massimiliano Mancini, Yiming Wang et al.CVPR 2024 · 57 citations
- Alert-CLIP: Abnormality-aware Latent-Enhanced Representation Tuning of CLIP for Video Anomaly DetectionYiyan Zhu, Menghao Zhang, Haifeng Sun, Pengfei Ren et al.CVPR 2026
