VigilanceNet: Decouple Intra- and Inter-Modality Learning for Multimodal Vigilance Estimation in RSVP-Based BCI
Xinyu Cheng, Wei Wei, Changde Du, Shuang Qiu, Sanli Tian, Xiaojun Ma, Huiguang He
Abstract
Recently, brain-computer interface (BCI) technology has made impressive progress and has been developed for many applications. Thereinto, the BCI system based on rapid serial visual presentation (RSVP) is a promising information detection technology. However, the use of RSVP is closely related to the user's performance, which can be influenced by their vigilance levels. Therefore it is crucial to detect vigilance levels in RSVP-based BCI. In this paper, we conducted a long-term RSVP target detection experiment to collect electroencephalography (EEG) and electrooculogram (EOG) data at different vigilance levels. In addition, to estimate vigilance levels in RSVP-based BCI, we propose a multimodal method named VigilanceNet using EEG and EOG. Firstly, we define the multiplicative relationships in conventional EOG features that can better describe the relationships between EOG features, and design an outer product embedding module to extract the multiplicative relationships. Secondly, we propose to decouple the learning of intra- and inter-modality to improve multimodal learning. Specifically, for intra-modality, we introduce an intra-modality representation learning (intra-RL) method to obtain effective representations of each modality by letting each modality independently predict vigilance levels during the multimodal training process. For inter-modality, we employ the cross-modal Transformer based on cross-attention to capture the complementary information between EEG and EOG, which only pays attention to the inter-modality relations. Extensive experiments and ablation studies are conducted on the RSVP and SEED-VIG public datasets. The results demonstrate the effectiveness of the method in terms of regression error and correlation.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Cited by top-tier papers5
- Decoding Natural Images from EEG for Object RecognitionYonghao Song, Bingchuan Liu, Xiang Li, Nanlin Shi et al.ICLR 2024 · 135 citations
- Multimodal Adaptive Emotion Transformer with Flexible Modality Inputs on A Novel Dataset with Continuous LabelsWei-Bang Jiang, Xuan-Hao Liu, Wei-Long Zheng, Bao-Liang LuACM MM 2023 · 44 citations
- Multi-to-Single: Reducing Multimodal Dependency in Emotion Recognition Through Contrastive LearningYan-Kai Liu, Jinyu Cai, Bao-Liang Lu, Wei-Long ZhengAAAI 2025 · 6 citations
- DHCM-CACL: Dynamic Hierarchical Cross-modal Mamba with Confidence-Adaptive Contrastive Learning for Multimodal Emotion RecognitionBaiqiang Wu, Yang LiAAAI 2026 · 1 citation
- A Multimodal EEG-Eye Movement Model for Automatic Depression DetectionHao-Long Yin, Jian-Ming Zhang, Ren-Jie Dai, Wei-Long Zheng et al.AAAI 2026
Related papers
- ElaSleepNet: Exploring an Elastic Multimodal Neural Network for Sleep Staging via Temporal and Contextual Consistency LearningQi Shen, Junchang Xin, Bing Tian Dai, Shudi Zhang et al.ACM MM 2025
- CognitionCapturer: Decoding Visual Stimuli from Human EEG Signal with Multimodal InformationKaifan Zhang, Lihuo He, Xin Jiang, Wen Lu et al.AAAI 2025 · 34 citations
- Adversarial Multimodal Representation Learning for Click-Through Rate PredictionXiang Li, Chao Wang, Jiwei Tan, Xiaoyi Zeng et al.WWW 2020 · 61 citations
- TFF-Former: Temporal-Frequency Fusion Transformer for Zero-training Decoding of Two BCI TasksXujin Li, Wei Wei, Shuang Qiu, Huiguang HeACM MM 2022 · 27 citations
- Robust Sleep Staging over Incomplete Multimodal Physiological Signals via Contrastive ImaginationQi Shen, Junchang Xin, Bing Tian Dai, Shudi Zhang et al.NeurIPS 2024 · 15 citations
