TITLE:
State-Aware Cross-Modal Contrastive Learning for Multimodal Vigilance Estimation
AUTHORS:
Wenfei Shu, Weidong Li
KEYWORDS:
Electroencephalography (EEG), Electrooculography (EOG), Multimodal Learning, Vigilance Estimation, Contrastive Learning
JOURNAL NAME:
Journal of Computer and Communications,
Vol.14 No.7,
July
23,
2026
ABSTRACT: To enhance multimodal representation learning for continuous vigilance estimation, this paper proposes a State-Aware Cross-Modal Contrastive Learning (SA-CL) framework for EEG-EOG collaborative modeling. In continuous vigilance regression tasks, samples from different temporal segments may correspond to similar vigilance states, while conventional alignment strategies tend to treat them as strictly negative pairs, thereby undermining the semantic consistency of cross-modal representations. To address this issue, a state-aware constraint is introduced into the alignment process between EEG and EOG representations, where negative sample relationships are selectively filtered or adaptively weighted according to inter-sample state similarity, thus alleviating the unreasonable repulsion between samples with similar vigilance states. Based on this strategy, an EEG-EOG multimodal collaborative framework is constructed. Specifically, the EEG branch extracts the spatial, topological, and temporal dynamic features of brain signals, while the EOG branch encodes eye-movement behavioral information. The complementary representations from both modalities are then fused to perform continuous vigilance regression prediction. Experimental results on the SEED-VIG dataset demonstrate that the proposed method effectively improves multimodal vigilance estimation performance while preserving the semantic consistency of cross-modal features.