TITLE:
Transformer-Based Multimodal Prediction of Depressive Episode Onset in Bipolar Disorder from Simulated 30-Day Digital Phenotyping Streams: A Proof-of-Concept Study of Actigraphy, HRV, and Smartphone Behaviour
AUTHORS:
Rocco de Filippis, Abdullah Al Foysal
KEYWORDS:
Digital Phenotyping, Multimodal Transformer, Bipolar Disorder, Depressive Episode, Actigraphy, Heart Rate Variability, Smartphone Sensing, Cross-Modal Attention, Passive Sensing, Early Warning
JOURNAL NAME:
Open Access Library Journal,
Vol.13 No.8,
August
27,
2026
ABSTRACT: Depressive episodes in bipolar disorder are preceded by a prodromal phase in which subtle physiological and behavioural changes reduced daytime activity, fragmented and lengthened sleep, blunted circadian rhythm, diminished heart rate variability, and social withdrawal emerge before the patient or clinician recognises a clinical episode. Digital phenotyping, the continuous passive collection of behavioural and physiological data from wearable sensors and smartphones, offers an unprecedented window onto this prodrome. Yet existing approaches typically analyse single sensor streams in isolation and fail to exploit the rich temporal structure and cross-modal interactions across the full multimodal data stream. No transformer-based multimodal architecture has been developed for depressive episode onset prediction in bipolar disorder from combined actigraphy, HRV, and smartphone behaviour streams. We developed a proof-of-concept multimodal transformer for depressive episode onset prediction using a fully synthetic digital phenotyping cohort of N = 700 simulated bipolar disorder patients (simulated onset prevalence 28.6%). Three modality streams were encoded: actigraphy (activity counts, sleep duration and efficiency, circadian amplitude, intradaily variability), heart rate variability (SDNN, RMSSD, LF/HF ratio, mean HR), and smartphone behaviour (screen time, typing speed, app-use entropy, GPS displacement, call and text frequency). Each 30-day modality stream was processed by a dedicated transformer encoder with a learnable CLS summary token and sinusoidal positional encoding; the three modality embeddings were combined via cross-modal multi-head attention fusion and concatenated with clinical meta-features for classification. The architecture was compared against single-modality transformers (ablation), a gradient-boosting model on summary statistics, and a multimodal LSTM. Cross-modal attention weights were extracted for interpretability. In this simulated proof-of-concept evaluation, the multimodal transformer achieved AUC = 0.981 (95% CI: 0.948 - 0.999), F1 = 0.901 (CI: 0.815 - 0.971), sensitivity = 0.933, and specificity = 0.947, with a Brier score of 0.039. Modality ablation demonstrated that no single modality achieved comparable performance (actigraphy 0.928, HRV 0.875, smartphone 0.910), and that multimodal fusion delivered a substantial AUC gain over the best single modality. Bayesian model comparison favoured the multimodal transformer over the evaluated baselines, including the multimodal LSTM. Simulated digital phenotyping trajectories diverged from approximately day 16 of the 30-day observation window, with reduced activity and HRV, blunted circadian amplitude, increased screen time, and social withdrawal in simulated onset cases. These simulation results support the technical feasibility of multimodal attention-based fusion for modelling a digitally expressed depressive prodrome. They do not establish clinical validity or a deployable two-week warning window; both require prospective validation in real-world bipolar disorder cohorts with naturally observed episode onset, missing data, device heterogeneity, and independent external evaluation.