<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.4 20241031//EN" "JATS-journalpublishing1-4.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article" dtd-version="1.4" xml:lang="en">
  <front>
    <journal-meta>
      <journal-id journal-id-type="publisher-id">Oalib</journal-id>
      <journal-title-group>
        <journal-title>Open Access Library Journal</journal-title>
      </journal-title-group>
      <issn pub-type="epub">2333-9721</issn>
      <issn pub-type="ppub">2333-9705</issn>
      <publisher>
        <publisher-name>Scientific Research Publishing</publisher-name>
      </publisher>
    </journal-meta>
    <article-meta>
      <article-id pub-id-type="doi">10.4236/oalib.1115346</article-id>
      <article-id pub-id-type="publisher-id">Oalib-151685</article-id>
      <article-categories>
        <subj-group>
          <subject>Article</subject>
        </subj-group>
        <subj-group>
          <subject>Biomedical</subject>
          <subject>Life Sciences</subject>
          <subject>Business</subject>
          <subject>Economics</subject>
          <subject>Chemistry</subject>
          <subject>Materials Science</subject>
          <subject>Computer Science</subject>
          <subject>Communications</subject>
          <subject>Earth</subject>
          <subject>Environmental Sciences</subject>
          <subject>Engineering</subject>
          <subject>Medicine</subject>
          <subject>Healthcare</subject>
          <subject>Physics</subject>
          <subject>Mathematics</subject>
          <subject>Social Sciences</subject>
          <subject>Humanities</subject>
        </subj-group>
      </article-categories>
      <title-group>
        <article-title>A Multimodal Deep Learning Framework for Continuous Mood Monitoring and Episode Prediction in Bipolar Disorder</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author" corresp="yes">
          <contrib-id contrib-id-type="orcid">0000-0001-9101-072X</contrib-id>
          <name name-style="western">
            <surname>Filippis</surname>
            <given-names>Rocco de</given-names>
          </name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <contrib-id contrib-id-type="orcid">0000-0002-5102-4999</contrib-id>
          <name name-style="western">
            <surname>Foysal</surname>
            <given-names>Abdullah Al</given-names>
          </name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
      </contrib-group>
      <aff id="aff1"><label>1</label> Department of Neuroscience, Institute of Psychopathology, Rome, Italy </aff>
      <aff id="aff2"><label>2</label> Department of Computer Engineering (AI), University of Genova, Genova, Italy </aff>
      <author-notes>
        <fn fn-type="conflict" id="fn-conflict">
          <p>The authors declare no conflicts of interest.</p>
        </fn>
      </author-notes>
      <pub-date pub-type="epub">
        <day>06</day>
        <month>05</month>
        <year>2026</year>
      </pub-date>
      <pub-date pub-type="collection">
        <month>05</month>
        <year>2026</year>
      </pub-date>
      <volume>13</volume>
      <issue>05</issue>
      <fpage>1</fpage>
      <lpage>22</lpage>
      <history>
        <date date-type="received">
          <day>14</day>
          <month>04</month>
          <year>2026</year>
        </date>
        <date date-type="accepted">
          <day>26</day>
          <month>05</month>
          <year>2026</year>
        </date>
        <date date-type="published">
          <day>29</day>
          <month>05</month>
          <year>2026</year>
        </date>
      </history>
      <permissions>
        <copyright-statement>© 2026 by the authors and Scientific Research Publishing Inc.</copyright-statement>
        <copyright-year>2026</copyright-year>
        <license license-type="open-access">
          <license-p> This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license ( <ext-link ext-link-type="uri" xlink:href="https://creativecommons.org/licenses/by/4.0/">https://creativecommons.org/licenses/by/4.0/</ext-link> ). </license-p>
        </license>
      </permissions>
      <self-uri content-type="doi" xlink:href="https://doi.org/10.4236/oalib.1115346">https://doi.org/10.4236/oalib.1115346</self-uri>
      <abstract>
        <p>Bipolar disorder does not announce itself with a clean clinical signal. It builds in sleep that fragments days before a manic break, in accelerating movement recorded through a wrist sensor, in speech that picks up tempo before the patient notices anything has changed. Capturing these signals passively and continuously is the promise of digital phenotyping. Delivering on it requires machine learning architectures capable of integrating heterogeneous, irregularly sampled, high-dimensional sensor streams with the contextual knowledge of self-reported mood and circadian biology. We introduce MoodSense-Net, an end-to-end multimodal deep learning framework that fuses smartphone accelerometery, sleep metrics, GPS mobility traces, speech acoustics, and ecological momentary assessment (EMA) data to continuously monitor and predict mood instability in bipolar disorder. The model integrates a Multi-Scale Temporal Convolutional Network (MS-TCN) for accelerometery, a Bidirectional LSTM for sleep staging, a Rhythm CNN for GPS circadian patterns, a Speech-BERT module for acoustic analysis, and a cross-modal transformer fusion layer with a Bayesian deep ensemble output for uncertainty-calibrated predictions. Trained and validated on a prospective cohort of 1847 participants monitored continuously for 12 months encompassing over 26 million sensor samples and 312,000 EMA responses, MoodSense-Net achieves 5-class mood state classification accuracy of 92.7%, AUC-ROC of 0.963, and macro-F1 of 0.891. Episode onset prediction at the 7-day horizon yields sensitivity of 89.1% and specificity of 90.3% for manic episodes, with a mean prediction lead time of 5.1 ± 1.9 days. The Bayesian ensemble achieves Expected Calibration Error (ECE) of 0.028. MoodSense-Net establishes a new methodological benchmark for passive monitoring in computational psychiatry, providing a validated, deployable architecture for continuous bipolar mood instability surveillance.</p>
      </abstract>
      <kwd-group kwd-group-type="author-generated" xml:lang="en">
        <kwd>Digital Phenotyping</kwd>
        <kwd>Bipolar Disorder</kwd>
        <kwd>Mood Instability</kwd>
        <kwd>Smartphone Sensing</kwd>
        <kwd>Deep Learning</kwd>
        <kwd>Temporal Convolutional Network</kwd>
        <kwd>Speech Analysis</kwd>
        <kwd>GPS Mobility</kwd>
        <kwd>Bayesian Uncertainty</kwd>
        <kwd>Ecological Momentary Assessment</kwd>
        <kwd>Circadian Rhythm</kwd>
        <kwd>Affective Computing</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec1">
      <title>1. Introduction</title>
      <p>Bipolar disorder does not announce itself with a clean clinical signal. It builds in sleep that contracts before a manic break, in motion that accelerates through a wearable sensor, in speech that picks up tempo before the patient notices anything is wrong. The emerging field of digital phenotyping [<xref ref-type="bibr" rid="B1">1</xref>][<xref ref-type="bibr" rid="B2">2</xref>] proposes to close the gap between what passive sensors can detect and what clinical services currently capture: by applying machine learning to continuously recorded smartphone data, it may be possible to construct a real-time computational phenotype of a person’s mental state that is richer, more continuous, and more objective than anything a weekly clinical interview can provide.</p>
      <p>This challenge is particularly urgent in bipolar disorder. The condition affects an estimated 1% - 4% of the global population across its spectrum subtypes [<xref ref-type="bibr" rid="B3">3</xref>][<xref ref-type="bibr" rid="B4">4</xref>], contributing over 9.9 million disability-adjusted life years annually [<xref ref-type="bibr" rid="B5">5</xref>]. The mean diagnostic delay exceeds seven years from symptom onset to confirmed diagnosis, a window in which patients accumulate functional impairment, undergo ineffective treatment trials, and face a lifetime suicide risk estimated at 15 - 20 times the general population rate [<xref ref-type="bibr" rid="B6">6</xref>].</p>
      <p>What makes BD tractable for digital phenotyping is that mood episodes are preceded by detectable behavioural and physiological signatures that emerge days before clinical threshold is crossed. Pharmacological management is highly phase-sensitive: lithium dose adjustments, quetiapine augmentation, and brief structured interventions show the greatest efficacy when deployed 3 - 7 days before episode onset [<xref ref-type="bibr" rid="B7">7</xref>][<xref ref-type="bibr" rid="B8">8</xref>]. Current reactive psychiatry systematically misses this window because it has no passive monitoring infrastructure to detect the prodromal period in real time.</p>
      <p>Evidence supports each sensing modality independently. Prodromal behavioural signatures for mood episodes, including sleep contraction, activity escalation, and social disengagement, have been documented in prospective studies [<xref ref-type="bibr" rid="B9">9</xref>]. Actigraphy-derived rest-activity rhythms have been linked to BD mood state transitions [<xref ref-type="bibr" rid="B10">10</xref>]. GPS mobility features have been associated with depressive episodes [<xref ref-type="bibr" rid="B11">11</xref>]. Speech acoustics have proven informative for mania and depression state detection [<xref ref-type="bibr" rid="B12">12</xref>]. Ecological momentary assessment captures fine-grained affective dynamics in daily life [<xref ref-type="bibr" rid="B13">13</xref>]. However, no prior system has jointly modelled all five modalities in a single end-to-end trainable architecture. The most competitive prior approach, a transformer-based ensemble combining actigraphy and EHR data, achieved 86.4% mood state accuracy [<xref ref-type="bibr" rid="B14">14</xref>] but did not integrate speech, GPS, or EMA. MoodSense-Net subsumes and extends all of these contributions.</p>
      <sec id="sec1dot1">
        <title>Summary of Contributions</title>
        <p>1) <bold>MoodSense-Net architecture:</bold>The first jointly trained, five-modality deep learning system for continuous bipolar mood monitoring, integrating accelerometery, sleep, GPS, speech, and EMA streams through a cross-modal transformer fusion layer with Bayesian ensemble output.</p>
        <p>2) <bold>MS-TCN for accelerometery:</bold>A Multi-Scale Temporal Convolutional Network with dilated causal convolutions and missingness-aware gating, designed for irregularly sampled wearable accelerometery from free-living bipolar disorder patients.</p>
        <p>3) <bold>Speech-BERT for acoustic phenotyping:</bold>A domain-adaptive BERT variant pre-trained on 2.8 million psychiatric consultation audio transcripts, achieving +3.8% accuracy over general audio BERT on BD mood state prediction.</p>
        <p>4) <bold>Circadian rhythm model:</bold>An explicit circadian phase estimation module operating on GPS and accelerometery, contributing an independent +3.4% accuracy improvement in ablation.</p>
        <p>5) <bold>State-of-the-art performance:</bold>92.7% accuracy, AUC-ROC 0.963, macro-F1 0.891; episode onset sensitivity 89.1%/specificity 90.3% at 7-day horizon; ECE = 0.028; mean prediction lead time 5.1 ± 1.9 days.</p>
        <p>6) <bold>Prospective cohort at scale:</bold>1847 participants monitored continuously for 12 months across four clinical sites, the largest prospective digital phenotyping cohort in bipolar disorder research reported to date.</p>
      </sec>
    </sec>
    <sec id="sec2">
      <title>2. Background and Related Work</title>
      <sec id="sec2dot1">
        <title>2.1. The Instability Problem in Bipolar Disorder</title>
        <p>Bipolar disorder is fundamentally a disorder of affective instability not simply of discrete episodes, but of the continuous oscillation between and through mood states that defines the lived experience of the condition. Clinical management built around scheduled appointments and patient-initiated contact captures almost none of this continuous dynamic. Patients spend approximately half their illness time in subsyndromal states neither euthymic nor meeting full episode criteria yet experiencing significant impairment, elevated relapse risk, and deteriorating functional capacity. The pharmacological sensitivity of BD means that interventions precisely timed to the prodromal period substantially outperform those initiated at episode onset. MoodSense-Net is designed to identify this window. For the purposes of this study, mood instability is operationalized as intra-individual variability in affective state across consecutive weekly assessments, quantified as the mean absolute difference in YMRS and HAMD-17 scores between adjacent weekly ratings. This umbrella construct encompasses two distinct prediction targets: 1) five-class mood state classification per week, and 2) binary episode onset prediction at the 7-day horizon, both of which are defined and evaluated independently in Sections 3.4 and 5.</p>
      </sec>
      <sec id="sec2dot2">
        <title>2.2. Digital Phenotyping in Bipolar Disorder: Prior Work</title>
        <p>Digital phenotyping has produced influential findings across individual sensing modalities. The Sleep Regularity Index [<xref ref-type="bibr" rid="B15">15</xref>] and actigraphy-based rest-activity rhythm analysis [<xref ref-type="bibr" rid="B16">16</xref>] have demonstrated associations between circadian disruption and mood episode transitions in bipolar disorder. GPS mobility studies have shown that radius of gyration, location entropy, and number of unique locations visited can distinguish depressive from euthymic periods with 78% accuracy in prospective cohorts [<xref ref-type="bibr" rid="B17">17</xref>]. A comprehensive review of physiological and behavioural monitoring technologies and their psychiatric applications [<xref ref-type="bibr" rid="B18">18</xref>] established the theoretical and empirical foundations on which MoodSense-Net builds. The most recent competitive system, a transformer-based ensemble of actigraphy and EHR features achieved 86.4% mood state accuracy but excluded speech, GPS, and EMA modalities, and did not provide Bayesian uncertainty quantification.</p>
      </sec>
      <sec id="sec2dot3">
        <title>2.3. Deep Learning Architectures for Temporal Sensing</title>
        <p>Temporal Convolutional Networks [<xref ref-type="bibr" rid="B19">19</xref>] demonstrated that dilated causal convolutions outperform LSTM [<xref ref-type="bibr" rid="B20">20</xref>] architectures on sequence modelling tasks, with substantially lower computational cost and freedom from gradient vanishing. WaveNet [<xref ref-type="bibr" rid="B21">21</xref>] extended this to audio generation, showing that hierarchical dilated convolutions can model dependencies spanning multiple timescales simultaneously. The Transformer architecture [<xref ref-type="bibr" rid="B22">22</xref>] introduced scaled dot-product attention, enabling superior long-range sequence modelling without recurrence and providing the foundational mechanism for cross-modal fusion. BERT [<xref ref-type="bibr" rid="B23">23</xref>] and its clinical adaptation ClinicalBERT [<xref ref-type="bibr" rid="B24">24</xref>] demonstrated that domain-adaptive pre-training substantially improves performance on specialized downstream tasks, a principle MoodSense-Net extends to speech acoustics through Speech-BERT.</p>
      </sec>
      <sec id="sec2dot4">
        <title>2.4. Bayesian Uncertainty in Clinical AI</title>
        <p>The deployment of AI systems in clinical mental health settings demands calibrated probabilistic outputs a well-calibrated model at 90% accuracy is safer than an overconfident model at 93%. Monte Carlo Dropout [<xref ref-type="bibr" rid="B25">25</xref>] and deep ensembles [<xref ref-type="bibr" rid="B26">26</xref>] provide principled approximations to Bayesian posterior inference; ensembles empirically produce superior calibration relative to single-model methods. The EU AI Act (Regulation EU 2024/1689) [<xref ref-type="bibr" rid="B27">27</xref>] classifies AI systems supporting psychiatric diagnosis as high-risk, requiring transparent uncertainty quantification making Bayesian calibration a regulatory necessity for European clinical deployment, not merely a scientific preference.</p>
      </sec>
    </sec>
    <sec id="sec3">
      <title>3. Dataset, Cohort Design, and Preprocessing</title>
      <sec id="sec3dot1">
        <title>3.1. Study Population and Recruitment</title>
        <p>The MoodSense cohort was assembled through a prospective multi-site observational study across four tertiary psychiatric centres in Italy, the United Kingdom, and Netherlands, with ethics approval from each site (Primary: IRB Ref. UNIGE-2024-MSN-02, GDPR Article 9 compliance). Inclusion criteria: adults aged 18 - 65 with DSM-5 confirmed BD-I or BD-II, minimum two documented mood episodes in the prior three years, capacity for written informed consent, and willingness to use the Mood Sense app for the 12-month monitoring period. Exclusion criteria: concurrent primary psychotic disorder, active substance use disorder with ongoing intoxication, neurological comorbidity, or inability to complete basic EMA.</p>
        <p>The final analytic cohort comprised 1847 participants (BD-I: n = 1041; BD-II: n = 806) monitored continuously over 12 months. The unit of analysis throughout this study is the participant-week-all sensor streams, EMA responses, and speech samples recorded within a given calendar week are aggregated into a single feature vector, aligned with the weekly clinician-assigned mood state label. With 1847 participants monitored over 52 weeks, the theoretical maximum is 96,044 participant-weeks; the realized dataset of 68,412 episode-weeks reflects exclusion of weeks with clinician assessment missing or sensor compliance below 40%. The dataset accumulated 26.4 million accelerometery samples, 14.2 million GPS location records, 8.7 million speech feature extracts, and 312,841 EMA responses (mean: 4.8 per participant per day). Clinician-rated assessments (YMRS and HAMD-17) were conducted fortnightly by trained raters blinded to sensor data. <bold>Table 1</bold> presents cohort characteristics.</p>
      </sec>
      <sec id="sec3dot2">
        <title>3.2. Cohort Characteristics</title>
        <p><bold>Table 1</bold> summarizes the sociodemographic and clinical profile of the final analytic cohort (n = 1847). The sample was approximately balanced by sex (52.9% female) and young-to-middle-aged in composition (mean age 34.4 ± 11.2 years), consistent with the peak burden period of bipolar disorder and with prior prospective digital phenotyping cohorts. BD-I participants (n = 1041) exhibited longer illness duration (10.2 ± 7.1 vs. 8.4 ± 6.2 years) and a higher mean prior episode count (6.8 ± 4.4 vs. 5.2 ± 3.6) relative to BD-II (n = 806), reflecting the characteristically more episodic and severe longitudinal trajectory of BD-I. Current mood stabilizer use was high across both subtypes (81.6% overall; BD-I: 83.2%, BD-II: 79.4%), reducing the confound of untreated illness on sensor signal interpretation, though the 3.8 percentage-point difference between subtypes was retained as a covariate in subgroup analyses. Comorbid anxiety disorder was present in 40.7% of the total cohort, with a higher prevalence in BD-II (44.7%) relative to BD-I (37.8%), consistent with established epidemiological patterns of affective comorbidity across bipolar subtypes. Smartphone compliance was high and comparable across groups (total: 88.1%; BD-I: 87.4%; BD-II: 89.1%), supporting robust passive monitoring coverage throughout the 12-month observation window. The realized dataset of 68,412 labelled episode-weeks reflects the natural heterogeneity of bipolar illness course, with a class distribution skewed toward euthymia (42%) and depression (26%), followed by hypomania (16%), mania (11%), and mixed features (5%) a distribution that mirrors real-world bipolar illness burden and was explicitly addressed through stratified sampling and weighted cross-entropy loss during model training.</p>
        <p><bold>Table 1.</bold> MoodSense cohort sociodemographic and clinical characteristics.</p>
        <table-wrap id="tbl1">
          <label>Table 1</label>
          <table>
            <tbody>
              <tr>
                <td>
                  <bold>Characteristic</bold>
                </td>
                <td>
                  <bold>BD-I</bold>
                  <bold>(n</bold>
                  <bold>=</bold>
                  <bold>1041)</bold>
                </td>
                <td>
                  <bold>BD-II</bold>
                  <bold>(n</bold>
                  <bold>=</bold>
                  <bold>806)</bold>
                </td>
                <td>
                  <bold>Total</bold>
                  <bold>(n</bold>
                  <bold>=</bold>
                  <bold>1847)</bold>
                </td>
              </tr>
              <tr>
                <td>Mean age (years ± SD)</td>
                <td>34.8 ± 11.4</td>
                <td>33.9 ± 10.8</td>
                <td>34.4 ± 11.2</td>
              </tr>
              <tr>
                <td>Female (%)</td>
                <td>51.3%</td>
                <td>55.1%</td>
                <td>52.9%</td>
              </tr>
              <tr>
                <td>Illness duration (years ± SD)</td>
                <td>10.2 ± 7.1</td>
                <td>8.4 ± 6.2</td>
                <td>9.4 ± 6.8</td>
              </tr>
              <tr>
                <td>Prior episodes (mean ± SD)</td>
                <td>6.8 ± 4.4</td>
                <td>5.2 ± 3.6</td>
                <td>6.1 ± 4.1</td>
              </tr>
              <tr>
                <td>Current mood stabilizer (%)</td>
                <td>83.2%</td>
                <td>79.4%</td>
                <td>81.6%</td>
              </tr>
              <tr>
                <td>Comorbid anxiety disorder (%)</td>
                <td>37.8%</td>
                <td>44.7%</td>
                <td>40.7%</td>
              </tr>
              <tr>
                <td>Smartphone compliance rate (%)</td>
                <td>
                  <bold>87.4%</bold>
                </td>
                <td>
                  <bold>89.1%</bold>
                </td>
                <td>
                  <bold>88.1%</bold>
                </td>
              </tr>
              <tr>
                <td>Total labelled episode-weeks</td>
                <td>
                </td>
                <td>
                </td>
                <td>
                  <bold>68</bold>
                  <bold>,</bold>
                  <bold>412</bold>
                </td>
              </tr>
              <tr>
                <td>Class distribution (Euth/Dep/Hypo/Mania/Mixed)</td>
                <td>
                </td>
                <td>
                </td>
                <td>42%/26%/16%/11%/5%</td>
              </tr>
            </tbody>
          </table>
        </table-wrap>
      </sec>
      <sec id="sec3dot3">
        <title>3.3. Sensing Modalities and Feature Engineering</title>
        <p>Each sensing modality was pre-processed and featured through a standardized, site-harmonized pipeline.</p>
        <p>3.3.1. Accelerometery Activity and Movement</p>
        <p>Raw triaxial accelerometery (32 Hz) was motion-artifact-corrected, bandpass filtered (0.1 - 15 Hz), and epoched into 1-minute windows. Features extracted per 24-hour day: mean activity intensity, activity fragmentation index, most active 10-hour window (M10), least active 5-hour window (L5), inter-daily stability (IS), intraday variability (IV), and spectral power in the circadian frequency band (0.032 - 0.042 Hz). Missing data from device non-wear was imputed using Gaussian process interpolation conditioned on adjacent valid epochs [<xref ref-type="bibr" rid="B28">28</xref>], with missingness masks propagated to the MS-TCN attention mechanism.</p>
        <p>3.3.2. Sleep Duration, Architecture, and Regularity</p>
        <p>Sleep episodes were identified from accelerometry using the validated Cole-Kripke algorithm [<xref ref-type="bibr" rid="B29">29</xref>], producing per-night estimates of total sleep time (TST), sleep efficiency (SE), sleep onset latency (SOL), and wake after sleep onset (WASO). Sleep regularity was quantified via the Sleep Regularity Index, the probability that the participant’s sleep-wake status at any two time points separated by 24 hours are concordant, a metric whose validity and associations with health outcomes have been established in prospective cohort research [<xref ref-type="bibr" rid="B30">30</xref>]. A sleep staging transformer [<xref ref-type="bibr" rid="B31">31</xref>] was applied to continuous accelerometery for automated NREM/REM/Wake estimation in participants with sufficient signal quality (76.3% of nights).</p>
        <p>3.3.3. GPS Mobility Spatial and Social Behaviour</p>
        <p>GPS location data (1-minute resolution) yielded per-day mobility features: radius of gyration, location entropy the Shannon entropy over time spent at distinct locations, capturing routine versus variability [<xref ref-type="bibr" rid="B32">32</xref>] number of unique locations visited, total distance travelled, home dwell time, and transition rate between locations. Social behaviour proxies derived from call metadata (duration, frequency, unique contacts) were included after participant consent.</p>
        <p>3.3.4. Speech Acoustics Prosody, Rate, and Coherence</p>
        <p>Passive speech feature extraction was conducted on audio from consented phone calls (minimum 60 seconds), using a privacy-preserving on-device pipeline that extracted acoustic features without storing raw audio. Features include fundamental frequency (F0) mean and variability, speech rate, pause duration distribution, voice tremor index, harmonics-to-noise ratio, MFCC coefficients, and a semantic coherence score from Speech-BERT. The clinical relevance of acoustic features for depression and mania detection is well established in the computational psychiatry literature [<xref ref-type="bibr" rid="B33">33</xref>][<xref ref-type="bibr" rid="B34">34</xref>].</p>
        <p>3.3.5. Ecological Momentary Assessment</p>
        <p>The MoodSense EMA protocol administered 4 brief surveys per day at semi-randomized times, capturing self-rated mood, energy level, sleep quality, social engagement, irritability, and medication adherence on validated visual analogue scales. Mean EMA compliance was 79.4% per participant-day, consistent with published compliance rates in smartphone EMA psychiatric research [<xref ref-type="bibr" rid="B35">35</xref>].</p>
        <p>3.3.6. Missing Data Rates and Imputation</p>
        <p><bold>Table 2</bold> reports per-modality missingness by site and study week. Briefly: accelerometery missingness (device non-wear) was 11.3% of participant-days; sleep feature missingness (insufficient accelerometery for Cole-Kripke estimation) was 14.7%; GPS missingness was 9.8%; speech missingness (insufficient call duration or no passive call in each week) was 22.1% of participant-weeks; EMA missingness was 20.6% of participant-days. For modality-level imputation, missing weekly feature vectors were replaced with the participant-specific modality mean computed exclusively from training-set observations. This mean was fixed at training time and applied identically during validation and test inference to prevent leakage. Missingness masks were propagated to all relevant encoder attention mechanisms (MS-TCN, Speech-BERT) so that imputed segments contribute attenuated gradients during fine-tuning.</p>
      </sec>
      <sec id="sec3dot4">
        <title>3.4. Ground Truth Labelling</title>
        <p>Mood state labels were assigned weekly by consensus of two senior psychiatrists using all available data sources: YMRS and HAMD-17 scores, EMA trend summaries, and clinical notes. Inter-rater reliability reached <italic>κ</italic> = 0.87 (95% CI: 0.84 - 0.90), indicating strong agreement. To prevent information leakage, EMA trend summaries used in label construction were derived exclusively from the preceding week’s self-report responses (days −14 to −8 relative to the label week), whereas the EMA encoder receives only the current week’s raw survey data. These input and label windows are strictly non-overlapping by design. Mood states followed DSM-5 definitions across five categories euthymia, depression, hypomania, mania, and mixed features. Final labelled dataset: 68,412 episode-weeks.</p>
      </sec>
    </sec>
    <sec id="sec4">
      <title>4. Mood Sense-Net: Architecture and Technical Specification</title>
      <p>Mood Sense-Net integrates five modality-specific encoders, a circadian rhythm model, a cross-modal transformer fusion layer, and a Bayesian deep ensemble output stage in a jointly trainable end-to-end architecture. <xref ref-type="fig" rid="fig1">Figure 1</xref><xref ref-type="fig" rid="fig1">Figure 1</xref> presents the complete system.</p>
      <sec id="sec4dot1">
        <title>4.1. Module 1: Multi-Scale Temporal Convolutional Network (MS-TCN)</title>
        <p>4.1.1. Design Rationale</p>
        <p>Accelerometery in free-living conditions presents three challenges that standard sequence models do not handle well: irregular sampling from non-wear gaps, multi-scale temporal structure spanning sub-minute autonomic rhythms to 24-hour circadian cycles, and non-stationarity from genuine behavioural changes. The MS-TCN addresses all three through hierarchical dilated convolutions, multi-head causal self-attention with missingness-aware gating, and multi-scale pooling.</p>
        <fig id="fig1">
          <label>Figure 1</label>
          <graphic xlink:href="https://html.scirp.org/file/1115346-rId16.jpeg?20260529050531" />
        </fig>
        <p><bold>Figure 1.</bold> MoodSense-Net end-to-end architecture. Five smartphone sensor streams are processed by dedicated modality encoders. A circadian rhythm model provides explicit phase estimation, enriching temporal representations across modalities. Cross-modal transformer fusion integrates all embeddings into a unified patient state representation. The Bayesian ensemble (<italic>M</italic> = 10, MC-Dropout) produces mood state classification, episode onset prediction, and calibrated uncertainty estimates.</p>
        <p>4.1.2. Formal Specification</p>
        <p>Let <inline-formula><mml:math display="inline"><mml:mrow><mml:mi> X </mml:mi><mml:mo> ∈ </mml:mo><mml:msup><mml:mi> ℝ </mml:mi><mml:mrow><mml:mi> T </mml:mi><mml:mo> × </mml:mo><mml:mi> C </mml:mi></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula> denote the accelerometery input with<italic>T</italic> timesteps and <italic>C</italic> = 7 feature channels. At each stack level <inline-formula><mml:math display="inline"><mml:mrow><mml:mi> l </mml:mi><mml:mo> ∈ </mml:mo><mml:mrow><mml:mo> { </mml:mo><mml:mrow><mml:mn> 1 </mml:mn><mml:mo> , </mml:mo><mml:mo> ⋯ </mml:mo><mml:mo> , </mml:mo><mml:mn> 6 </mml:mn></mml:mrow><mml:mo> } </mml:mo></mml:mrow></mml:mrow></mml:math></inline-formula> :</p>
        <disp-formula id="FD1">
          <label>(1)</label>
          <mml:math display="inline">
            <mml:mrow>
              <mml:msub>
                <mml:mi>H</mml:mi>
                <mml:mi>l</mml:mi>
              </mml:msub>
              <mml:mo>=</mml:mo>
              <mml:mtext>LayerNorm</mml:mtext>
              <mml:mrow>
                <mml:mo>(</mml:mo>
                <mml:mrow>
                  <mml:msub>
                    <mml:mrow>
                      <mml:mtext>TCN</mml:mtext>
                    </mml:mrow>
                    <mml:mrow>
                      <mml:mrow>
                        <mml:mo>{</mml:mo>
                        <mml:mrow>
                          <mml:msub>
                            <mml:mi>d</mml:mi>
                            <mml:mi>l</mml:mi>
                          </mml:msub>
                        </mml:mrow>
                        <mml:mo>}</mml:mo>
                      </mml:mrow>
                    </mml:mrow>
                  </mml:msub>
                  <mml:mrow>
                    <mml:mo>(</mml:mo>
                    <mml:mrow>
                      <mml:msub>
                        <mml:mi>H</mml:mi>
                        <mml:mrow>
                          <mml:mi>l</mml:mi>
                          <mml:mo>−</mml:mo>
                          <mml:mn>1</mml:mn>
                        </mml:mrow>
                      </mml:msub>
                    </mml:mrow>
                    <mml:mo>)</mml:mo>
                  </mml:mrow>
                  <mml:mo>⊙</mml:mo>
                  <mml:msub>
                    <mml:mrow>
                      <mml:mtext>Attn</mml:mtext>
                    </mml:mrow>
                    <mml:mi>l</mml:mi>
                  </mml:msub>
                  <mml:mrow>
                    <mml:mo>(</mml:mo>
                    <mml:mrow>
                      <mml:msub>
                        <mml:mi>H</mml:mi>
                        <mml:mrow>
                          <mml:mi>l</mml:mi>
                          <mml:mo>−</mml:mo>
                          <mml:mn>1</mml:mn>
                        </mml:mrow>
                      </mml:msub>
                    </mml:mrow>
                    <mml:mo>)</mml:mo>
                  </mml:mrow>
                  <mml:mo>⊙</mml:mo>
                  <mml:mtext>Gate</mml:mtext>
                  <mml:mrow>
                    <mml:mo>(</mml:mo>
                    <mml:mrow>
                      <mml:msub>
                        <mml:mi>H</mml:mi>
                        <mml:mrow>
                          <mml:mi>l</mml:mi>
                          <mml:mo>−</mml:mo>
                          <mml:mn>1</mml:mn>
                        </mml:mrow>
                      </mml:msub>
                      <mml:mo>,</mml:mo>
                      <mml:mi>M</mml:mi>
                    </mml:mrow>
                    <mml:mo>)</mml:mo>
                  </mml:mrow>
                  <mml:mo>+</mml:mo>
                  <mml:msub>
                    <mml:mi>H</mml:mi>
                    <mml:mrow>
                      <mml:mi>l</mml:mi>
                      <mml:mo>−</mml:mo>
                      <mml:mn>1</mml:mn>
                    </mml:mrow>
                  </mml:msub>
                </mml:mrow>
                <mml:mo>)</mml:mo>
              </mml:mrow>
            </mml:mrow>
          </mml:math>
        </disp-formula>
        <p>where dilation <inline-formula><mml:math display="inline"><mml:mrow><mml:msub><mml:mi> d </mml:mi><mml:mi> l </mml:mi></mml:msub><mml:mo> = </mml:mo><mml:msup><mml:mn> 2 </mml:mn><mml:mrow><mml:mi> l </mml:mi><mml:mo> − </mml:mo><mml:mn> 1 </mml:mn></mml:mrow></mml:msup><mml:mo> ∈ </mml:mo><mml:mrow><mml:mo> { </mml:mo><mml:mrow><mml:mn> 1 </mml:mn><mml:mo> , </mml:mo><mml:mn> 2 </mml:mn><mml:mo> , </mml:mo><mml:mn> 4 </mml:mn><mml:mo> , </mml:mo><mml:mn> 8 </mml:mn><mml:mo> , </mml:mo><mml:mn> 16 </mml:mn><mml:mo> , </mml:mo><mml:mn> 32 </mml:mn></mml:mrow><mml:mo> } </mml:mo></mml:mrow></mml:mrow></mml:math></inline-formula> ; <inline-formula><mml:math display="inline"><mml:mo> ⊙ </mml:mo></mml:math></inline-formula> is element-wise multiplication; <inline-formula><mml:math display="inline"><mml:mrow><mml:mi> M </mml:mi><mml:mo> ∈ </mml:mo><mml:msup><mml:mrow><mml:mrow><mml:mo> { </mml:mo><mml:mrow><mml:mn> 0 </mml:mn><mml:mo> , </mml:mo><mml:mn> 1 </mml:mn></mml:mrow><mml:mo> } </mml:mo></mml:mrow></mml:mrow><mml:mi> T </mml:mi></mml:msup></mml:mrow></mml:math></inline-formula> is the missingness mask from non-wear detection; Attn<italic><sub>l</sub></italic> is 4-head causal masked self-attention; and Gate(·, <italic>M</italic>) is a sigmoid-gated linear unit conditioned on <italic>M</italic>, ensuring missing segments contribute zero gradient. Multi-scale pooling at <inline-formula><mml:math display="inline"><mml:mrow><mml:mi> l </mml:mi><mml:mo> ∈ </mml:mo><mml:mrow><mml:mo> { </mml:mo><mml:mrow><mml:mn> 2 </mml:mn><mml:mo> , </mml:mo><mml:mn> 4 </mml:mn><mml:mo> , </mml:mo><mml:mn> 6 </mml:mn></mml:mrow><mml:mo> } </mml:mo></mml:mrow></mml:mrow></mml:math></inline-formula> before MLP projection yields <inline-formula><mml:math display="inline"><mml:mrow><mml:msub><mml:mi> h </mml:mi><mml:mrow><mml:mtext> accel </mml:mtext></mml:mrow></mml:msub><mml:mo> ∈ </mml:mo><mml:msup><mml:mi> ℝ </mml:mi><mml:mrow><mml:mn> 256 </mml:mn></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula> .</p>
      </sec>
      <sec id="sec4dot2">
        <title>4.2. Module 2: Bidirectional LSTM for Sleep Architecture</title>
        <p>Sleep variables exhibit strong temporal autocorrelation tonight’s sleep efficiency predicts tomorrow’s mood state more strongly than any single night’s reading in isolation making bidirectional sequence modelling appropriate. A 2-layer Bidirectional LSTM with hidden dimension 128 (bidirectional: 256 total) processes a 14-night rolling window of sleep feature vectors (12 features per night). Attention pooling over hidden states produces <inline-formula><mml:math display="inline"><mml:mrow><mml:msub><mml:mi> h </mml:mi><mml:mrow><mml:mtext> sleep </mml:mtext></mml:mrow></mml:msub><mml:mo> ∈ </mml:mo><mml:msup><mml:mi> ℝ </mml:mi><mml:mrow><mml:mn> 128 </mml:mn></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula> . Layer normalization and dropout (p = 0.25) are applied between LSTM layers. The BiLSTM is initialized with weights pre-trained on the Montreal Archive of Sleep Studies before fine-tuning on the MoodSense cohort.</p>
      </sec>
      <sec id="sec4dot3">
        <title>4.3. Module 3: Rhythm CNN for GPS Circadian Patterns</title>
        <p>GPS mobility traces encode circadian structure in their temporal distribution, the timing of location transitions, home departure and return rhythms, and regularity of social location visits. A specialized Rhythm CNN processes GPS feature vectors through 1D convolutions with kernels of size 24 (one per hour of day), learning time-of-day sensitivity explicitly. Three convolutional blocks with max pooling, followed by a fully connected layer, produce <inline-formula><mml:math display="inline"><mml:mrow><mml:msub><mml:mi> h </mml:mi><mml:mrow><mml:mtext> gps </mml:mtext></mml:mrow></mml:msub><mml:mo> ∈ </mml:mo><mml:msup><mml:mi> ℝ </mml:mi><mml:mrow><mml:mn> 128 </mml:mn></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula> . The Rhythm CNN receives explicit circadian phase embeddings from the circadian rhythm module (Section 4.5) as positional context.</p>
      </sec>
      <sec id="sec4dot4">
        <title>4.4. Module 4: Speech-BERT for Acoustic Phenotyping</title>
        <p>4.4.1. Domain-Adaptive Pre-Training</p>
        <p>General-domain audio transformers misrepresent psychiatric speech by underweighting prosodic features clinically associated with mania (elevated F0, increased rate, reduced pause duration) and depression (flattened F0, prolonged pauses, reduced rate). Speech-BERT was developed through two-stage pre-training. Stage 1 domain-adaptive pre-training: acoustic features (MFCCs [13 coefficients], fundamental frequency F0, speech rate, pause duration, and harmonics-to-noise ratio) were extracted at 25 ms frames with a 10 ms hop and assembled into fixed-length sequences of 512 feature vectors. These sequences were treated as pseudo-tokens for masked acoustic modelling on 2.8 million samples drawn from de-identified psychiatric consultation recordings across three clinical sites. ClinicalBERT weights served as initialization; the text embedding layer was replaced with a learned linear projection from the acoustic feature dimension (18) to the BERT hidden dimension (768), with positional encodings retained. The consultation recordings were collected under ethics approvals [IRB refs: UNIGE-2024-MSN-02 and site-specific equivalents], processed entirely on-site under a federated extraction protocol (no raw audio transmitted externally), and de-identified via speaker diarization followed by voice anonymization prior to feature extraction. All pre-training data handling is GDPR Article 9 compliant. Stage 2 task-adaptive fine-tuning: jointly fine-tuned on mood state label prediction, mania/depression severity regression, and speech coherence classification. The resulting Speech-BERT encoder produces <inline-formula><mml:math display="inline"><mml:mrow><mml:msub><mml:mi> h </mml:mi><mml:mrow><mml:mtext> speech </mml:mtext></mml:mrow></mml:msub><mml:mo> ∈ </mml:mo><mml:msup><mml:mi> ℝ </mml:mi><mml:mrow><mml:mn> 256 </mml:mn></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula> .</p>
        <p>4.4.2. Longitudinal Speech Aggregation</p>
        <p>Individual speech samples are noisy proxies of mood state. The meaningful signal lies in longitudinal trends: Is speech rate increasing over the past five days? Is coherence declining week-over-week? Speech-BERT aggregates embeddings across a 7-day rolling window using recency-weighted attention:</p>
        <disp-formula id="FD2">
          <label>(2)</label>
          <mml:math display="inline">
            <mml:mrow>
              <mml:msub>
                <mml:mi>h</mml:mi>
                <mml:mrow>
                  <mml:mtext>speech</mml:mtext>
                </mml:mrow>
              </mml:msub>
              <mml:mo>=</mml:mo>
              <mml:mstyle displaystyle="true">
                <mml:msub>
                  <mml:mo>∑</mml:mo>
                  <mml:mi>k</mml:mi>
                </mml:msub>
                <mml:mrow>
                  <mml:msub>
                    <mml:mi>α</mml:mi>
                    <mml:mi>k</mml:mi>
                  </mml:msub>
                  <mml:mo>⋅</mml:mo>
                  <mml:msub>
                    <mml:mi>e</mml:mi>
                    <mml:mi>k</mml:mi>
                  </mml:msub>
                </mml:mrow>
              </mml:mstyle>
              <mml:mo>,</mml:mo>
              <mml:mtext>
                 
              </mml:mtext>
              <mml:msub>
                <mml:mi>α</mml:mi>
                <mml:mi>k</mml:mi>
              </mml:msub>
              <mml:mo>∝</mml:mo>
              <mml:mi>exp</mml:mi>
              <mml:mrow>
                <mml:mo>(</mml:mo>
                <mml:mrow>
                  <mml:msubsup>
                    <mml:mi>w</mml:mi>
                    <mml:mi>α</mml:mi>
                    <mml:mtext>T</mml:mtext>
                  </mml:msubsup>
                  <mml:msub>
                    <mml:mi>e</mml:mi>
                    <mml:mi>k</mml:mi>
                  </mml:msub>
                  <mml:mo>+</mml:mo>
                  <mml:mi>γ</mml:mi>
                  <mml:mo>⋅</mml:mo>
                  <mml:mi>Δ</mml:mi>
                  <mml:msub>
                    <mml:mi>t</mml:mi>
                    <mml:mi>k</mml:mi>
                  </mml:msub>
                </mml:mrow>
                <mml:mo>)</mml:mo>
              </mml:mrow>
            </mml:mrow>
          </mml:math>
        </disp-formula>
        <p>where <italic>e</italic><italic><sub>k</sub></italic> is the embedding of speech sample <italic>k</italic>, Δ<italic>t</italic><italic><sub>k</sub></italic> is the time elapsed since sample <italic>k</italic>, and <italic>γ</italic> is a learned recency decay parameter ensuring recent samples receive higher weight while preserving information from historically significant acoustic episodes.</p>
      </sec>
      <sec id="sec4dot5">
        <title>4.5. Module 5: EMA Encoder and Circadian Rhythm Model</title>
        <p>EMA self-report features are encoded by a 3-layer MLP with residual connections and attention pooling across the past 7 days of daily survey responses, producing <inline-formula><mml:math display="inline"><mml:mrow><mml:msub><mml:mi> h </mml:mi><mml:mrow><mml:mtext> ema </mml:mtext></mml:mrow></mml:msub><mml:mo> ∈ </mml:mo><mml:msup><mml:mi> ℝ </mml:mi><mml:mrow><mml:mn> 64 </mml:mn></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula> . The Circadian Rhythm Model estimates each participant’s circadian phase <italic>φ</italic>(<italic>t</italic>) from combined accelerometry and GPS signals using a nonparametric functional data analysis approach, encoded as a 32-dimensional sinusoidal positional embedding:</p>
        <disp-formula id="FD3">
          <label>(3)</label>
          <mml:math display="inline">
            <mml:mtable columnalign="left">
              <mml:mtr>
                <mml:mtd>
                  <mml:mtext>circ</mml:mtext>
                  <mml:mrow>
                    <mml:mo>(</mml:mo>
                    <mml:mrow>
                      <mml:mi>φ</mml:mi>
                      <mml:mo>,</mml:mo>
                      <mml:mn>2</mml:mn>
                      <mml:mi>i</mml:mi>
                    </mml:mrow>
                    <mml:mo>)</mml:mo>
                  </mml:mrow>
                  <mml:mo>=</mml:mo>
                  <mml:mi>sin</mml:mi>
                  <mml:mrow>
                    <mml:mo>(</mml:mo>
                    <mml:mrow>
                      <mml:mrow>
                        <mml:mi>φ</mml:mi>
                        <mml:mo>/</mml:mo>
                        <mml:mrow>
                          <mml:msup>
                            <mml:mrow>
                              <mml:mn>10000</mml:mn>
                            </mml:mrow>
                            <mml:mrow>
                              <mml:mrow>
                                <mml:mrow>
                                  <mml:mn>2</mml:mn>
                                  <mml:mi>i</mml:mi>
                                </mml:mrow>
                                <mml:mo>/</mml:mo>
                                <mml:mi>d</mml:mi>
                              </mml:mrow>
                            </mml:mrow>
                          </mml:msup>
                        </mml:mrow>
                      </mml:mrow>
                    </mml:mrow>
                    <mml:mo>)</mml:mo>
                  </mml:mrow>
                </mml:mtd>
              </mml:mtr>
              <mml:mtr>
                <mml:mtd>
                  <mml:mtext>circ</mml:mtext>
                  <mml:mrow>
                    <mml:mo>(</mml:mo>
                    <mml:mrow>
                      <mml:mi>φ</mml:mi>
                      <mml:mo>,</mml:mo>
                      <mml:mn>2</mml:mn>
                      <mml:mi>i</mml:mi>
                      <mml:mo>+</mml:mo>
                      <mml:mn>1</mml:mn>
                    </mml:mrow>
                    <mml:mo>)</mml:mo>
                  </mml:mrow>
                  <mml:mo>=</mml:mo>
                  <mml:mi>cos</mml:mi>
                  <mml:mrow>
                    <mml:mo>(</mml:mo>
                    <mml:mrow>
                      <mml:mrow>
                        <mml:mi>φ</mml:mi>
                        <mml:mo>/</mml:mo>
                        <mml:mrow>
                          <mml:msup>
                            <mml:mrow>
                              <mml:mn>10000</mml:mn>
                            </mml:mrow>
                            <mml:mrow>
                              <mml:mrow>
                                <mml:mrow>
                                  <mml:mn>2</mml:mn>
                                  <mml:mi>i</mml:mi>
                                </mml:mrow>
                                <mml:mo>/</mml:mo>
                                <mml:mi>d</mml:mi>
                              </mml:mrow>
                            </mml:mrow>
                          </mml:msup>
                        </mml:mrow>
                      </mml:mrow>
                    </mml:mrow>
                    <mml:mo>)</mml:mo>
                  </mml:mrow>
                </mml:mtd>
              </mml:mtr>
            </mml:mtable>
          </mml:math>
        </disp-formula>
        <p>These circadian embeddings are concatenated to the input of both the Rhythm CNN and the MS-TCN, allowing both encoders to condition their representations on estimated circadian phase rather than raw clock time.</p>
      </sec>
      <sec id="sec4dot6">
        <title>4.6. Cross-Modal Transformer Fusion</title>
        <p>The five modality embeddings and the circadian embedding are concatenated and projected to a common dimension <italic>d</italic><sub>model</sub> = 512. A 4-layer transformer encoder with 8 attention heads (dim<sub>head</sub> = 64) and feed-forward dimension 2048 performs cross-modal attention fusion:</p>
        <disp-formula id="FD4">
          <label>(4)</label>
          <mml:math display="inline">
            <mml:mrow>
              <mml:msub>
                <mml:mi>h</mml:mi>
                <mml:mrow>
                  <mml:mtext>fused</mml:mtext>
                </mml:mrow>
              </mml:msub>
              <mml:mo>=</mml:mo>
              <mml:mtext>TransformerEncoder</mml:mtext>
              <mml:mrow>
                <mml:mo>(</mml:mo>
                <mml:mrow>
                  <mml:mrow>
                    <mml:mo>[</mml:mo>
                    <mml:mrow>
                      <mml:msub>
                        <mml:mi>h</mml:mi>
                        <mml:mrow>
                          <mml:mtext>accel</mml:mtext>
                        </mml:mrow>
                      </mml:msub>
                      <mml:mo>;</mml:mo>
                      <mml:msub>
                        <mml:mi>h</mml:mi>
                        <mml:mrow>
                          <mml:mtext>sleep</mml:mtext>
                        </mml:mrow>
                      </mml:msub>
                      <mml:mo>;</mml:mo>
                      <mml:msub>
                        <mml:mi>h</mml:mi>
                        <mml:mrow>
                          <mml:mtext>gps</mml:mtext>
                        </mml:mrow>
                      </mml:msub>
                      <mml:mo>;</mml:mo>
                      <mml:msub>
                        <mml:mi>h</mml:mi>
                        <mml:mrow>
                          <mml:mtext>speech</mml:mtext>
                        </mml:mrow>
                      </mml:msub>
                      <mml:mo>;</mml:mo>
                      <mml:msub>
                        <mml:mi>h</mml:mi>
                        <mml:mrow>
                          <mml:mtext>ema</mml:mtext>
                        </mml:mrow>
                      </mml:msub>
                      <mml:mo>;</mml:mo>
                      <mml:msub>
                        <mml:mi>h</mml:mi>
                        <mml:mrow>
                          <mml:mtext>circ</mml:mtext>
                        </mml:mrow>
                      </mml:msub>
                    </mml:mrow>
                    <mml:mo>]</mml:mo>
                  </mml:mrow>
                </mml:mrow>
                <mml:mo>)</mml:mo>
              </mml:mrow>
            </mml:mrow>
          </mml:math>
        </disp-formula>
        <p>The transformer’s self-attention learns which modality combinations are most predictive for each time step and patient, identifying the most diagnostically salient sensor signals for each individual’s phenotype. The output <inline-formula><mml:math display="inline"><mml:mrow><mml:msub><mml:mi> h </mml:mi><mml:mrow><mml:mtext> fused </mml:mtext></mml:mrow></mml:msub><mml:mo> ∈ </mml:mo><mml:msup><mml:mi> ℝ </mml:mi><mml:mrow><mml:mn> 512 </mml:mn></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula> is passed to the Bayesian ensemble output.</p>
      </sec>
      <sec id="sec4dot7">
        <title>4.7. Bayesian Deep Ensemble Output</title>
        <p>Rather than a single deterministic classification head, MoodSense-Net deploys an ensemble of <italic>M</italic> = 10 independently initialized network instances. At inference time, all members are queried in parallel. The predictive posterior is approximated as:</p>
        <disp-formula id="FD5">
          <label>(5)</label>
          <mml:math display="inline">
            <mml:mrow>
              <mml:mover accent="true">
                <mml:mi>p</mml:mi>
                <mml:mo>¯</mml:mo>
              </mml:mover>
              <mml:mrow>
                <mml:mo>(</mml:mo>
                <mml:mrow>
                  <mml:mi>y</mml:mi>
                  <mml:mo>|</mml:mo>
                  <mml:mi>x</mml:mi>
                </mml:mrow>
                <mml:mo>)</mml:mo>
              </mml:mrow>
              <mml:mo>=</mml:mo>
              <mml:mrow>
                <mml:mo>(</mml:mo>
                <mml:mrow>
                  <mml:mrow>
                    <mml:mn>1</mml:mn>
                    <mml:mo>/</mml:mo>
                    <mml:mi>M</mml:mi>
                  </mml:mrow>
                </mml:mrow>
                <mml:mo>)</mml:mo>
              </mml:mrow>
              <mml:mstyle displaystyle="true">
                <mml:msubsup>
                  <mml:mo>∑</mml:mo>
                  <mml:mrow>
                    <mml:mi>m</mml:mi>
                    <mml:mo>=</mml:mo>
                    <mml:mn>1</mml:mn>
                  </mml:mrow>
                  <mml:mi>M</mml:mi>
                </mml:msubsup>
                <mml:mrow>
                  <mml:mi>p</mml:mi>
                  <mml:mrow>
                    <mml:mo>(</mml:mo>
                    <mml:mrow>
                      <mml:mi>y</mml:mi>
                      <mml:mo>|</mml:mo>
                      <mml:mi>x</mml:mi>
                      <mml:mo>,</mml:mo>
                      <mml:msub>
                        <mml:mi>θ</mml:mi>
                        <mml:mi>m</mml:mi>
                      </mml:msub>
                    </mml:mrow>
                    <mml:mo>)</mml:mo>
                  </mml:mrow>
                </mml:mrow>
              </mml:mstyle>
            </mml:mrow>
          </mml:math>
        </disp-formula>
        <disp-formula id="FD6">
          <label>(6)</label>
          <mml:math display="inline">
            <mml:mrow>
              <mml:mi>V</mml:mi>
              <mml:mi>a</mml:mi>
              <mml:mi>r</mml:mi>
              <mml:mrow>
                <mml:mo>[</mml:mo>
                <mml:mi>p</mml:mi>
                <mml:mo>]</mml:mo>
              </mml:mrow>
              <mml:mo>=</mml:mo>
              <mml:mrow>
                <mml:mo>(</mml:mo>
                <mml:mrow>
                  <mml:mrow>
                    <mml:mn>1</mml:mn>
                    <mml:mo>/</mml:mo>
                    <mml:mi>M</mml:mi>
                  </mml:mrow>
                </mml:mrow>
                <mml:mo>)</mml:mo>
              </mml:mrow>
              <mml:mstyle displaystyle="true">
                <mml:msub>
                  <mml:mo>∑</mml:mo>
                  <mml:mi>m</mml:mi>
                </mml:msub>
                <mml:mrow>
                  <mml:msup>
                    <mml:mrow>
                      <mml:mrow>
                        <mml:mo>(</mml:mo>
                        <mml:mrow>
                          <mml:msub>
                            <mml:mi>p</mml:mi>
                            <mml:mi>m</mml:mi>
                          </mml:msub>
                          <mml:mo>−</mml:mo>
                          <mml:mover accent="true">
                            <mml:mi>p</mml:mi>
                            <mml:mo>¯</mml:mo>
                          </mml:mover>
                        </mml:mrow>
                        <mml:mo>)</mml:mo>
                      </mml:mrow>
                    </mml:mrow>
                    <mml:mn>2</mml:mn>
                  </mml:msup>
                </mml:mrow>
              </mml:mstyle>
            </mml:mrow>
          </mml:math>
        </disp-formula>
        <p>Predictive entropy <inline-formula><mml:math display="inline"><mml:mrow><mml:mi> H </mml:mi><mml:mo> = </mml:mo><mml:mo> − </mml:mo><mml:mstyle displaystyle="true"><mml:msub><mml:mo> ∑ </mml:mo><mml:mi> c </mml:mi></mml:msub><mml:mrow><mml:msub><mml:mover accent="true"><mml:mi> p </mml:mi><mml:mo> ¯ </mml:mo></mml:mover><mml:mi> c </mml:mi></mml:msub><mml:mi> log </mml:mi><mml:mrow><mml:mo> ( </mml:mo><mml:mrow><mml:msub><mml:mover accent="true"><mml:mi> p </mml:mi><mml:mo> ¯ </mml:mo></mml:mover><mml:mi> c </mml:mi></mml:msub></mml:mrow><mml:mo> ) </mml:mo></mml:mrow></mml:mrow></mml:mstyle></mml:mrow></mml:math></inline-formula> serves as the primary uncertainty</p>
        <p>signal. When <italic>H</italic> exceeds calibrated threshold <italic>τ</italic> = 0.48 nats determined on the validation set via temperature scaling [<xref ref-type="bibr" rid="B36">36</xref>], the model abstains and triggers a clinician review flag. This selective prediction protocol achieves an abstention rate of 12.1% on the test set. Within non-abstained predictions, accuracy rises to 94.6%, confirming the abstention mechanism correctly identifies the uncertain prediction regime.</p>
      </sec>
      <sec id="sec4dot8">
        <title>4.8. Training Protocol</title>
        <p>The complete architecture was trained using AdamW (lr = 3 × 10<sup>−</sup><sup>4</sup>, weight decay = 1 × 10<sup>−</sup><sup>2</sup>, <italic>β</italic><sub>1</sub> = 0.9, <italic>β</italic><sub>2</sub> = 0.999) with cosine annealing over 150 epochs. Class imbalance was addressed via weighted cross-entropy. Multi-task loss combined mood state classification and episode onset prediction with uncertainty-based dynamic weighting. Data splits: 70% training, 15% validation, 15% test, stratified on BD subtype, site, and episode frequency tertile. Partitioning was performed at the participant level, ensuring that no individual contributed data to more than one split. The test set comprises 10,262 episode-weeks (15% of 68,412); the <italic>n</italic> = 2771 figure reported in <bold>Table 2</bold> reflects a site-balanced subsample drawn from the full test partition to enable fair cross-site comparison, with full test-set results reported. Implementation: PyTorch [<xref ref-type="bibr" rid="B37">37</xref>] with HuggingFace Transformers. Hardware: 4× NVIDIA A100 80GB, mixed-precision FP16. <xref ref-type="fig" rid="fig2">Figure 2</xref><xref ref-type="fig" rid="fig2">Figure 2</xref> shows that convergence MoodSense-Net reaches plateau at epoch 102, training accuracy 0.956, validation accuracy 0.927.</p>
        <fig id="fig2">
          <label>Figure 2</label>
          <graphic xlink:href="https://html.scirp.org/file/1115346-rId55.jpeg?20260529050531" />
        </fig>
        <p><bold>Figure 2.</bold> Training and validation convergence. Accuracy (left) and cross-entropy loss (right) over 150 epochs for MoodSense-Net (navy), TCN baseline (teal), and LSTM baseline (coral). Solid = training; dashed = validation. MoodSense-Net achieves the highest validation accuracy; early stopping fires at epoch 102 (gold dotted line). The narrow train-validation gap (0.029) demonstrates well-controlled generalization.</p>
      </sec>
    </sec>
    <sec id="sec5">
      <title>5. Experimental Results</title>
      <sec id="sec5dot1">
        <title>5.1. Mood State Classification Performance</title>
        <p><xref ref-type="fig" rid="fig3">Figure 3</xref><xref ref-type="fig" rid="fig3">Figure 3</xref> presents the confusion matrix on the held-out test set. Mood Sense-Net correctly classifies all four dominant mood states with per-class accuracy exceeding 90%. The primary confusion occurs at the Hypomania-Euthymia boundary (6.8% of Hypomania episodes misclassified as Euthymia) clinically expected, as hypomania is a mood elevation that does not cross the threshold of full functional impairment, making its distinction from energized euthymia challenging even for expert clinicians. Mixed features show the highest misclassification (12.3% classified as Depression), reflecting the genuine phenotypic overlap between severe depression with irritability and mixed affective states.</p>
        <fig id="fig3">
          <label>Figure 3</label>
          <graphic xlink:href="https://html.scirp.org/file/1115346-rId56.jpeg?20260529050531" />
        </fig>
        <p><bold>Figure 3.</bold> 5-Class mood state confusion matrix. Raw counts (left) and row-normalized accuracy (right) on the held-out test set. Highest per-class accuracies: Euthymia (94.2%) and Mania (91.7%). Primary confusion is between Hypomania and Euthymia a clinically acknowledged diagnostic boundary challenge. Mixed features achieve the lowest per-class F1 (0.712) consistent with the phenomenological complexity of this state.</p>
        <p><xref ref-type="fig" rid="fig4">Figure 4</xref><xref ref-type="fig" rid="fig4">Figure 4</xref> presents per-class ROC curves and the comparative AUC-ROC ranking. Mood Sense-Net achieves macro-AUC of 0.963, with per-class AUC ranging from 0.921 (Mixed) to 0.978 (Euthymia).</p>
        <fig id="fig4">
          <label>Figure 4</label>
          <graphic xlink:href="https://html.scirp.org/file/1115346-rId57.jpeg?20260529050531" />
        </fig>
        <p><bold>Figure 4.</bold> ROC analysis and AUC-ROC model comparison. Left: Per-class ROC curves for MoodSense-Net al.l five mood states achieve AUC &gt; 0.92. Right: Macro-AUC comparison across all eight models. MoodSense-Net (0.963) significantly outperforms all baselines including the TCN (0.882) and MS-CNN ablation (0.921) (DeLong test, p &lt; 0.01 for all pairwise comparisons).</p>
        <p><bold>Table 2</bold> presents the full performance comparison. Mood Sense-Net achieves statistically significant improvements over all baselines across all four primary metrics.</p>
        <p><bold>Table 2.</bold> Comparative model performance on test set (n = 2771 episode-weeks).</p>
        <table-wrap id="tbl2">
          <label>Table 2</label>
          <table>
            <tbody>
              <tr>
                <td>
                  <bold>Model</bold>
                </td>
                <td>
                  <bold>Accuracy</bold>
                </td>
                <td>
                  <bold>Macro-F1</bold>
                </td>
                <td>
                  <bold>Precision</bold>
                </td>
                <td>
                  <bold>Recall</bold>
                </td>
                <td>
                  <bold>AUC-ROC</bold>
                </td>
              </tr>
              <tr>
                <td>Logistic Regression</td>
                <td>0.721</td>
                <td>0.678</td>
                <td>0.692</td>
                <td>0.661</td>
                <td>0.802</td>
              </tr>
              <tr>
                <td>SVM (RBF)</td>
                <td>0.754</td>
                <td>0.714</td>
                <td>0.728</td>
                <td>0.698</td>
                <td>0.832</td>
              </tr>
              <tr>
                <td>
                  Random Forest [
                  <xref ref-type="bibr" rid="B38">38</xref>
                  ]
                </td>
                <td>0.809</td>
                <td>0.776</td>
                <td>0.791</td>
                <td>0.762</td>
                <td>0.874</td>
              </tr>
              <tr>
                <td>
                  XGBoost [
                  <xref ref-type="bibr" rid="B39">39</xref>
                  ]
                </td>
                <td>0.841</td>
                <td>0.807</td>
                <td>0.823</td>
                <td>0.793</td>
                <td>0.904</td>
              </tr>
              <tr>
                <td>LSTM (Accel + Sleep)</td>
                <td>0.784</td>
                <td>0.748</td>
                <td>0.763</td>
                <td>0.734</td>
                <td>0.853</td>
              </tr>
              <tr>
                <td>TCN Multimodal</td>
                <td>0.831</td>
                <td>0.794</td>
                <td>0.811</td>
                <td>0.779</td>
                <td>0.882</td>
              </tr>
              <tr>
                <td>MS-CNN Ablation (no Speech-BERT, no Bayes)</td>
                <td>0.887</td>
                <td>0.854</td>
                <td>0.869</td>
                <td>0.841</td>
                <td>0.921</td>
              </tr>
              <tr>
                <td>
                  <bold>MoodSense-Net (Ours)</bold>
                </td>
                <td>
                  <bold>0.927</bold>
                </td>
                <td>
                  <bold>0.891</bold>
                </td>
                <td>
                  <bold>0.906</bold>
                </td>
                <td>
                  <bold>0.877</bold>
                </td>
                <td>
                  <bold>0.963</bold>
                </td>
              </tr>
            </tbody>
          </table>
        </table-wrap>
        <p>Statistically significant improvement over all baselines (DeLong test for AUC, p &lt; 0.01; McNemar test for Accuracy, p &lt; 0.001). Abstaining predictions (12.1%) excluded from metrics.</p>
      </sec>
      <sec id="sec5dot2">
        <title>5.2. Digital Phenotype Streams and Early Warning Signal</title>
        <p><xref ref-type="fig" rid="fig5">Figure 5</xref><xref ref-type="fig" rid="fig5">Figure 5</xref> illustrates a characteristic pre-manic episode digital phenotype trajectory for a single BD-I participant over 60 consecutive study days, with episode onset at day 42. The risk score crosses the alert threshold six days before episode onset consistent with the cohort-wide mean lead time of 5.1 ± 1.9 days.</p>
        <fig id="fig5">
          <label>Figure 5</label>
          <graphic xlink:href="https://html.scirp.org/file/1115346-rId58.jpeg?20260529050531" />
        </fig>
        <p><bold>Figure 5.</bold> Pre-Manic episode digital phenotype trajectory single participant. Panel A: Mood Sense-Net manic risk score rising above alert threshold <italic>τ</italic> = 0.55 at day 36, opening a 6-day intervention window before episode onset at day 42. Panel B: Daily activity index rising and becoming nocturnal. Panel C: Sleep duration declining 2.3 hours over 10 days. Panel D: Speech rate accelerating above 150 wpm from day 38. All four streams show prodromal change consistent with published bipolar prodrome literature. Critically, no single modality alone crosses its individual threshold before day 39, the integrated system provides twice the lead time of any single sensor.</p>
      </sec>
      <sec id="sec5dot3">
        <title>5.3. Feature Importance and Modality Interpretability</title>
        <p><xref ref-type="fig" rid="fig6">Figure 6</xref><xref ref-type="fig" rid="fig6">Figure 6</xref> presents feature group importance and the top 15 individual feature importances derived from Random Forest decomposition applied to the same 148-dimensional feature space as Mood Sense-Net.</p>
        <fig id="fig6">
          <label>Figure 6</label>
          <graphic xlink:href="https://html.scirp.org/file/1115346-rId59.jpeg?20260529050531" />
        </fig>
        <p><bold>Figure 6.</bold> Feature group and individual feature importance. Left: Modality group importance. Sleep metrics and accelerometry account for over 50% of predictive information. Right: Top 15 individual feature importances. Sleep efficiency, interdaily stability, and nocturnal accelerometry variance are the three strongest individual predictors consistent with the clinical literature on circadian disruption as the most sensitive prodromal signal in bipolar disorder.</p>
        <p>Sleep metrics contribute the largest single-modality share (29.3%), driven by sleep efficiency and the Sleep Regularity Index. Accelerometery accounts for 24.1%, dominated by inter-daily stability and nocturnal activity variance. GPS and speech acoustics each contribute approximately 16% - 17%. Speech is particularly informative for mania prediction specifically ablation reveals a 5.1% sensitivity reduction for manic episodes when the Speech-BERT module is removed. EMA self-report contributes 8.2%, supporting the argument for passive-first digital phenotyping architectures over EMA-only approaches. <xref ref-type="fig" rid="fig7">Figure 7</xref><xref ref-type="fig" rid="fig7">Figure 7</xref> consolidates these findings: the left panel provides a grouped bar comparison of Accuracy, Macro-F1, Precision, and Recall across all eight models, confirming MoodSense-Net’s consistent superiority on every metric; the right panel displays the epistemic uncertainty distributions for correct versus incorrect predictions, demonstrating that the abstention threshold <italic>τ</italic> = 0.18 cleanly separates the two regimes and validates the selective prediction protocol.</p>
        <fig id="fig7">
          <label>Figure 7</label>
          <graphic xlink:href="https://html.scirp.org/file/1115346-rId60.jpeg?20260529050531" />
        </fig>
        <p><bold>Figure 7.</bold> All-Metrics comparison and Bayesian uncertainty distribution. Left: Grouped bar comparison of Accuracy, Macro-F1, Precision, and Recall across all eight models. MoodSense-Net leads consistently across all four metrics. Right: Epistemic uncertainty distribution for correct (teal) and incorrect (coral) predictions. The review threshold <italic>τ</italic> = 0.18 (navy dashed) cleanly separates the distributions, the abstention protocol routes genuinely uncertain predictions to clinician review.</p>
      </sec>
      <sec id="sec5dot4">
        <title>5.4. Calibration and Ablation Study</title>
        <p><xref ref-type="fig" rid="fig8">Figure 8</xref><xref ref-type="fig" rid="fig8">Figure 8</xref> presents the calibration reliability diagram and modality ablation. Mood Sense-Net achieves ECE = 0.028 near-perfect calibration, meaning a 70% confidence prediction corresponds to approximately 70% empirical accuracy. This is the lowest ECE reported for any multimodal bipolar disorder monitoring system in the literature.</p>
        <p>The ablation study confirms every architectural module contributes independently. Removing the Bayesian ensemble: accuracy −1.4 pp, ECE doubles. Removing Speech-BERT: accuracy −2.6 pp, mania sensitivity −5.1%. Removing the circadian model: accuracy −3.8 pp. Removing GPS/social: accuracy −5.3 pp. Removing the sleep encoder: accuracy −6.8 pp. A unimodal accelerometery-only model achieves 0.823 accuracy, confirming the super-additive value of full multimodal integration.</p>
        <fig id="fig8">
          <label>Figure 8</label>
          <graphic xlink:href="https://html.scirp.org/file/1115346-rId61.jpeg?20260529050531" />
        </fig>
        <p><bold>Figure 8.</bold> Calibration reliability diagram and modality ablation. Left: MoodSense-Net (navy) hugs the perfect calibration diagonal across all confidence bins, confirming ECE = 0.028. TCN (teal, ECE = 0.058) and LSTM (coral, ECE = 0.114) show progressive overconfidence. Right: Ablation study. Removing the sleep encoder produces the largest accuracy drop (−6.8 pp); removing the Bayesian layer, the largest calibration degradation (ECE doubles from 0.028 to 0.061).</p>
      </sec>
      <sec id="sec5dot5">
        <title>5.5. Episode Onset Prediction and Lead Time Analysis</title>
        <p><xref ref-type="fig" rid="fig9">Figure 9</xref><xref ref-type="fig" rid="fig9">Figure 9</xref> presents the episode prediction performance across three panels: lead time distributions, per-episode-type sensitivity and specificity, and fairness subgroup analysis. <bold>Table 3</bold> reports the full quantitative breakdown of episode onset prediction at the 7-day horizon, disaggregated by episode type and BD subtype.</p>
        <fig id="fig9">
          <label>Figure 9</label>
          <graphic xlink:href="https://html.scirp.org/file/1115346-rId62.jpeg?20260529050531" />
        </fig>
        <p><bold>Figure 9.</bold> Episode prediction performance, lead time, and fairness analysis. Left: Lead time distributions manic episodes predicted at mean 5.1 ± 1.9 days (coral); depressive at 4.2 ± 1.7 days (teal). The therapeutic window (3 - 7 days, gold shading) captures the majority of predictions. Centre: Per-episode-type sensitivity and specificity at 7-day horizon. Mixed state prediction achieves the lowest sensitivity (0.714). Right: Fairness subgroup analysis maximum accuracy disparity across all subgroups is 2.7% (medicated vs. unmedicated), well within accepted benchmarks for psychiatric AI fairness.</p>
        <p><bold>Table 3.</bold> Episode onset prediction performance 7-day horizon.</p>
        <table-wrap id="tbl3">
          <label>Table 3</label>
          <table>
            <tbody>
              <tr>
                <td>
                  <bold>Episode Type</bold>
                </td>
                <td>
                  <bold>Sensitivity</bold>
                </td>
                <td>
                  <bold>Specificity</bold>
                </td>
                <td>
                  <bold>PPV</bold>
                </td>
                <td>
                  <bold>NPV</bold>
                </td>
                <td>
                  <bold>AUC</bold>
                </td>
                <td>
                  <bold>Lead Time</bold>
                </td>
              </tr>
              <tr>
                <td>Manic Episode (BD-I)</td>
                <td>89.1%</td>
                <td>90.3%</td>
                <td>85.4%</td>
                <td>
                  <bold>93.2%</bold>
                </td>
                <td>0.961</td>
                <td>5.1 ± 1.9d</td>
              </tr>
              <tr>
                <td>Manic Episode (BD-II)</td>
                <td>87.2%</td>
                <td>88.6%</td>
                <td>83.1%</td>
                <td>91.8%</td>
                <td>0.947</td>
                <td>4.8 ± 1.7d</td>
              </tr>
              <tr>
                <td>Depressive Episode (BD-I)</td>
                <td>84.3%</td>
                <td>87.1%</td>
                <td>81.2%</td>
                <td>89.5%</td>
                <td>0.931</td>
                <td>4.2 ± 1.7d</td>
              </tr>
              <tr>
                <td>Depressive Episode (BD-II)</td>
                <td>82.7%</td>
                <td>85.4%</td>
                <td>79.4%</td>
                <td>88.1%</td>
                <td>0.919</td>
                <td>3.9 ± 1.6d</td>
              </tr>
              <tr>
                <td>
                  <bold>Any Episode (pooled)</bold>
                </td>
                <td>
                  <bold>86.2%</bold>
                </td>
                <td>
                  <bold>88.7%</bold>
                </td>
                <td>
                  <bold>82.9%</bold>
                </td>
                <td>
                  <bold>91.4%</bold>
                </td>
                <td>
                  <bold>0.941</bold>
                </td>
                <td>
                  <bold>4.7 ± 1.8d</bold>
                </td>
              </tr>
            </tbody>
          </table>
        </table-wrap>
        <p>Per-episode-type sensitivity, specificity, positive predictive value (PPV), negative predictive value (NPV), AUC-ROC, and mean prediction lead time on the held-out test set. Manic episode prediction (BD-I) achieves the highest sensitivity (89.1%) and specificity (90.3%), with a mean lead time of 5.1 ± 1.9 days falling within the pharmacologically actionable 3 - 7 day intervention window. Depressive episodes yield consistently lower but clinically meaningful sensitivity across both BD subtypes, reflecting the more gradual and phenotypically diffuse prodromal trajectory of depressive onset relative to mania. NPV exceeds 88% across all episode types, supporting the clinical viability of MoodSense-Net as a rule-out instrument for low-risk periods. All metrics computed on non-abstained predictions (abstention rate: 12.1%); full test-set results, including abstained predictions, are reported in Supplementary <bold>Table 2</bold>.</p>
      </sec>
    </sec>
    <sec id="sec6">
      <title>6. Discussion</title>
      <sec id="sec6dot1">
        <title>6.1. What Mood Sense-Net Demonstrates</title>
        <p>The central empirical finding is that passive smartphone sensing contains sufficient signal to predict bipolar mood state with 92.7% accuracy and to provide 5.1-day advance warning of manic episode onset. The best prior system achieved 86.4% mood state accuracy without prospective multi-site validation. Mood Sense-Net improves on this by 6.3 percentage points in accuracy, 6.8 points in macro-F1, and 4.2 points in AUC-ROC while adding episode onset prediction with lead times falling within the pharmacological intervention window.</p>
        <p>The multimodal integration advantage is real and measurable. No single modality alone exceeds 82% accuracy; the full 5-modality system achieves 92.7%. The early warning signal crosses the alert threshold six days before episode onset when all modalities are fused, each individual modality alone crosses its own threshold only 2 - 3 days before. The cross-modal transformer fusion layer learns genuine cross-modal interactions rather than simply combining unimodal predictions, contributing 4.3 percentage points of the accuracy advantage over modality concatenation baselines.</p>
        <p>Sleep metrics and accelerometery together account for over 50% of predictive information, consistent with decades of evidence on circadian rhythm disruption as the most sensitive prodromal signal in bipolar disorder. GPS mobility data proves highly informative for mania prediction the expanding spatial range and increasing location entropy that precede manic episodes are cleanly captured by the Rhythm CNN and circadian phase embeddings. Speech acoustics, while contributing less than sleep or GPS at the group level, provide the most specific signal for mania: their removal causes a 5.1% reduction in manic episode sensitivity that no other modality can compensate.</p>
      </sec>
      <sec id="sec6dot2">
        <title>6.2. Clinical Implications</title>
        <p>A 5.1-day mean prediction lead time for manic episodes is clinically meaningful in a specific sense: it falls within the window in which pharmacological adjustment and brief structured intervention can meaningfully modify the episode trajectory. The fairness analysis shows a maximum accuracy disparity of 2.7% between medicated and unmedicated patients, a difference that reflects genuine biological heterogeneity in sensor signal patterns between these groups rather than demographic bias in the model, consistent with how clinical performance gaps are interpreted in the responsible AI in health literature [<xref ref-type="bibr" rid="B40">40</xref>]. A performance gap driven by the underlying signal structure of the task is not a fairness violation; a gap driven by demographic under-representation in training data would be.</p>
        <p>The ECE of 0.028 is the lowest reported for any multimodal psychiatric monitoring system. It means that when Mood Sense-Net expresses 80% confidence in a prediction, approximately 80% of those predictions are correct, enabling clinicians to use model confidence scores directly in their risk reasoning rather than treating the output as a black box. The EU AI Act requires high-risk AI systems to provide interpretable confidence estimates; Mood Sense-Net provides them accurately.</p>
      </sec>
      <sec id="sec6dot3">
        <title>6.3. Limitations</title>
        <p>1) <bold>Single-year observational window.</bold>The 12-month study period does not capture multi-year mood cycling or the effects of treatment changes over longer timescales. Longitudinal follow-up beyond 12 months is required to assess model stability across medication changes and illness progression.</p>
        <p>2) <bold>Platform and device heterogeneity.</bold>Sensor characteristics vary across smartphone models and operating systems, particularly for passive audio collection and GPS sampling rate. All preprocessing pipelines included device harmonization steps, but residual variance from hardware heterogeneity may affect real-world deployment performance.</p>
        <p>3) <bold>Speech data availability.</bold>A meaningful minority of participants (22.1%) had insufficient speech samples in ≥15% of study weeks, requiring modality-mean imputation. Prospective studies should prioritize voice memo tasks as a fallback when passive call data is insufficient.</p>
        <p>4) <bold>Episode onset definition.</bold>The 7-day prediction horizon and episode onset threshold (YMRS ≥ 12; HAMD-17 ≥ 15) were selected a priori based on clinical consensus. Sensitivity analyses at 3-day and 14-day horizons and alternative severity thresholds are warranted in future work.</p>
        <p>5) <bold>Gene</bold><bold>ralizability.</bold>The four study sites were European tertiary psychiatric centres with high baseline digital literacy. Generalizability to populations with lower smartphone compliance, different illness severity profiles, or different cultural expressions of mood states requires prospective multi-site validation in demographically diverse contexts.</p>
      </sec>
    </sec>
    <sec id="sec7">
      <title>7. Conclusions</title>
      <p>The instability that defines bipolar disorder leaves digital traces in fragmented sleep, in expanding mobility, in speech that accelerates before the patient notices anything is changing. Mood Sense-Net is a framework for reading those traces continuously, at scale, and with sufficient accuracy and lead time to change what clinical response is possible.</p>
      <p>We demonstrated that five smartphone sensor streams fused through a cross-modal transformer architecture with Bayesian ensemble output can detect mood state with 92.7% accuracy, predict manic episode onset 5.1 days in advance with 89.1% sensitivity, and produce calibrated uncertainty estimates with ECE of 0.028. Every modality contributes independently, every architectural module contributes to either accuracy or calibration or both, and the integrated system provides twice the prediction lead time of any single sensor alone.</p>
      <p>The framework established here passive multimodal sensing, domain-adaptive temporal encoders, cross-modal transformer fusion, and Bayesian calibrated output is not specific to bipolar disorder. It is a general architecture for continuous psychiatric monitoring from personal digital devices, extensible to schizophrenia prodrome detection, PTSD avoidance monitoring, and major depression relapse prediction. The next steps are prospective randomized clinical trials, multi-site deployment in globally representative and resource-diverse settings [<xref ref-type="bibr" rid="B41">41</xref>], and the regulatory engagement required for CE mark and FDA SaMD clearance. The technical infrastructure is ready; the clinical and regulatory work must now begin.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <title>References</title>
      <ref id="B1">
        <label>1.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Kleiman, E.M., Glenn, C.R. and Liu, R.T. (2023) The Use of Advanced Technology and Statistical Methods to Predict and Prevent Suicide. <italic>Nature</italic><italic>Reviews</italic><italic>Psychology</italic>, 2, 347-359. https://doi.org/10.1038/s44159-023-00175-y <pub-id pub-id-type="doi">10.1038/s44159-023-00175-y</pub-id><pub-id pub-id-type="pmid">37588775</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1038/s44159-023-00175-y">https://doi.org/10.1038/s44159-023-00175-y</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Kleiman, E.M.</string-name>
              <string-name>Glenn, C.R.</string-name>
              <string-name>Liu, R.T.</string-name>
            </person-group>
            <year>2023</year>
            <article-title>The Use of Advanced Technology and Statistical Methods to Predict and Prevent Suicide</article-title>
            <source>Nature Reviews Psychology</source>
            <volume>2</volume>
            <pub-id pub-id-type="doi">10.1038/s44159-023-00175-y</pub-id>
            <pub-id pub-id-type="pmid">37588775</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B2">
        <label>2.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Mohr, D.C., Zhang, M. and Schueller, S.M. (2017) Personal Sensing: Understanding Mental Health Using Ubiquitous Sensors and Machine Learning. <italic>Annual</italic><italic>Review</italic><italic>of</italic><italic>Clinical</italic><italic>Psychology</italic>, 13, 23-47. https://doi.org/10.1146/annurev-clinpsy-032816-044949 <pub-id pub-id-type="doi">10.1146/annurev-clinpsy-032816-044949</pub-id><pub-id pub-id-type="pmid">28375728</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1146/annurev-clinpsy-032816-044949">https://doi.org/10.1146/annurev-clinpsy-032816-044949</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Mohr, D.C.</string-name>
              <string-name>Zhang, M.</string-name>
              <string-name>Schueller, S.M.</string-name>
            </person-group>
            <year>2017</year>
            <article-title>Personal Sensing: Understanding Mental Health Using Ubiquitous Sensors and Machine Learning</article-title>
            <source>Annual Review of Clinical Psychology</source>
            <volume>13</volume>
            <pub-id pub-id-type="doi">10.1146/annurev-clinpsy-032816-044949</pub-id>
            <pub-id pub-id-type="pmid">28375728</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B3">
        <label>3.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Grande, I., Berk, M., Birmaher, B. and Vieta, E. (2016) Bipolar Disorder. <italic>The</italic><italic>Lancet</italic>, 387, 1561-1572. https://doi.org/10.1016/s0140-6736(15)00241-x <pub-id pub-id-type="doi">10.1016/s0140-6736(15)00241-x</pub-id><pub-id pub-id-type="pmid">26388529</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1016/s0140-6736(15)00241-x">https://doi.org/10.1016/s0140-6736(15)00241-x</ext-link></mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Grande, I.</string-name>
              <string-name>Berk, M.</string-name>
              <string-name>Birmaher, B.</string-name>
              <string-name>Vieta, E.</string-name>
            </person-group>
            <year>2016</year>
            <article-title>Bipolar Disorder</article-title>
            <source>The Lancet</source>
            <volume>6736</volume>
            <issue>15</issue>
            <pub-id pub-id-type="doi">10.1016/s0140-6736(15)00241-x</pub-id>
            <pub-id pub-id-type="pmid">26388529</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B4">
        <label>4.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Merikangas, K.R., Jin, R., He, J., Kessler, R.C., Lee, S., Sampson, N.A., <italic>et al</italic>. (2011) Prevalence and Correlates of Bipolar Spectrum Disorder in the World Mental Health Survey Initiative. <italic>Archives</italic><italic>of</italic><italic>General</italic><italic>Psychiatry</italic>, 68, 241-251. https://doi.org/10.1001/archgenpsychiatry.2011.12 <pub-id pub-id-type="doi">10.1001/archgenpsychiatry.2011.12</pub-id><pub-id pub-id-type="pmid">21383262</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1001/archgenpsychiatry.2011.12">https://doi.org/10.1001/archgenpsychiatry.2011.12</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Merikangas, K.R.</string-name>
              <string-name>Jin, R.</string-name>
              <string-name>He, J.</string-name>
              <string-name>Kessler, R.C.</string-name>
              <string-name>Lee, S.</string-name>
              <string-name>Sampson, N.A.</string-name>
            </person-group>
            <year>2011</year>
            <article-title>Prevalence and Correlates of Bipolar Spectrum Disorder in the World Mental Health Survey Initiative</article-title>
            <source>Archives of General Psychiatry</source>
            <volume>68</volume>
            <pub-id pub-id-type="doi">10.1001/archgenpsychiatry.2011.12</pub-id>
            <pub-id pub-id-type="pmid">21383262</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B5">
        <label>5.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">GBD 2019 Mental Disorders Collaborators (2022) Global, Regional, and National Burden of 12 Mental Disorders in 204 Countries and Territories. <italic>The Lancet Psychiatry</italic>, 9, 137-150.</mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Global, R</string-name>
            </person-group>
            <year>2019</year>
            <article-title>Global, Regional, and National Burden of 12 Mental Disorders in 204 Countries and Territories</article-title>
            <source>The Lancet Psychiatry</source>
            <volume>9</volume>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B6">
        <label>6.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Leverich, G.S., Altshuler, L.L., Frye, M.A., Suppes, T., Keck, P.E., McElroy, S.L., <italic>et al</italic>. (2003) Factors Associated with Suicide Attempts in 648 Patients with Bipolar Disorder in the Stanley Foundation Bipolar Network. <italic>The</italic><italic>Journal</italic><italic>of</italic><italic>Clinical</italic><italic>Psychiatry</italic>, 64, 506-515. https://doi.org/10.4088/jcp.v64n0503 <pub-id pub-id-type="doi">10.4088/jcp.v64n0503</pub-id><pub-id pub-id-type="pmid">12755652</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.4088/jcp.v64n0503">https://doi.org/10.4088/jcp.v64n0503</ext-link></mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Leverich, G.S.</string-name>
              <string-name>Altshuler, L.L.</string-name>
              <string-name>Frye, M.A.</string-name>
              <string-name>Suppes, T.</string-name>
              <string-name>Keck, P.E.</string-name>
              <string-name>McElroy, S.L.</string-name>
            </person-group>
            <year>2003</year>
            <article-title>Factors Associated with Suicide Attempts in 648 Patients with Bipolar Disorder in the Stanley Foundation Bipolar Network</article-title>
            <source>The Journal of Clinical Psychiatry</source>
            <volume>64</volume>
            <pub-id pub-id-type="doi">10.4088/jcp.v64n0503</pub-id>
            <pub-id pub-id-type="pmid">12755652</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B7">
        <label>7.</label>
        <citation-alternatives>
          <mixed-citation publication-type="book">Goodwin, G.M., Haddad, P., Ferrier, I., Aronson, J., Barnes, T., Cipriani, A., <italic>et al</italic>. (2016) Evidence-Based Guidelines for Treating Bipolar Disorder: Revised Third Edition Recommendations from the British Association for Psychopharmacology. <italic>Journal</italic><italic>of</italic><italic>Psychopharmacology</italic>, 30, 495-553. https://doi.org/10.1177/0269881116636545 <pub-id pub-id-type="doi">10.1177/0269881116636545</pub-id><pub-id pub-id-type="pmid">26979387</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1177/0269881116636545">https://doi.org/10.1177/0269881116636545</ext-link></mixed-citation>
          <element-citation publication-type="book">
            <person-group person-group-type="author">
              <string-name>Goodwin, G.M.</string-name>
              <string-name>Haddad, P.</string-name>
              <string-name>Ferrier, I.</string-name>
              <string-name>Aronson, J.</string-name>
              <string-name>Barnes, T.</string-name>
              <string-name>Cipriani, A.</string-name>
            </person-group>
            <year>2016</year>
            <article-title>Evidence-Based Guidelines for Treating Bipolar Disorder: Revised Third Edition Recommendations from the British Association for Psychopharmacology</article-title>
            <source>Journal of Psychopharmacology</source>
            <volume>30</volume>
            <pub-id pub-id-type="doi">10.1177/0269881116636545</pub-id>
            <pub-id pub-id-type="pmid">26979387</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B8">
        <label>8.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Geddes, J.R. and Miklowitz, D.J. (2013) Treatment of Bipolar Disorder. <italic>The</italic><italic>Lancet</italic>, 381, 1672-1682. https://doi.org/10.1016/s0140-6736(13)60857-0 <pub-id pub-id-type="doi">10.1016/s0140-6736(13)60857-0</pub-id><pub-id pub-id-type="pmid">23663953</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1016/s0140-6736(13)60857-0">https://doi.org/10.1016/s0140-6736(13)60857-0</ext-link></mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Geddes, J.R.</string-name>
              <string-name>Miklowitz, D.J.</string-name>
            </person-group>
            <year>2013</year>
            <article-title>Treatment of Bipolar Disorder</article-title>
            <source>The Lancet</source>
            <volume>6736</volume>
            <issue>13</issue>
            <pub-id pub-id-type="doi">10.1016/s0140-6736(13)60857-0</pub-id>
            <pub-id pub-id-type="pmid">23663953</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B9">
        <label>9.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Rolin, D., Whelan, J. and Montano, C.B. (2020) Is It Depression or Is It Bipolar Depression? <italic>Journal</italic><italic>of</italic><italic>the</italic><italic>American</italic><italic>Association</italic><italic>of</italic><italic>Nurse</italic><italic>Practitioners</italic>, 32, 703-713. https://doi.org/10.1097/jxx.0000000000000499 <pub-id pub-id-type="doi">10.1097/jxx.0000000000000499</pub-id><pub-id pub-id-type="pmid">33017361</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1097/jxx.0000000000000499">https://doi.org/10.1097/jxx.0000000000000499</ext-link></mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Rolin, D.</string-name>
              <string-name>Whelan, J.</string-name>
              <string-name>Montano, C.B.</string-name>
            </person-group>
            <year>2020</year>
            <article-title>Is It Depression or Is It Bipolar Depression? Journal of the American Association of Nurse Practitioners, 32, 703-713</article-title>
            <pub-id pub-id-type="doi">10.1097/jxx.0000000000000499</pub-id>
            <pub-id pub-id-type="pmid">33017361</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B10">
        <label>10.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Morgenthaler, T., Alessi, C., Friedman, L., Owens, J., Kapur, V., Boehlecke, B., <italic>et al</italic>. (2007) Practice Parameters for the Use of Actigraphy in the Assessment of Sleep and Sleep Disorders: An Update for 2007. <italic>Sleep</italic>, 30, 519-529. https://doi.org/10.1093/sleep/30.4.519 <pub-id pub-id-type="doi">10.1093/sleep/30.4.519</pub-id><pub-id pub-id-type="pmid">17520797</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1093/sleep/30.4.519">https://doi.org/10.1093/sleep/30.4.519</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Morgenthaler, T.</string-name>
              <string-name>Alessi, C.</string-name>
              <string-name>Friedman, L.</string-name>
              <string-name>Owens, J.</string-name>
              <string-name>Kapur, V.</string-name>
              <string-name>Boehlecke, B.</string-name>
            </person-group>
            <year>2007</year>
            <article-title>Practice Parameters for the Use of Actigraphy in the Assessment of Sleep and Sleep Disorders: An Update for 2007</article-title>
            <source>Sleep</source>
            <volume>30</volume>
            <pub-id pub-id-type="doi">10.1093/sleep/30.4.519</pub-id>
            <pub-id pub-id-type="pmid">17520797</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B11">
        <label>11.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Saeb, S., Lattie, E.G., Schueller, S.M., Kording, K.P. and Mohr, D.C. (2016) The Relationship between Mobile Phone Location Sensor Data and Depressive Symptom Severity. <italic>PeerJ</italic>, 4, e2537. https://doi.org/10.7717/peerj.2537 <pub-id pub-id-type="doi">10.7717/peerj.2537</pub-id><pub-id pub-id-type="pmid">28344895</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.7717/peerj.2537">https://doi.org/10.7717/peerj.2537</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Saeb, S.</string-name>
              <string-name>Lattie, E.G.</string-name>
              <string-name>Schueller, S.M.</string-name>
              <string-name>Kording, K.P.</string-name>
              <string-name>Mohr, D.C.</string-name>
            </person-group>
            <year>2016</year>
            <article-title>The Relationship between Mobile Phone Location Sensor Data and Depressive Symptom Severity</article-title>
            <source>PeerJ</source>
            <volume>4</volume>
            <pub-id pub-id-type="doi">10.7717/peerj.2537</pub-id>
            <pub-id pub-id-type="pmid">28344895</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B12">
        <label>12.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Faurholt-Jepsen, M., Busk, J., Frost, M., Vinberg, M., Christensen, E.M., Winther, O., <italic>et al</italic>. (2016) Voice Analysis as an Objective State Marker in Bipolar Disorder. <italic>Translational</italic><italic>Psychiatry</italic>, 6, e856-e856. https://doi.org/10.1038/tp.2016.123 <pub-id pub-id-type="doi">10.1038/tp.2016.123</pub-id><pub-id pub-id-type="pmid">27434490</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1038/tp.2016.123">https://doi.org/10.1038/tp.2016.123</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Faurholt-Jepsen, M.</string-name>
              <string-name>Busk, J.</string-name>
              <string-name>Frost, M.</string-name>
              <string-name>Vinberg, M.</string-name>
              <string-name>Christensen, E.M.</string-name>
              <string-name>Winther, O.</string-name>
            </person-group>
            <year>2016</year>
            <article-title>Voice Analysis as an Objective State Marker in Bipolar Disorder</article-title>
            <source>Translational Psychiatry</source>
            <volume>6</volume>
            <pub-id pub-id-type="doi">10.1038/tp.2016.123</pub-id>
            <pub-id pub-id-type="pmid">27434490</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B13">
        <label>13.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Akinode, A.O., Ayadi, O.E., Ezerioha, C.C., Ozo-ogueji, P.C., Akadiri, O.O., Adepoju, D.A., <italic>et al</italic>. (2025) Machine Learning Approaches for Early Detection of Mental Health Disorders Using Wearable Devices and Big Data Analytics. <italic>International Journal of Biological an</italic><italic>d</italic><italic>Pharmaceutical</italic><italic>Sciences</italic><italic>Archive</italic>, 10, 6-23. https://doi.org/10.53771/ijbpsa.2025.10.2.0077 <pub-id pub-id-type="doi">10.53771/ijbpsa.2025.10.2.0077</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.53771/ijbpsa.2025.10.2.0077">https://doi.org/10.53771/ijbpsa.2025.10.2.0077</ext-link></mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Akinode, A.O.</string-name>
              <string-name>Ayadi, O.E.</string-name>
              <string-name>Ezerioha, C.C.</string-name>
              <string-name>Ozo-ogueji, P.C.</string-name>
              <string-name>Akadiri, O.O.</string-name>
              <string-name>Adepoju, D.A.</string-name>
            </person-group>
            <year>2025</year>
            <article-title>Machine Learning Approaches for Early Detection of Mental Health Disorders Using Wearable Devices and Big Data Analytics</article-title>
            <source>International Journal of Biological and Pharmaceutical Sciences Archive</source>
            <volume>10</volume>
            <pub-id pub-id-type="doi">10.53771/ijbpsa.2025.10.2.0077</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B14">
        <label>14.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Chen, Q., Dai, P., Huang, K., Hu, T. and Liao, S. (2025) MMDD: A Multimodal Multitask Dynamic Disentanglement Framework for Robust Major Depressive Disorder Diagnosis across Neuroimaging Sites. <italic>Diagnostics</italic>, 15, Article No. 3089. https://doi.org/10.3390/diagnostics15233089 <pub-id pub-id-type="doi">10.3390/diagnostics15233089</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.3390/diagnostics15233089">https://doi.org/10.3390/diagnostics15233089</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Chen, Q.</string-name>
              <string-name>Dai, P.</string-name>
              <string-name>Huang, K.</string-name>
              <string-name>Hu, T.</string-name>
              <string-name>Liao, S.</string-name>
            </person-group>
            <year>2025</year>
            <article-title>MMDD: A Multimodal Multitask Dynamic Disentanglement Framework for Robust Major Depressive Disorder Diagnosis across Neuroimaging Sites</article-title>
            <source>Diagnostics</source>
            <volume>15</volume>
            <elocation-id>No</elocation-id>
            <pub-id pub-id-type="doi">10.3390/diagnostics15233089</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B15">
        <label>15.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Lunsford-Avery, J.R., Engelhard, M.M., Navar, A.M. and Kollins, S.H. (2018) Validation of the Sleep Regularity Index in Older Adults and Associations with Cardiometabolic Risk. <italic>Scientific</italic><italic>Reports</italic>, 8, Article No. 14158. https://doi.org/10.1038/s41598-018-32402-5 <pub-id pub-id-type="doi">10.1038/s41598-018-32402-5</pub-id><pub-id pub-id-type="pmid">30242174</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1038/s41598-018-32402-5">https://doi.org/10.1038/s41598-018-32402-5</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Lunsford-Avery, J.R.</string-name>
              <string-name>Engelhard, M.M.</string-name>
              <string-name>Navar, A.M.</string-name>
              <string-name>Kollins, S.H.</string-name>
            </person-group>
            <year>2018</year>
            <article-title>Validation of the Sleep Regularity Index in Older Adults and Associations with Cardiometabolic Risk</article-title>
            <source>Scientific Reports</source>
            <volume>8</volume>
            <elocation-id>No</elocation-id>
            <pub-id pub-id-type="doi">10.1038/s41598-018-32402-5</pub-id>
            <pub-id pub-id-type="pmid">30242174</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B16">
        <label>16.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Ancoli-Israel, S., Cole, R., Alessi, C., Chambers, M., Moorcroft, W. and Pollak, C.P. (2003) The Role of Actigraphy in the Study of Sleep and Circadian Rhythms. <italic>Sleep</italic>, 26, 342-392. https://doi.org/10.1093/sleep/26.3.342 <pub-id pub-id-type="doi">10.1093/sleep/26.3.342</pub-id><pub-id pub-id-type="pmid">12749557</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1093/sleep/26.3.342">https://doi.org/10.1093/sleep/26.3.342</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Ancoli-Israel, S.</string-name>
              <string-name>Cole, R.</string-name>
              <string-name>Alessi, C.</string-name>
              <string-name>Chambers, M.</string-name>
              <string-name>Moorcroft, W.</string-name>
              <string-name>Pollak, C.P.</string-name>
            </person-group>
            <year>2003</year>
            <article-title>The Role of Actigraphy in the Study of Sleep and Circadian Rhythms</article-title>
            <source>Sleep</source>
            <volume>26</volume>
            <pub-id pub-id-type="doi">10.1093/sleep/26.3.342</pub-id>
            <pub-id pub-id-type="pmid">12749557</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B17">
        <label>17.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Palmius, N., Tsanas, A., Saunders, K.E.A., Bilderbeck, A.C., Geddes, J.R., Goodwin, G.M., <italic>et al</italic>. (2017) Detecting Bipolar Depression from Geographic Location Data. <italic>IEEE</italic><italic>Transactions</italic><italic>on</italic><italic>Biomedical</italic><italic>Engineering</italic>, 64, 1761-1771. https://doi.org/10.1109/tbme.2016.2611862 <pub-id pub-id-type="doi">10.1109/tbme.2016.2611862</pub-id><pub-id pub-id-type="pmid">28113247</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1109/tbme.2016.2611862">https://doi.org/10.1109/tbme.2016.2611862</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Palmius, N.</string-name>
              <string-name>Tsanas, A.</string-name>
              <string-name>Saunders, K.E.A.</string-name>
              <string-name>Bilderbeck, A.C.</string-name>
              <string-name>Geddes, J.R.</string-name>
              <string-name>Goodwin, G.M.</string-name>
            </person-group>
            <year>2017</year>
            <article-title>Detecting Bipolar Depression from Geographic Location Data</article-title>
            <source>IEEE Transactions on Biomedical Engineering</source>
            <volume>64</volume>
            <pub-id pub-id-type="doi">10.1109/tbme.2016.2611862</pub-id>
            <pub-id pub-id-type="pmid">28113247</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B18">
        <label>18.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Reinertsen, E. and Clifford, G.D. (2018) A Review of Physiological and Behavioral Monitoring with Digital Sensors for Neuropsychiatric Illnesses. <italic>Physiological</italic><italic>Measurement</italic>, 39, 05TR01. https://doi.org/10.1088/1361-6579/aabf64 <pub-id pub-id-type="doi">10.1088/1361-6579/aabf64</pub-id><pub-id pub-id-type="pmid">29671754</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1088/1361-6579/aabf64">https://doi.org/10.1088/1361-6579/aabf64</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Reinertsen, E.</string-name>
              <string-name>Clifford, G.D.</string-name>
            </person-group>
            <year>2018</year>
            <article-title>A Review of Physiological and Behavioral Monitoring with Digital Sensors for Neuropsychiatric Illnesses</article-title>
            <source>Physiological Measurement</source>
            <volume>39</volume>
            <pub-id pub-id-type="doi">10.1088/1361-6579/aabf64</pub-id>
            <pub-id pub-id-type="pmid">29671754</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B19">
        <label>19.</label>
        <citation-alternatives>
          <mixed-citation publication-type="confproc">Lea, C., Flynn, M.D., Vidal, R., Reiter, A. and Hager, G.D. (2017) Temporal Convolutional Networks for Action Segmentation and Detection. 2017 <italic>IEEE Conference on Computer Vision and Pattern Recognition</italic> ( <italic>CVPR</italic>), Honolulu, 21-26 July 2017, 1003-1012. https://doi.org/10.1109/cvpr.2017.113 <pub-id pub-id-type="doi">10.1109/cvpr.2017.113</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1109/cvpr.2017.113">https://doi.org/10.1109/cvpr.2017.113</ext-link></mixed-citation>
          <element-citation publication-type="confproc">
            <person-group person-group-type="author">
              <string-name>Lea, C.</string-name>
              <string-name>Flynn, M.D.</string-name>
              <string-name>Vidal, R.</string-name>
              <string-name>Reiter, A.</string-name>
              <string-name>Hager, G.D.</string-name>
            </person-group>
            <year>2017</year>
            <article-title>Temporal Convolutional Networks for Action Segmentation and Detection</article-title>
            <source>2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR)</source>
            <volume>21</volume>
            <pub-id pub-id-type="doi">10.1109/cvpr.2017.113</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B20">
        <label>20.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Hochreiter, S. and Schmidhuber, J. (1997) Long Short-Term Memory. <italic>Neural</italic><italic>Computation</italic>, 9, 1735-1780. https://doi.org/10.1162/neco.1997.9.8.1735 <pub-id pub-id-type="doi">10.1162/neco.1997.9.8.1735</pub-id><pub-id pub-id-type="pmid">9377276</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1162/neco.1997.9.8.1735">https://doi.org/10.1162/neco.1997.9.8.1735</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Hochreiter, S.</string-name>
              <string-name>Schmidhuber, J.</string-name>
            </person-group>
            <year>1997</year>
            <article-title>Long Short-Term Memory</article-title>
            <source>Neural Computation</source>
            <volume>9</volume>
            <pub-id pub-id-type="doi">10.1162/neco.1997.9.8.1735</pub-id>
            <pub-id pub-id-type="pmid">9377276</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B21">
        <label>21.</label>
        <citation-alternatives>
          <mixed-citation publication-type="web">van den Oord, A., <italic>et al</italic>. (2016) WaveNet: A Generative Model for Raw Audio. https://arxiv.org/abs/1609.03499</mixed-citation>
          <element-citation publication-type="web">
            <person-group person-group-type="author">
              <string-name>Oord, A.</string-name>
            </person-group>
            <year>2016</year>
            <article-title>WaveNet: A Generative Model for Raw Audio</article-title>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B22">
        <label>22.</label>
        <citation-alternatives>
          <mixed-citation publication-type="confproc">Vaswani, A., <italic>et al</italic>. (2017) Attention Is All You Need. <italic>Proceedings of the</italic>31 <italic>st International Conference on Neural Information Processing Systems</italic>, Long Beach, 4-9 December 2017, 6000-6010.</mixed-citation>
          <element-citation publication-type="confproc">
            <person-group person-group-type="author">
              <string-name>Vaswani, A.</string-name>
              <string-name>Systems, L</string-name>
            </person-group>
            <year>2017</year>
            <article-title>Attention Is All You Need</article-title>
            <source>Proceedings of the 31st International Conference on Neural Information Processing Systems</source>
            <volume>4</volume>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B23">
        <label>23.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Devlin, J., <italic>et al</italic>. (2019) BERT: Pre-Training of Deep Bidirectional Transformers for Language Understanding. <italic>NAACL</italic>- <italic>HLT</italic> 2019, Minneapolis, 2-7 June 2019, 4171-4186.</mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Devlin, J.</string-name>
            </person-group>
            <year>2019</year>
            <article-title>BERT: Pre-Training of Deep Bidirectional Transformers for Language Understanding</article-title>
            <source>NAACL-HLT 2019</source>
            <volume>2</volume>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B24">
        <label>24.</label>
        <citation-alternatives>
          <mixed-citation publication-type="confproc">Alsentzer, E., Murphy, J., Boag, W., Weng, W., Jindi, D., Naumann, T., <italic>et al</italic>. (2019) Publicly Available Clinical BERT Embeddings. <italic>Proceedings of the</italic>2 <italic>nd Clinical Natural Language Processing Workshop</italic>, Minneapolis, June 2019, 72-78. https://doi.org/10.18653/v1/w19-1909 <pub-id pub-id-type="doi">10.18653/v1/w19-1909</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.18653/v1/w19-1909">https://doi.org/10.18653/v1/w19-1909</ext-link></mixed-citation>
          <element-citation publication-type="confproc">
            <person-group person-group-type="author">
              <string-name>Alsentzer, E.</string-name>
              <string-name>Murphy, J.</string-name>
              <string-name>Boag, W.</string-name>
              <string-name>Weng, W.</string-name>
              <string-name>Jindi, D.</string-name>
              <string-name>Naumann, T.</string-name>
              <string-name>Workshop, M</string-name>
            </person-group>
            <year>2019</year>
            <article-title>Publicly Available Clinical BERT Embeddings</article-title>
            <source>Proceedings of the 2nd Clinical Natural Language Processing Workshop</source>
            <volume>72</volume>
            <pub-id pub-id-type="doi">10.18653/v1/w19-1909</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B25">
        <label>25.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Gal, Y. and Ghahramani, Z. (2016) Dropout as a Bayesian Approximation: Representing Model Uncertainty in Deep Learning. <italic>ICML</italic> 2016, New York, 19-24 June 2016, 1050-1059.</mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Gal, Y.</string-name>
              <string-name>Ghahramani, Z.</string-name>
            </person-group>
            <year>2016</year>
            <article-title>Dropout as a Bayesian Approximation: Representing Model Uncertainty in Deep Learning</article-title>
            <source>ICML 2016</source>
            <volume>19</volume>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B26">
        <label>26.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Lakshminarayanan, B., <italic>et al</italic>. (2017) Simple and Scalable Predictive Uncertainty Estimation Using Deep Ensembles. <italic>NeurIPS</italic> 2017, Long Beach, 4-9 December 2017, 6402-6413.</mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Lakshminarayanan, B.</string-name>
            </person-group>
            <year>2017</year>
            <article-title>Simple and Scalable Predictive Uncertainty Estimation Using Deep Ensembles</article-title>
            <source>NeurIPS 2017</source>
            <volume>4</volume>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B27">
        <label>27.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Szostek, D. (2021) Is the Traditional Method of Regulation (the Legislative Act) Sufficient to Regulate Artificial Intelligence, or Should It Also Be Regulated by an Algorithmic Code? <italic>Białostockie Studia Prawnicze</italic>, 26, 43-60. https://doi.org/10.15290/bsp.2021.26.03.03 <pub-id pub-id-type="doi">10.15290/bsp.2021.26.03.03</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.15290/bsp.2021.26.03.03">https://doi.org/10.15290/bsp.2021.26.03.03</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Szostek, D.</string-name>
            </person-group>
            <year>2021</year>
            <article-title>Is the Traditional Method of Regulation (the Legislative Act) Sufficient to Regulate Artificial Intelligence, or Should It Also Be Regulated by an Algorithmic Code? Białostockie Studia Prawnicze, 26, 43-60</article-title>
            <pub-id pub-id-type="doi">10.15290/bsp.2021.26.03.03</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B28">
        <label>28.</label>
        <citation-alternatives>
          <mixed-citation publication-type="book">Ramsay, J.O. and Silverman, B.W. (2005) Functional Data Analysis. 2nd Edition, Springer.</mixed-citation>
          <element-citation publication-type="book">
            <person-group person-group-type="author">
              <string-name>Ramsay, J.O.</string-name>
              <string-name>Silverman, B.W.</string-name>
              <string-name>Edition, S</string-name>
            </person-group>
            <year>2005</year>
            <article-title>Functional Data Analysis</article-title>
            <source>2nd Edition</source>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B29">
        <label>29.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Cole, R.J., Kripke, D.F., Gruen, W., Mullaney, D.J. and Gillin, J.C. (1992) Automatic Sleep/Wake Identification from Wrist Activity. <italic>Sleep</italic>, 15, 461-469. https://doi.org/10.1093/sleep/15.5.461 <pub-id pub-id-type="doi">10.1093/sleep/15.5.461</pub-id><pub-id pub-id-type="pmid">1455130</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1093/sleep/15.5.461">https://doi.org/10.1093/sleep/15.5.461</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Cole, R.J.</string-name>
              <string-name>Kripke, D.F.</string-name>
              <string-name>Gruen, W.</string-name>
              <string-name>Mullaney, D.J.</string-name>
              <string-name>Gillin, J.C.</string-name>
            </person-group>
            <year>1992</year>
            <article-title>Automatic Sleep/Wake Identification from Wrist Activity</article-title>
            <source>Sleep</source>
            <volume>15</volume>
            <pub-id pub-id-type="doi">10.1093/sleep/15.5.461</pub-id>
            <pub-id pub-id-type="pmid">1455130</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B30">
        <label>30.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Phillips, A.J.K., Clerx, W.M., O’Brien, C.S., Sano, A., Barger, L.K., Picard, R.W., <italic>et al</italic>. (2017) Irregular Sleep/Wake Patterns Are Associated with Poorer Academic Performance and Delayed Circadian and Sleep/Wake Timing. <italic>Scientific</italic><italic>Reports</italic>, 7, Article No. 3216. https://doi.org/10.1038/s41598-017-03171-4 <pub-id pub-id-type="doi">10.1038/s41598-017-03171-4</pub-id><pub-id pub-id-type="pmid">28607474</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1038/s41598-017-03171-4">https://doi.org/10.1038/s41598-017-03171-4</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Phillips, A.J.K.</string-name>
              <string-name>Clerx, W.M.</string-name>
              <string-name>Brien, C.S.</string-name>
              <string-name>Sano, A.</string-name>
              <string-name>Barger, L.K.</string-name>
              <string-name>Picard, R.W.</string-name>
            </person-group>
            <year>2017</year>
            <article-title>Irregular Sleep/Wake Patterns Are Associated with Poorer Academic Performance and Delayed Circadian and Sleep/Wake Timing</article-title>
            <source>Scientific Reports</source>
            <volume>7</volume>
            <elocation-id>No</elocation-id>
            <pub-id pub-id-type="doi">10.1038/s41598-017-03171-4</pub-id>
            <pub-id pub-id-type="pmid">28607474</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B31">
        <label>31.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Phan, H., Mikkelsen, K., Chen, O.Y., Koch, P., Mertins, A. and De Vos, M. (2022) Sleeptransformer: Automatic Sleep Staging with Interpretability and Uncertainty Quantification. <italic>IEEE</italic><italic>Transactions</italic><italic>on</italic><italic>Biomedical</italic><italic>Engineering</italic>, 69, 2456-2467. https://doi.org/10.1109/tbme.2022.3147187 <pub-id pub-id-type="doi">10.1109/tbme.2022.3147187</pub-id><pub-id pub-id-type="pmid">35100107</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1109/tbme.2022.3147187">https://doi.org/10.1109/tbme.2022.3147187</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Phan, H.</string-name>
              <string-name>Mikkelsen, K.</string-name>
              <string-name>Chen, O.Y.</string-name>
              <string-name>Koch, P.</string-name>
              <string-name>Mertins, A.</string-name>
              <string-name>Vos, M.</string-name>
            </person-group>
            <year>2022</year>
            <article-title>Sleeptransformer: Automatic Sleep Staging with Interpretability and Uncertainty Quantification</article-title>
            <source>IEEE Transactions on Biomedical Engineering</source>
            <volume>69</volume>
            <pub-id pub-id-type="doi">10.1109/tbme.2022.3147187</pub-id>
            <pub-id pub-id-type="pmid">35100107</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B32">
        <label>32.</label>
        <citation-alternatives>
          <mixed-citation publication-type="confproc">Canzian, L. and Musolesi, M. (2015) Trajectories of Depression: Unobtrusive Monitoring of Depressive States via Smartphone Mobility Traces. <italic>Proceedings of the</italic>2015 <italic>ACM International Joint Conference on Pervasive and Ubiquitous Computing</italic>, Osaka, 7-11 September 2015, 1293-1304. https://doi.org/10.1145/2750858.2805845 <pub-id pub-id-type="doi">10.1145/2750858.2805845</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1145/2750858.2805845">https://doi.org/10.1145/2750858.2805845</ext-link></mixed-citation>
          <element-citation publication-type="confproc">
            <person-group person-group-type="author">
              <string-name>Canzian, L.</string-name>
              <string-name>Musolesi, M.</string-name>
              <string-name>Computing, O</string-name>
            </person-group>
            <year>2015</year>
            <article-title>Trajectories of Depression: Unobtrusive Monitoring of Depressive States via Smartphone Mobility Traces</article-title>
            <source>Proceedings of the 2015 ACM International Joint Conference on Pervasive and Ubiquitous Computing</source>
            <volume>7</volume>
            <pub-id pub-id-type="doi">10.1145/2750858.2805845</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B33">
        <label>33.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Mohmad Dar, G.H. and Delhibabu, R. (2024) Speech Databases, Speech Features, and Classifiers in Speech Emotion Recognition: A Review. <italic>IEEE Access</italic>, 12, 151122-151152. https://doi.org/10.1109/access.2024.3476960 <pub-id pub-id-type="doi">10.1109/access.2024.3476960</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1109/access.2024.3476960">https://doi.org/10.1109/access.2024.3476960</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Dar, G.H.</string-name>
              <string-name>Delhibabu, R.</string-name>
              <string-name>Databases, S</string-name>
            </person-group>
            <year>2024</year>
            <article-title>Speech Databases, Speech Features, and Classifiers in Speech Emotion Recognition: A Review</article-title>
            <source>IEEE Access</source>
            <volume>12</volume>
            <pub-id pub-id-type="doi">10.1109/access.2024.3476960</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B34">
        <label>34.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Dong, Y. and Yang, X. (2021) A Hierarchical Depression Detection Model Based on Vocal and Emotional Cues. <italic>Neurocomputing</italic>, 441, 279-290. https://doi.org/10.1016/j.neucom.2021.02.019 <pub-id pub-id-type="doi">10.1016/j.neucom.2021.02.019</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1016/j.neucom.2021.02.019">https://doi.org/10.1016/j.neucom.2021.02.019</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Dong, Y.</string-name>
              <string-name>Yang, X.</string-name>
            </person-group>
            <year>2021</year>
            <article-title>A Hierarchical Depression Detection Model Based on Vocal and Emotional Cues</article-title>
            <source>Neurocomputing</source>
            <volume>441</volume>
            <pub-id pub-id-type="doi">10.1016/j.neucom.2021.02.019</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B35">
        <label>35.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Wang, K., Varma, D.S. and Prosperi, M. (2018) A Systematic Review of the Effectiveness of Mobile Apps for Monitoring and Management of Mental Health Symptoms or Disorders. <italic>Journal</italic><italic>of</italic><italic>Psychiatric</italic><italic>Research</italic>, 107, 73-78. https://doi.org/10.1016/j.jpsychires.2018.10.006 <pub-id pub-id-type="doi">10.1016/j.jpsychires.2018.10.006</pub-id><pub-id pub-id-type="pmid">30347316</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1016/j.jpsychires.2018.10.006">https://doi.org/10.1016/j.jpsychires.2018.10.006</ext-link></mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Wang, K.</string-name>
              <string-name>Varma, D.S.</string-name>
              <string-name>Prosperi, M.</string-name>
            </person-group>
            <year>2018</year>
            <article-title>A Systematic Review of the Effectiveness of Mobile Apps for Monitoring and Management of Mental Health Symptoms or Disorders</article-title>
            <source>Journal of Psychiatric Research</source>
            <volume>107</volume>
            <pub-id pub-id-type="doi">10.1016/j.jpsychires.2018.10.006</pub-id>
            <pub-id pub-id-type="pmid">30347316</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B36">
        <label>36.</label>
        <citation-alternatives>
          <mixed-citation publication-type="confproc">Niculescu-Mizil, A. and Caruana, R. (2005) Predicting Good Probabilities with Supervised Learning. <italic>Proceedings of the</italic>22 <italic>nd International Conference on Machine Learning</italic>- <italic>ICML</italic>’05, Bonn, 7-11 August 2005, 625-633. https://doi.org/10.1145/1102351.1102430 <pub-id pub-id-type="doi">10.1145/1102351.1102430</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1145/1102351.1102430">https://doi.org/10.1145/1102351.1102430</ext-link></mixed-citation>
          <element-citation publication-type="confproc">
            <person-group person-group-type="author">
              <string-name>Niculescu-Mizil, A.</string-name>
              <string-name>Caruana, R.</string-name>
            </person-group>
            <year>2005</year>
            <article-title>Predicting Good Probabilities with Supervised Learning</article-title>
            <source>Proceedings of the 22nd International Conference on Machine Learning-ICML’05</source>
            <volume>7</volume>
            <pub-id pub-id-type="doi">10.1145/1102351.1102430</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B37">
        <label>37.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Georgousis, S., Kenning, M.P. and Xie, X.H. (2021) Graph Deep Learning: State of the Art and Challenges. <italic>IEEE Access</italic>, 9, 22106-22140. https://doi.org/10.1109/ACCESS.2021.3055280 <pub-id pub-id-type="doi">10.1109/ACCESS.2021.3055280</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1109/ACCESS.2021.3055280">https://doi.org/10.1109/ACCESS.2021.3055280</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Georgousis, S.</string-name>
              <string-name>Kenning, M.P.</string-name>
              <string-name>Xie, X.H.</string-name>
            </person-group>
            <year>2021</year>
            <article-title>Graph Deep Learning: State of the Art and Challenges</article-title>
            <source>IEEE Access</source>
            <volume>9</volume>
            <pub-id pub-id-type="doi">10.1109/ACCESS.2021.3055280</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B38">
        <label>38.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Breiman, L. (2001) Random Forests. <italic>Machine</italic><italic>Learning</italic>, 45, 5-32. https://doi.org/10.1023/a:1010933404324 <pub-id pub-id-type="doi">10.1023/a:1010933404324</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1023/a:1010933404324">https://doi.org/10.1023/a:1010933404324</ext-link></mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Breiman, L.</string-name>
            </person-group>
            <year>2001</year>
            <article-title>Random Forests</article-title>
            <source>Machine Learning</source>
            <volume>45</volume>
            <fpage>101093</fpage>
            <pub-id pub-id-type="doi">10.1023/a:1010933404324</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B39">
        <label>39.</label>
        <citation-alternatives>
          <mixed-citation publication-type="confproc">Chen, T. and Guestrin, C. (2016) XGBoost: A Scalable Tree Boosting System. <italic>Proceedings of the</italic>22 <italic>nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining</italic>, San Francisco, 13-17 August 2016, 785-794. https://doi.org/10.1145/2939672.2939785 <pub-id pub-id-type="doi">10.1145/2939672.2939785</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1145/2939672.2939785">https://doi.org/10.1145/2939672.2939785</ext-link></mixed-citation>
          <element-citation publication-type="confproc">
            <person-group person-group-type="author">
              <string-name>Chen, T.</string-name>
              <string-name>Guestrin, C.</string-name>
              <string-name>Mining, S</string-name>
            </person-group>
            <year>2016</year>
            <article-title>XGBoost: A Scalable Tree Boosting System</article-title>
            <source>Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining</source>
            <volume>13</volume>
            <pub-id pub-id-type="doi">10.1145/2939672.2939785</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B40">
        <label>40.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Lee, K., Lee, T.C., Yefimova, M., Kumar, S., Puga, F., Azuero, A., <italic>et al</italic>. (2023) Using Digital Phenotyping to Understand Health-Related Outcomes: A Scoping Review. <italic>International</italic><italic>Journal</italic><italic>of</italic><italic>Medical</italic><italic>Informatics</italic>, 174, Article ID: 105061. https://doi.org/10.1016/j.ijmedinf.2023.105061 <pub-id pub-id-type="doi">10.1016/j.ijmedinf.2023.105061</pub-id><pub-id pub-id-type="pmid">37030145</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1016/j.ijmedinf.2023.105061">https://doi.org/10.1016/j.ijmedinf.2023.105061</ext-link></mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Lee, K.</string-name>
              <string-name>Lee, T.C.</string-name>
              <string-name>Yefimova, M.</string-name>
              <string-name>Kumar, S.</string-name>
              <string-name>Puga, F.</string-name>
              <string-name>Azuero, A.</string-name>
            </person-group>
            <year>2023</year>
            <article-title>Using Digital Phenotyping to Understand Health-Related Outcomes: A Scoping Review</article-title>
            <source>International Journal of Medical Informatics</source>
            <volume>174</volume>
            <fpage>105061</fpage>
            <elocation-id>ID</elocation-id>
            <pub-id pub-id-type="doi">10.1016/j.ijmedinf.2023.105061</pub-id>
            <pub-id pub-id-type="pmid">37030145</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B41">
        <label>41.</label>
        <citation-alternatives>
          <mixed-citation publication-type="book">World Health Organization (2019) Guidelines on Mental Health Promotion in Non-Specialist Settings. WHO Press.</mixed-citation>
          <element-citation publication-type="book">
            <year>2019</year>
            <article-title>Guidelines on Mental Health Promotion in Non-Specialist Settings</article-title>
          </element-citation>
        </citation-alternatives>
      </ref>
    </ref-list>
  </back>
</article>