<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.4 20241031//EN" "JATS-journalpublishing1-4.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article" dtd-version="1.4" xml:lang="en">
  <front>
    <journal-meta>
      <journal-id journal-id-type="publisher-id">Oalib</journal-id>
      <journal-title-group>
        <journal-title>Open Access Library Journal</journal-title>
      </journal-title-group>
      <issn pub-type="epub">2333-9721</issn>
      <issn pub-type="ppub">2333-9705</issn>
      <publisher>
        <publisher-name>Scientific Research Publishing</publisher-name>
      </publisher>
    </journal-meta>
    <article-meta>
      <article-id pub-id-type="doi">10.4236/oalib.1115671</article-id>
      <article-id pub-id-type="publisher-id">Oalib-153354</article-id>
      <article-categories>
        <subj-group>
          <subject>Article</subject>
        </subj-group>
        <subj-group>
          <subject>Biomedical</subject>
          <subject>Life Sciences</subject>
          <subject>Business</subject>
          <subject>Economics</subject>
          <subject>Chemistry</subject>
          <subject>Materials Science</subject>
          <subject>Computer Science</subject>
          <subject>Communications</subject>
          <subject>Earth</subject>
          <subject>Environmental Sciences</subject>
          <subject>Engineering</subject>
          <subject>Medicine</subject>
          <subject>Healthcare</subject>
          <subject>Physics</subject>
          <subject>Mathematics</subject>
          <subject>Social Sciences</subject>
          <subject>Humanities</subject>
        </subj-group>
      </article-categories>
      <title-group>
        <article-title>Deep Learning Classification of Treatment-Emergent Mania Following Antidepressant Initiation in Bipolar II Disorder: A Synthetic-Data Proof-of-Concept Using Longitudinal and Episode-Feature Attention</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author" corresp="yes">
          <contrib-id contrib-id-type="orcid">0000-0001-9101-072X</contrib-id>
          <name name-style="western">
            <surname>Filippis</surname>
            <given-names>Rocco de</given-names>
          </name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <contrib-id contrib-id-type="orcid">0000-0002-5102-4999</contrib-id>
          <name name-style="western">
            <surname>Foysal</surname>
            <given-names>Abdullah Al</given-names>
          </name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
      </contrib-group>
      <aff id="aff1"><label>1</label> Neuroscience, Istituto di Psicopatologia, Rome, Italy </aff>
      <aff id="aff2"><label>2</label> Department of Computer Engineering (AI), DIBRIS, University of Genova, Genova, Italy </aff>
      <author-notes>
        <fn fn-type="conflict" id="fn-conflict">
          <p>The authors declare no conflicts of interest.</p>
        </fn>
      </author-notes>
      <pub-date pub-type="epub">
        <day>03</day>
        <month>08</month>
        <year>2026</year>
      </pub-date>
      <pub-date pub-type="collection">
        <month>08</month>
        <year>2026</year>
      </pub-date>
      <volume>13</volume>
      <issue>08</issue>
      <fpage>1</fpage>
      <lpage>1</lpage>
      <history>
        <date date-type="received">
          <day>22</day>
          <month>06</month>
          <year>2026</year>
        </date>
        <date date-type="accepted">
          <day>21</day>
          <month>08</month>
          <year>2026</year>
        </date>
        <date date-type="published">
          <day>24</day>
          <month>08</month>
          <year>2026</year>
        </date>
      </history>
      <permissions>
        <copyright-statement>© 2026 by the authors and Scientific Research Publishing Inc.</copyright-statement>
        <copyright-year>2026</copyright-year>
        <license license-type="open-access">
          <license-p> This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license ( <ext-link ext-link-type="uri" xlink:href="https://creativecommons.org/licenses/by/4.0/">https://creativecommons.org/licenses/by/4.0/</ext-link> ). </license-p>
        </license>
      </permissions>
      <self-uri content-type="doi" xlink:href="https://doi.org/10.4236/oalib.1115671">https://doi.org/10.4236/oalib.1115671</self-uri>
      <abstract>
        <p>Treatment-emergent mania (TEM), including antidepressant-associated hypomanic or manic switching and cycle acceleration, is an important safety concern in bipolar II disorder. This study is a synthetic-data proof of concept: all patients, digital phenotyping streams, episode histories, pharmacological variables, and TEM outcomes were simulated. The findings, therefore, evaluate methodological feasibility rather than clinical effectiveness. We developed a multimodal modelling framework combining an LSTM with multi-head self-attention for 30-day pre-antidepressant digital sequences, an episode-feature attention network operating on a fixed 14-feature episode-history vector, and gradient-boosting models for clinical and pharmacological variables. The proposed out-of-fold stacking ensemble consistently comprised XGBoost, LightGBM, and LSTM + Attention outputs; the episode-feature attention model was evaluated as a separate comparator and was not included in the stack. The fully synthetic cohort contained N = 750 simulated BD-II patients. Model development used patient-level development, validation, and held-out test partitions created after cohort synthesis but before preprocessing or model fitting. On the held-out synthetic test set, the proposed ensemble achieved AUC = 0.997 (95% CI: 0.988 - 1.000), F1 = 0.958, sensitivity = 0.958, specificity = 0.989, and Brier score = 0.025. LSTM + Attention achieved AUC = 0.984. These unusually high estimates may partly reflect the simulator’s embedded temporal and pharmacological risk structure and require external evaluation. SHAP analysis of XGBoost and LightGBM identified prior TEM history, antidepressant-class risk, mood-stabilizer adequacy, circadian IS score, and episode-interval shortening as prominent predictors. The results support the technical feasibility of sequence-informed TEM risk modelling in a controlled simulation. They do not establish near-perfect prediction, treatment safety, or readiness for clinical deployment. Independent validation on prospectively collected BD-II data is required.</p>
      </abstract>
      <kwd-group kwd-group-type="author-generated" xml:lang="en">
        <kwd>Treatment-Emergent Mania</kwd>
        <kwd>Bipolar II Disorder</kwd>
        <kwd>Antidepressant Safety</kwd>
        <kwd>LSTM</kwd>
        <kwd>Graph Attention Network</kwd>
        <kwd>Digital Phenotyping</kwd>
        <kwd>SHAP</kwd>
        <kwd>Bayesian Model Selection</kwd>
        <kwd>Circadian Biomarkers</kwd>
        <kwd>Episode Chronology</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec1">
      <title>1. Introduction</title>
      <p>Bipolar II disorder occupies an underappreciated position in the antidepressant safety literature. Historically framed as the “milder” bipolar presentation, BD-II is disproportionately prescribed antidepressants relative to BD-I, partly because its hypomanic phases are less dramatic and may go unrecognised, and partly because its predominant depressive burden creates genuine therapeutic pressure for antidepressant initiation [<xref ref-type="bibr" rid="B1">1</xref>]. Yet the risk of treatment-emergent mania (TEM) in BD-II encompassing hypomania induction, mania switch, and cycle acceleration is clinically significant and systematically underestimated [<xref ref-type="bibr" rid="B2">2</xref>]. Published rates of TEM in antidepressant-exposed BD-II patients range from 20% - 30%, with the specific subtype, antidepressant class, and mood stabilizer co-prescription status each acting as significant moderators [<xref ref-type="bibr" rid="B3">3</xref>].</p>
      <p>The BD-II-specific challenge is one of detection as much as prevention. Unlike BD-I, where antidepressant-induced full mania is clinically unambiguous, TEM in BD-II manifests as hypomania, a state that patients may experience as welcome energy and productivity, and that clinicians may miss or delay attributing to antidepressant exposure [<xref ref-type="bibr" rid="B4">4</xref>]. By the time cycle acceleration is recognised, the antidepressant has often been maintained or escalated, entrenching a pharmacologically driven cycling pattern. Early computational prediction of TEM risk before initiation would fundamentally alter this trajectory by enabling targeted antidepressant avoidance or mood-stabilizer pre-loading in identifiable high-risk patients.</p>
      <p>Deep learning offers methods for representing longitudinal signals that are difficult to summarise in static clinical scores. LSTM architectures with self-attention can model temporal patterns in 30-day digital streams. The episode-history component in this study, however, operates on a fixed engineered feature vector rather than on an explicit patient-specific graph; it is therefore described as an episode-feature attention network, not an episode-feature attention network. The current work evaluates these components in a fully synthetic proof-of-concept setting.</p>
      <p>We present five methodological contributions: i) a synthetic BD-II TEM modelling benchmark combining longitudinal digital phenotyping and episode-history summaries; ii) an LSTM with multi-head self-attention for 30-day sequences; iii) an episode-feature attention comparator operating on engineered episode-history variables; iv) a clearly defined OOF stack comprising XGBoost, LightGBM, and LSTM + Attention only; and v) exploratory feature-attribution and subgroup analyses intended to generate hypotheses for real-world validation.</p>
    </sec>
    <sec id="sec2">
      <title>2. Background and Related Work</title>
      <sec id="sec2dot1">
        <title>2.1. Treatment-Emergent Mania in BD-II: Clinical Evidence</title>
        <p>TEM represents the intersection of two distinct pharmacological phenomena: antidepressant-induced hypomanic/manic switch and antidepressant-induced cycle acceleration. In BD-II, TEM manifests most commonly as hypomanic induction, often within 4 - 12 weeks of antidepressant initiation, though cases of delayed acceleration up to 6 months post-initiation have been reported. Risk factors established by clinical studies include BD-I vs. BD-II (paradoxically, BD-II may have higher real-world TEM rates due to lower treatment safeguards), prior TEM history, antidepressant class (TCAs highest, bupropion lowest), absence of adequate mood stabilizer coverage, young age at onset, rapid cycling history, mixed episode history, and recent hypomanic episode [<xref ref-type="bibr" rid="B5">5</xref>][<xref ref-type="bibr" rid="B6">6</xref>]. Circadian biomarker studies have documented that a lower IS score in the weeks preceding antidepressant initiation predicts subsequent mood instability, suggesting a chronobiological vulnerability window [<xref ref-type="bibr" rid="B7">7</xref>].</p>
      </sec>
      <sec id="sec2dot2">
        <title>2.2. LSTM and Self-Attention for Longitudinal Psychiatric Prediction</title>
        <p>Long Short-Term Memory networks [<xref ref-type="bibr" rid="B8">8</xref>] have demonstrated strong performance on longitudinal clinical prediction tasks, including sepsis onset, readmission prediction, and medication adherence modelling. The addition of multi-head self-attention [<xref ref-type="bibr" rid="B9">9</xref>] to LSTM outputs enables the model to selectively weight time points within the sequence based on their relevance to the target, complementing the LSTM’s recurrent memory with a global receptive field over the 30-day monitoring window. For mood state modelling, this is particularly relevant: the hypomanic prodrome may be concentrated in the final 5 - 10 days of the pre-antidepressant window, and attention weights over time steps should reflect this.</p>
      </sec>
      <sec id="sec2dot3">
        <title>2.3. Attention-Based Modelling of Episode-History Features</title>
        <p>Episode-feature attention networks require explicit nodes, edges, and neighbourhood aggregation [<xref ref-type="bibr" rid="B10">10</xref>][<xref ref-type="bibr" rid="B11">11</xref>]. The present episode-history module does not satisfy that definition because it receives a fixed patient-level feature vector. It is therefore treated as a feed-forward feature-attention model. True graph modelling of bipolar episode histories would require episode nodes, temporally ordered edges, optional pharmacological or similarity edges, and message passing over each patient’s adjacency structure.</p>
      </sec>
      <sec id="sec2dot4">
        <title>2.4. Digital Phenotyping of Bipolar II Hypomanic Prodrome</title>
        <p>Hypomanic activation in BD-II produces characteristic digital signatures: increased step count and physical activity, reduced sleep duration with preserved energy, increased communication frequency, greater smartphone usage entropy, and rising circadian amplitude (RA) with declining IS score [<xref ref-type="bibr" rid="B12">12</xref>][<xref ref-type="bibr" rid="B13">13</xref>]. These signatures, detectable through passive smartphone sensing and wrist actigraphy in the weeks preceding antidepressant initiation, constitute a potentially actionable prodromal risk window for TEM prediction.</p>
      </sec>
    </sec>
    <sec id="sec3">
      <title>3. Methods</title>
      <sec id="sec3dot1">
        <title>3.1. Synthetic Cohort Design</title>
        <p>The cohort comprised N = 750 fully synthetic BD-II patients with simulated antidepressant exposure and TEM prevalence of 22.0%. No real patient records were used. Marginal ranges and relative-risk directions were informed by antidepressant-switch literature, bipolar treatment guidance, episode-recurrence theory, and digital phenotyping studies [<xref ref-type="bibr" rid="B14">14</xref>]-[<xref ref-type="bibr" rid="B17">17</xref>]. Clinical and pharmacological continuous variables were sampled from truncated normal, log-normal, or beta distributions; binary variables from Bernoulli distributions; and antidepressant and mood-stabilizer classes from categorical distributions.</p>
        <p>Cross-variable dependence was imposed through correlated latent Gaussian factors transformed to each required marginal distribution. Prior TEM, rapid cycling, mixed-state history, recent hypomania, shorter episode intervals, and higher-risk antidepressant classes were positively correlated. Mood-stabilizer adequacy was negatively correlated with TEM risk. Circadian instability linked lower IS, higher IV, shorter and more variable sleep, increased activity, and rising YMRS sub-scores. Patient-specific random intercepts, day-level autoregressive noise, and measurement noise were added to the 30-day sequences.</p>
        <p>TEM labels were sampled from a probabilistic function of prior TEM, antidepressant-class risk, mood-stabilizer adequacy, recent hypomania, rapid cycling, episode-interval shortening, circadian instability, and selected interactions. The intercept was calibrated to 22.0% prevalence. Because related signals occur in both predictors and the label generator, performance may partly reflect recovery of the simulation rules; stochastic noise and overlapping class distributions reduce but do not eliminate this structural optimism.</p>
        <p>The complete cohort was synthesized first and then split once at the patient level into development, validation, and held-out test sets. The test set contained 113 patients and was not used for preprocessing, synthetic oversampling, hyperparameter selection, threshold selection, early stopping, or ensemble fitting. All partitions were stratified by TEM outcome [<xref ref-type="bibr" rid="B18">18</xref>].</p>
      </sec>
      <sec id="sec3dot2">
        <title>3.2. Feature Architecture</title>
        <p>Episode-history features (14) encoded depressive and hypomanic episode counts and severity summaries, mixed episode count and fraction, mean inter-episode interval and shortening trend, sequence entropy, depression-to-hypomania ratio, and recent hypomanic episode status. These were patient-level engineered variables, not nodes or edges in an explicit graph. Pharmacological, clinical, circadian, and 30-day longitudinal features were retained as originally specified.</p>
      </sec>
      <sec id="sec3dot3">
        <title>3.3. LSTM + Multi-Head Self-Attention Architecture</title>
        <p>The sequence encoder comprised a 2-layer LSTM (hidden dimension 128, dropout 0.2) processing 30  × 8 input sequences, followed by a 4-head multi-head self-attention module (attention dropout 0.1) over the full 30-day LSTM output sequence. Mean pooling over the attended sequence produced a 128-dimensional sequence embedding, projected to 64 dimensions and classified by a 2-layer head with GELU activation and dropout of 0.3. The self-attention mechanism enables the model to selectively upweight time steps corresponding to the hypomanic prodromal window rather than treating all 30 days equally.</p>
      </sec>
      <sec id="sec3dot4">
        <title>3.4. Episode-Feature Attention Network</title>
        <p>The episode-feature attention network operated on a fixed 14-dimensional patient-level episode-history vector. A two-layer fully connected encoder projected the vector to 128 dimensions, followed by a learned gating/attention function that reweighted latent feature dimensions before classification. There were no episode nodes, adjacency matrix, edge weights, neighbourhood sets, or graph message-passing steps. Accordingly, this component is a feed-forward attention block and is reported as such throughout the manuscript.</p>
      </sec>
      <sec id="sec3dot5">
        <title>3.5. Ensemble and Evaluation</title>
        <p>Five-fold stratified OOF stacking combined exactly three Level-1 learners: XGBoost, LightGBM, and LSTM + Attention. The episode-feature attention network was a standalone comparator and did not contribute to the stack. For each development fold, preprocessing and SMOTE were fit only on that fold’s training partition; predictions for the untouched fold formed the OOF meta-learner inputs. After hyperparameters were selected using development/validation data, base learners were refit on the full development set, and the logistic-regression meta-learner was applied to the held-out test predictions. The classification threshold was 0.27, selected on the validation set and locked before test evaluation. Leakage control was implemented at the patient level. No patient appeared in more than one partition; test observations were never used for feature scaling, SMOTE, early stopping, model selection, SHAP fitting, or threshold optimisation. The simulator used common population-level rules but independent patient-level random draws, so no duplicated templates or trajectories were shared across partitions. Nevertheless, the shared generative mechanism remains a source of optimism and is explicitly acknowledged.</p>
      </sec>
    </sec>
    <sec id="sec4">
      <title>4. Results</title>
      <sec id="sec4dot1">
        <title>4.1. Calibration and Clinical Utility</title>
        <p><xref ref-type="fig" rid="fig1">Figure 1</xref><xref ref-type="fig" rid="fig1">Figure 1</xref> presents calibration curves on the held-out synthetic test set. The proposed ensemble had the lowest Brier score (0.025), followed by LSTM + Attention (0.036). Given the modest number of TEM events and the synthetic label mechanism, these values indicate good apparent calibration within this simulation rather than near-perfect real-world calibration. <xref ref-type="fig" rid="fig2">Figure 2</xref><xref ref-type="fig" rid="fig2">Figure 2</xref> reports simulated decision-curve results and should not be interpreted as evidence of clinical utility.</p>
        <fig id="fig1">
          <label>Figure 1</label>
          <graphic xlink:href="https://html.scirp.org/file/1115671-rId16.jpeg?20260824033316" />
        </fig>
        <p><bold>Figure 1.</bold> Calibration curves (reliability diagrams) for all seven models. Brier scores: Proposed Ensemble 0.025 (lowest), LSTM + Attention 0.036, LightGBM 0.068, XGBoost 0.071, LR 0.093, RF 0.110, GAT 0.110. The ensemble and LSTM + Attention show the nearest tracking to perfect calibration. Lower Brier score = better.</p>
        <fig id="fig2">
          <label>Figure 2</label>
          <graphic xlink:href="https://html.scirp.org/file/1115671-rId17.jpeg?20260824033316" />
        </fig>
        <p><bold>Figure 2.</bold> Decision curve analysis: net benefit versus threshold probability (0.0 - 0.70). The Proposed Ensemble and LSTM + Attention module achieves the highest net clinical benefit across threshold probabilities 0.15 - 0.55, confirming the superior clinical utility of longitudinal sequence-informed TEM risk stratification over tabular models alone.</p>
      </sec>
      <sec id="sec4dot2">
        <title>4.2. Discriminative Performance</title>
        <p><bold>Table 1</bold> presents performance on the held-out synthetic test set. The proposed three-model stack (XGBoost + LightGBM + LSTM + Attention) achieved AUC = 0.997, F1 = 0.958, sensitivity = 0.958, and specificity = 0.989. LSTM + Attention achieved AUC = 0.984. The episode-feature attention comparator achieved AUC = 0.914. The high estimates are proof-of-concept results and may be optimistic because the simulator embeds predictive temporal and pharmacological structure.</p>
        <p><bold>Table 1.</bold>Comparative model performance—test set (n = 113).</p>
        <table-wrap id="tbl1">
          <label>Table 1</label>
          <table>
            <tbody>
              <tr>
                <td>
                  <bold>Model</bold>
                </td>
                <td>
                  <bold>AUC</bold>
                </td>
                <td>
                  <bold>F1</bold>
                </td>
                <td>
                  <bold>Sensitivity</bold>
                </td>
                <td>
                  <bold>Specificity</bold>
                </td>
                <td>
                  <bold>Precision</bold>
                </td>
                <td>
                  <bold>Brier</bold>
                </td>
              </tr>
              <tr>
                <td>Logistic Regression</td>
                <td>0.953</td>
                <td>0.766</td>
                <td>0.802</td>
                <td>0.919</td>
                <td>0.730</td>
                <td>0.081</td>
              </tr>
              <tr>
                <td>Random Forest</td>
                <td>0.935</td>
                <td>0.770</td>
                <td>0.681</td>
                <td>0.977</td>
                <td>0.875</td>
                <td>0.097</td>
              </tr>
              <tr>
                <td>XGBoost</td>
                <td>0.943</td>
                <td>0.762</td>
                <td>0.719</td>
                <td>0.955</td>
                <td>0.810</td>
                <td>0.071</td>
              </tr>
              <tr>
                <td>LightGBM</td>
                <td>0.941</td>
                <td>0.790</td>
                <td>0.762</td>
                <td>0.955</td>
                <td>0.821</td>
                <td>0.068</td>
              </tr>
              <tr>
                <td>LSTM + Attention</td>
                <td>0.984</td>
                <td>0.913</td>
                <td>0.876</td>
                <td>0.989</td>
                <td>0.952</td>
                <td>0.036</td>
              </tr>
              <tr>
                <td>Episode-Feature Attention</td>
                <td>0.914</td>
                <td>0.665</td>
                <td>0.723</td>
                <td>0.874</td>
                <td>0.614</td>
                <td>0.110</td>
              </tr>
              <tr>
                <td>
                  <bold>Proposed Ensemble (</bold>
                  <bold>Proposed</bold>
                  <bold>)</bold>
                </td>
                <td>
                  <bold>0.997</bold>
                </td>
                <td>
                  <bold>0.958</bold>
                </td>
                <td>
                  <bold>0.958</bold>
                </td>
                <td>
                  <bold>0.989</bold>
                </td>
                <td>
                  <bold>0.958</bold>
                </td>
                <td>
                  <bold>0.025</bold>
                </td>
              </tr>
            </tbody>
          </table>
        </table-wrap>
        <p>Note: 95% bootstrap CI (1000 resamples): AUC [0.988 - 1.000], F1 [0.885 - 1.000], sensitivity [0.863 - 1.000], specificity [0.963 - 1.000]. Threshold = 0.27 (F1-optimised on validation set).</p>
      </sec>
      <sec id="sec4dot3">
        <title>4.3. SHAP Interpretability</title>
        <p><xref ref-type="fig" rid="fig3">Figure 3</xref><xref ref-type="fig" rid="fig3">Figure 3</xref> presents SHAP values from the XGBoost and LightGBM base learners after feature alignment. These attributions do not directly explain the LSTM sequence encoder or the logistic stacking meta-learner and, therefore, should not be described as unified explanations of the complete ensemble. Prior TEM history, antidepressant-class risk, mood-stabilizer adequacy, IS score, and episode-interval shortening were the largest tree-model attributions.</p>
        <fig id="fig3">
          <label>Figure 3</label>
          <graphic xlink:href="https://html.scirp.org/file/1115671-rId18.jpeg?20260824033316" />
        </fig>
        <p><bold>Figure 3.</bold> SHAP beeswarm—top 20 predictors of TEM in BD-II. Combined XGBoost + LightGBM TreeExplainer SHAP values. Each point represents one test patient. Red: high feature value; blue: low feature value. Prior TEM history, AD class risk, mood stabilizer adequacy, IS score, and episode interval shortening are the five dominant predictors.</p>
      </sec>
      <sec id="sec4dot4">
        <title>4.4. ROC Curves</title>
        <p><xref ref-type="fig" rid="fig4">Figure 4</xref><xref ref-type="fig" rid="fig4">Figure 4</xref> presents ROC curves with bootstrap confidence bands. The LSTM + Attention point-estimate AUC exceeded the strongest tabular baseline by 0.041 in the synthetic test set. This is consistent with the simulator assigning predictive temporal structure to the 30-day sequence, but it does not establish that the same gain will occur in prospectively collected digital phenotyping data. The episode-feature attention comparator performed below the tabular models.</p>
        <fig id="fig4">
          <label>Figure 4</label>
          <graphic xlink:href="https://html.scirp.org/file/1115671-rId19.jpeg?20260824033316" />
        </fig>
        <p><bold>Figure 4.</bold> ROC curves with 95% bootstrap CI bands (400 resamples). LSTM + Attention (purple) substantially outperforms all tabular baselines (ΔAUC = 0.041 vs. XGBoost). Proposed ensemble (dark bold) achieves AUC = 0.997. Episode-Feature Attention (teal) shows the widest CI, consistent with its limited episode-only feature scope.</p>
      </sec>
      <sec id="sec4dot5">
        <title>4.5. Confusion Matrix</title>
        <fig id="fig5">
          <label>Figure 5</label>
          <graphic xlink:href="https://html.scirp.org/file/1115671-rId20.jpeg?20260824033316" />
        </fig>
        <p><bold>Figure 5.</bold> Confusion matrix—proposed ensemble (threshold = 0.27). True TEM: 24 correctly classified (sensitivity = 0.958). True non-TEM: 87 correctly classified (specificity = 0.989). False negatives = 1; false positives = 1.</p>
        <p><xref ref-type="fig" rid="fig5">Figure 5</xref><xref ref-type="fig" rid="fig5">Figure 5</xref> presents the confusion matrix at the validation-selected threshold of 0.27. Of 113 synthetic test patients, 111 were correctly classified. Among 25 TEM cases, 24 were identified; among 88 non-TEM cases, 87 were identified. These counts are descriptive of one held-out simulated test set.</p>
      </sec>
      <sec id="sec4dot6">
        <title>4.6. SHAP Dependence Analysis</title>
        <p><xref ref-type="fig" rid="fig6">Figure 6</xref><xref ref-type="fig" rid="fig6">Figure 6</xref> presents SHAP dependence plots for the three highest-ranked predictors. The prior TEM history dependence plot confirms a bimodal SHAP distribution (positive history  →  strong positive contributions; no history  →  consistent negative contributions). The antidepressant class risk plot shows a monotonically increasing SHAP-to-risk relationship across the continuous risk index (0.12 for bupropion to 0.45 for TCAs), with an inflection above risk index 0.30 suggesting a non-linear acceleration of TEM probability at higher-risk class exposures. The mood stabilizer adequacy plot confirms a monotonically protective dose-response, with the steepest SHAP gradient in the 0.3 - 0.7 adequacy range, suggesting that partial mood stabilizer coverage provides meaningful but incomplete protection.</p>
        <fig id="fig6">
          <label>Figure 6</label>
          <graphic xlink:href="https://html.scirp.org/file/1115671-rId21.jpeg?20260824033316" />
        </fig>
        <p><bold>Figure 6.</bold> SHAP dependence plots for the three dominant predictors. Left: prior TEM history (bimodal distribution). Centre: AD class pharmacological risk index (monotonically increasing, inflection at risk &gt; 0.30). Right: mood stabilizer adequacy (monotonically protective, steepest gradient 0.3 - 0.7). Curves: second-degree polynomial fits.</p>
      </sec>
      <sec id="sec4dot7">
        <title>4.7. Bayesian Model Comparison</title>
        <p><bold>Table 2.</bold>Bayesian model comparison: BIC, WAIC, and Bayes factors.</p>
        <table-wrap id="tbl2">
          <label>Table 2</label>
          <table>
            <tbody>
              <tr>
                <td>
                  <bold>Model</bold>
                </td>
                <td>
                  <bold>Log</bold>
                  <bold>-</bold>
                  <bold>Lik.</bold>
                </td>
                <td>
                  <bold>k</bold>
                </td>
                <td>
                  <bold>BIC</bold>
                </td>
                <td>
                  <bold>WAIC</bold>
                </td>
                <td>
                  <bold>log₁₀(BF)</bold>
                </td>
                <td>
                  <bold>Evidence</bold>
                </td>
              </tr>
              <tr>
                <td>Logistic Regression</td>
                <td>−30.2</td>
                <td>57</td>
                <td>301.5</td>
                <td>60.9</td>
                <td>57.7</td>
                <td>Decisive</td>
              </tr>
              <tr>
                <td>Random Forest</td>
                <td>−39.0</td>
                <td>40</td>
                <td>267.1</td>
                <td>78.0</td>
                <td>50.2</td>
                <td>Decisive</td>
              </tr>
              <tr>
                <td>XGBoost</td>
                <td>−27.2</td>
                <td>60</td>
                <td>338.0</td>
                <td>54.8</td>
                <td>65.6</td>
                <td>Decisive</td>
              </tr>
              <tr>
                <td>LightGBM</td>
                <td>−26.7</td>
                <td>60</td>
                <td>337.0</td>
                <td>53.6</td>
                <td>65.4</td>
                <td>Decisive</td>
              </tr>
              <tr>
                <td>LSTM + Attention</td>
                <td>−23.7</td>
                <td>50</td>
                <td>283.7</td>
                <td>50.8</td>
                <td>53.8</td>
                <td>Decisive</td>
              </tr>
              <tr>
                <td>Episode + Feature Attention</td>
                <td>−39.0</td>
                <td>35</td>
                <td>243.4</td>
                <td>77.9</td>
                <td>45.1</td>
                <td>Decisive</td>
              </tr>
              <tr>
                <td>
                  <bold>Proposed Ensemble (</bold>
                  <bold>Proposed</bold>
                  <bold>)</bold>
                </td>
                <td>
                  <bold>−</bold>
                  <bold>8.5</bold>
                </td>
                <td>
                  <bold>4</bold>
                </td>
                <td>
                  <bold>35.9</bold>
                </td>
                <td>
                  <bold>17.4</bold>
                </td>
                <td>
                  <bold>0.0 (ref.)</bold>
                </td>
                <td>
                  <bold>Reference</bold>
                </td>
              </tr>
            </tbody>
          </table>
        </table-wrap>
        <p>Note: log<sub>10</sub>(BF): originally reported base-10 Bayes-factor calculation relative to the proposed ensemble. Because the ensemble complexity count excludes its base learners, these values are exploratory and should not be interpreted as decisive posterior evidence.</p>
        <p><bold>Table 2</bold> and <xref ref-type="fig" rid="fig7">Figure 7</xref><xref ref-type="fig" rid="fig7">Figure 7</xref> report the original exploratory BIC, WAIC, and Bayes-factor calculations, with the ensemble BIC consistently reported as 35.9. For neural networks, boosting models, and stacked ensembles, likelihoods and effective parameter counts are not uniquely defined. Counting only four meta-learner coefficients for the stack omits the complexity of its three fitted base learners and mechanically favours the ensemble. Accordingly, these criteria are retained for transparency but are not treated as decisive evidence of superiority.</p>
        <fig id="fig7">
          <label>Figure 7</label>
          <graphic xlink:href="https://html.scirp.org/file/1115671-rId22.jpeg?20260824033316" />
        </fig>
        <p><bold>Figure 7.</bold>Exploratory Bayesian model comparison. Left: reported BIC values (proposed ensemble = 35.9, consistent with <bold>Table 2</bold>). Right: reported Bayes-factor calculations. Because the ensemble parameter count excludes base-learner complexity, these comparisons should be interpreted cautiously.</p>
      </sec>
      <sec id="sec4dot8">
        <title>4.8. Subgroup Analysis</title>
        <p><bold>Table 3</bold> presents exploratory subgroup point estimates. Several groups are small, and estimates near 1.00 are unstable and may reflect the synthetic label mechanism. Subgroup-level prediction vectors were not available in the manuscript package, so valid bootstrap confidence intervals could not be reconstructed without inventing results. The table, therefore, explicitly marks confidence intervals as not reported (NR); they should be added from the completed analysis outputs before resubmission.</p>
        <p><bold>Table 3.</bold>Subgroup analysis—proposed ensemble performance.</p>
        <table-wrap id="tbl3">
          <label>Table 3</label>
          <table>
            <tbody>
              <tr>
                <td>
                  <bold>Subgroup</bold>
                </td>
                <td>
                  <bold>N</bold>
                </td>
                <td>
                  <bold>AUC</bold>
                </td>
                <td>
                  <bold>F1</bold>
                </td>
                <td>
                  <bold>Sensitivity</bold>
                </td>
                <td>
                  <bold>95% CI</bold>
                </td>
              </tr>
              <tr>
                <td>Prior TEM History</td>
                <td>17</td>
                <td>1.000</td>
                <td>0.968</td>
                <td>0.938</td>
                <td>NR</td>
              </tr>
              <tr>
                <td>No Adequate Mood Stabilizer</td>
                <td>77</td>
                <td>0.995</td>
                <td>0.947</td>
                <td>0.947</td>
                <td>NR</td>
              </tr>
              <tr>
                <td>High-Risk AD Class (TCA/SNRI/MAOI)</td>
                <td>39</td>
                <td>0.991</td>
                <td>0.960</td>
                <td>1.000</td>
                <td>NR</td>
              </tr>
              <tr>
                <td>Recent Hypomanic Episode</td>
                <td>20</td>
                <td>1.000</td>
                <td>1.000</td>
                <td>1.000</td>
                <td>NR</td>
              </tr>
              <tr>
                <td>Low IS Score (Circadian Disruption)</td>
                <td>56</td>
                <td>0.992</td>
                <td>0.917</td>
                <td>0.917</td>
                <td>NR</td>
              </tr>
              <tr>
                <td>Female Sex</td>
                <td>65</td>
                <td>0.996</td>
                <td>0.938</td>
                <td>0.938</td>
                <td>NR</td>
              </tr>
            </tbody>
          </table>
        </table-wrap>
        <p>Note: Test set (n = 113). Low IS score: below test-set median. High-risk AD class: TCA, SNRI, or MAOI. All subgroup analyses are exploratory. NR = confidence interval not reconstructable from the available aggregate outputs; subgroup-level bootstrap intervals must be inserted from the analysis files.</p>
      </sec>
      <sec id="sec4dot9">
        <title>4.9. SHAP Waterfall-Representative Patients</title>
        <fig id="fig8">
          <label>Figure 8</label>
          <graphic xlink:href="https://html.scirp.org/file/1115671-rId23.jpeg?20260824033316" />
        </fig>
        <p><bold>Figure 8.</bold> SHAP waterfall for the highest-confidence TEM prediction (left) and highest-confidence non-TEM prediction (right).</p>
        <p>The TEM case is characterised by prior TEM history, high-risk AD class, low IS score, and recent hypomanic episode. The non-TEM case is protected by adequate mood stabilizer coverage, a high IS score, and a low-risk AD class. Red: features increasing TEM risk; Blue: features reducing TEM risk (See <xref ref-type="fig" rid="fig8">Figure 8</xref><xref ref-type="fig" rid="fig8">Figure 8</xref>).</p>
      </sec>
    </sec>
    <sec id="sec5">
      <title>5. Discussion</title>
      <p>The principal methodological observation is a 0.041 point-estimate AUC difference between LSTM + Attention and the strongest tabular baseline in one synthetic test set. Because the simulated 30-day trajectories were generated to contain circadian and behavioural risk signals, this difference demonstrates recovery of the designed temporal signal. It is a hypothesis for prospective evaluation, not evidence supporting a mandatory monitoring protocol [<xref ref-type="bibr" rid="B19">19</xref>]-[<xref ref-type="bibr" rid="B21">21</xref>].</p>
      <p>The IS-score attribution is consistent with the simulated circadian-risk structure and with prior bipolar rhythm research. The apparent non-linear threshold near 0.45 is model- and simulator-specific and should not be used to defer treatment or prescribe a chronobiological intervention without clinical validation.</p>
      <p>The high sensitivity in the recent-hypomania subgroup is exploratory. Recent hypomania was included in the synthetic TEM-generating process, so the result does not constitute independent construct validation [<xref ref-type="bibr" rid="B22">22</xref>]-[<xref ref-type="bibr" rid="B25">25</xref>].</p>
    </sec>
    <sec id="sec6">
      <title>6. Limitations</title>
      <p>The cohort, longitudinal streams, and labels are fully synthetic, and performance may primarily reflect recovery of the shared generative mechanism. The episode-history attention component is not a graph neural network and does not perform message passing. The proposed ensemble excludes that component and contains only XGBoost, LightGBM, and LSTM + Attention. Bootstrap intervals from a single test set do not quantify variation across independent simulations or seeds. Subgroup confidence intervals require the original patient-level prediction files. The BIC/Bayes-factor comparison uses an ensemble parameter count that omits base-learner complexity. Real-world digital streams will introduce missingness, device heterogeneity, non-wear, ascertainment error, and distribution shift. Recent reviews and reporting guidance emphasise external evaluation, transparent model description, and caution when translating digital phenotyping models into clinical care [<xref ref-type="bibr" rid="B26">26</xref>]-[<xref ref-type="bibr" rid="B29">29</xref>].</p>
    </sec>
    <sec id="sec7">
      <title>7. Conclusion</title>
      <p>This synthetic-data proof-of-concept evaluates longitudinal and tabular modelling for TEM risk in BD-II. The proposed OOF stack consistently comprises XGBoost, LightGBM, and LSTM + Attention and achieved AUC = 0.997 in one held-out simulated test set. The separate episode-feature attention comparator is not a graph neural network. The high performance, feature attributions, decision curves, Bayesian comparisons, and subgroup estimates require cautious interpretation because they are conditioned by the simulator. Next steps are repeated simulations across independent seeds, reporting of subgroup confidence intervals from patient-level outputs, evaluation under realistic missingness and device shift, and external prospective validation before any prescribing workflow integration.</p>
    </sec>
    <sec id="sec8">
      <title>Author Contributions</title>
      <p>Conceptualization: RDF and AAF; Methodology: AAF and RDF; Software: AAF; Validation: RDF and AAF; Formal analysis: AAF; Investigation: RDF and AAF; Resources: RDF; Data curation: AAF; Writing original draft preparation: AAF; Writing review and editing: RDF and AAF; Visualization: AAF; Supervision: RDF; Project administration: RDF. All authors have read and agreed to the published version of the manuscript.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <title>References</title>
      <ref id="B1">
        <label>1.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Altshuler, L.L., Post, R.M., Leverich, G.S., <italic>et al</italic>. (1995) Antidepressant-Induced Mania and Cycle Acceleration: A Controversy Revisited. <italic>American Journal of Psychia</italic><italic>try</italic>, 152, 1130-1138.</mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Altshuler, L.L.</string-name>
              <string-name>Post, R.M.</string-name>
              <string-name>Leverich, G.S.</string-name>
            </person-group>
            <year>1995</year>
            <article-title>Antidepressant-Induced Mania and Cycle Acceleration: A Controversy Revisited</article-title>
            <source>American Journal of Psychiatry</source>
            <volume>152</volume>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B2">
        <label>2.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Wehr, T.A. and Goodwin, F.K. (1987) Can Antidepressants Cause Mania and Worsen the Course of Affective Illness? <italic>American Journal of Psychiatry</italic>, 144, 1403-1411.</mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Wehr, T.A.</string-name>
              <string-name>Goodwin, F.K.</string-name>
            </person-group>
            <year>1987</year>
            <article-title>Can Antidepressants Cause Mania and Worsen the Course of Affective Illness? American Journal of Psychiatry, 144, 1403-1411</article-title>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B3">
        <label>3.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Gijsman, H.J., Geddes, J.R., Rendell, J.M., Nolen, W.A. and Goodwin, G.M. (2004) Antidepressants for Bipolar Depression: A Systematic Review of Randomized, Controlled Trials. <italic>American Journal of Psychiatry</italic>, 161, 1537-1547. https://doi.org/10.1176/appi.ajp.161.9.1537 <pub-id pub-id-type="doi">10.1176/appi.ajp.161.9.1537</pub-id><pub-id pub-id-type="pmid">15337640</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1176/appi.ajp.161.9.1537">https://doi.org/10.1176/appi.ajp.161.9.1537</ext-link></mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Gijsman, H.J.</string-name>
              <string-name>Geddes, J.R.</string-name>
              <string-name>Rendell, J.M.</string-name>
              <string-name>Nolen, W.A.</string-name>
              <string-name>Goodwin, G.M.</string-name>
              <string-name>Randomized, C</string-name>
            </person-group>
            <year>2004</year>
            <article-title>Antidepressants for Bipolar Depression: A Systematic Review of Randomized, Controlled Trials</article-title>
            <source>American Journal of Psychiatry</source>
            <volume>161</volume>
            <pub-id pub-id-type="doi">10.1176/appi.ajp.161.9.1537</pub-id>
            <pub-id pub-id-type="pmid">15337640</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B4">
        <label>4.</label>
        <citation-alternatives>
          <mixed-citation publication-type="confproc">Suppes, T., Dennehy, E.B., Swann, A.C., Bowden, C.L., Calabrese, J.R., Hirschfeld, R.M.A., <italic>et al</italic>. (2002) Report of the Texas Consensus Conference Panel on Medication Treatment of Bipolar Disorder 2000. <italic>The Journal of Clinical Psychiatry</italic>, 63, 288-299. https://doi.org/10.4088/jcp.v63n0404 <pub-id pub-id-type="doi">10.4088/jcp.v63n0404</pub-id><pub-id pub-id-type="pmid">12004801</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.4088/jcp.v63n0404">https://doi.org/10.4088/jcp.v63n0404</ext-link></mixed-citation>
          <element-citation publication-type="confproc">
            <person-group person-group-type="author">
              <string-name>Suppes, T.</string-name>
              <string-name>Dennehy, E.B.</string-name>
              <string-name>Swann, A.C.</string-name>
              <string-name>Bowden, C.L.</string-name>
              <string-name>Calabrese, J.R.</string-name>
              <string-name>Hirschfeld, R.M.A.</string-name>
            </person-group>
            <year>2002</year>
            <article-title>Report of the Texas Consensus Conference Panel on Medication Treatment of Bipolar Disorder 2000</article-title>
            <source>The Journal of Clinical Psychiatry</source>
            <volume>63</volume>
            <pub-id pub-id-type="doi">10.4088/jcp.v63n0404</pub-id>
            <pub-id pub-id-type="pmid">12004801</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B5">
        <label>5.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Tondo, L. and Baldessarini, R.J. (1998) Rapid Cycling in Women and Men with Bipolar Manic-Depressive Disorders. <italic>American Journal of Psychiatry</italic>, 155, 1434-1436. https://doi.org/10.1176/ajp.155.10.1434 <pub-id pub-id-type="doi">10.1176/ajp.155.10.1434</pub-id><pub-id pub-id-type="pmid">9766777</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1176/ajp.155.10.1434">https://doi.org/10.1176/ajp.155.10.1434</ext-link></mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Tondo, L.</string-name>
              <string-name>Baldessarini, R.J.</string-name>
            </person-group>
            <year>1998</year>
            <article-title>Rapid Cycling in Women and Men with Bipolar Manic-Depressive Disorders</article-title>
            <source>American Journal of Psychiatry</source>
            <volume>155</volume>
            <pub-id pub-id-type="doi">10.1176/ajp.155.10.1434</pub-id>
            <pub-id pub-id-type="pmid">9766777</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B6">
        <label>6.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Yatham, L.N., Kennedy, S.H., Parikh, S.V., Schaffer, A., Bond, D.J., Frey, B.N., <italic>et</italic><italic>al</italic>. (2018) Canadian Network for Mood and Anxiety Treatments (CANMAT) and International Society for Bipolar Disorders (ISBD) 2018 Guidelines for the Management of Patients with Bipolar Disorder. <italic>Bipolar Disorders</italic>, 20, 97-170. https://doi.org/10.1111/bdi.12609 <pub-id pub-id-type="doi">10.1111/bdi.12609</pub-id><pub-id pub-id-type="pmid">29536616</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1111/bdi.12609">https://doi.org/10.1111/bdi.12609</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Yatham, L.N.</string-name>
              <string-name>Kennedy, S.H.</string-name>
              <string-name>Parikh, S.V.</string-name>
              <string-name>Schaffer, A.</string-name>
              <string-name>Bond, D.J.</string-name>
              <string-name>Frey, B.N.</string-name>
            </person-group>
            <year>2018</year>
            <article-title>Canadian Network for Mood and Anxiety Treatments (CANMAT) and International Society for Bipolar Disorders (ISBD) 2018 Guidelines for the Management of Patients with Bipolar Disorder</article-title>
            <source>Bipolar Disorders</source>
            <volume>20</volume>
            <pub-id pub-id-type="doi">10.1111/bdi.12609</pub-id>
            <pub-id pub-id-type="pmid">29536616</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B7">
        <label>7.</label>
        <citation-alternatives>
          <mixed-citation publication-type="book">Frank, E. (2005) Treating Bipolar Disorder: A Clinician’s Guide to Interpersonal and Social Rhythm Therapy. Guilford Press.</mixed-citation>
          <element-citation publication-type="book">
            <person-group person-group-type="author">
              <string-name>Frank, E.</string-name>
            </person-group>
            <year>2005</year>
            <article-title>Treating Bipolar Disorder: A Clinician’s Guide to Interpersonal and Social Rhythm Therapy</article-title>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B8">
        <label>8.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Hochreiter, S. and Schmidhuber, J. (1997) Long Short-Term Memory. <italic>Neural Computation</italic>, 9, 1735-1780. https://doi.org/10.1162/neco.1997.9.8.1735 <pub-id pub-id-type="doi">10.1162/neco.1997.9.8.1735</pub-id><pub-id pub-id-type="pmid">9377276</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1162/neco.1997.9.8.1735">https://doi.org/10.1162/neco.1997.9.8.1735</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Hochreiter, S.</string-name>
              <string-name>Schmidhuber, J.</string-name>
            </person-group>
            <year>1997</year>
            <article-title>Long Short-Term Memory</article-title>
            <source>Neural Computation</source>
            <volume>9</volume>
            <pub-id pub-id-type="doi">10.1162/neco.1997.9.8.1735</pub-id>
            <pub-id pub-id-type="pmid">9377276</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B9">
        <label>9.</label>
        <citation-alternatives>
          <mixed-citation publication-type="confproc">Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., <italic>et al</italic>. (2017) Attention Is All You Need. 31 <italic>st Conference on Neural Information Processing Systems</italic>( <italic>NIPS</italic> 2017), Long Beach, 4-9 December 2017, 1-11.</mixed-citation>
          <element-citation publication-type="confproc">
            <person-group person-group-type="author">
              <string-name>Vaswani, A.</string-name>
              <string-name>Shazeer, N.</string-name>
              <string-name>Parmar, N.</string-name>
              <string-name>Uszkoreit, J.</string-name>
            </person-group>
            <year>2017</year>
            <article-title>Attention Is All You Need</article-title>
            <source>31st Conference on Neural Information Processing Systems (NIPS 2017)</source>
            <volume>4</volume>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B10">
        <label>10.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Su, G., Wang, H., Zhang, Y., Zhang, W. and Lin, X. (2024) Simple and Deep Graph Attention Networks. <italic>Knowledge</italic>- <italic>Based Systems</italic>, 293, Article 111649. https://doi.org/10.1016/j.knosys.2024.111649 <pub-id pub-id-type="doi">10.1016/j.knosys.2024.111649</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1016/j.knosys.2024.111649">https://doi.org/10.1016/j.knosys.2024.111649</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Su, G.</string-name>
              <string-name>Wang, H.</string-name>
              <string-name>Zhang, Y.</string-name>
              <string-name>Zhang, W.</string-name>
              <string-name>Lin, X.</string-name>
            </person-group>
            <year>2024</year>
            <article-title>Simple and Deep Graph Attention Networks</article-title>
            <source>Knowledge-Based Systems</source>
            <volume>293</volume>
            <elocation-id>111649</elocation-id>
            <pub-id pub-id-type="doi">10.1016/j.knosys.2024.111649</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B11">
        <label>11.</label>
        <citation-alternatives>
          <mixed-citation publication-type="confproc">Wang, X., Ji, H.Y., Shi, C., Wang, B., <italic>et al</italic>. (2019) Heterogeneous Graph Attention Network. <italic>The World Wide Web Conference</italic>, San Francisco, 13-17 May 2019, 2022-2032. https://doi.org/10.1145/3308558.3313562 <pub-id pub-id-type="doi">10.1145/3308558.3313562</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1145/3308558.3313562">https://doi.org/10.1145/3308558.3313562</ext-link></mixed-citation>
          <element-citation publication-type="confproc">
            <person-group person-group-type="author">
              <string-name>Wang, X.</string-name>
              <string-name>Ji, H.Y.</string-name>
              <string-name>Shi, C.</string-name>
              <string-name>Wang, B.</string-name>
              <string-name>Conference, S</string-name>
            </person-group>
            <year>2019</year>
            <article-title>Heterogeneous Graph Attention Network</article-title>
            <source>The World Wide Web Conference</source>
            <volume>13</volume>
            <pub-id pub-id-type="doi">10.1145/3308558.3313562</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B12">
        <label>12.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Post, R.M. (2007) Kindling and Sensitization as Models for Affective Episode Recurrence, Cyclicity, and Tolerance Phenomena. <italic>Neuroscience &amp; Biobehavioral Reviews</italic>, 31, 858-873. https://doi.org/10.1016/j.neubiorev.2007.04.003 <pub-id pub-id-type="doi">10.1016/j.neubiorev.2007.04.003</pub-id><pub-id pub-id-type="pmid">17555817</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1016/j.neubiorev.2007.04.003">https://doi.org/10.1016/j.neubiorev.2007.04.003</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Post, R.M.</string-name>
              <string-name>Recurrence, C</string-name>
            </person-group>
            <year>2007</year>
            <article-title>Kindling and Sensitization as Models for Affective Episode Recurrence, Cyclicity, and Tolerance Phenomena</article-title>
            <source>Neuroscience &amp; Biobehavioral Reviews</source>
            <volume>31</volume>
            <pub-id pub-id-type="doi">10.1016/j.neubiorev.2007.04.003</pub-id>
            <pub-id pub-id-type="pmid">17555817</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B13">
        <label>13.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Boudebesse, C., Geoffroy, F. and Bellivier, F. (2014) Correlations between Sleep and Circadian Rhythms, and the Recurrences in Bipolar Disorder: A Systematic Review. <italic>Chronobiology International</italic>, 31, 696-712.</mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Boudebesse, C.</string-name>
              <string-name>Geoffroy, F.</string-name>
              <string-name>Bellivier, F.</string-name>
            </person-group>
            <year>2014</year>
            <article-title>Correlations between Sleep and Circadian Rhythms, and the Recurrences in Bipolar Disorder: A Systematic Review</article-title>
            <source>Chronobiology International</source>
            <volume>31</volume>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B14">
        <label>14.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">DeLong, E.R., DeLong, D.M. and Clarke-Pearson, D.L. (1988) Comparing the Areas under Two or More Correlated Receiver Operating Characteristic Curves: A Nonparametric Approach. <italic>Biometrics</italic>, 44, 837-844. https://doi.org/10.2307/2531595 <pub-id pub-id-type="doi">10.2307/2531595</pub-id><pub-id pub-id-type="pmid">3203132</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.2307/2531595">https://doi.org/10.2307/2531595</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>DeLong, E.R.</string-name>
              <string-name>DeLong, D.M.</string-name>
              <string-name>Clarke-Pearson, D.L.</string-name>
            </person-group>
            <year>1988</year>
            <article-title>Comparing the Areas under Two or More Correlated Receiver Operating Characteristic Curves: A Nonparametric Approach</article-title>
            <source>Biometrics</source>
            <volume>44</volume>
            <pub-id pub-id-type="doi">10.2307/2531595</pub-id>
            <pub-id pub-id-type="pmid">3203132</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B15">
        <label>15.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Kass, R.E. and Raftery, A.E. (1995) Bayes Factors. <italic>Journal of the American Statistical Association</italic>, 90, 773-795. https://doi.org/10.1080/01621459.1995.10476572 <pub-id pub-id-type="doi">10.1080/01621459.1995.10476572</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1080/01621459.1995.10476572">https://doi.org/10.1080/01621459.1995.10476572</ext-link></mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Kass, R.E.</string-name>
              <string-name>Raftery, A.E.</string-name>
            </person-group>
            <year>1995</year>
            <article-title>Bayes Factors</article-title>
            <source>Journal of the American Statistical Association</source>
            <volume>90</volume>
            <pub-id pub-id-type="doi">10.1080/01621459.1995.10476572</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B16">
        <label>16.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Vickers, A.J. and Elkin, E.B. (2006) Decision Curve Analysis: A Novel Method for Evaluating Prediction Models. <italic>Medical Decision Making</italic>, 26, 565-574. https://doi.org/10.1177/0272989x06295361 <pub-id pub-id-type="doi">10.1177/0272989x06295361</pub-id><pub-id pub-id-type="pmid">17099194</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1177/0272989x06295361">https://doi.org/10.1177/0272989x06295361</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Vickers, A.J.</string-name>
              <string-name>Elkin, E.B.</string-name>
            </person-group>
            <year>2006</year>
            <article-title>Decision Curve Analysis: A Novel Method for Evaluating Prediction Models</article-title>
            <source>Medical Decision Making</source>
            <volume>26</volume>
            <pub-id pub-id-type="doi">10.1177/0272989x06295361</pub-id>
            <pub-id pub-id-type="pmid">17099194</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B17">
        <label>17.</label>
        <citation-alternatives>
          <mixed-citation publication-type="confproc">Chen, T.Q. and Guestrin, C. (2016) XGBoost: A Scalable Tree Boosting System. <italic>Proceedings of the</italic>22 <italic>nd ACM SIGKDD International Conference on Knowledge Dis</italic><italic>covery and Data Mining</italic>, San Francisco, 13-17 August 2016, 785-794. https://doi.org/10.1145/2939672.2939785 <pub-id pub-id-type="doi">10.1145/2939672.2939785</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1145/2939672.2939785">https://doi.org/10.1145/2939672.2939785</ext-link></mixed-citation>
          <element-citation publication-type="confproc">
            <person-group person-group-type="author">
              <string-name>Chen, T.Q.</string-name>
              <string-name>Guestrin, C.</string-name>
              <string-name>Mining, S</string-name>
            </person-group>
            <year>2016</year>
            <article-title>XGBoost: A Scalable Tree Boosting System</article-title>
            <source>Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining</source>
            <volume>13</volume>
            <pub-id pub-id-type="doi">10.1145/2939672.2939785</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B18">
        <label>18.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Ke, G.L., Qi, M., Finley, T., Wang, T.F., <italic>et al</italic>. (2017) LightGBM: A Highly Efficient Gradient Boosting Decision Tree. <italic>Advances in Neural Information Processing Systems</italic>, 30, 3146-3154.</mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Ke, G.L.</string-name>
              <string-name>Qi, M.</string-name>
              <string-name>Finley, T.</string-name>
              <string-name>Wang, T.F.</string-name>
            </person-group>
            <year>2017</year>
            <article-title>LightGBM: A Highly Efficient Gradient Boosting Decision Tree</article-title>
            <source>Advances in Neural Information Processing Systems</source>
            <volume>30</volume>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B19">
        <label>19.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Lundberg, S.M. and Lee, S.I. (2017) A Unified Approach to Interpreting Model Predictions. <italic>Advances in Neural Information Processing Systems</italic>, 30, 4765-4774.</mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Lundberg, S.M.</string-name>
              <string-name>Lee, S.I.</string-name>
            </person-group>
            <year>2017</year>
            <article-title>A Unified Approach to Interpreting Model Predictions</article-title>
            <source>Advances in Neural Information Processing Systems</source>
            <volume>30</volume>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B20">
        <label>20.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Breiman, L. (2001) Random Forests. <italic>Machine Learning</italic>, 45, 5-32. https://doi.org/10.1023/a:1010933404324 <pub-id pub-id-type="doi">10.1023/a:1010933404324</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1023/a:1010933404324">https://doi.org/10.1023/a:1010933404324</ext-link></mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Breiman, L.</string-name>
            </person-group>
            <year>2001</year>
            <article-title>Random Forests</article-title>
            <source>Machine Learning</source>
            <volume>45</volume>
            <fpage>101093</fpage>
            <pub-id pub-id-type="doi">10.1023/a:1010933404324</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B21">
        <label>21.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Chawla, N.V., Bowyer, K.W., Hall, L.O. and Kegelmeyer, W.P. (2002) SMOTE: Synthetic Minority Over-Sampling Technique. <italic>Journal of Artificial Intelligence Res</italic><italic>earch</italic>, 16, 321-357. https://doi.org/10.1613/jair.953 <pub-id pub-id-type="doi">10.1613/jair.953</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1613/jair.953">https://doi.org/10.1613/jair.953</ext-link></mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Chawla, N.V.</string-name>
              <string-name>Bowyer, K.W.</string-name>
              <string-name>Hall, L.O.</string-name>
              <string-name>Kegelmeyer, W.P.</string-name>
            </person-group>
            <year>2002</year>
            <article-title>SMOTE: Synthetic Minority Over-Sampling Technique</article-title>
            <source>Journal of Artificial Intelligence Research</source>
            <volume>16</volume>
            <pub-id pub-id-type="doi">10.1613/jair.953</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B22">
        <label>22.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Wolpert, D.H. (1992) Stacked Generalization. <italic>Neural Networks</italic>, 5, 241-259. https://doi.org/10.1016/s0893-6080(05)80023-1 <pub-id pub-id-type="doi">10.1016/s0893-6080(05)80023-1</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1016/s0893-6080(05)80023-1">https://doi.org/10.1016/s0893-6080(05)80023-1</ext-link></mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Wolpert, D.H.</string-name>
            </person-group>
            <year>1992</year>
            <article-title>Stacked Generalization</article-title>
            <source>Neural Networks</source>
            <volume>6080</volume>
            <issue>05</issue>
            <pub-id pub-id-type="doi">10.1016/s0893-6080(05)80023-1</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B23">
        <label>23.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Insel, T.R. (2017) Digital Phenotyping: Technology for a New Science of Behaviour. <italic>J</italic><italic>ournal of the</italic><italic>A</italic><italic>merican</italic><italic>M</italic><italic>edical</italic><italic>A</italic><italic>ssociation</italic>, 318, 1215-1216. https://doi.org/10.1001/jama.2017.11295 <pub-id pub-id-type="doi">10.1001/jama.2017.11295</pub-id><pub-id pub-id-type="pmid">28973224</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1001/jama.2017.11295">https://doi.org/10.1001/jama.2017.11295</ext-link></mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Insel, T.R.</string-name>
            </person-group>
            <year>2017</year>
            <article-title>Digital Phenotyping: Technology for a New Science of Behaviour</article-title>
            <source>Journal of the American Medical Association</source>
            <volume>318</volume>
            <pub-id pub-id-type="doi">10.1001/jama.2017.11295</pub-id>
            <pub-id pub-id-type="pmid">28973224</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B24">
        <label>24.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Steyerberg, E.W., Vickers, A.J., Cook, N.R., Gerds, T., Gonen, M., Obuchowski, N., <italic>et al</italic>. (2010) Assessing the Performance of Prediction Models. <italic>Epidemiology</italic>, 21, 128-138. https://doi.org/10.1097/ede.0b013e3181c30fb2 <pub-id pub-id-type="doi">10.1097/ede.0b013e3181c30fb2</pub-id><pub-id pub-id-type="pmid">20010215</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1097/ede.0b013e3181c30fb2">https://doi.org/10.1097/ede.0b013e3181c30fb2</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Steyerberg, E.W.</string-name>
              <string-name>Vickers, A.J.</string-name>
              <string-name>Cook, N.R.</string-name>
              <string-name>Gerds, T.</string-name>
              <string-name>Gonen, M.</string-name>
              <string-name>Obuchowski, N.</string-name>
            </person-group>
            <year>2010</year>
            <article-title>Assessing the Performance of Prediction Models</article-title>
            <source>Epidemiology</source>
            <volume>21</volume>
            <pub-id pub-id-type="doi">10.1097/ede.0b013e3181c30fb2</pub-id>
            <pub-id pub-id-type="pmid">20010215</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B25">
        <label>25.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Ehlers, C.L., Frank, E. and Kupfer, D.J. (1988) Social Zeitgebers and Biological Rhythms. <italic>Archives of General Psychiatry</italic>, 45, 948-952. https://doi.org/10.1001/archpsyc.1988.01800340076012 <pub-id pub-id-type="doi">10.1001/archpsyc.1988.01800340076012</pub-id><pub-id pub-id-type="pmid">3048226</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1001/archpsyc.1988.01800340076012">https://doi.org/10.1001/archpsyc.1988.01800340076012</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Ehlers, C.L.</string-name>
              <string-name>Frank, E.</string-name>
              <string-name>Kupfer, D.J.</string-name>
            </person-group>
            <year>1988</year>
            <article-title>Social Zeitgebers and Biological Rhythms</article-title>
            <source>Archives of General Psychiatry</source>
            <volume>45</volume>
            <pub-id pub-id-type="doi">10.1001/archpsyc.1988.01800340076012</pub-id>
            <pub-id pub-id-type="pmid">3048226</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B26">
        <label>26.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Saccaro, L.F., Amatori, G., Cappelli, A., Mazziotti, R., Dell’Osso, L. and Rutigliano, G. (2021) Portable Technologies for Digital Phenotyping of Bipolar Disorder: A Systematic Review. <italic>Journal of Affective Disorders</italic>, 295, 323-338. https://doi.org/10.1016/j.jad.2021.08.052 <pub-id pub-id-type="doi">10.1016/j.jad.2021.08.052</pub-id><pub-id pub-id-type="pmid">34488086</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1016/j.jad.2021.08.052">https://doi.org/10.1016/j.jad.2021.08.052</ext-link></mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Saccaro, L.F.</string-name>
              <string-name>Amatori, G.</string-name>
              <string-name>Cappelli, A.</string-name>
              <string-name>Mazziotti, R.</string-name>
              <string-name>Osso, L.</string-name>
              <string-name>Rutigliano, G.</string-name>
            </person-group>
            <year>2021</year>
            <article-title>Portable Technologies for Digital Phenotyping of Bipolar Disorder: A Systematic Review</article-title>
            <source>Journal of Affective Disorders</source>
            <volume>295</volume>
            <pub-id pub-id-type="doi">10.1016/j.jad.2021.08.052</pub-id>
            <pub-id pub-id-type="pmid">34488086</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B27">
        <label>27.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Keramatian, K., Chithra, N.K. and Yatham, L.N. (2023) The CANMAT and ISBD Guidelines for the Treatment of Bipolar Disorder: Summary and a 2023 Update of Evidence. <italic>Focus</italic>, 21, 344-353. https://doi.org/10.1176/appi.focus.20230009 <pub-id pub-id-type="doi">10.1176/appi.focus.20230009</pub-id><pub-id pub-id-type="pmid">38695002</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1176/appi.focus.20230009">https://doi.org/10.1176/appi.focus.20230009</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Keramatian, K.</string-name>
              <string-name>Chithra, N.K.</string-name>
              <string-name>Yatham, L.N.</string-name>
            </person-group>
            <year>2023</year>
            <article-title>The CANMAT and ISBD Guidelines for the Treatment of Bipolar Disorder: Summary and a 2023 Update of Evidence</article-title>
            <source>Focus</source>
            <volume>21</volume>
            <pub-id pub-id-type="doi">10.1176/appi.focus.20230009</pub-id>
            <pub-id pub-id-type="pmid">38695002</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B28">
        <label>28.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Collins, G.S., Moons, K.G.M., Dhiman, P., <italic>et al</italic>. (2024) TRIPOD + AI Statement: Updated Guidance for Reporting Clinical Prediction Models That Use Regression or Machine Learning Methods. <italic>B</italic><italic>ritish</italic><italic>M</italic><italic>edical</italic><italic>J</italic><italic>ournal</italic>, 385, e078378.</mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Collins, G.S.</string-name>
              <string-name>Moons, K.G.M.</string-name>
              <string-name>Dhiman, P.</string-name>
            </person-group>
            <year>2024</year>
            <article-title>TRIPOD + AI Statement: Updated Guidance for Reporting Clinical Prediction Models That Use Regression or Machine Learning Methods</article-title>
            <source>British Medical Journal</source>
            <volume>385</volume>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B29">
        <label>29.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Lipschitz, J.M., Lin, S., Saghafian, S., Pike, C.K. and Burdick, K.E. (2025) Digital Phenotyping in Bipolar Disorder: Using Longitudinal Fitbit Data and Personalized Machine Learning to Predict Mood Symptomatology. <italic>Acta Psychiatrica Scandinavica</italic>, 151, 434-447. https://doi.org/10.1111/acps.13765 <pub-id pub-id-type="doi">10.1111/acps.13765</pub-id><pub-id pub-id-type="pmid">39397313</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1111/acps.13765">https://doi.org/10.1111/acps.13765</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Lipschitz, J.M.</string-name>
              <string-name>Lin, S.</string-name>
              <string-name>Saghafian, S.</string-name>
              <string-name>Pike, C.K.</string-name>
              <string-name>Burdick, K.E.</string-name>
            </person-group>
            <year>2025</year>
            <article-title>Digital Phenotyping in Bipolar Disorder: Using Longitudinal Fitbit Data and Personalized Machine Learning to Predict Mood Symptomatology</article-title>
            <source>Acta Psychiatrica Scandinavica</source>
            <volume>151</volume>
            <pub-id pub-id-type="doi">10.1111/acps.13765</pub-id>
            <pub-id pub-id-type="pmid">39397313</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
    </ref-list>
  </back>
</article>