<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.4 20241031//EN" "JATS-journalpublishing1-4.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article" dtd-version="1.4" xml:lang="en">
  <front>
    <journal-meta>
      <journal-id journal-id-type="publisher-id">Oalib</journal-id>
      <journal-title-group>
        <journal-title>Open Access Library Journal</journal-title>
      </journal-title-group>
      <issn pub-type="epub">2333-9721</issn>
      <issn pub-type="ppub">2333-9705</issn>
      <publisher>
        <publisher-name>Scientific Research Publishing</publisher-name>
      </publisher>
    </journal-meta>
    <article-meta>
      <article-id pub-id-type="doi">10.4236/oalib.1115137</article-id>
      <article-id pub-id-type="publisher-id">Oalib-151575</article-id>
      <article-categories>
        <subj-group>
          <subject>Article</subject>
        </subj-group>
        <subj-group>
          <subject>Biomedical</subject>
          <subject>Life Sciences</subject>
          <subject>Business</subject>
          <subject>Economics</subject>
          <subject>Chemistry</subject>
          <subject>Materials Science</subject>
          <subject>Computer Science</subject>
          <subject>Communications</subject>
          <subject>Earth</subject>
          <subject>Environmental Sciences</subject>
          <subject>Engineering</subject>
          <subject>Medicine</subject>
          <subject>Healthcare</subject>
          <subject>Physics</subject>
          <subject>Mathematics</subject>
          <subject>Social Sciences</subject>
          <subject>Humanities</subject>
        </subj-group>
      </article-categories>
      <title-group>
        <article-title>Deep Reinforcement Learning for Personalized Antidepressant Decision Support in Bipolar Spectrum Disorders: Simulated Randomized Trial Framework</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author" corresp="yes">
          <contrib-id contrib-id-type="orcid">0000-0001-9101-072X</contrib-id>
          <name name-style="western">
            <surname>Filippis</surname>
            <given-names>Rocco de</given-names>
          </name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <contrib-id contrib-id-type="orcid">0000-0002-5102-4999</contrib-id>
          <name name-style="western">
            <surname>Foysal</surname>
            <given-names>Abdullah Al</given-names>
          </name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
      </contrib-group>
      <aff id="aff1"><label>1</label> Department of Neuroscience, Institute of Psychopathology, Rome, Italy </aff>
      <aff id="aff2"><label>2</label> Department of Computer Engineering (AI), University of Genova, Genova, Italy </aff>
      <author-notes>
        <fn fn-type="conflict" id="fn-conflict">
          <p>The authors declare no conflicts of interest.</p>
        </fn>
      </author-notes>
      <pub-date pub-type="epub">
        <day>06</day>
        <month>05</month>
        <year>2026</year>
      </pub-date>
      <pub-date pub-type="collection">
        <month>05</month>
        <year>2026</year>
      </pub-date>
      <volume>13</volume>
      <issue>05</issue>
      <fpage>1</fpage>
      <lpage>19</lpage>
      <history>
        <date date-type="received">
          <day>10</day>
          <month>03</month>
          <year>2026</year>
        </date>
        <date date-type="accepted">
          <day>25</day>
          <month>05</month>
          <year>2026</year>
        </date>
        <date date-type="published">
          <day>28</day>
          <month>05</month>
          <year>2026</year>
        </date>
      </history>
      <permissions>
        <copyright-statement>© 2026 by the authors and Scientific Research Publishing Inc.</copyright-statement>
        <copyright-year>2026</copyright-year>
        <license license-type="open-access">
          <license-p> This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license ( <ext-link ext-link-type="uri" xlink:href="https://creativecommons.org/licenses/by/4.0/">https://creativecommons.org/licenses/by/4.0/</ext-link> ). </license-p>
        </license>
      </permissions>
      <self-uri content-type="doi" xlink:href="https://doi.org/10.4236/oalib.1115137">https://doi.org/10.4236/oalib.1115137</self-uri>
      <abstract>
        <p>Treatment selection for bipolar depression remains largely trial-and-error, with substantial non-response to first-line strategies and clinically meaningful risk of mood destabilization. We developed a deep reinforcement learning (RL) framework to optimize treatment selection while explicitly penalizing destabilization events. We implemented RL-CADENCE, a simulated multi-centre experimental framework designed to emulate a parallel-group randomized trial across 12 virtual psychiatric centres. Using a combination of publicly available online data sources and clinically informed synthetic generation, we constructed a cohort of 2500 virtual participants representing bipolar spectrum disorders. Virtual participants were algorithmically allocated (3:3:3:1) to four treatment strategies: 1) lithium + SSRI, 2) quetiapine + lamotrigine, 3) lurasidone + mood stabilizer (lithium or valproate), or 4) RL-personalized treatment selection. The primary endpoint was the simulated change in Montgomery Åsberg Depression Rating Scale (MADRS) score over 12 months. Secondary outcomes included response, mood destabilization events, and quality-adjusted life years (QALYs). A causal machine learning pipeline estimated conditional average treatment effects (CATE) to characterize heterogeneity across subgroups within the synthetic cohort. In simulation, the RL-personalized strategy achieved greater MADRS improvement than pooled standard protocols (mean difference: −5.6 points; 95% CI: −7.6 to −3.6; Cohen’s <italic>d</italic> = 0.78). Simulated response rates (≥50% MADRS reduction) were 95.7% versus 58.9%, and mood destabilization occurred in 4.8% versus 10.8% of synthetic patient-months. The RL policy network achieved an AUC-ROC of 0.89 for predicting the optimal treatment strategy under the simulated counterfactual evaluation. Heterogeneous effects were largest in mixed features (CATE = 15.2; 95% CI: 8.9 - 21.5) and bipolar I subtype (CATE = 12.3; 95% CI: 7.1 - 17.5). Within a simulated, synthetic-data evaluation, deep RL showed strong potential to personalize antidepressant-related treatment selection in bipolar spectrum disorders, improving depressive symptom outcomes while reducing destabilization risk. These findings provide proof-of-concept for RL-based precision psychiatry and motivate prospective validation in real-world clinical cohorts.</p>
      </abstract>
      <kwd-group kwd-group-type="author-generated" xml:lang="en">
        <kwd>Bipolar Disorder</kwd>
        <kwd>Reinforcement Learning</kwd>
        <kwd>Precision Psychiatry</kwd>
        <kwd>Treatment Optimization</kwd>
        <kwd>Causal Inference</kwd>
        <kwd>Machine Learning</kwd>
        <kwd>Antidepressant</kwd>
        <kwd>Mood Destabilization</kwd>
        <kwd>Deep Learning</kwd>
        <kwd>Personalized Medicine</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec1">
      <title>1. Introduction</title>
      <p>Bipolar spectrum disorders affect approximately 2.4% of the global population and represent a leading cause of disability among young adults [<xref ref-type="bibr" rid="B1">1</xref>][<xref ref-type="bibr" rid="B2">2</xref>]. Despite the availability of numerous pharmacological interventions, treatment selection remains predominantly guided by clinical intuition and trial-and-error approaches. This conventional paradigm yields suboptimal outcomes: nearly 60% of patients with bipolar depression with bipolar depression fail to achieve remission with first-line treatments, and approximately 20% experience antidepressant-associated mood destabilization, including switches to mania or rapid cycling [<xref ref-type="bibr" rid="B3">3</xref>][<xref ref-type="bibr" rid="B4">4</xref>]. The heterogeneity of bipolar spectrum disorders presents a fundamental challenge to traditional treatment approaches. Synthetic patients vary substantially in clinical presentation (bipolar I vs. II vs. mixed features), comorbidity profiles, pharmacogenomic markers, and treatment history [<xref ref-type="bibr" rid="B5">5</xref>]. Current guidelines provide limited personalization beyond broad categorical distinctions, failing to capitalize on the multidimensional data increasingly available in contemporary psychiatric practice [<xref ref-type="bibr" rid="B6">6</xref>].</p>
      <p>Artificial intelligence (AI) and machine learning (ML) have emerged as promising tools for precision medicine, with applications ranging from diagnostic imaging to drug discovery [<xref ref-type="bibr" rid="B7">7</xref>][<xref ref-type="bibr" rid="B8">8</xref>]. In psychiatry, ML approaches have demonstrated potential for predicting treatment response in major depressive disorder [<xref ref-type="bibr" rid="B9">9</xref>][<xref ref-type="bibr" rid="B10">10</xref>]. However, several critical limitations have hindered clinical translation: 1) most models rely on static prediction rather than sequential decision-making; 2) they inadequately account for delayed rewards and long-term outcomes; 3) they rarely incorporate explicit safety constraints to prevent simulated adverse events; and 4) they often lack causal validity for treatment recommendation tasks [<xref ref-type="bibr" rid="B11">11</xref>][<xref ref-type="bibr" rid="B12">12</xref>]. Reinforcement learning (RL) provides a mathematical framework well suited to addressing these limitations. Unlike supervised learning, RL optimizes sequential decision-making through interaction with an environment, maximizing cumulative rewards while accounting for delayed consequences [<xref ref-type="bibr" rid="B13">13</xref>]. In healthcare-oriented settings, RL can model the dynamic nature of treatment response, learn policies from observational or simulated trajectories, and explicitly incorporate safety constraints to reduce destabilization risk [<xref ref-type="bibr" rid="B14">14</xref>][<xref ref-type="bibr" rid="B15">15</xref>]. Recent advances in deep RL, combining neural network function approximation with policy-gradient methods, have enabled successful applications in complex domains such as robotics, game playing, and resource allocation [<xref ref-type="bibr" rid="B16">16</xref>][<xref ref-type="bibr" rid="B17">17</xref>]. In medicine, deep RL has shown promise for sepsis management, mechanical ventilation, and treatment sequencing in oncology [<xref ref-type="bibr" rid="B18">18</xref>]-[<xref ref-type="bibr" rid="B20">20</xref>]. However, applications to psychiatric treatment optimization remain limited, with existing studies largely focusing on static prediction rather than dynamic decision-making [<xref ref-type="bibr" rid="B21">21</xref>][<xref ref-type="bibr" rid="B22">22</xref>].</p>
      <p>We hypothesized that a deep reinforcement learning framework integrating clinical, demographic, and pharmacogenomic information could: 1) learn optimal treatment policies from sequential simulated synthetic patient trajectories; 2) personalize treatment recommendations based on individual synthetic patient characteristics; 3) explicitly minimize mood destabilization risk through constrained optimization; and 4) provide interpretable decision-support insights relevant to clinical decision-making [<xref ref-type="bibr" rid="B23">23</xref>]-[<xref ref-type="bibr" rid="B26">26</xref>]. To evaluate this hypothesis, we implemented the Reinforcement Learning for Clinical Antidepressant Decision-making in Bipolar Spectrum Disorders (RL-CADENCE) framework within a simulated trial-emulation environment constructed using publicly available online datasets and clinically informed synthetic patient trajectories [<xref ref-type="bibr" rid="B27">27</xref>]-[<xref ref-type="bibr" rid="B30">30</xref>].</p>
    </sec>
    <sec id="sec2">
      <title>2. Methods</title>
      <sec id="sec2dot1">
        <title>2.1. Study Design and Virtual Participants</title>
        <p>A simulated multi-centre trial framework was implemented using a combination of publicly available online datasets and clinically informed synthetic data generation. Three publicly available sources were used: 1) the STEP-BD (Systematic Treatment Enhancement Program for Bipolar Disorder) publicly available summary statistics provided distributions for baseline MADRS, YMRS, prior episode counts, and bipolar subtype prevalence used to calibrate the synthetic cohort generator; 2) the CANMAT 2018 Bipolar Guidelines supplementary tables provided treatment response rates by arm and subtype used to parameterize outcome functions; 3) the UK Biobank publicly released aggregate phenotypic statistics for age, sex, illness duration, and comorbidity rates. No individual-level patient data from any of these sources were used; only published summary statistics (means, standard deviations, proportions) were imported to set generative parameters. All individual synthetic patient records were computationally generated as described in Section 2.2. The pre-training dataset (n = 8432 trajectories) was generated using the same synthetic pipeline with parameters calibrated from the STEP-BD and CANMAT sources; no external real patient trajectories were used for pre-training. Virtual participants represented adults aged 18 - 75 years with a primary diagnosis of bipolar spectrum disorder (bipolar I, bipolar II, or cyclothymic disorder), operationalized according to DSM-5 diagnostic criteria and modelled to reflect distributions observed in clinical cohorts. Inclusion criteria simulated virtual participants experiencing a current major depressive episode, defined by a Montgomery Åsberg Depression Rating Scale (MADRS) score ≥ 20. Simulated exclusion criteria mirrored standard psychiatric trial protocols and included: 1) current manic or mixed episodes; 2) active psychotic symptoms requiring recent medication changes; 3) recent substance use disorder (excluding nicotine or caffeine); 4) pregnancy or lactation; 5) contraindications to study medications; and 6) inability to provide informed consent, represented through synthetic eligibility constraints within the data generation pipeline.</p>
      </sec>
      <sec id="sec2dot2">
        <title>2.2. Randomization and Masking</title>
        <p>Virtual participants were algorithmically assigned in a 3:3:3:1 ratio to four treatment strategies: 1) Lithium plus selective serotonin reuptake inhibitor (SSRI); 2) Quetiapine plus Lamotrigine; 3) Lurasidone plus mood stabilizer (lithium or valproate); or 4) RL-personalized treatment selection. Treatment allocation was implemented within the simulation framework using stratified randomization procedures based on bipolar subtype (I vs. II vs. other), baseline symptom severity (MADRS &lt; 30 vs. ≥30), and presence of mixed features, ensuring balanced subgroup representation across arms.</p>
        <p>As the study was conducted within a computational simulation environment using online and synthetic data sources, traditional virtual participant and clinician masking was not applicable. However, to preserve methodological consistency with clinical trial standards, outcome evaluation pipelines were designed to remain independent of treatment assignment during metric computation, and statistical analyses were performed using pre-specified blinded scripts prior to final model evaluation.</p>
        <p>Synthetic Data Generation Details</p>
        <p>Joint feature distributions were generated as follows. Continuous features (age, MADRS, YMRS, prior failed trials) were drawn from multivariate normal distributions with covariance matrices derived from published correlation tables in STEP-BD: MADRS and YMRS were correlated at r = 0.31; prior failed trials and illness duration at r = 0.48. Binary features (bipolar I, mixed features, CYP2D6 poor metabolizer) were drawn from Bernoulli distributions with site-specific prevalences. Treatment response trajectories were generated using a linear mixed-effects outcome model: <italic>MADRS</italic><italic><sub>t</sub></italic> = <italic>MADRS</italic><sub>0</sub> + <italic>β</italic><sub>arm</sub>·<italic>t</italic> + <italic>β</italic><sub>interaction</sub>·(arm × subtype)·<italic>t</italic> + <italic>u</italic><italic><sub>i</sub></italic> + <italic>ε</italic><italic><sub>it</sub></italic>, where <italic>β</italic><sub>arm</sub> coefficients were set to −0.93 (Li + SSRI), −1.05 (QTP + LTG), −1.10 (LUR + MS), and −1.40 (RL-arm) points/month, derived from published meta-analytic effect sizes [Sidor &amp; MacQueen 2011]; <italic>u</italic><italic><sub>i</sub></italic> ~ N(0, 3.2<sup>2</sup>) is a random patient intercept; <italic>ε</italic><italic><sub>it</sub></italic> ~ N(0, 1.5<sup>2</sup>). Mood destabilization events were generated as Bernoulli draws with monthly probabilities: 1.2% (Li + SSRI), 0.9% (QTP + LTG), 0.8% (LUR + MS), and 0.4% (RL-arm) the RL-arm rate was set based on the hard constraint excluding antidepressant monotherapy in high-YMRS patients, not independently estimated. Medication adherence was modelled as beta-distributed (<italic>α</italic> = 5, <italic>β</italic> = 2) per arm, declining by 3% per failed prior trial. Missing data (10% MAR missingness at months 3 and 6) were introduced by randomly setting MADRS and YMRS to missing with probability proportional to side-effect burden. All parameters not directly from published evidence were flagged as expert assumptions; these include the RL-arm treatment response slope, the RL destabilization rate, and the adherence decay coefficient.</p>
      </sec>
      <sec id="sec2dot3">
        <title>2.3. Interventions</title>
        <p><bold>Standard Treatment Arms:</bold> Virtual participants in arms 1 - 3 received guideline-concordant pharmacotherapy as specified by protocol. Medications were titrated according to standardized algorithms targeting therapeutic blood levels (for lithium: 0.6 - 1.0 mEq/L; for valproate: 50 - 100 μg/mL) or maximum tolerated doses. Concomitant medications were permitted for anxiety or sleep but restricted to non-study antidepressants or mood stabilizers.</p>
        <p><bold>RL-Personalized Arm:</bold> Virtual participants in the RL arm received treatment recommendations generated by the deep RL policy network (described in Section 2.4). Decision rules emulating clinician override behaviour were incorporated and could override recommendations based on clinical judgment. Overrides were documented for secondary analysis. The RL system provided monthly recommendations based on updated clinical data.</p>
        <p><bold>Dynamic vs. Fixed Policy Comparison</bold><bold>:</bold> An important structural asymmetry exists between the RL arm and the three standard arms, the RL policy selected treatment actions monthly based on updated state observations, while the standard arms assigned fixed treatment protocols at baseline without adaptive switching. This means the comparison is between a dynamic, state-adaptive policy and three fixed protocols, not between four equally adaptive strategies. This structural difference confers an inherent advantage to the RL arm independent of the quality of the learned policy, because any adaptive system can exploit trajectory information unavailable to fixed-protocol arms. All comparisons between the RL arm and standard arms should therefore be interpreted as evaluating the value of dynamic adaptation relative to fixed guideline protocols, not as a head-to-head comparison of equivalent decision architectures. In clinical practice, clinicians do adapt treatments over time; future comparisons should include a dynamic-clinician-judgment arm to isolate the incremental value of the RL policy above and beyond human adaptive decision-making.</p>
      </sec>
      <sec id="sec2dot4">
        <title>2.4. Deep Reinforcement Learning Framework</title>
        <p>2.4.1. State Space</p>
        <p>The reinforcement learning (RL) environment was defined through synthetic patient states <inline-formula><mml:math><mml:mrow><mml:msub><mml:mi> s </mml:mi><mml:mi> t </mml:mi></mml:msub><mml:mo> ∈ </mml:mo><mml:mi> S </mml:mi></mml:mrow></mml:math></inline-formula> , representing multidimensional clinical information at each decision step:</p>
        <disp-formula id="FD1">
          <mml:math>
            <mml:mrow>
              <mml:msub>
                <mml:mi>s</mml:mi>
                <mml:mi>t</mml:mi>
              </mml:msub>
              <mml:mo>=</mml:mo>
              <mml:mrow>
                <mml:mo>[</mml:mo>
                <mml:mrow>
                  <mml:msub>
                    <mml:mi>X</mml:mi>
                    <mml:mrow>
                      <mml:mtext>demo</mml:mtext>
                    </mml:mrow>
                  </mml:msub>
                  <mml:mo>,</mml:mo>
                  <mml:msub>
                    <mml:mi>X</mml:mi>
                    <mml:mrow>
                      <mml:mtext>clinical</mml:mtext>
                    </mml:mrow>
                  </mml:msub>
                  <mml:mo>,</mml:mo>
                  <mml:msub>
                    <mml:mi>X</mml:mi>
                    <mml:mrow>
                      <mml:mtext>genomic</mml:mtext>
                    </mml:mrow>
                  </mml:msub>
                  <mml:mo>,</mml:mo>
                  <mml:msub>
                    <mml:mi>X</mml:mi>
                    <mml:mrow>
                      <mml:mtext>history</mml:mtext>
                    </mml:mrow>
                  </mml:msub>
                </mml:mrow>
                <mml:mo>]</mml:mo>
              </mml:mrow>
            </mml:mrow>
          </mml:math>
        </disp-formula>
        <p>where <inline-formula><mml:math><mml:mrow><mml:msub><mml:mi> X </mml:mi><mml:mrow><mml:mtext> demo </mml:mtext></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula> includes demographic attributes such as age, sex, and socioeconomic indicators; <inline-formula><mml:math><mml:mrow><mml:msub><mml:mi> X </mml:mi><mml:mrow><mml:mtext> clinical </mml:mtext></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula> comprises symptom severity measures (YMRS, MADRS), bipolar subtype, comorbidities, and side-effect burden; <inline-formula><mml:math><mml:mrow><mml:msub><mml:mi> X </mml:mi><mml:mrow><mml:mtext> genomic </mml:mtext></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula> includes pharmacogenomic markers such as CYP2D6 metabolizer status, COMT Val158Met, and BDNF Val66Met polymorphisms; and <inline-formula><mml:math><mml:mrow><mml:msub><mml:mi> X </mml:mi><mml:mrow><mml:mtext> history </mml:mtext></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula> captures previous treatment trials and observed response trajectories.</p>
        <p>2.4.2. Action Space</p>
        <p>The action space was defined as:</p>
        <disp-formula id="FD2">
          <mml:math>
            <mml:mrow>
              <mml:mi>A</mml:mi>
              <mml:mo>=</mml:mo>
              <mml:mrow>
                <mml:mo>{</mml:mo>
                <mml:mrow>
                  <mml:msub>
                    <mml:mi>A</mml:mi>
                    <mml:mn>1</mml:mn>
                  </mml:msub>
                  <mml:mo>,</mml:mo>
                  <mml:msub>
                    <mml:mi>A</mml:mi>
                    <mml:mn>2</mml:mn>
                  </mml:msub>
                  <mml:mo>,</mml:mo>
                  <mml:msub>
                    <mml:mi>A</mml:mi>
                    <mml:mn>3</mml:mn>
                  </mml:msub>
                  <mml:mo>,</mml:mo>
                  <mml:msub>
                    <mml:mi>A</mml:mi>
                    <mml:mn>4</mml:mn>
                  </mml:msub>
                </mml:mrow>
                <mml:mo>}</mml:mo>
              </mml:mrow>
            </mml:mrow>
          </mml:math>
        </disp-formula>
        <p>corresponding to four predefined treatment strategies. Actions were selected at monthly intervals based on current state observations, enabling adaptive treatment selection over time.</p>
        <p>2.4.3. Reward Function</p>
        <p>The reward function was formulated to balance symptom improvement against safety and tolerability constraints:</p>
        <disp-formula id="FD3">
          <mml:math>
            <mml:mrow>
              <mml:msub>
                <mml:mi>r</mml:mi>
                <mml:mi>t</mml:mi>
              </mml:msub>
              <mml:mo>=</mml:mo>
              <mml:mo>−</mml:mo>
              <mml:mrow>
                <mml:mo>[</mml:mo>
                <mml:mrow>
                  <mml:mi>α</mml:mi>
                  <mml:mo>⋅</mml:mo>
                  <mml:mi>Y</mml:mi>
                  <mml:mi>M</mml:mi>
                  <mml:mi>R</mml:mi>
                  <mml:msub>
                    <mml:mi>S</mml:mi>
                    <mml:mi>t</mml:mi>
                  </mml:msub>
                  <mml:mo>+</mml:mo>
                  <mml:mi>β</mml:mi>
                  <mml:mo>⋅</mml:mo>
                  <mml:mi>M</mml:mi>
                  <mml:mi>A</mml:mi>
                  <mml:mi>D</mml:mi>
                  <mml:mi>R</mml:mi>
                  <mml:msub>
                    <mml:mi>S</mml:mi>
                    <mml:mi>t</mml:mi>
                  </mml:msub>
                </mml:mrow>
                <mml:mo>]</mml:mo>
              </mml:mrow>
              <mml:mo>−</mml:mo>
              <mml:mi>λ</mml:mi>
              <mml:mo>⋅</mml:mo>
              <mml:mn>1</mml:mn>
              <mml:mrow>
                <mml:mo>(</mml:mo>
                <mml:mrow>
                  <mml:msub>
                    <mml:mrow>
                      <mml:mtext>destabilization</mml:mtext>
                    </mml:mrow>
                    <mml:mi>t</mml:mi>
                  </mml:msub>
                </mml:mrow>
                <mml:mo>)</mml:mo>
              </mml:mrow>
              <mml:mo>−</mml:mo>
              <mml:mi>γ</mml:mi>
              <mml:mo>⋅</mml:mo>
              <mml:msub>
                <mml:mrow>
                  <mml:mtext>side_effects</mml:mtext>
                </mml:mrow>
                <mml:mi>t</mml:mi>
              </mml:msub>
            </mml:mrow>
          </mml:math>
        </disp-formula>
        <p>where <inline-formula><mml:math><mml:mrow><mml:mi> α </mml:mi><mml:mo> = </mml:mo><mml:mn> 0.4 </mml:mn></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math><mml:mrow><mml:mi> β </mml:mi><mml:mo> = </mml:mo><mml:mn> 0.6 </mml:mn></mml:mrow></mml:math></inline-formula> weight manic and depressive symptom severity, respectively; <inline-formula><mml:math><mml:mrow><mml:mi> λ </mml:mi><mml:mo> = </mml:mo><mml:mn> 15 </mml:mn></mml:mrow></mml:math></inline-formula> imposes a strong penalty for mood destabilization events; and <inline-formula><mml:math><mml:mrow><mml:mi> γ </mml:mi><mml:mo> = </mml:mo><mml:mn> 0.1 </mml:mn></mml:mrow></mml:math></inline-formula> penalizes treatment-related adverse effects. A discount factor of <inline-formula><mml:math><mml:mrow><mml:mi> γ </mml:mi><mml:mo> = </mml:mo><mml:mn> 0.95 </mml:mn></mml:mrow></mml:math></inline-formula> was employed to emphasize long-term outcomes.</p>
        <p>The cumulative discounted return was defined as:</p>
        <disp-formula id="FD4">
          <mml:math>
            <mml:mrow>
              <mml:msub>
                <mml:mi>G</mml:mi>
                <mml:mi>t</mml:mi>
              </mml:msub>
              <mml:mo>=</mml:mo>
              <mml:munderover>
                <mml:mstyle mathsize="140%" displaystyle="true">
                  <mml:mo>∑</mml:mo>
                </mml:mstyle>
                <mml:mrow>
                  <mml:mi>k</mml:mi>
                  <mml:mo>=</mml:mo>
                  <mml:mn>0</mml:mn>
                </mml:mrow>
                <mml:mi>∞</mml:mi>
              </mml:munderover>
              <mml:mtext>
                 
              </mml:mtext>
              <mml:msup>
                <mml:mi>γ</mml:mi>
                <mml:mi>k</mml:mi>
              </mml:msup>
              <mml:msub>
                <mml:mi>r</mml:mi>
                <mml:mrow>
                  <mml:mi>t</mml:mi>
                  <mml:mo>+</mml:mo>
                  <mml:mi>k</mml:mi>
                  <mml:mo>+</mml:mo>
                  <mml:mn>1</mml:mn>
                </mml:mrow>
              </mml:msub>
            </mml:mrow>
          </mml:math>
        </disp-formula>
        <p>2.4.4. Policy Network Architecture</p>
        <p>An actor–critic framework based on proximal policy optimization (PPO) was employed. The policy network <inline-formula><mml:math><mml:mrow><mml:msub><mml:mi> π </mml:mi><mml:mi> θ </mml:mi></mml:msub><mml:mrow><mml:mo> ( </mml:mo><mml:mrow><mml:mrow><mml:mrow><mml:mi> a </mml:mi><mml:mtext>   </mml:mtext></mml:mrow><mml:mo> | </mml:mo></mml:mrow><mml:mi> s </mml:mi></mml:mrow><mml:mo> ) </mml:mo></mml:mrow></mml:mrow></mml:math></inline-formula> , parameterized by <inline-formula><mml:math><mml:mi> θ </mml:mi></mml:math></inline-formula> , produced treatment probabilities, while the value network <inline-formula><mml:math><mml:mrow><mml:msub><mml:mi> V </mml:mi><mml:mi> ψ </mml:mi></mml:msub><mml:mrow><mml:mo> ( </mml:mo><mml:mi> s </mml:mi><mml:mo> ) </mml:mo></mml:mrow></mml:mrow></mml:math></inline-formula> estimated expected future rewards. Both networks were implemented as three-layer multilayer perceptrons containing 256 hidden units per layer, ReLU activations, and dropout regularization (rate = 0.3). A shared representation layer learned state embeddings <inline-formula><mml:math><mml:mrow><mml:mi> ϕ </mml:mi><mml:mrow><mml:mo> ( </mml:mo><mml:mi> s </mml:mi><mml:mo> ) </mml:mo></mml:mrow></mml:mrow></mml:math></inline-formula> .</p>
        <p>The PPO objective used a clipped surrogate loss:</p>
        <disp-formula id="FD5">
          <mml:math display="inline">
            <mml:mrow>
              <mml:msub>
                <mml:mi>L</mml:mi>
                <mml:mrow>
                  <mml:mtext>CLIP</mml:mtext>
                </mml:mrow>
              </mml:msub>
              <mml:mrow>
                <mml:mo>(</mml:mo>
                <mml:mi>θ</mml:mi>
                <mml:mo>)</mml:mo>
              </mml:mrow>
              <mml:mo>=</mml:mo>
              <mml:msub>
                <mml:mover accent="true">
                  <mml:mi mathvariant="double-struck">E</mml:mi>
                  <mml:mo>^</mml:mo>
                </mml:mover>
                <mml:mi>t</mml:mi>
              </mml:msub>
              <mml:mrow>
                <mml:mo>[</mml:mo>
                <mml:mrow>
                  <mml:mi>min</mml:mi>
                  <mml:mrow>
                    <mml:mo>(</mml:mo>
                    <mml:mrow>
                      <mml:mi>r</mml:mi>
                      <mml:mi>t</mml:mi>
                      <mml:mrow>
                        <mml:mo>(</mml:mo>
                        <mml:mi>θ</mml:mi>
                        <mml:mo>)</mml:mo>
                      </mml:mrow>
                      <mml:msup>
                        <mml:mi>A</mml:mi>
                        <mml:mi>t</mml:mi>
                      </mml:msup>
                      <mml:mo>,</mml:mo>
                      <mml:mi>c</mml:mi>
                      <mml:mi>l</mml:mi>
                      <mml:mi>i</mml:mi>
                      <mml:mi>p</mml:mi>
                      <mml:mrow>
                        <mml:mo>(</mml:mo>
                        <mml:mrow>
                          <mml:mi>r</mml:mi>
                          <mml:mi>t</mml:mi>
                          <mml:mrow>
                            <mml:mo>(</mml:mo>
                            <mml:mi>θ</mml:mi>
                            <mml:mo>)</mml:mo>
                          </mml:mrow>
                          <mml:mo>,</mml:mo>
                          <mml:mn>1</mml:mn>
                          <mml:mo>−</mml:mo>
                          <mml:mi>ϵ</mml:mi>
                          <mml:mo>,</mml:mo>
                          <mml:mn>1</mml:mn>
                          <mml:mo>+</mml:mo>
                          <mml:mi>ϵ</mml:mi>
                        </mml:mrow>
                        <mml:mo>)</mml:mo>
                      </mml:mrow>
                      <mml:msup>
                        <mml:mi>A</mml:mi>
                        <mml:mrow>
                          <mml:mi>t</mml:mi>
                          <mml:mtext>
                          </mml:mtext>
                        </mml:mrow>
                      </mml:msup>
                    </mml:mrow>
                    <mml:mo>)</mml:mo>
                  </mml:mrow>
                </mml:mrow>
                <mml:mo>]</mml:mo>
              </mml:mrow>
            </mml:mrow>
          </mml:math>
        </disp-formula>
        <p>where,</p>
        <disp-formula id="FD6">
          <mml:math>
            <mml:mrow>
              <mml:msub>
                <mml:mi>r</mml:mi>
                <mml:mi>t</mml:mi>
              </mml:msub>
              <mml:mrow>
                <mml:mo>(</mml:mo>
                <mml:mi>θ</mml:mi>
                <mml:mo>)</mml:mo>
              </mml:mrow>
              <mml:mo>=</mml:mo>
              <mml:mfrac>
                <mml:mrow>
                  <mml:msub>
                    <mml:mi>π</mml:mi>
                    <mml:mi>θ</mml:mi>
                  </mml:msub>
                  <mml:mrow>
                    <mml:mo>(</mml:mo>
                    <mml:mrow>
                      <mml:mrow>
                        <mml:mrow>
                          <mml:msub>
                            <mml:mi>a</mml:mi>
                            <mml:mi>i</mml:mi>
                          </mml:msub>
                        </mml:mrow>
                        <mml:mo>|</mml:mo>
                      </mml:mrow>
                      <mml:msub>
                        <mml:mi>s</mml:mi>
                        <mml:mi>i</mml:mi>
                      </mml:msub>
                    </mml:mrow>
                    <mml:mo>)</mml:mo>
                  </mml:mrow>
                </mml:mrow>
                <mml:mrow>
                  <mml:msub>
                    <mml:mi>π</mml:mi>
                    <mml:mrow>
                      <mml:msub>
                        <mml:mi>θ</mml:mi>
                        <mml:mrow>
                          <mml:mtext>old</mml:mtext>
                        </mml:mrow>
                      </mml:msub>
                    </mml:mrow>
                  </mml:msub>
                  <mml:mrow>
                    <mml:mo>(</mml:mo>
                    <mml:mrow>
                      <mml:mrow>
                        <mml:mrow>
                          <mml:msub>
                            <mml:mi>a</mml:mi>
                            <mml:mi>i</mml:mi>
                          </mml:msub>
                        </mml:mrow>
                        <mml:mo>|</mml:mo>
                      </mml:mrow>
                      <mml:msub>
                        <mml:mi>s</mml:mi>
                        <mml:mi>i</mml:mi>
                      </mml:msub>
                    </mml:mrow>
                    <mml:mo>)</mml:mo>
                  </mml:mrow>
                </mml:mrow>
              </mml:mfrac>
            </mml:mrow>
          </mml:math>
        </disp-formula>
        <p>is the probability ratio, <inline-formula><mml:math><mml:mrow><mml:msub><mml:mover accent="true"><mml:mi> A </mml:mi><mml:mo> ^ </mml:mo></mml:mover><mml:mi> t </mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> denotes the estimated advantage function, and <inline-formula><mml:math><mml:mrow><mml:mi> ϵ </mml:mi><mml:mo> = </mml:mo><mml:mn> 0.2 </mml:mn></mml:mrow></mml:math></inline-formula> defines the clipping threshold.</p>
        <p>The value function loss minimized the squared prediction error:</p>
        <disp-formula id="FD7">
          <mml:math>
            <mml:mrow>
              <mml:msub>
                <mml:mi>L</mml:mi>
                <mml:mrow>
                  <mml:mtext>VF</mml:mtext>
                </mml:mrow>
              </mml:msub>
              <mml:mrow>
                <mml:mo>(</mml:mo>
                <mml:mi>ψ</mml:mi>
                <mml:mo>)</mml:mo>
              </mml:mrow>
              <mml:mo>=</mml:mo>
              <mml:msub>
                <mml:mover accent="true">
                  <mml:mi mathvariant="double-struck">E</mml:mi>
                  <mml:mo>^</mml:mo>
                </mml:mover>
                <mml:mi>t</mml:mi>
              </mml:msub>
              <mml:mrow>
                <mml:mo>[</mml:mo>
                <mml:mrow>
                  <mml:msup>
                    <mml:mrow>
                      <mml:mrow>
                        <mml:mo>(</mml:mo>
                        <mml:mrow>
                          <mml:msub>
                            <mml:mi>V</mml:mi>
                            <mml:mi>ψ</mml:mi>
                          </mml:msub>
                          <mml:mrow>
                            <mml:mo>(</mml:mo>
                            <mml:mrow>
                              <mml:msub>
                                <mml:mi>s</mml:mi>
                                <mml:mi>t</mml:mi>
                              </mml:msub>
                            </mml:mrow>
                            <mml:mo>)</mml:mo>
                          </mml:mrow>
                          <mml:mo>−</mml:mo>
                          <mml:msubsup>
                            <mml:mi>V</mml:mi>
                            <mml:mi>t</mml:mi>
                            <mml:mrow>
                              <mml:mtext>target</mml:mtext>
                            </mml:mrow>
                          </mml:msubsup>
                        </mml:mrow>
                        <mml:mo>)</mml:mo>
                      </mml:mrow>
                    </mml:mrow>
                    <mml:mn>2</mml:mn>
                  </mml:msup>
                </mml:mrow>
                <mml:mo>]</mml:mo>
              </mml:mrow>
            </mml:mrow>
          </mml:math>
        </disp-formula>
        <p>The total optimization objective combined policy loss, value loss, and entropy regularization:</p>
        <disp-formula id="FD8">
          <mml:math>
            <mml:mrow>
              <mml:msub>
                <mml:mi>L</mml:mi>
                <mml:mrow>
                  <mml:mtext>TOTAL</mml:mtext>
                </mml:mrow>
              </mml:msub>
              <mml:mrow>
                <mml:mo>(</mml:mo>
                <mml:mrow>
                  <mml:mi>θ</mml:mi>
                  <mml:mo>,</mml:mo>
                  <mml:mi>ψ</mml:mi>
                </mml:mrow>
                <mml:mo>)</mml:mo>
              </mml:mrow>
              <mml:mo>=</mml:mo>
              <mml:msub>
                <mml:mi>L</mml:mi>
                <mml:mrow>
                  <mml:mtext>CLIP</mml:mtext>
                </mml:mrow>
              </mml:msub>
              <mml:mrow>
                <mml:mo>(</mml:mo>
                <mml:mi>θ</mml:mi>
                <mml:mo>)</mml:mo>
              </mml:mrow>
              <mml:mo>−</mml:mo>
              <mml:msub>
                <mml:mi>c</mml:mi>
                <mml:mn>1</mml:mn>
              </mml:msub>
              <mml:msub>
                <mml:mi>L</mml:mi>
                <mml:mrow>
                  <mml:mtext>VF</mml:mtext>
                </mml:mrow>
              </mml:msub>
              <mml:mrow>
                <mml:mo>(</mml:mo>
                <mml:mi>ψ</mml:mi>
                <mml:mo>)</mml:mo>
              </mml:mrow>
              <mml:mo>+</mml:mo>
              <mml:msub>
                <mml:mi>c</mml:mi>
                <mml:mn>2</mml:mn>
              </mml:msub>
              <mml:mi>H</mml:mi>
              <mml:mrow>
                <mml:mo>(</mml:mo>
                <mml:mrow>
                  <mml:msub>
                    <mml:mi>π</mml:mi>
                    <mml:mi>θ</mml:mi>
                  </mml:msub>
                </mml:mrow>
                <mml:mo>)</mml:mo>
              </mml:mrow>
            </mml:mrow>
          </mml:math>
        </disp-formula>
        <p>where <inline-formula><mml:math><mml:mrow><mml:mi> H </mml:mi><mml:mrow><mml:mo> ( </mml:mo><mml:mrow><mml:msub><mml:mi> π </mml:mi><mml:mi> θ </mml:mi></mml:msub></mml:mrow><mml:mo> ) </mml:mo></mml:mrow></mml:mrow></mml:math></inline-formula> represents the entropy bonus encouraging exploration, with coefficients <inline-formula><mml:math><mml:mrow><mml:msub><mml:mi> c </mml:mi><mml:mn> 1 </mml:mn></mml:msub><mml:mo> = </mml:mo><mml:mn> 0.5 </mml:mn></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math><mml:mrow><mml:msub><mml:mi> c </mml:mi><mml:mn> 2 </mml:mn></mml:msub><mml:mo> = </mml:mo><mml:mn> 0.01 </mml:mn></mml:mrow></mml:math></inline-formula> .</p>
        <p>2.4.5. Training Procedure</p>
        <p>The policy network was initially pre-trained using historical online datasets and synthetic trajectories (<italic>n</italic> = 8432) generated to reflect distributions reported in clinical cohorts (2015-2021). Off-policy evaluation was performed using importance sampling:</p>
        <disp-formula id="FD9">
          <mml:math>
            <mml:mrow>
              <mml:msub>
                <mml:mover accent="true">
                  <mml:mi>V</mml:mi>
                  <mml:mo>^</mml:mo>
                </mml:mover>
                <mml:mrow>
                  <mml:mi>I</mml:mi>
                  <mml:mi>S</mml:mi>
                </mml:mrow>
              </mml:msub>
              <mml:mo>=</mml:mo>
              <mml:mfrac>
                <mml:mn>1</mml:mn>
                <mml:mi>n</mml:mi>
              </mml:mfrac>
              <mml:munderover>
                <mml:mstyle mathsize="140%" displaystyle="true">
                  <mml:mo>∑</mml:mo>
                </mml:mstyle>
                <mml:mrow>
                  <mml:mi>i</mml:mi>
                  <mml:mo>=</mml:mo>
                  <mml:mn>1</mml:mn>
                </mml:mrow>
                <mml:mi>n</mml:mi>
              </mml:munderover>
              <mml:mfrac>
                <mml:mrow>
                  <mml:msub>
                    <mml:mi>π</mml:mi>
                    <mml:mi>θ</mml:mi>
                  </mml:msub>
                  <mml:mrow>
                    <mml:mo>(</mml:mo>
                    <mml:mrow>
                      <mml:mrow>
                        <mml:mrow>
                          <mml:msub>
                            <mml:mi>a</mml:mi>
                            <mml:mi>i</mml:mi>
                          </mml:msub>
                        </mml:mrow>
                        <mml:mo>|</mml:mo>
                      </mml:mrow>
                      <mml:msub>
                        <mml:mi>s</mml:mi>
                        <mml:mi>i</mml:mi>
                      </mml:msub>
                    </mml:mrow>
                    <mml:mo>)</mml:mo>
                  </mml:mrow>
                </mml:mrow>
                <mml:mrow>
                  <mml:msub>
                    <mml:mi>π</mml:mi>
                    <mml:mi>b</mml:mi>
                  </mml:msub>
                  <mml:mrow>
                    <mml:mo>(</mml:mo>
                    <mml:mrow>
                      <mml:mrow>
                        <mml:mrow>
                          <mml:msub>
                            <mml:mi>a</mml:mi>
                            <mml:mi>i</mml:mi>
                          </mml:msub>
                        </mml:mrow>
                        <mml:mo>|</mml:mo>
                      </mml:mrow>
                      <mml:msub>
                        <mml:mi>s</mml:mi>
                        <mml:mi>i</mml:mi>
                      </mml:msub>
                    </mml:mrow>
                    <mml:mo>)</mml:mo>
                  </mml:mrow>
                </mml:mrow>
              </mml:mfrac>
              <mml:msub>
                <mml:mi>R</mml:mi>
                <mml:mi>i</mml:mi>
              </mml:msub>
            </mml:mrow>
          </mml:math>
        </disp-formula>
        <p>where <inline-formula><mml:math><mml:mrow><mml:msub><mml:mi> π </mml:mi><mml:mi> b </mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> denotes the behavior policy and <inline-formula><mml:math><mml:mrow><mml:msub><mml:mi> R </mml:mi><mml:mi> i </mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> represents observed returns.</p>
        <p>During the clinical trial, online learning was employed with conservative policy updates to ensure stability. A trust-region constraint limited divergence between successive policies:</p>
        <disp-formula id="FD10">
          <mml:math display="inline">
            <mml:mrow>
              <mml:msub>
                <mml:mi mathvariant="double-struck">E</mml:mi>
                <mml:mi>t</mml:mi>
              </mml:msub>
              <mml:mrow>
                <mml:mo>[</mml:mo>
                <mml:mrow>
                  <mml:mi>K</mml:mi>
                  <mml:mi>L</mml:mi>
                  <mml:mrow>
                    <mml:mo>(</mml:mo>
                    <mml:mrow>
                      <mml:msub>
                        <mml:mi>π</mml:mi>
                        <mml:mrow>
                          <mml:msub>
                            <mml:mi>θ</mml:mi>
                            <mml:mrow>
                              <mml:mtext>old</mml:mtext>
                            </mml:mrow>
                          </mml:msub>
                        </mml:mrow>
                      </mml:msub>
                      <mml:mrow>
                        <mml:mo>(</mml:mo>
                        <mml:mrow>
                          <mml:mrow>
                            <mml:mrow>
                              <mml:mo>⋅</mml:mo>
                              <mml:mtext>
                                 
                              </mml:mtext>
                            </mml:mrow>
                            <mml:mo>|</mml:mo>
                          </mml:mrow>
                          <mml:msub>
                            <mml:mi>s</mml:mi>
                            <mml:mi>t</mml:mi>
                          </mml:msub>
                        </mml:mrow>
                        <mml:mo>)</mml:mo>
                      </mml:mrow>
                      <mml:mo>,</mml:mo>
                      <mml:msub>
                        <mml:mi>π</mml:mi>
                        <mml:mi>θ</mml:mi>
                      </mml:msub>
                      <mml:mrow>
                        <mml:mo>(</mml:mo>
                        <mml:mrow>
                          <mml:mrow>
                            <mml:mrow>
                              <mml:mo>⋅</mml:mo>
                              <mml:mtext>
                                 
                              </mml:mtext>
                            </mml:mrow>
                            <mml:mo>|</mml:mo>
                          </mml:mrow>
                          <mml:msub>
                            <mml:mi>s</mml:mi>
                            <mml:mi>t</mml:mi>
                          </mml:msub>
                        </mml:mrow>
                        <mml:mo>)</mml:mo>
                      </mml:mrow>
                    </mml:mrow>
                    <mml:mo>)</mml:mo>
                  </mml:mrow>
                </mml:mrow>
                <mml:mo>]</mml:mo>
              </mml:mrow>
              <mml:mo>≤</mml:mo>
              <mml:mi>δ</mml:mi>
            </mml:mrow>
          </mml:math>
        </disp-formula>
      </sec>
      <sec id="sec2dot5">
        <title>2.5. Causal Machine Learning Analysis</title>
        <p>To estimate heterogeneous treatment effects, a doubly robust estimator combining outcome regression and inverse probability weighting was applied. The conditional average treatment effect (CATE) for synthetic patient <inline-formula><mml:math><mml:mi> i </mml:mi></mml:math></inline-formula> was defined as:</p>
        <disp-formula id="FD11">
          <mml:math>
            <mml:mrow>
              <mml:mi>τ</mml:mi>
              <mml:mrow>
                <mml:mo>(</mml:mo>
                <mml:mrow>
                  <mml:msub>
                    <mml:mi>X</mml:mi>
                    <mml:mi>i</mml:mi>
                  </mml:msub>
                </mml:mrow>
                <mml:mo>)</mml:mo>
              </mml:mrow>
              <mml:mo>=</mml:mo>
              <mml:mi>E</mml:mi>
              <mml:mrow>
                <mml:mo>[</mml:mo>
                <mml:mrow>
                  <mml:mrow>
                    <mml:mrow>
                      <mml:msub>
                        <mml:mi>Y</mml:mi>
                        <mml:mi>i</mml:mi>
                      </mml:msub>
                      <mml:mrow>
                        <mml:mo>(</mml:mo>
                        <mml:mn>1</mml:mn>
                        <mml:mo>)</mml:mo>
                      </mml:mrow>
                      <mml:mo>−</mml:mo>
                      <mml:msub>
                        <mml:mi>Y</mml:mi>
                        <mml:mi>i</mml:mi>
                      </mml:msub>
                      <mml:mrow>
                        <mml:mo>(</mml:mo>
                        <mml:mn>0</mml:mn>
                        <mml:mo>)</mml:mo>
                      </mml:mrow>
                      <mml:mtext>
                         
                      </mml:mtext>
                    </mml:mrow>
                    <mml:mo>|</mml:mo>
                  </mml:mrow>
                  <mml:msub>
                    <mml:mi>X</mml:mi>
                    <mml:mi>i</mml:mi>
                  </mml:msub>
                </mml:mrow>
                <mml:mo>]</mml:mo>
              </mml:mrow>
            </mml:mrow>
          </mml:math>
        </disp-formula>
        <p>where <inline-formula><mml:math><mml:mrow><mml:msub><mml:mi> Y </mml:mi><mml:mi> i </mml:mi></mml:msub><mml:mrow><mml:mo> ( </mml:mo><mml:mi> a </mml:mi><mml:mo> ) </mml:mo></mml:mrow></mml:mrow></mml:math></inline-formula> denotes the potential outcome under treatment <inline-formula><mml:math><mml:mi> a </mml:mi></mml:math></inline-formula> .</p>
        <p>The doubly robust estimator is:</p>
        <disp-formula id="FD12">
          <mml:math>
            <mml:mrow>
              <mml:msub>
                <mml:mover accent="true">
                  <mml:mi>τ</mml:mi>
                  <mml:mo>^</mml:mo>
                </mml:mover>
                <mml:mrow>
                  <mml:mi>D</mml:mi>
                  <mml:mi>R</mml:mi>
                </mml:mrow>
              </mml:msub>
              <mml:mo>=</mml:mo>
              <mml:mfrac>
                <mml:mn>1</mml:mn>
                <mml:mi>n</mml:mi>
              </mml:mfrac>
              <mml:munderover>
                <mml:mstyle mathsize="140%" displaystyle="true">
                  <mml:mo>∑</mml:mo>
                </mml:mstyle>
                <mml:mrow>
                  <mml:mi>i</mml:mi>
                  <mml:mo>=</mml:mo>
                  <mml:mn>1</mml:mn>
                </mml:mrow>
                <mml:mi>n</mml:mi>
              </mml:munderover>
              <mml:mrow>
                <mml:mo>[</mml:mo>
                <mml:mrow>
                  <mml:mover accent="true">
                    <mml:mi>μ</mml:mi>
                    <mml:mo>^</mml:mo>
                  </mml:mover>
                  <mml:mrow>
                    <mml:mo>(</mml:mo>
                    <mml:mrow>
                      <mml:msub>
                        <mml:mi>X</mml:mi>
                        <mml:mi>i</mml:mi>
                      </mml:msub>
                      <mml:mo>,</mml:mo>
                      <mml:mn>1</mml:mn>
                    </mml:mrow>
                    <mml:mo>)</mml:mo>
                  </mml:mrow>
                  <mml:mo>−</mml:mo>
                  <mml:mover accent="true">
                    <mml:mi>μ</mml:mi>
                    <mml:mo>^</mml:mo>
                  </mml:mover>
                  <mml:mrow>
                    <mml:mo>(</mml:mo>
                    <mml:mrow>
                      <mml:msub>
                        <mml:mi>X</mml:mi>
                        <mml:mi>i</mml:mi>
                      </mml:msub>
                      <mml:mo>,</mml:mo>
                      <mml:mn>0</mml:mn>
                    </mml:mrow>
                    <mml:mo>)</mml:mo>
                  </mml:mrow>
                  <mml:mo>+</mml:mo>
                  <mml:mfrac>
                    <mml:mrow>
                      <mml:msub>
                        <mml:mi>A</mml:mi>
                        <mml:mi>i</mml:mi>
                      </mml:msub>
                      <mml:mrow>
                        <mml:mo>(</mml:mo>
                        <mml:mrow>
                          <mml:msub>
                            <mml:mi>Y</mml:mi>
                            <mml:mi>i</mml:mi>
                          </mml:msub>
                          <mml:mo>−</mml:mo>
                          <mml:mover accent="true">
                            <mml:mi>μ</mml:mi>
                            <mml:mo>^</mml:mo>
                          </mml:mover>
                          <mml:mrow>
                            <mml:mo>(</mml:mo>
                            <mml:mrow>
                              <mml:msub>
                                <mml:mi>X</mml:mi>
                                <mml:mi>i</mml:mi>
                              </mml:msub>
                              <mml:mo>,</mml:mo>
                              <mml:mn>1</mml:mn>
                            </mml:mrow>
                            <mml:mo>)</mml:mo>
                          </mml:mrow>
                        </mml:mrow>
                        <mml:mo>)</mml:mo>
                      </mml:mrow>
                    </mml:mrow>
                    <mml:mrow>
                      <mml:mover accent="true">
                        <mml:mi>e</mml:mi>
                        <mml:mo>^</mml:mo>
                      </mml:mover>
                      <mml:mrow>
                        <mml:mo>(</mml:mo>
                        <mml:mrow>
                          <mml:msub>
                            <mml:mi>X</mml:mi>
                            <mml:mi>i</mml:mi>
                          </mml:msub>
                        </mml:mrow>
                        <mml:mo>)</mml:mo>
                      </mml:mrow>
                    </mml:mrow>
                  </mml:mfrac>
                  <mml:mo>−</mml:mo>
                  <mml:mfrac>
                    <mml:mrow>
                      <mml:mrow>
                        <mml:mo>(</mml:mo>
                        <mml:mrow>
                          <mml:mn>1</mml:mn>
                          <mml:mo>−</mml:mo>
                          <mml:msub>
                            <mml:mi>A</mml:mi>
                            <mml:mi>i</mml:mi>
                          </mml:msub>
                        </mml:mrow>
                        <mml:mo>)</mml:mo>
                      </mml:mrow>
                      <mml:mrow>
                        <mml:mo>(</mml:mo>
                        <mml:mrow>
                          <mml:msub>
                            <mml:mi>Y</mml:mi>
                            <mml:mi>i</mml:mi>
                          </mml:msub>
                          <mml:mo>−</mml:mo>
                          <mml:mover accent="true">
                            <mml:mi>μ</mml:mi>
                            <mml:mo>^</mml:mo>
                          </mml:mover>
                          <mml:mrow>
                            <mml:mo>(</mml:mo>
                            <mml:mrow>
                              <mml:msub>
                                <mml:mi>X</mml:mi>
                                <mml:mi>i</mml:mi>
                              </mml:msub>
                              <mml:mo>,</mml:mo>
                              <mml:mn>0</mml:mn>
                            </mml:mrow>
                            <mml:mo>)</mml:mo>
                          </mml:mrow>
                        </mml:mrow>
                        <mml:mo>)</mml:mo>
                      </mml:mrow>
                    </mml:mrow>
                    <mml:mrow>
                      <mml:mn>1</mml:mn>
                      <mml:mo>−</mml:mo>
                      <mml:mover accent="true">
                        <mml:mi>e</mml:mi>
                        <mml:mo>^</mml:mo>
                      </mml:mover>
                      <mml:mrow>
                        <mml:mo>(</mml:mo>
                        <mml:mrow>
                          <mml:msub>
                            <mml:mi>X</mml:mi>
                            <mml:mi>i</mml:mi>
                          </mml:msub>
                        </mml:mrow>
                        <mml:mo>)</mml:mo>
                      </mml:mrow>
                    </mml:mrow>
                  </mml:mfrac>
                </mml:mrow>
                <mml:mo>]</mml:mo>
              </mml:mrow>
            </mml:mrow>
          </mml:math>
        </disp-formula>
        <p>where <inline-formula><mml:math><mml:mrow><mml:mover accent="true"><mml:mi> μ </mml:mi><mml:mo> ^ </mml:mo></mml:mover><mml:mrow><mml:mo> ( </mml:mo><mml:mrow><mml:mi> X </mml:mi><mml:mo> , </mml:mo><mml:mi> A </mml:mi></mml:mrow><mml:mo> ) </mml:mo></mml:mrow><mml:mo> = </mml:mo><mml:mi> E </mml:mi><mml:mrow><mml:mo> [ </mml:mo><mml:mrow><mml:mrow><mml:mrow><mml:mi> Y </mml:mi><mml:mtext>   </mml:mtext></mml:mrow><mml:mo> | </mml:mo></mml:mrow><mml:mi> X </mml:mi><mml:mo> , </mml:mo><mml:mi> A </mml:mi></mml:mrow><mml:mo> ] </mml:mo></mml:mrow></mml:mrow></mml:math></inline-formula> is the outcome model and <inline-formula><mml:math><mml:mrow><mml:mover accent="true"><mml:mi> e </mml:mi><mml:mo> ^ </mml:mo></mml:mover><mml:mrow><mml:mo> ( </mml:mo><mml:mi> X </mml:mi><mml:mo> ) </mml:mo></mml:mrow><mml:mo> = </mml:mo><mml:mi> P </mml:mi><mml:mrow><mml:mo> ( </mml:mo><mml:mrow><mml:mrow><mml:mrow><mml:mi> A </mml:mi><mml:mo> = </mml:mo><mml:mn> 1 </mml:mn><mml:mtext>   </mml:mtext></mml:mrow><mml:mo> | </mml:mo></mml:mrow><mml:mi> X </mml:mi></mml:mrow><mml:mo> ) </mml:mo></mml:mrow></mml:mrow></mml:math></inline-formula> denotes the propensity score.</p>
        <p>Gradient boosting machines were used for both propensity estimation and outcome regression. Subgroup analyses evaluated treatment effect modification by bipolar subtype, mixed features, rapid cycling, anxiety comorbidity, and CYP2D6 status.</p>
      </sec>
      <sec id="sec2dot6">
        <title>2.6. Outcomes</title>
        <p><bold>Primary outcome:</bold> Simulated change in Montgomery–Åsberg Depression Rating Scale (MADRS) score from baseline to 12 months, computed using predefined scoring functions within the simulation framework independent of treatment assignment.</p>
        <p><bold>Secondary outcomes:</bold> Simulated treatment response (≥50% reduction in MADRS), remission (MADRS ≤ 10), time-to-response, mood destabilization events, quality-of-life indices (SF-36), functional impairment metrics (WHO Disability Assessment Schedule), and treatment-emergent simulated adverse event frequencies generated within the synthetic cohort.</p>
        <p><bold>Exploratory outcomes:</bold> Simulation-based estimates of cost-effectiveness (cost per QALY), medication adherence patterns, and synthetic patient satisfaction indicators derived from modelled behavioural trajectories.</p>
      </sec>
      <sec id="sec2dot7">
        <title>2.7. Statistical Analysis</title>
        <p><bold>Primary analyses:</bold> Differences in simulated MADRS change between the RL-personalized strategy and pooled standard protocols were evaluated using ANCOVA, adjusting for predefined baseline covariates within the simulation framework.</p>
        <disp-formula id="FD13">
          <mml:math>
            <mml:mrow>
              <mml:msub>
                <mml:mi>Y</mml:mi>
                <mml:mi>i</mml:mi>
              </mml:msub>
              <mml:mo>=</mml:mo>
              <mml:msub>
                <mml:mi>β</mml:mi>
                <mml:mn>0</mml:mn>
              </mml:msub>
              <mml:mo>+</mml:mo>
              <mml:msub>
                <mml:mi>β</mml:mi>
                <mml:mn>1</mml:mn>
              </mml:msub>
              <mml:msub>
                <mml:mi>T</mml:mi>
                <mml:mi>i</mml:mi>
              </mml:msub>
              <mml:mo>+</mml:mo>
              <mml:msub>
                <mml:mi>β</mml:mi>
                <mml:mn>2</mml:mn>
              </mml:msub>
              <mml:msub>
                <mml:mi>Y</mml:mi>
                <mml:mrow>
                  <mml:mn>0</mml:mn>
                  <mml:mi>i</mml:mi>
                </mml:mrow>
              </mml:msub>
              <mml:mo>+</mml:mo>
              <mml:msub>
                <mml:mi>β</mml:mi>
                <mml:mn>3</mml:mn>
              </mml:msub>
              <mml:msub>
                <mml:mi>S</mml:mi>
                <mml:mi>i</mml:mi>
              </mml:msub>
              <mml:mo>+</mml:mo>
              <mml:msub>
                <mml:mi>ϵ</mml:mi>
                <mml:mi>i</mml:mi>
              </mml:msub>
            </mml:mrow>
          </mml:math>
        </disp-formula>
        <p>where <inline-formula><mml:math><mml:mrow><mml:msub><mml:mi> Y </mml:mi><mml:mi> i </mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> represents the 12-month MADRS score, <inline-formula><mml:math><mml:mrow><mml:msub><mml:mi> T </mml:mi><mml:mi> i </mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> treatment assignment, <inline-formula><mml:math><mml:mrow><mml:msub><mml:mi> Y </mml:mi><mml:mrow><mml:mn> 0 </mml:mn><mml:mi> i </mml:mi></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula> baseline MADRS, and <inline-formula><mml:math><mml:mrow><mml:msub><mml:mi> S </mml:mi><mml:mi> i </mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> stratification factors.</p>
        <p>Binary outcomes were modelled via logistic regression:</p>
        <disp-formula id="FD14">
          <mml:math>
            <mml:mrow>
              <mml:mtext>logit</mml:mtext>
              <mml:mrow>
                <mml:mo>(</mml:mo>
                <mml:mrow>
                  <mml:mi>P</mml:mi>
                  <mml:mrow>
                    <mml:mo>(</mml:mo>
                    <mml:mrow>
                      <mml:msub>
                        <mml:mi>Y</mml:mi>
                        <mml:mi>i</mml:mi>
                      </mml:msub>
                      <mml:mo>=</mml:mo>
                      <mml:mn>1</mml:mn>
                    </mml:mrow>
                    <mml:mo>)</mml:mo>
                  </mml:mrow>
                </mml:mrow>
                <mml:mo>)</mml:mo>
              </mml:mrow>
              <mml:mo>=</mml:mo>
              <mml:msub>
                <mml:mi>β</mml:mi>
                <mml:mn>0</mml:mn>
              </mml:msub>
              <mml:mo>+</mml:mo>
              <mml:msub>
                <mml:mi>β</mml:mi>
                <mml:mn>1</mml:mn>
              </mml:msub>
              <mml:msub>
                <mml:mi>T</mml:mi>
                <mml:mi>i</mml:mi>
              </mml:msub>
              <mml:mo>+</mml:mo>
              <mml:msub>
                <mml:mi>β</mml:mi>
                <mml:mn>2</mml:mn>
              </mml:msub>
              <mml:msub>
                <mml:mi>X</mml:mi>
                <mml:mi>i</mml:mi>
              </mml:msub>
            </mml:mrow>
          </mml:math>
        </disp-formula>
        <p>Time-to-event outcomes were analysed using Cox proportional hazards models:</p>
        <disp-formula id="FD15">
          <mml:math>
            <mml:mrow>
              <mml:mi>h</mml:mi>
              <mml:mrow>
                <mml:mo>(</mml:mo>
                <mml:mrow>
                  <mml:mrow>
                    <mml:mi>t</mml:mi>
                    <mml:mo>|</mml:mo>
                  </mml:mrow>
                  <mml:mi>X</mml:mi>
                </mml:mrow>
                <mml:mo>)</mml:mo>
              </mml:mrow>
              <mml:mo>=</mml:mo>
              <mml:msub>
                <mml:mi>h</mml:mi>
                <mml:mn>0</mml:mn>
              </mml:msub>
              <mml:mrow>
                <mml:mo>(</mml:mo>
                <mml:mi>t</mml:mi>
                <mml:mo>)</mml:mo>
              </mml:mrow>
              <mml:mtext>exp</mml:mtext>
              <mml:mrow>
                <mml:mo>(</mml:mo>
                <mml:mrow>
                  <mml:msup>
                    <mml:mi>β</mml:mi>
                    <mml:mtext>T</mml:mtext>
                  </mml:msup>
                  <mml:mi>X</mml:mi>
                </mml:mrow>
                <mml:mo>)</mml:mo>
              </mml:mrow>
            </mml:mrow>
          </mml:math>
        </disp-formula>
        <p>Multiple imputation addressed missing data under missing-at-random assumptions, with sensitivity analyses for missing-not-at-random scenarios. All analyses followed the intention-to-treat principle.</p>
      </sec>
      <sec id="sec2dot8">
        <title>2.8. Model Interpretability and Safety</title>
        <p>Model interpretability was assessed using SHAP values, defined as:</p>
        <disp-formula id="FD16">
          <mml:math display="inline">
            <mml:mrow>
              <mml:msub>
                <mml:mi>ϕ</mml:mi>
                <mml:mi>j</mml:mi>
              </mml:msub>
              <mml:mo>=</mml:mo>
              <mml:munder>
                <mml:mstyle mathsize="140%" displaystyle="true">
                  <mml:mo>∑</mml:mo>
                </mml:mstyle>
                <mml:mrow>
                  <mml:mi>S</mml:mi>
                  <mml:mo>⊆</mml:mo>
                  <mml:mi>N</mml:mi>
                  <mml:mo>∖</mml:mo>
                  <mml:mrow>
                    <mml:mo>{</mml:mo>
                    <mml:mi>j</mml:mi>
                    <mml:mo>}</mml:mo>
                  </mml:mrow>
                </mml:mrow>
              </mml:munder>
              <mml:mfrac>
                <mml:mrow>
                  <mml:mrow>
                    <mml:mo>|</mml:mo>
                    <mml:mi>S</mml:mi>
                    <mml:mo>|</mml:mo>
                  </mml:mrow>
                  <mml:mo>!</mml:mo>
                  <mml:mrow>
                    <mml:mo>(</mml:mo>
                    <mml:mrow>
                      <mml:mrow>
                        <mml:mo>|</mml:mo>
                        <mml:mi>N</mml:mi>
                        <mml:mo>|</mml:mo>
                      </mml:mrow>
                      <mml:mo>−</mml:mo>
                      <mml:mrow>
                        <mml:mo>|</mml:mo>
                        <mml:mi>S</mml:mi>
                        <mml:mo>|</mml:mo>
                      </mml:mrow>
                      <mml:mo>−</mml:mo>
                      <mml:mn>1</mml:mn>
                    </mml:mrow>
                    <mml:mo>)</mml:mo>
                  </mml:mrow>
                  <mml:mo>!</mml:mo>
                </mml:mrow>
                <mml:mrow>
                  <mml:mo>∣</mml:mo>
                  <mml:mi>N</mml:mi>
                  <mml:mo>∣</mml:mo>
                  <mml:mo>!</mml:mo>
                </mml:mrow>
              </mml:mfrac>
              <mml:mrow>
                <mml:mo>[</mml:mo>
                <mml:mrow>
                  <mml:msub>
                    <mml:mi>f</mml:mi>
                    <mml:mrow>
                      <mml:mi>S</mml:mi>
                      <mml:mo>∪</mml:mo>
                      <mml:mrow>
                        <mml:mo>{</mml:mo>
                        <mml:mi>j</mml:mi>
                        <mml:mo>}</mml:mo>
                      </mml:mrow>
                    </mml:mrow>
                  </mml:msub>
                  <mml:mrow>
                    <mml:mo>(</mml:mo>
                    <mml:mrow>
                      <mml:msub>
                        <mml:mi>x</mml:mi>
                        <mml:mrow>
                          <mml:mi>S</mml:mi>
                          <mml:mo>∪</mml:mo>
                          <mml:mrow>
                            <mml:mo>{</mml:mo>
                            <mml:mi>j</mml:mi>
                            <mml:mo>}</mml:mo>
                          </mml:mrow>
                        </mml:mrow>
                      </mml:msub>
                    </mml:mrow>
                    <mml:mo>)</mml:mo>
                  </mml:mrow>
                  <mml:mo>−</mml:mo>
                  <mml:mo>
                  </mml:mo>
                  <mml:msub>
                    <mml:mi>f</mml:mi>
                    <mml:mi>S</mml:mi>
                  </mml:msub>
                  <mml:mrow>
                    <mml:mo>(</mml:mo>
                    <mml:mrow>
                      <mml:msub>
                        <mml:mi>x</mml:mi>
                        <mml:mi>S</mml:mi>
                      </mml:msub>
                    </mml:mrow>
                    <mml:mo>)</mml:mo>
                  </mml:mrow>
                </mml:mrow>
                <mml:mo>]</mml:mo>
              </mml:mrow>
            </mml:mrow>
          </mml:math>
        </disp-formula>
        <p>Clinical utility was evaluated using decision curve analysis:</p>
        <disp-formula id="FD17">
          <mml:math>
            <mml:mrow>
              <mml:mtext>Net Benefit</mml:mtext>
              <mml:mo>=</mml:mo>
              <mml:mfrac>
                <mml:mrow>
                  <mml:mi>T</mml:mi>
                  <mml:mi>P</mml:mi>
                </mml:mrow>
                <mml:mi>N</mml:mi>
              </mml:mfrac>
              <mml:mo>−</mml:mo>
              <mml:mfrac>
                <mml:mrow>
                  <mml:mi>F</mml:mi>
                  <mml:mi>P</mml:mi>
                </mml:mrow>
                <mml:mi>N</mml:mi>
              </mml:mfrac>
              <mml:mo>⋅</mml:mo>
              <mml:mfrac>
                <mml:mrow>
                  <mml:mn>1</mml:mn>
                  <mml:mo>−</mml:mo>
                  <mml:msub>
                    <mml:mi>p</mml:mi>
                    <mml:mi>t</mml:mi>
                  </mml:msub>
                </mml:mrow>
                <mml:mrow>
                  <mml:msub>
                    <mml:mi>p</mml:mi>
                    <mml:mi>t</mml:mi>
                  </mml:msub>
                </mml:mrow>
              </mml:mfrac>
            </mml:mrow>
          </mml:math>
        </disp-formula>
        <p>where <inline-formula><mml:math><mml:mrow><mml:msub><mml:mi> p </mml:mi><mml:mi> t </mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> denotes the decision threshold probability.</p>
        <p>Safety monitoring included real-time surveillance for destabilization events. Hard constraints within the RL policy prohibited antidepressant monotherapy in synthetic patients with high baseline YMRS scores or mixed features. Clinician overrides were systematically analysed to identify policy limitations and inform iterative refinement.</p>
      </sec>
    </sec>
    <sec id="sec3">
      <title>3. Results</title>
      <sec id="sec3dot1">
        <title>3.1. Virtual Participant Characteristics</title>
        <p>Between March 2022 and August 2023, 4856 virtual candidate profiles were generated and filtered according to eligibility rules, with 2500 randomized (<xref ref-type="fig" rid="fig1">Figure 1</xref><xref ref-type="fig" rid="fig1">Figure 1</xref>). Baseline characteristics were well-balanced across arms (<bold>Table 1</bold>). Mean age was 38.6 years (SD = 12.3), 52% were female, and 45% had bipolar I disorder. Mean baseline MADRS was 24.1 (SD = 11.6) and mean YMRS was 14.5 (SD = 7.9). Mixed features were present in 55% of virtual participants.</p>
        <fig id="fig1">
          <label>Figure 1</label>
          <graphic xlink:href="https://html.scirp.org/file/1115137-rId112.jpeg?20260528022033" />
        </fig>
        <p><bold>Figure 1.</bold> CONSORT-style simulation flow diagram showing virtual participant allocation within the RL-CADENCE framework. Through the RL-CADENCE simulation trial of 4856 individuals screened, 2500 were randomized to four treatment arms. Completion rates at 12 months were similar across arms (87% - 95%).</p>
        <p><bold>Table 1.</bold> Baseline characteristics of randomized virtual participants.</p>
        <table-wrap id="tbl1">
          <label>Table 1</label>
          <table>
            <tbody>
              <tr>
                <td>
                  <bold>Characteristic</bold>
                </td>
                <td>
                  <bold>Li + SSRI</bold>
                  <bold>(n = 773)</bold>
                </td>
                <td>
                  <bold>QTP + LTG</bold>
                  <bold>(n = 711)</bold>
                </td>
                <td>
                  <bold>LUR + MS</bold>
                  <bold>(n = 762)</bold>
                </td>
                <td>
                  <bold>RL-Pers</bold>
                  <bold>(n = 254)</bold>
                </td>
                <td>
                  <bold>Total</bold>
                  <bold>(n = 2500)</bold>
                </td>
              </tr>
              <tr>
                <td>Age, years</td>
                <td>38.4 ± 12.1</td>
                <td>38.7 ± 12.5</td>
                <td>38.5 ± 12.2</td>
                <td>38.9 ± 12.4</td>
                <td>38.6 ± 12.3</td>
              </tr>
              <tr>
                <td>Female, %</td>
                <td>52.0</td>
                <td>51.8</td>
                <td>52.1</td>
                <td>51.6</td>
                <td>51.9</td>
              </tr>
              <tr>
                <td>Bipolar I, %</td>
                <td>45.0</td>
                <td>45.0</td>
                <td>45.0</td>
                <td>44.9</td>
                <td>45.0</td>
              </tr>
              <tr>
                <td>Mixed features, %</td>
                <td>55.0</td>
                <td>55.0</td>
                <td>55.0</td>
                <td>55.1</td>
                <td>55.0</td>
              </tr>
              <tr>
                <td>Baseline MADRS</td>
                <td>24.0 ± 11.5</td>
                <td>24.2 ± 11.7</td>
                <td>24.1 ± 11.6</td>
                <td>24.3 ± 11.8</td>
                <td>24.1 ± 11.6</td>
              </tr>
              <tr>
                <td>Baseline YMRS</td>
                <td>14.4 ± 7.8</td>
                <td>14.6 ± 8.0</td>
                <td>14.5 ± 7.9</td>
                <td>14.7 ± 8.1</td>
                <td>14.5 ± 7.9</td>
              </tr>
              <tr>
                <td>Prior failed trials</td>
                <td>1.8 ± 1.2</td>
                <td>1.7 ± 1.1</td>
                <td>1.8 ± 1.2</td>
                <td>1.8 ± 1.3</td>
                <td>1.8 ± 1.2</td>
              </tr>
              <tr>
                <td>CYP2D6 poor metabolizer, %</td>
                <td>8.0</td>
                <td>8.0</td>
                <td>8.0</td>
                <td>7.9</td>
                <td>8.0</td>
              </tr>
            </tbody>
          </table>
        </table-wrap>
      </sec>
      <sec id="sec3dot2">
        <title>3.2. Primary Outcome</title>
        <p>At 12 months, the RL-personalized strategy demonstrated greater simulated MADRS reduction compared with pooled standard treatment strategies (mean difference: −5.6 points, 95% CI: −7.6 to −3.6; F (1, 2495) = 28.4, <italic>p</italic> &lt; 0.001; Cohen’s <italic>d</italic> = 0.78, indicating a large effect size). The least-squares mean change from baseline was −16.8 points (95% CI: −18.2 to −15.4) for the RL-personalized strategy versus −11.2 points (95% CI: −12.1 to −10.3) for pooled standard strategies within the simulated cohort.</p>
        <fig id="fig2">
          <label>Figure 2</label>
          <graphic xlink:href="https://html.scirp.org/file/1115137-rId113.jpeg?20260528022033" />
        </fig>
        <p><bold>Figure 2.</bold> Longitudinal clinical outcomes and predictive features. (A) Montgomery-Åsberg Depression Rating Scale (MADRS) trajectories by treatment arm, showing superior symptom reduction in the RL-personalized arm. (B) Young Mania Rating Scale (YMRS) trajectories demonstrating maintained mood stability. (C) Cumulative mood destabilization events over follow-up. (D) Twelve-month outcome summary including response and remission rates. (E) Feature importance for RL treatment selection policy, with clinical severity measures (MADRS, YMRS baseline) as the strongest predictors.</p>
        <p>Longitudinal outcome trajectories showed divergence beginning around month 2, with the RL-based strategy maintaining superior simulated outcomes throughout follow-up (<xref ref-type="fig" rid="fig2">Figure 2(A)</xref>,<xref ref-type="fig" rid="fig2">Figure 2(B)</xref>). Sensitivity analyses using mixed models for repeated measures and alternative simulation assumptions yielded consistent results.</p>
      </sec>
      <sec id="sec3dot3">
        <title>3.3. Secondary Outcomes</title>
        <p><bold>Treatment Response:</bold> Response rates (≥50% MADRS reduction) were significantly higher in the RL-personalized arm (95.7% vs. 58.9%; odds ratio [OR] = 14.8, 95% CI: 8.9 - 24.6; p &lt; 0.001; number needed to treat [NNT] = 2.2) (<bold>Table 2</bold>, <xref ref-type="fig" rid="fig2">Figure 2(D)</xref>). Time to response was shorter (median 6 weeks vs. 10 weeks; hazard ratio [HR] = 1.68, 95% CI: 1.52 - 1.86; p &lt; 0.001).</p>
        <p><bold>Remission:</bold> Remission rates (MADRS ≤ 10) were 89.4% vs. 47.2% (OR = 9.6, 95% CI: 6.8 - 13.6; p &lt; 0.001; NNT = 2.4).</p>
        <p><bold>Mood Destabilization:</bold> The RL-personalized arm experienced significantly fewer mood destabilization events (4.8% vs. 10.8% of synthetic patient-months; incidence rate ratio [IRR] = 0.45, 95% CI: 0.32 - 0.63; p &lt; 0.001; number needed to harm [NNH] = 16.8) (<bold>Table 2</bold>, <xref ref-type="fig" rid="fig2">Figure 2(C)</xref>). No virtual participant in the RL arm experienced antidepressant-induced mania requiring hospitalization, compared to 12 (0.6%) in standard arms.</p>
        <p><bold>Table 2.</bold> Primary and secondary outcomes at 12 months.</p>
        <table-wrap id="tbl2">
          <label>Table 2</label>
          <table>
            <tbody>
              <tr>
                <td>
                  <bold>Outcome</bold>
                </td>
                <td>
                  <bold>RL-Personalized</bold>
                </td>
                <td>
                  <bold>Standard Protocols</bold>
                </td>
                <td>
                  <bold>Difference (95% CI)</bold>
                </td>
                <td>
                  <bold>p-value</bold>
                </td>
              </tr>
              <tr>
                <td>MADRS change, mean</td>
                <td>−16.8</td>
                <td>−11.2</td>
                <td>−5.6 (−7.6, −3.6)</td>
                <td>&lt;0.001</td>
              </tr>
              <tr>
                <td>Response rate, %</td>
                <td>95.7</td>
                <td>58.9</td>
                <td>36.8 (31.2, 42.4)</td>
                <td>&lt;0.001</td>
              </tr>
              <tr>
                <td>Remission rate, %</td>
                <td>89.4</td>
                <td>47.2</td>
                <td>42.2 (36.8, 47.6)</td>
                <td>&lt;0.001</td>
              </tr>
              <tr>
                <td>Destabilization rate, %*</td>
                <td>4.82</td>
                <td>10.77</td>
                <td>−5.95 (−8.2, −3.7)</td>
                <td>&lt;0.001</td>
              </tr>
              <tr>
                <td>Quality of life (SF-36)</td>
                <td>68.4 ± 12.3</td>
                <td>58.2 ± 14.1</td>
                <td>10.2 (8.1, 12.3)</td>
                <td>&lt;0.001</td>
              </tr>
              <tr>
                <td>Functional impairment</td>
                <td>12.4 ± 8.2</td>
                <td>18.9 ± 10.5</td>
                <td>−6.5 (−8.1, −4.9)</td>
                <td>&lt;0.001</td>
              </tr>
            </tbody>
          </table>
        </table-wrap>
      </sec>
      <sec id="sec3dot4">
        <title>3.4. Heterogeneous Treatment Effects</title>
        <p>Causal machine learning analysis revealed substantial heterogeneity in treatment effects (<xref ref-type="fig" rid="fig3">Figure 3</xref>). The conditional average treatment effect (CATE) of RL-personalization versus standard care varied significantly across subgroups (<xref ref-type="fig" rid="fig3">Figure 3(D)</xref>, <xref ref-type="fig" rid="fig4">Figure 4(B)</xref>).</p>
        <p>Synthetic patients with mixed features demonstrated the largest benefit (CATE = 15.2, 95% CI: 8.9 - 21.5), followed by those with bipolar I disorder (CATE = 12.3, 95% CI: 7.1 - 17.5). Conversely, synthetic patients with substance use disorders showed smaller but still significant benefits (CATE = 7.2, 95% CI: 1.2 - 13.2). Individual treatment effects (ITE) followed a bimodal distribution, with 62.6% of synthetic patients expected to benefit from RL-personalization (ITE &gt; 0) and mean benefit of 9.6 units (95% CI: 5.2 - 14.0) (<xref ref-type="fig" rid="fig3">Figure 3(C)</xref>).</p>
        <fig id="fig3">
          <label>Figure 3</label>
          <graphic xlink:href="https://html.scirp.org/file/1115137-rId114.jpeg?20260528022033" />
        </fig>
        <p><bold>Figure 3.</bold> Causal inference and counterfactual analysis. (A) Causal directed acyclic graph (DAG) depicting relationships between baseline confounders (X, G, C), treatment assignment (A), unobserved confounders (U), and outcome (Y). (B) Counterfactual outcome distributions for each treatment arm, with RL-personalized showing the highest expected reward. (C) Distribution of individual treatment effects (ITE), with 62.6% of synthetic patients expected to benefit from RL-personalization. (D) Conditional average treatment effects (CATE) across clinical subgroups. (E) Calibration plot demonstrating excellent agreement between predicted and observed probabilities of optimal treatment assignment.</p>
      </sec>
      <sec id="sec3dot5">
        <title>3.5. Model Performance and Interpretability</title>
        <p>The RL policy network achieved an AUC-ROC of 0.89 (95% CI: 0.87 - 0.91) for predicting optimal treatment assignment on held-out test data. Calibration was excellent, with predicted probabilities closely matching observed outcomes (<xref ref-type="fig" rid="fig3">Figure 3(E)</xref>).</p>
        <p>Feature importance analysis (SHAP values) identified baseline MADRS (mean |SHAP| = 0.245), baseline YMRS (0.198), age (0.156), and mixed features (0.134) as the strongest predictors of treatment selection (<xref ref-type="fig" rid="fig2">Figure 2(E)</xref>, <xref ref-type="fig" rid="fig4">Figure 4(D)</xref>). Pharmacogenomic markers (CYP2D6 metabolizer status) contributed modestly (0.067) but showed significant interaction effects with medication class. Decision curve analysis demonstrated superior clinical utility of the RL model across clinically relevant threshold probabilities (<xref ref-type="fig" rid="fig4">Figure 4(C)</xref>). At a threshold of 15% (indicating willingness to treat if probability of response exceeds 15%), the net benefit was 0.28 compared to 0.15 for standard protocols.</p>
        <fig id="fig4">
          <label>Figure 4</label>
          <graphic xlink:href="https://html.scirp.org/file/1115137-rId115.jpeg?20260528022033" />
        </fig>
        <p><bold>Figure 4.</bold> Deep reinforcement learning architecture and clinical decision support. (A) Neural architecture comprising synthetic patient state observation, state embedding, policy and value networks, treatment action selection, clinical environment interaction, and reward calculation. (B) Heterogeneous treatment effects by synthetic patient subgroups, with the largest benefits in mixed states and Type I bipolar disorder. (C) Decision curve analysis showing superior clinical utility of RL-personalized approach across threshold probabilities. (D) SHAP summary of feature importance for treatment decisions. (E) Distribution of mood destabilization risk categories, with RL-personalized approach shifting synthetic patients toward lower-risk categories.</p>
      </sec>
      <sec id="sec3dot6">
        <title>3.6. Safety and Tolerability</title>
        <p>Treatment-emergent simulated adverse events were reported by 68% of virtual participants in the RL arm versus 72% in standard arms (p = 0.18). Serious simulated adverse events occurred in 3.9% vs. 5.2% (p = 0.34). Discontinuation due to simulated adverse events was lower in the RL arm (8.3% vs. 14.2%; p = 0.008). The RL policy avoided high-risk recommendations: among 140 virtual participants with mixed features and high baseline YMRS (&gt;15), the system recommended combination therapy or atypical antipsychotics in 94% of cases, avoiding antidepressant monotherapy.</p>
      </sec>
    </sec>
    <sec id="sec4">
      <title>4. Discussion</title>
      <sec id="sec4dot1">
        <title>4.1. Principal Findings</title>
        <p>The simulation framework suggests that deep reinforcement learning significantly improves antidepressant selection in bipolar spectrum disorders. The RL-personalized approach achieved a 5.6-point greater reduction in depression severity compared to standard protocols, representing a large effect size (Cohen’s d = 0.78) that exceeds thresholds for clinical meaningfulness [<xref ref-type="bibr" rid="B31">31</xref>]. Notably, this improvement was accompanied by a 55% relative reduction in mood destabilization events, addressing the central safety concern in bipolar depression treatment. The NNT of 2.2 indicates that for every two synthetic patients treated with RL-personalization rather than standard care, one additional synthetic patient achieves treatment response. This compares favourably to NNTs of 4 - 7 reported for FDA-approved treatments in bipolar depression [<xref ref-type="bibr" rid="B32">32</xref>]. The NNH of 16.8 suggests a favourable benefit-risk profile, with mood destabilization events occurring less than half as frequently in the RL arm.</p>
      </sec>
      <sec id="sec4dot2">
        <title>4.2. Heterogeneity and Precision Medicine</title>
        <p>Our causal machine learning analysis reveals that treatment effects are highly heterogeneous across synthetic patient subgroups. The finding that synthetic patients with mixed features derive the greatest benefit (CATE = 15.2) is clinically significant, as mixed states represent a treatment challenge with limited evidence-based options [<xref ref-type="bibr" rid="B33">33</xref>]. The RL system’s ability to identify and appropriately treat these synthetic patients avoiding antidepressant monotherapy while optimizing combination strategies likely contributes to the superior outcomes. The modest contribution of pharmacogenomic markers to treatment selection (SHAP importance = 0.067) suggests that clinical features remain primary drivers of treatment response, consistent with recent polygenic score studies in psychiatric disorders [<xref ref-type="bibr" rid="B34">34</xref>]. However, significant gene-treatment interactions indicate that pharmacogenomic data may become more informative as sample sizes increase and genetic architectures are better characterized.</p>
      </sec>
      <sec id="sec4dot3">
        <title>4.3. Comparison with Previous Work</title>
        <p>Prior machine learning studies in bipolar disorder have focused primarily on diagnosis or prognosis prediction [<xref ref-type="bibr" rid="B35">35</xref>][<xref ref-type="bibr" rid="B36">36</xref>]. To our knowledge, large-scale simulated randomized trial framework of deep RL for treatment optimization in psychiatry. Our findings extend prior observational studies suggesting that algorithmic treatment selection can improve outcomes in depression [<xref ref-type="bibr" rid="B37">37</xref>][<xref ref-type="bibr" rid="B38">38</xref>] by demonstrating causality through randomization and addressing safety-critical constraints specific to bipolar disorder. The performance of our RL policy (AUC-ROC = 0.89) compares favourably to predictive models in other medical domains, such as sepsis (AUC 0.70 - 0.80) or acute kidney injury (AUC 0.75 - 0.85) [<xref ref-type="bibr" rid="B39">39</xref>][<xref ref-type="bibr" rid="B40">40</xref>]. The explicit incorporation of safety constraints penalizing destabilization events with represents a methodological advance over standard RL approaches, aligning with principles of safe reinforcement learning [<xref ref-type="bibr" rid="B41">41</xref>].</p>
      </sec>
      <sec id="sec4dot4">
        <title>4.4. Clinical and Policy Implications</title>
        <p>These findings support the integration of AI-driven decision support into clinical practice for bipolar disorder. The RL system functions not as a replacement for clinical judgment but as a tool augmenting evidence-based decision-making. The high override rate observed in other AI implementations [<xref ref-type="bibr" rid="B42">42</xref>] was notably low in our study (8.3%), suggesting good alignment between algorithmic recommendations and clinician preferences.</p>
        <p>From a health systems perspective, the RL approach may reduce costs through improved efficiency (shorter time to response, fewer failed trials) and reduced simulated adverse events (fewer hospitalizations for mood destabilization) [<xref ref-type="bibr" rid="B43">43</xref>]-[<xref ref-type="bibr" rid="B45">45</xref>]. Cost-effectiveness analyses are underway and will inform reimbursement and implementation decisions.</p>
      </sec>
      <sec id="sec4dot5">
        <title>4.5. Limitations</title>
        <p>Several limitations should be considered. First, the trial duration (12 months) captures medium-term but not long-term outcomes. Durability of benefits and potential late-emerging adverse effects require extended follow-up. Second, the saredy population was restricted to synthetic patients with access to academic medical centres; generalizability to community settings or resource-limited environments is uncertain. Third, the RL system was trained primarily on synthetic patients with European ancestry; performance in diverse racial and ethnic groups requires validation. Fourth, while we employed causal inference methods to estimate heterogeneous effects, residual confounding may persist in subgroup analyses. The randomized design ensures unbiased estimation of average treatment effects, but subgroup analyses are inherently observational and should be interpreted cautiously. Fifth, the relatively small sample size in the RL arm (<italic>n</italic> = 254) limits precision for rare simulated adverse events and subgroup analyses. Because outcomes are generated by a synthetic environment designed from prior assumptions, performance estimates may partially reflect modelling choices rather than real clinical complexity.</p>
      </sec>
      <sec id="sec4dot6">
        <title>4.6. Future Directions</title>
        <p>First, extending the reinforcement learning (RL) paradigm toward sequential treatment optimization represents an important next step. Rather than selecting a single treatment at baseline, future systems could dynamically adapt medication strategies based on longitudinal synthetic patient responses, thereby reflecting real clinical workflows. In this context, contextual multi-armed bandit algorithms offer an attractive solution by balancing exploration of alternative treatments with exploitation of known effective strategies under safety constraints. A commonly used approach is the Upper Confidence Bound (UCB) strategy, where the action selected at time step <inline-formula><mml:math><mml:mi> t </mml:mi></mml:math></inline-formula> is defined as:</p>
        <disp-formula id="FD18">
          <mml:math>
            <mml:mrow>
              <mml:msub>
                <mml:mi>A</mml:mi>
                <mml:mi>t</mml:mi>
              </mml:msub>
              <mml:mo>=</mml:mo>
              <mml:mtext>arg</mml:mtext>
              <mml:munder>
                <mml:mrow>
                  <mml:mtext>max</mml:mtext>
                </mml:mrow>
                <mml:mi>a</mml:mi>
              </mml:munder>
              <mml:mrow>
                <mml:mo>[</mml:mo>
                <mml:mrow>
                  <mml:msub>
                    <mml:mover accent="true">
                      <mml:mi>μ</mml:mi>
                      <mml:mo>^</mml:mo>
                    </mml:mover>
                    <mml:mi>a</mml:mi>
                  </mml:msub>
                  <mml:mrow>
                    <mml:mo>(</mml:mo>
                    <mml:mrow>
                      <mml:msub>
                        <mml:mi>x</mml:mi>
                        <mml:mi>t</mml:mi>
                      </mml:msub>
                    </mml:mrow>
                    <mml:mo>)</mml:mo>
                  </mml:mrow>
                  <mml:mo>+</mml:mo>
                  <mml:mi>c</mml:mi>
                  <mml:msqrt>
                    <mml:mrow>
                      <mml:mfrac>
                        <mml:mrow>
                          <mml:mi>ln</mml:mi>
                          <mml:mi>t</mml:mi>
                        </mml:mrow>
                        <mml:mrow>
                          <mml:msub>
                            <mml:mi>N</mml:mi>
                            <mml:mi>a</mml:mi>
                          </mml:msub>
                          <mml:mrow>
                            <mml:mo>(</mml:mo>
                            <mml:mi>t</mml:mi>
                            <mml:mo>)</mml:mo>
                          </mml:mrow>
                        </mml:mrow>
                      </mml:mfrac>
                    </mml:mrow>
                  </mml:msqrt>
                </mml:mrow>
                <mml:mo>]</mml:mo>
              </mml:mrow>
            </mml:mrow>
          </mml:math>
        </disp-formula>
        <p>where <inline-formula><mml:math><mml:mrow><mml:msub><mml:mover accent="true"><mml:mi> μ </mml:mi><mml:mo> ^ </mml:mo></mml:mover><mml:mi> a </mml:mi></mml:msub><mml:mrow><mml:mo> ( </mml:mo><mml:mrow><mml:msub><mml:mi> x </mml:mi><mml:mi> t </mml:mi></mml:msub></mml:mrow><mml:mo> ) </mml:mo></mml:mrow></mml:mrow></mml:math></inline-formula> denotes the estimated expected reward for treatment action <inline-formula><mml:math><mml:mi> a </mml:mi></mml:math></inline-formula> given synthetic patient context <inline-formula><mml:math><mml:mrow><mml:msub><mml:mi> x </mml:mi><mml:mi> t </mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> , <inline-formula><mml:math><mml:mrow><mml:msub><mml:mi> N </mml:mi><mml:mi> a </mml:mi></mml:msub><mml:mrow><mml:mo> ( </mml:mo><mml:mi> t </mml:mi><mml:mo> ) </mml:mo></mml:mrow></mml:mrow></mml:math></inline-formula> represents the number of times action <inline-formula><mml:math><mml:mi> a </mml:mi></mml:math></inline-formula> has been previously selected, and the constant <inline-formula><mml:math><mml:mi> c </mml:mi></mml:math></inline-formula> controls the exploration-exploitation trade-off. Such formulations allow adaptive learning while limiting excessive exploration that could compromise synthetic patient safety.</p>
        <p>Second, expanding the multimodal data sources used for prediction may substantially improve model robustness and clinical relevance. Future frameworks should incorporate digital biomarkers such as actigraphy and voice-based features, alongside neuroimaging and large-scale electronic health record (EHR) data. In parallel, federated learning approaches present a practical solution for cross-institutional model training, enabling collaborative learning while preserving privacy and regulatory compliance by keeping synthetic patient data localized. Third, improving interpretability remains critical for real-world adoption. Although feature attribution methods such as SHAP provide useful insights, future research should focus on more clinically intuitive explainability paradigms. Natural language generation (NLG) systems capable of transforming model outputs into structured narrative explanations may bridge the gap between complex AI reasoning and clinician decision-making by summarizing evidence, uncertainty, and synthetic patient-specific risk factors in an interpretable manner.</p>
      </sec>
    </sec>
    <sec id="sec5">
      <title>5. Conclusion</title>
      <p>Deep reinforcement learning showed strong performance for optimizing antidepressant selection in bipolar spectrum disorders within a simulated evaluation framework, achieving improved symptom reduction while minimizing mood destabilization risk. The RL-CADENCE framework serves as a proof-of-concept for AI-driven precision psychiatry, demonstrating that algorithmic treatment optimization can be modelled within clinically inspired workflows using online and synthetically generated data. These results support further methodological development and motivate future real-world validation, as well as the design of regulatory and implementation frameworks for clinical AI systems in precision mental healthcare.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <title>References</title>
      <ref id="B1">
        <label>1.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Merikangas, K.R., Jin, R., He, J., Kessler, R.C., Lee, S., Sampson, N.A., <italic>et al</italic>. (2011) Prevalence and Correlates of Bipolar Spectrum Disorder in the World Mental Health Survey Initiative. <italic>Archives</italic><italic>of</italic><italic>General</italic><italic>Psychiatry</italic>, 68, 241-251. https://doi.org/10.1001/archgenpsychiatry.2011.12 <pub-id pub-id-type="doi">10.1001/archgenpsychiatry.2011.12</pub-id><pub-id pub-id-type="pmid">21383262</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1001/archgenpsychiatry.2011.12">https://doi.org/10.1001/archgenpsychiatry.2011.12</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Merikangas, K.R.</string-name>
              <string-name>Jin, R.</string-name>
              <string-name>He, J.</string-name>
              <string-name>Kessler, R.C.</string-name>
              <string-name>Lee, S.</string-name>
              <string-name>Sampson, N.A.</string-name>
            </person-group>
            <year>2011</year>
            <article-title>Prevalence and Correlates of Bipolar Spectrum Disorder in the World Mental Health Survey Initiative</article-title>
            <source>Archives of General Psychiatry</source>
            <volume>68</volume>
            <pub-id pub-id-type="doi">10.1001/archgenpsychiatry.2011.12</pub-id>
            <pub-id pub-id-type="pmid">21383262</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B2">
        <label>2.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Vos, T., Lim, S.S., Abbafati, C., Abbas, K.M., Abbasi, M., Abbasifard, M., <italic>et al</italic>. (2020) Global Burden of 369 Diseases and Injuries in 204 Countries and Territories, 1990-2019: A Systematic Analysis for the Global Burden of Disease Study 2019. <italic>The</italic><italic>Lancet</italic>, 396, 1204-1222. https://doi.org/10.1016/s0140-6736(20)30925-9 <pub-id pub-id-type="doi">10.1016/s0140-6736(20)30925-9</pub-id><pub-id pub-id-type="pmid">33069326</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1016/s0140-6736(20)30925-9">https://doi.org/10.1016/s0140-6736(20)30925-9</ext-link></mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Vos, T.</string-name>
              <string-name>Lim, S.S.</string-name>
              <string-name>Abbafati, C.</string-name>
              <string-name>Abbas, K.M.</string-name>
              <string-name>Abbasi, M.</string-name>
              <string-name>Abbasifard, M.</string-name>
            </person-group>
            <year>2020</year>
            <article-title>Global Burden of 369 Diseases and Injuries in 204 Countries and Territories, 1990-2019: A Systematic Analysis for the Global Burden of Disease Study 2019</article-title>
            <source>The Lancet</source>
            <volume>6736</volume>
            <issue>20</issue>
            <pub-id pub-id-type="doi">10.1016/s0140-6736(20)30925-9</pub-id>
            <pub-id pub-id-type="pmid">33069326</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B3">
        <label>3.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Sidor, M.M. and MacQueen, G.M. (2011) Antidepressants for the Acute Treatment of Bipolar Depression: A Systematic Review and Meta-Analysis. <italic>The</italic><italic>Journal</italic><italic>of</italic><italic>Clin</italic><italic>ical</italic><italic>Psychiatry</italic>, 72, 156-167. https://doi.org/10.4088/jcp.09r05385gre <pub-id pub-id-type="doi">10.4088/jcp.09r05385gre</pub-id><pub-id pub-id-type="pmid">21034686</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.4088/jcp.09r05385gre">https://doi.org/10.4088/jcp.09r05385gre</ext-link></mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Sidor, M.M.</string-name>
              <string-name>MacQueen, G.M.</string-name>
            </person-group>
            <year>2011</year>
            <article-title>Antidepressants for the Acute Treatment of Bipolar Depression: A Systematic Review and Meta-Analysis</article-title>
            <source>The Journal of Clinical Psychiatry</source>
            <volume>72</volume>
            <pub-id pub-id-type="doi">10.4088/jcp.09r05385gre</pub-id>
            <pub-id pub-id-type="pmid">21034686</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B4">
        <label>4.</label>
        <citation-alternatives>
          <mixed-citation publication-type="report">Pacchiarotti, I., Bond, D.J., Baldessarini, R.J., Nolen, W.A., Grunze, H., Licht, R.W., <italic>et al</italic>. (2013) The International Society for Bipolar Disorders (ISBD) Task Force Report on Antidepressant Use in Bipolar Disorders. <italic>American</italic><italic>Journal</italic><italic>of</italic><italic>Psychiatry</italic>, 170, 1249-1262. https://doi.org/10.1176/appi.ajp.2013.13020185 <pub-id pub-id-type="doi">10.1176/appi.ajp.2013.13020185</pub-id><pub-id pub-id-type="pmid">24030475</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1176/appi.ajp.2013.13020185">https://doi.org/10.1176/appi.ajp.2013.13020185</ext-link></mixed-citation>
          <element-citation publication-type="report">
            <person-group person-group-type="author">
              <string-name>Pacchiarotti, I.</string-name>
              <string-name>Bond, D.J.</string-name>
              <string-name>Baldessarini, R.J.</string-name>
              <string-name>Nolen, W.A.</string-name>
              <string-name>Grunze, H.</string-name>
              <string-name>Licht, R.W.</string-name>
            </person-group>
            <year>2013</year>
            <article-title>The International Society for Bipolar Disorders (ISBD) Task Force Report on Antidepressant Use in Bipolar Disorders</article-title>
            <source>American Journal of Psychiatry</source>
            <volume>170</volume>
            <pub-id pub-id-type="doi">10.1176/appi.ajp.2013.13020185</pub-id>
            <pub-id pub-id-type="pmid">24030475</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B5">
        <label>5.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Phillips, M.L. and Kupfer, D.J. (2013) Bipolar Disorder Diagnosis: Challenges and Future Directions. <italic>The Lancet</italic>, 381, 1663-1671. https://doi.org/10.1016/s0140-6736(13)60989-7 <pub-id pub-id-type="doi">10.1016/s0140-6736(13)60989-7</pub-id><pub-id pub-id-type="pmid">23663952</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1016/s0140-6736(13)60989-7">https://doi.org/10.1016/s0140-6736(13)60989-7</ext-link></mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Phillips, M.L.</string-name>
              <string-name>Kupfer, D.J.</string-name>
            </person-group>
            <year>2013</year>
            <article-title>Bipolar Disorder Diagnosis: Challenges and Future Directions</article-title>
            <source>The Lancet</source>
            <volume>6736</volume>
            <issue>13</issue>
            <pub-id pub-id-type="doi">10.1016/s0140-6736(13)60989-7</pub-id>
            <pub-id pub-id-type="pmid">23663952</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B6">
        <label>6.</label>
        <citation-alternatives>
          <mixed-citation publication-type="book">Goodwin, G., Haddad, P., Ferrier, I., Aronson, J., Barnes, T., Cipriani, A., <italic>et al</italic>. (2016) Evidence-Based Guidelines for Treating Bipolar Disorder: Revised Third Edition Recommendations from the British Association for Psychopharmacology. <italic>Journal</italic><italic>of</italic><italic>Psychopharmacology</italic>, 30, 495-553. https://doi.org/10.1177/0269881116636545 <pub-id pub-id-type="doi">10.1177/0269881116636545</pub-id><pub-id pub-id-type="pmid">26979387</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1177/0269881116636545">https://doi.org/10.1177/0269881116636545</ext-link></mixed-citation>
          <element-citation publication-type="book">
            <person-group person-group-type="author">
              <string-name>Goodwin, G.</string-name>
              <string-name>Haddad, P.</string-name>
              <string-name>Ferrier, I.</string-name>
              <string-name>Aronson, J.</string-name>
              <string-name>Barnes, T.</string-name>
              <string-name>Cipriani, A.</string-name>
            </person-group>
            <year>2016</year>
            <article-title>Evidence-Based Guidelines for Treating Bipolar Disorder: Revised Third Edition Recommendations from the British Association for Psychopharmacology</article-title>
            <source>Journal of Psychopharmacology</source>
            <volume>30</volume>
            <pub-id pub-id-type="doi">10.1177/0269881116636545</pub-id>
            <pub-id pub-id-type="pmid">26979387</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B7">
        <label>7.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Topol, E.J. (2019) High-Performance Medicine: The Convergence of Human and Artificial Intelligence. <italic>Nature</italic><italic>Medicine</italic>, 25, 44-56. https://doi.org/10.1038/s41591-018-0300-7 <pub-id pub-id-type="doi">10.1038/s41591-018-0300-7</pub-id><pub-id pub-id-type="pmid">30617339</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1038/s41591-018-0300-7">https://doi.org/10.1038/s41591-018-0300-7</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Topol, E.J.</string-name>
            </person-group>
            <year>2019</year>
            <article-title>High-Performance Medicine: The Convergence of Human and Artificial Intelligence</article-title>
            <source>Nature Medicine</source>
            <volume>25</volume>
            <pub-id pub-id-type="doi">10.1038/s41591-018-0300-7</pub-id>
            <pub-id pub-id-type="pmid">30617339</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B8">
        <label>8.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Rajpurkar, P., Chen, E., Banerjee, O. and Topol, E.J. (2022) AI in Health and Medicine. <italic>Nature</italic><italic>Medicine</italic>, 28, 31-38. https://doi.org/10.1038/s41591-021-01614-0 <pub-id pub-id-type="doi">10.1038/s41591-021-01614-0</pub-id><pub-id pub-id-type="pmid">35058619</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1038/s41591-021-01614-0">https://doi.org/10.1038/s41591-021-01614-0</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Rajpurkar, P.</string-name>
              <string-name>Chen, E.</string-name>
              <string-name>Banerjee, O.</string-name>
              <string-name>Topol, E.J.</string-name>
            </person-group>
            <year>2022</year>
            <article-title>AI in Health and Medicine</article-title>
            <source>Nature Medicine</source>
            <volume>28</volume>
            <pub-id pub-id-type="doi">10.1038/s41591-021-01614-0</pub-id>
            <pub-id pub-id-type="pmid">35058619</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B9">
        <label>9.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Chekroud, A.M., Zotti, R.J., Shehzad, Z., Gueorguieva, R., Johnson, M.K., Trivedi, M.H., <italic>et al</italic>. (2016) Cross-Trial Prediction of Treatment Outcome in Depression: A Machine Learning Approach. <italic>The</italic><italic>Lancet</italic><italic>Psychiatry</italic>, 3, 243-250. https://doi.org/10.1016/s2215-0366(15)00471-x <pub-id pub-id-type="doi">10.1016/s2215-0366(15)00471-x</pub-id><pub-id pub-id-type="pmid">26803397</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1016/s2215-0366(15)00471-x">https://doi.org/10.1016/s2215-0366(15)00471-x</ext-link></mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Chekroud, A.M.</string-name>
              <string-name>Zotti, R.J.</string-name>
              <string-name>Shehzad, Z.</string-name>
              <string-name>Gueorguieva, R.</string-name>
              <string-name>Johnson, M.K.</string-name>
              <string-name>Trivedi, M.H.</string-name>
            </person-group>
            <year>2016</year>
            <article-title>Cross-Trial Prediction of Treatment Outcome in Depression: A Machine Learning Approach</article-title>
            <source>The Lancet Psychiatry</source>
            <volume>0366</volume>
            <issue>15</issue>
            <pub-id pub-id-type="doi">10.1016/s2215-0366(15)00471-x</pub-id>
            <pub-id pub-id-type="pmid">26803397</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B10">
        <label>10.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Kessler, R.C., Warner, C.H., Ivany, C., Petukhova, M.V., Rose, S., Bromet, E.J., <italic>et al</italic>. (2015) Predicting Suicides after Psychiatric Hospitalization in US Army Soldiers: The Army Study to Assess Risk and Resilience in Servicemembers (Army STARRS). <italic>JAMA</italic><italic>Psychiatry</italic>, 72, 49-57. https://doi.org/10.1001/jamapsychiatry.2014.1754 <pub-id pub-id-type="doi">10.1001/jamapsychiatry.2014.1754</pub-id><pub-id pub-id-type="pmid">25390793</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1001/jamapsychiatry.2014.1754">https://doi.org/10.1001/jamapsychiatry.2014.1754</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Kessler, R.C.</string-name>
              <string-name>Warner, C.H.</string-name>
              <string-name>Ivany, C.</string-name>
              <string-name>Petukhova, M.V.</string-name>
              <string-name>Rose, S.</string-name>
              <string-name>Bromet, E.J.</string-name>
            </person-group>
            <year>2015</year>
            <article-title>Predicting Suicides after Psychiatric Hospitalization in US Army Soldiers: The Army Study to Assess Risk and Resilience in Servicemembers (Army STARRS)</article-title>
            <source>JAMA Psychiatry</source>
            <volume>72</volume>
            <pub-id pub-id-type="doi">10.1001/jamapsychiatry.2014.1754</pub-id>
            <pub-id pub-id-type="pmid">25390793</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B11">
        <label>11.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Chen, J.H. and Asch, S.M. (2017) Machine Learning and Prediction in Medicine—Beyond the Peak of Inflated Expectations. <italic>New</italic><italic>England</italic><italic>Journal</italic><italic>of</italic><italic>Medicine</italic>, 376, 2507-2509. https://doi.org/10.1056/nejmp1702071 <pub-id pub-id-type="doi">10.1056/nejmp1702071</pub-id><pub-id pub-id-type="pmid">28657867</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1056/nejmp1702071">https://doi.org/10.1056/nejmp1702071</ext-link></mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Chen, J.H.</string-name>
              <string-name>Asch, S.M.</string-name>
            </person-group>
            <year>2017</year>
            <article-title>Machine Learning and Prediction in Medicine—Beyond the Peak of Inflated Expectations</article-title>
            <source>New England Journal of Medicine</source>
            <volume>376</volume>
            <pub-id pub-id-type="doi">10.1056/nejmp1702071</pub-id>
            <pub-id pub-id-type="pmid">28657867</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B12">
        <label>12.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Sendak, M., Gao, M., Nichols, C., <italic>et al</italic>. (2020) “Human-Compatible” Machine Learning as a Step toward Safe Clinical AI. <italic>NPJ Digital Medicine</italic>, 3, Article No. 141.</mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Sendak, M.</string-name>
              <string-name>Gao, M.</string-name>
              <string-name>Nichols, C.</string-name>
            </person-group>
            <year>2020</year>
            <article-title>“Human-Compatible” Machine Learning as a Step toward Safe Clinical AI</article-title>
            <source>NPJ Digital Medicine</source>
            <volume>3</volume>
            <elocation-id>No</elocation-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B13">
        <label>13.</label>
        <citation-alternatives>
          <mixed-citation publication-type="book">Sutton, R.S. and Barto, A.G. (2018) Reinforcement Learning: An Introduction. 2nd Edition, MIT Press.</mixed-citation>
          <element-citation publication-type="book">
            <person-group person-group-type="author">
              <string-name>Sutton, R.S.</string-name>
              <string-name>Barto, A.G.</string-name>
              <string-name>Edition, M</string-name>
            </person-group>
            <year>2018</year>
            <article-title>Reinforcement Learning: An Introduction</article-title>
            <source>2nd Edition</source>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B14">
        <label>14.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Gottesman, O., Johansson, F., Komorowski, M., Faisal, A., Sontag, D., Doshi-Velez, F., <italic>et al</italic>. (2019) Guidelines for Reinforcement Learning in Healthcare. <italic>Nature</italic><italic>Medicine</italic>, 25, 16-18. https://doi.org/10.1038/s41591-018-0310-5 <pub-id pub-id-type="doi">10.1038/s41591-018-0310-5</pub-id><pub-id pub-id-type="pmid">30617332</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1038/s41591-018-0310-5">https://doi.org/10.1038/s41591-018-0310-5</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Gottesman, O.</string-name>
              <string-name>Johansson, F.</string-name>
              <string-name>Komorowski, M.</string-name>
              <string-name>Faisal, A.</string-name>
              <string-name>Sontag, D.</string-name>
              <string-name>Doshi-Velez, F.</string-name>
            </person-group>
            <year>2019</year>
            <article-title>Guidelines for Reinforcement Learning in Healthcare</article-title>
            <source>Nature Medicine</source>
            <volume>25</volume>
            <pub-id pub-id-type="doi">10.1038/s41591-018-0310-5</pub-id>
            <pub-id pub-id-type="pmid">30617332</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B15">
        <label>15.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Yu, C., Liu, J., Nemati, S. and Yin, G. (2021) Reinforcement Learning in Healthcare: A Survey. <italic>ACM Computing Surveys</italic>, 55, 1-36.</mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Yu, C.</string-name>
              <string-name>Liu, J.</string-name>
              <string-name>Nemati, S.</string-name>
              <string-name>Yin, G.</string-name>
            </person-group>
            <year>2021</year>
            <article-title>Reinforcement Learning in Healthcare: A Survey</article-title>
            <source>ACM Computing Surveys</source>
            <volume>55</volume>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B16">
        <label>16.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Silver, D., Huang, A., Maddison, C.J., Guez, A., Sifre, L., van den Driessche, G., <italic>et al</italic>. (2016) Mastering the Game of Go with Deep Neural Networks and Tree Search. <italic>Nature</italic>, 529, 484-489. https://doi.org/10.1038/nature16961 <pub-id pub-id-type="doi">10.1038/nature16961</pub-id><pub-id pub-id-type="pmid">26819042</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1038/nature16961">https://doi.org/10.1038/nature16961</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Silver, D.</string-name>
              <string-name>Huang, A.</string-name>
              <string-name>Maddison, C.J.</string-name>
              <string-name>Guez, A.</string-name>
              <string-name>Sifre, L.</string-name>
              <string-name>Driessche, G.</string-name>
            </person-group>
            <year>2016</year>
            <article-title>Mastering the Game of Go with Deep Neural Networks and Tree Search</article-title>
            <source>Nature</source>
            <volume>529</volume>
            <pub-id pub-id-type="doi">10.1038/nature16961</pub-id>
            <pub-id pub-id-type="pmid">26819042</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B17">
        <label>17.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Vinyals, O., Babuschkin, I., Czarnecki, W.M., Mathieu, M., Dudzik, A., Chung, J., <italic>et al</italic>. (2019) Grandmaster Level in StarCraft II Using Multi-Agent Reinforcement Learning. <italic>Nature</italic>, 575, 350-354. https://doi.org/10.1038/s41586-019-1724-z <pub-id pub-id-type="doi">10.1038/s41586-019-1724-z</pub-id><pub-id pub-id-type="pmid">31666705</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1038/s41586-019-1724-z">https://doi.org/10.1038/s41586-019-1724-z</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Vinyals, O.</string-name>
              <string-name>Babuschkin, I.</string-name>
              <string-name>Czarnecki, W.M.</string-name>
              <string-name>Mathieu, M.</string-name>
              <string-name>Dudzik, A.</string-name>
              <string-name>Chung, J.</string-name>
            </person-group>
            <year>2019</year>
            <article-title>Grandmaster Level in StarCraft II Using Multi-Agent Reinforcement Learning</article-title>
            <source>Nature</source>
            <volume>575</volume>
            <pub-id pub-id-type="doi">10.1038/s41586-019-1724-z</pub-id>
            <pub-id pub-id-type="pmid">31666705</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B18">
        <label>18.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Komorowski, M., Celi, L.A., Badawi, O., Gordon, A.C. and Faisal, A.A. (2018) The Artificial Intelligence Clinician Learns Optimal Treatment Strategies for Sepsis in Intensive Care. <italic>Nature</italic><italic>Medicine</italic>, 24, 1716-1720. https://doi.org/10.1038/s41591-018-0213-5 <pub-id pub-id-type="doi">10.1038/s41591-018-0213-5</pub-id><pub-id pub-id-type="pmid">30349085</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1038/s41591-018-0213-5">https://doi.org/10.1038/s41591-018-0213-5</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Komorowski, M.</string-name>
              <string-name>Celi, L.A.</string-name>
              <string-name>Badawi, O.</string-name>
              <string-name>Gordon, A.C.</string-name>
              <string-name>Faisal, A.A.</string-name>
            </person-group>
            <year>2018</year>
            <article-title>The Artificial Intelligence Clinician Learns Optimal Treatment Strategies for Sepsis in Intensive Care</article-title>
            <source>Nature Medicine</source>
            <volume>24</volume>
            <pub-id pub-id-type="doi">10.1038/s41591-018-0213-5</pub-id>
            <pub-id pub-id-type="pmid">30349085</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B19">
        <label>19.</label>
        <citation-alternatives>
          <mixed-citation publication-type="confproc">Peng, X., Ding, Y., Wirsching, W., <italic>et al</italic>. (2018) Improving Sepsis Treatment Strategies by Combining Deep and Kernel-Based Reinforcement Learning. <italic>AMIA Annual Symposium Proceedings</italic>, San Francisco, 3-7 November 2018, 887-896.</mixed-citation>
          <element-citation publication-type="confproc">
            <person-group person-group-type="author">
              <string-name>Peng, X.</string-name>
              <string-name>Ding, Y.</string-name>
              <string-name>Wirsching, W.</string-name>
              <string-name>Proceedings, S</string-name>
            </person-group>
            <year>2018</year>
            <article-title>Improving Sepsis Treatment Strategies by Combining Deep and Kernel-Based Reinforcement Learning</article-title>
            <source>AMIA Annual Symposium Proceedings</source>
            <volume>3</volume>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B20">
        <label>20.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Zhao, R., Pacella, M., Sanmugarajah, J., <italic>et al</italic>. (2022) Deep Reinforcement Learning for Treatment Duration Decision Making in Acute Lymphoblastic Leukemia. <italic>IEEE Journal of Biomedical and Health Informatics</italic>, 26, 4623-4634.</mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Zhao, R.</string-name>
              <string-name>Pacella, M.</string-name>
              <string-name>Sanmugarajah, J.</string-name>
            </person-group>
            <year>2022</year>
            <article-title>Deep Reinforcement Learning for Treatment Duration Decision Making in Acute Lymphoblastic Leukemia</article-title>
            <source>IEEE Journal of Biomedical and Health Informatics</source>
            <volume>26</volume>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B21">
        <label>21.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Colombo, F., Calesella, F., Mazza, M.G., Melloni, E.M.T., Morelli, M.J., Scotti, G.M., <italic>et al</italic>. (2022) Machine Learning Approaches for Prediction of Bipolar Disorder Based on Biological, Clinical and Neuropsychological Markers: A Systematic Review and Meta-Analysis. <italic>Neuroscience &amp; Biobehavioral Reviews</italic>, 135, Article 104552. https://doi.org/10.1016/j.neubiorev.2022.104552 <pub-id pub-id-type="doi">10.1016/j.neubiorev.2022.104552</pub-id><pub-id pub-id-type="pmid">35120970</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1016/j.neubiorev.2022.104552">https://doi.org/10.1016/j.neubiorev.2022.104552</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Colombo, F.</string-name>
              <string-name>Calesella, F.</string-name>
              <string-name>Mazza, M.G.</string-name>
              <string-name>Melloni, E.M.T.</string-name>
              <string-name>Morelli, M.J.</string-name>
              <string-name>Scotti, G.M.</string-name>
              <string-name>Biological, C</string-name>
            </person-group>
            <year>2022</year>
            <article-title>Machine Learning Approaches for Prediction of Bipolar Disorder Based on Biological, Clinical and Neuropsychological Markers: A Systematic Review and Meta-Analysis</article-title>
            <source>Neuroscience &amp; Biobehavioral Reviews</source>
            <volume>135</volume>
            <elocation-id>104552</elocation-id>
            <pub-id pub-id-type="doi">10.1016/j.neubiorev.2022.104552</pub-id>
            <pub-id pub-id-type="pmid">35120970</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B22">
        <label>22.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">He, M., Bakker, E.M. and Lew, M.S. (2024) DPD (Depression Detection) Net: A Deep Neural Network for Multimodal Depression Detection. <italic>Health Information Science and Systems</italic>, 12, Article No. 53. https://doi.org/10.1007/s13755-024-00311-9 <pub-id pub-id-type="doi">10.1007/s13755-024-00311-9</pub-id><pub-id pub-id-type="pmid">39544256</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1007/s13755-024-00311-9">https://doi.org/10.1007/s13755-024-00311-9</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>He, M.</string-name>
              <string-name>Bakker, E.M.</string-name>
              <string-name>Lew, M.S.</string-name>
            </person-group>
            <year>2024</year>
            <article-title>DPD (Depression Detection) Net: A Deep Neural Network for Multimodal Depression Detection</article-title>
            <source>Health Information Science and Systems</source>
            <volume>12</volume>
            <elocation-id>No</elocation-id>
            <pub-id pub-id-type="doi">10.1007/s13755-024-00311-9</pub-id>
            <pub-id pub-id-type="pmid">39544256</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B23">
        <label>23.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">First, M.B., Williams, J.B.W., Karg, R.S. and Spitzer, R.L. (2015) Structured Clinical Interview for DSM-5 Research Version (SCID-5 for DSM-5, Research Version; SCID-5-RV). American Psychiatric Association.</mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>First, M.B.</string-name>
              <string-name>Williams, J.B.W.</string-name>
              <string-name>Karg, R.S.</string-name>
              <string-name>Spitzer, R.L.</string-name>
            </person-group>
            <year>2015</year>
            <article-title>Structured Clinical Interview for DSM-5 Research Version (SCID-5 for DSM-5, Research Version; SCID-5-RV)</article-title>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B24">
        <label>24.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Montgomery, S.A. and Åsberg, M. (1979) A New Depression Scale Designed to Be Sensitive to Change. <italic>British Journal of Psychiatry</italic>, 134, 382-389. https://doi.org/10.1192/bjp.134.4.382 <pub-id pub-id-type="doi">10.1192/bjp.134.4.382</pub-id><pub-id pub-id-type="pmid">444788</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1192/bjp.134.4.382">https://doi.org/10.1192/bjp.134.4.382</ext-link></mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Montgomery, S.A.</string-name>
            </person-group>
            <year>1979</year>
            <article-title>A New Depression Scale Designed to Be Sensitive to Change</article-title>
            <source>British Journal of Psychiatry</source>
            <volume>134</volume>
            <pub-id pub-id-type="doi">10.1192/bjp.134.4.382</pub-id>
            <pub-id pub-id-type="pmid">444788</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B25">
        <label>25.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Schulman, J., Wolski, F., Dhariwal, P., Radford, A. and Klimov, O. (2017) Proximal Policy Optimization Algorithms.</mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Schulman, J.</string-name>
              <string-name>Wolski, F.</string-name>
              <string-name>Dhariwal, P.</string-name>
              <string-name>Radford, A.</string-name>
              <string-name>Klimov, O.</string-name>
            </person-group>
            <year>2017</year>
            <article-title>Proximal Policy Optimization Algorithms</article-title>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B26">
        <label>26.</label>
        <citation-alternatives>
          <mixed-citation publication-type="confproc">Thomas, P., Theocharous, G. and Ghavamzadeh, M. (2015) High-Confidence Off-Policy Evaluation. <italic>Proceedings of the AAAI Conference on Artificial Intelligence</italic>, 29, 3000-3006. https://doi.org/10.1609/aaai.v29i1.9541 <pub-id pub-id-type="doi">10.1609/aaai.v29i1.9541</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1609/aaai.v29i1.9541">https://doi.org/10.1609/aaai.v29i1.9541</ext-link></mixed-citation>
          <element-citation publication-type="confproc">
            <person-group person-group-type="author">
              <string-name>Thomas, P.</string-name>
              <string-name>Theocharous, G.</string-name>
              <string-name>Ghavamzadeh, M.</string-name>
            </person-group>
            <year>2015</year>
            <article-title>High-Confidence Off-Policy Evaluation</article-title>
            <source>Proceedings of the AAAI Conference on Artificial Intelligence</source>
            <volume>29</volume>
            <pub-id pub-id-type="doi">10.1609/aaai.v29i1.9541</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B27">
        <label>27.</label>
        <citation-alternatives>
          <mixed-citation publication-type="confproc">Schulman, J., Levine, S., Abbeel, P., Jordan, M. and Moritz, P. (2015) Trust Region Policy Optimization. <italic>International Conference on Machine Learning</italic>, Lille, 7-9 July 2015, 1889-1897.</mixed-citation>
          <element-citation publication-type="confproc">
            <person-group person-group-type="author">
              <string-name>Schulman, J.</string-name>
              <string-name>Levine, S.</string-name>
              <string-name>Abbeel, P.</string-name>
              <string-name>Jordan, M.</string-name>
              <string-name>Moritz, P.</string-name>
              <string-name>Learning, L</string-name>
            </person-group>
            <year>2015</year>
            <article-title>Trust Region Policy Optimization</article-title>
            <source>International Conference on Machine Learning</source>
            <volume>7</volume>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B28">
        <label>28.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Chernozhukov, V., Chetverikov, D., Demirer, M., Duflo, E., Hansen, C., Newey, W., <italic>et al</italic>. (2018) Double/Debiased Machine Learning for Treatment and Structural Parameters. <italic>The Econometrics Journal</italic>, 21, C1-C68. https://doi.org/10.1111/ectj.12097 <pub-id pub-id-type="doi">10.1111/ectj.12097</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1111/ectj.12097">https://doi.org/10.1111/ectj.12097</ext-link></mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Chernozhukov, V.</string-name>
              <string-name>Chetverikov, D.</string-name>
              <string-name>Demirer, M.</string-name>
              <string-name>Duflo, E.</string-name>
              <string-name>Hansen, C.</string-name>
              <string-name>Newey, W.</string-name>
            </person-group>
            <year>2018</year>
            <article-title>Double/Debiased Machine Learning for Treatment and Structural Parameters</article-title>
            <source>The Econometrics Journal</source>
            <volume>21</volume>
            <pub-id pub-id-type="doi">10.1111/ectj.12097</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B29">
        <label>29.</label>
        <citation-alternatives>
          <mixed-citation publication-type="confproc">Lundberg, S.M. and Lee, S.I. (2017) A Unified Approach to Interpreting Model Predictions. <italic>Advances in Neural Information Processing Systems</italic>30: <italic>Annual Conference on Neural Information Processing Systems</italic> 2017, Long Beach, 4-9 December 2017, 4765-4774.</mixed-citation>
          <element-citation publication-type="confproc">
            <person-group person-group-type="author">
              <string-name>Lundberg, S.M.</string-name>
              <string-name>Lee, S.I.</string-name>
            </person-group>
            <year>2017</year>
            <article-title>A Unified Approach to Interpreting Model Predictions</article-title>
            <source>Advances in Neural Information Processing Systems 30: Annual Conference on Neural Information Processing Systems 2017</source>
            <volume>4</volume>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B30">
        <label>30.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Vickers, A.J. and Elkin, E.B. (2006) Decision Curve Analysis: A Novel Method for Evaluating Prediction Models. <italic>Medical Decision Making</italic>, 26, 565-574. https://doi.org/10.1177/0272989x06295361 <pub-id pub-id-type="doi">10.1177/0272989x06295361</pub-id><pub-id pub-id-type="pmid">17099194</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1177/0272989x06295361">https://doi.org/10.1177/0272989x06295361</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Vickers, A.J.</string-name>
              <string-name>Elkin, E.B.</string-name>
            </person-group>
            <year>2006</year>
            <article-title>Decision Curve Analysis: A Novel Method for Evaluating Prediction Models</article-title>
            <source>Medical Decision Making</source>
            <volume>26</volume>
            <pub-id pub-id-type="doi">10.1177/0272989x06295361</pub-id>
            <pub-id pub-id-type="pmid">17099194</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B31">
        <label>31.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Cuijpers, P., Turner, E.H., Mohr, D.C., Hofmann, S.G., Andersson, G., Berking, M., <italic>et al</italic>. (2014) Comparison of Psychotherapies for Adult Depression to Pill Placebo Control Groups: A Meta-Analysis. <italic>Psychological Medicine</italic>, 44, 685-695. https://doi.org/10.1017/s0033291713000457 <pub-id pub-id-type="doi">10.1017/s0033291713000457</pub-id><pub-id pub-id-type="pmid">23552610</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1017/s0033291713000457">https://doi.org/10.1017/s0033291713000457</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Cuijpers, P.</string-name>
              <string-name>Turner, E.H.</string-name>
              <string-name>Mohr, D.C.</string-name>
              <string-name>Hofmann, S.G.</string-name>
              <string-name>Andersson, G.</string-name>
              <string-name>Berking, M.</string-name>
            </person-group>
            <year>2014</year>
            <article-title>Comparison of Psychotherapies for Adult Depression to Pill Placebo Control Groups: A Meta-Analysis</article-title>
            <source>Psychological Medicine</source>
            <volume>44</volume>
            <pub-id pub-id-type="doi">10.1017/s0033291713000457</pub-id>
            <pub-id pub-id-type="pmid">23552610</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B32">
        <label>32.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Geddes, J.R., Gardiner, A., Rendell, J., Voysey, M., Tunbridge, E., Hinds, A., <italic>et al</italic>. (2022) Comparative Evaluation of Quetiapine plus Lamotrigine Combination versus Quetiapine Monotherapy in Bipolar Depression: A Randomized, Double-Blind, Placebo-Controlled Trial. <italic>The</italic><italic>Lancet Psychiatry</italic>, 9, 883-894.</mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Geddes, J.R.</string-name>
              <string-name>Gardiner, A.</string-name>
              <string-name>Rendell, J.</string-name>
              <string-name>Voysey, M.</string-name>
              <string-name>Tunbridge, E.</string-name>
              <string-name>Hinds, A.</string-name>
              <string-name>Randomized, D</string-name>
              <string-name>Blind, P</string-name>
            </person-group>
            <year>2022</year>
            <article-title>Comparative Evaluation of Quetiapine plus Lamotrigine Combination versus Quetiapine Monotherapy in Bipolar Depression: A Randomized, Double-Blind, Placebo-Controlled Trial</article-title>
            <source>The Lancet Psychiatry</source>
            <volume>9</volume>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B33">
        <label>33.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">McIntyre, R.S., Berk, M., Brietzke, E., Goldstein, B.I., López-Jaramillo, C., Kessing, L.V., <italic>et al</italic>. (2020) Bipolar Disorders. <italic>The Lancet</italic>, 396, 1841-1856. https://doi.org/10.1016/s0140-6736(20)31544-0 <pub-id pub-id-type="doi">10.1016/s0140-6736(20)31544-0</pub-id><pub-id pub-id-type="pmid">33278937</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1016/s0140-6736(20)31544-0">https://doi.org/10.1016/s0140-6736(20)31544-0</ext-link></mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>McIntyre, R.S.</string-name>
              <string-name>Berk, M.</string-name>
              <string-name>Brietzke, E.</string-name>
              <string-name>Goldstein, B.I.</string-name>
              <string-name>Jaramillo, C.</string-name>
              <string-name>Kessing, L.V.</string-name>
            </person-group>
            <year>2020</year>
            <article-title>Bipolar Disorders</article-title>
            <source>The Lancet</source>
            <volume>6736</volume>
            <issue>20</issue>
            <pub-id pub-id-type="doi">10.1016/s0140-6736(20)31544-0</pub-id>
            <pub-id pub-id-type="pmid">33278937</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B34">
        <label>34.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Wray, N.R., Ripke, S., Mattheisen, M., Trzaskowski, M., Byrne, E.M., Abdellaoui, A., <italic>et al</italic>. (2018) Genome-Wide Association Analyses Identify 44 Risk Variants and Refine the Genetic Architecture of Major Depression. <italic>Nature Genetics</italic>, 50, 668-681. https://doi.org/10.1038/s41588-018-0090-3 <pub-id pub-id-type="doi">10.1038/s41588-018-0090-3</pub-id><pub-id pub-id-type="pmid">29700475</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1038/s41588-018-0090-3">https://doi.org/10.1038/s41588-018-0090-3</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Wray, N.R.</string-name>
              <string-name>Ripke, S.</string-name>
              <string-name>Mattheisen, M.</string-name>
              <string-name>Trzaskowski, M.</string-name>
              <string-name>Byrne, E.M.</string-name>
              <string-name>Abdellaoui, A.</string-name>
            </person-group>
            <year>2018</year>
            <article-title>Genome-Wide Association Analyses Identify 44 Risk Variants and Refine the Genetic Architecture of Major Depression</article-title>
            <source>Nature Genetics</source>
            <volume>50</volume>
            <pub-id pub-id-type="doi">10.1038/s41588-018-0090-3</pub-id>
            <pub-id pub-id-type="pmid">29700475</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B35">
        <label>35.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Zeng, J., Zhang, Y., Xiang, Y., Liang, S., Xue, C., Zhang, J., <italic>et al</italic>. (2023) Optimizing Multi-Domain Hematologic Biomarkers and Clinical Features for the Differential Diagnosis of Unipolar Depression and Bipolar Depression. <italic>NPJ Mental Health Research</italic>, 2, Article No. 4. https://doi.org/10.1038/s44184-023-00024-z <pub-id pub-id-type="doi">10.1038/s44184-023-00024-z</pub-id><pub-id pub-id-type="pmid">38609642</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1038/s44184-023-00024-z">https://doi.org/10.1038/s44184-023-00024-z</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Zeng, J.</string-name>
              <string-name>Zhang, Y.</string-name>
              <string-name>Xiang, Y.</string-name>
              <string-name>Liang, S.</string-name>
              <string-name>Xue, C.</string-name>
              <string-name>Zhang, J.</string-name>
            </person-group>
            <year>2023</year>
            <article-title>Optimizing Multi-Domain Hematologic Biomarkers and Clinical Features for the Differential Diagnosis of Unipolar Depression and Bipolar Depression</article-title>
            <source>NPJ Mental Health Research</source>
            <volume>2</volume>
            <elocation-id>No</elocation-id>
            <pub-id pub-id-type="doi">10.1038/s44184-023-00024-z</pub-id>
            <pub-id pub-id-type="pmid">38609642</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B36">
        <label>36.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Kanchapogu, N.R. and Mohanty, S.N. (2025) Deep Learning with Ensemble-Based Hybrid AI Model for Bipolar and Unipolar Depression Detection Using Demographic and Behavioral Based on Time-Series Data. <italic>Dialogues</italic><italic>in</italic><italic>Clinical</italic><italic>Neuroscience</italic>, 27, 16-35. https://doi.org/10.1080/19585969.2025.2524337 <pub-id pub-id-type="doi">10.1080/19585969.2025.2524337</pub-id><pub-id pub-id-type="pmid">40588165</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1080/19585969.2025.2524337">https://doi.org/10.1080/19585969.2025.2524337</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Kanchapogu, N.R.</string-name>
              <string-name>Mohanty, S.N.</string-name>
            </person-group>
            <year>2025</year>
            <article-title>Deep Learning with Ensemble-Based Hybrid AI Model for Bipolar and Unipolar Depression Detection Using Demographic and Behavioral Based on Time-Series Data</article-title>
            <source>Dialogues in Clinical Neuroscience</source>
            <volume>27</volume>
            <pub-id pub-id-type="doi">10.1080/19585969.2025.2524337</pub-id>
            <pub-id pub-id-type="pmid">40588165</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B37">
        <label>37.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Kessler, R.C., Bossarte, R.M., Luedtke, A., <italic>et al</italic>. (2023) Evaluation of a Machine Learning-Based Prediction Model for Benefit and Harm from Antidepressant Treatment in the EM-BARC Randomized Clinical Trial. <italic>JAMA Network Open</italic>, 6, e2327755.</mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Kessler, R.C.</string-name>
              <string-name>Bossarte, R.M.</string-name>
              <string-name>Luedtke, A.</string-name>
            </person-group>
            <year>2023</year>
            <article-title>Evaluation of a Machine Learning-Based Prediction Model for Benefit and Harm from Antidepressant Treatment in the EM-BARC Randomized Clinical Trial</article-title>
            <source>JAMA Network Open</source>
            <volume>6</volume>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B38">
        <label>38.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Iniesta, R., Hodgson, K., Stahl, D., Malki, K., Maier, W., Rietschel, M., <italic>et al</italic>. (2018) Antidepressant Drug-Specific Prediction of Depression Treatment Outcomes from Genetic and Clinical Variables. <italic>Scientific</italic><italic>Reports</italic>, 8, Article No. 5380. https://doi.org/10.1038/s41598-018-23584-z <pub-id pub-id-type="doi">10.1038/s41598-018-23584-z</pub-id><pub-id pub-id-type="pmid">29615645</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1038/s41598-018-23584-z">https://doi.org/10.1038/s41598-018-23584-z</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Iniesta, R.</string-name>
              <string-name>Hodgson, K.</string-name>
              <string-name>Stahl, D.</string-name>
              <string-name>Malki, K.</string-name>
              <string-name>Maier, W.</string-name>
              <string-name>Rietschel, M.</string-name>
            </person-group>
            <year>2018</year>
            <article-title>Antidepressant Drug-Specific Prediction of Depression Treatment Outcomes from Genetic and Clinical Variables</article-title>
            <source>Scientific Reports</source>
            <volume>8</volume>
            <elocation-id>No</elocation-id>
            <pub-id pub-id-type="doi">10.1038/s41598-018-23584-z</pub-id>
            <pub-id pub-id-type="pmid">29615645</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B39">
        <label>39.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Henry, K.E., Hager, D.N., Pronovost, P.J. and Saria, S. (2015) A Targeted Real-Time Early Warning Score (TREWScore) for Septic Shock. <italic>Science</italic><italic>Translational</italic><italic>Medicine</italic>, 7, 299ra122. https://doi.org/10.1126/scitranslmed.aab3719 <pub-id pub-id-type="doi">10.1126/scitranslmed.aab3719</pub-id><pub-id pub-id-type="pmid">26246167</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1126/scitranslmed.aab3719">https://doi.org/10.1126/scitranslmed.aab3719</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Henry, K.E.</string-name>
              <string-name>Hager, D.N.</string-name>
              <string-name>Pronovost, P.J.</string-name>
              <string-name>Saria, S.</string-name>
            </person-group>
            <year>2015</year>
            <article-title>A Targeted Real-Time Early Warning Score (TREWScore) for Septic Shock</article-title>
            <source>Science Translational Medicine</source>
            <volume>7</volume>
            <pub-id pub-id-type="doi">10.1126/scitranslmed.aab3719</pub-id>
            <pub-id pub-id-type="pmid">26246167</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B40">
        <label>40.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Tomašev, N., Glorot, X., Rae, J.W., Zielinski, M., Askham, H., Saraiva, A., <italic>et al</italic>. (2019) A Clinically Applicable Approach to Continuous Prediction of Future Acute Kidney Injury. <italic>Nature</italic>, 572, 116-119. https://doi.org/10.1038/s41586-019-1390-1 <pub-id pub-id-type="doi">10.1038/s41586-019-1390-1</pub-id><pub-id pub-id-type="pmid">31367026</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1038/s41586-019-1390-1">https://doi.org/10.1038/s41586-019-1390-1</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Glorot, X.</string-name>
              <string-name>Rae, J.W.</string-name>
              <string-name>Zielinski, M.</string-name>
              <string-name>Askham, H.</string-name>
              <string-name>Saraiva, A.</string-name>
            </person-group>
            <year>2019</year>
            <article-title>A Clinically Applicable Approach to Continuous Prediction of Future Acute Kidney Injury</article-title>
            <source>Nature</source>
            <volume>572</volume>
            <pub-id pub-id-type="doi">10.1038/s41586-019-1390-1</pub-id>
            <pub-id pub-id-type="pmid">31367026</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B41">
        <label>41.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Amodei, D., Olah, C., Steinhardt, J., Christiano, P., Schulman, J. and Mané, D. (2016) Concrete Problems in AI Safety. https://doi.org/10.48550/arXiv.1606.06565 <pub-id pub-id-type="doi">10.48550/arXiv.1606.06565</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.48550/arXiv.1606.06565">https://doi.org/10.48550/arXiv.1606.06565</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Amodei, D.</string-name>
              <string-name>Olah, C.</string-name>
              <string-name>Steinhardt, J.</string-name>
              <string-name>Christiano, P.</string-name>
              <string-name>Schulman, J.</string-name>
            </person-group>
            <year>2016</year>
            <article-title>Concrete Problems in AI Safety</article-title>
            <pub-id pub-id-type="doi">10.48550/arXiv.1606.06565</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B42">
        <label>42.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Knevel, R. and Liao, K.P. (2023) From Real-World Electronic Health Record Data to Real-World Results Using Artificial Intelligence. <italic>Annals of the Rheumatic Diseases</italic>, 82, 306-311. https://doi.org/10.1136/ard-2022-222626 <pub-id pub-id-type="doi">10.1136/ard-2022-222626</pub-id><pub-id pub-id-type="pmid">36150748</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1136/ard-2022-222626">https://doi.org/10.1136/ard-2022-222626</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Knevel, R.</string-name>
              <string-name>Liao, K.P.</string-name>
            </person-group>
            <year>2023</year>
            <article-title>From Real-World Electronic Health Record Data to Real-World Results Using Artificial Intelligence</article-title>
            <source>Annals of the Rheumatic Diseases</source>
            <volume>82</volume>
            <pub-id pub-id-type="doi">10.1136/ard-2022-222626</pub-id>
            <pub-id pub-id-type="pmid">36150748</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B43">
        <label>43.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Kosorok, M.R. and Laber, E.B. (2019) Precision Medicine. <italic>Annual Review of Statistics and Its Application</italic>, 6, 263-286. https://doi.org/10.1146/annurev-statistics-030718-105251 <pub-id pub-id-type="doi">10.1146/annurev-statistics-030718-105251</pub-id><pub-id pub-id-type="pmid">31073534</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1146/annurev-statistics-030718-105251">https://doi.org/10.1146/annurev-statistics-030718-105251</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Kosorok, M.R.</string-name>
              <string-name>Laber, E.B.</string-name>
            </person-group>
            <year>2019</year>
            <article-title>Precision Medicine</article-title>
            <source>Annual Review of Statistics and Its Application</source>
            <volume>6</volume>
            <pub-id pub-id-type="doi">10.1146/annurev-statistics-030718-105251</pub-id>
            <pub-id pub-id-type="pmid">31073534</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B44">
        <label>44.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Rieke, N., Hancox, J., Li, W., Milletarì, F., Roth, H.R., Albarqouni, S., <italic>et al</italic>. (2020) The Future of Digital Health with Federated Learning. <italic>NPJ</italic><italic>Digital</italic><italic>Medicine</italic>, 3, Article No. 119. https://doi.org/10.1038/s41746-020-00323-1 <pub-id pub-id-type="doi">10.1038/s41746-020-00323-1</pub-id><pub-id pub-id-type="pmid">33015372</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1038/s41746-020-00323-1">https://doi.org/10.1038/s41746-020-00323-1</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Rieke, N.</string-name>
              <string-name>Hancox, J.</string-name>
              <string-name>Li, W.</string-name>
              <string-name>Roth, H.R.</string-name>
              <string-name>Albarqouni, S.</string-name>
            </person-group>
            <year>2020</year>
            <article-title>The Future of Digital Health with Federated Learning</article-title>
            <source>NPJ Digital Medicine</source>
            <volume>3</volume>
            <elocation-id>No</elocation-id>
            <pub-id pub-id-type="doi">10.1038/s41746-020-00323-1</pub-id>
            <pub-id pub-id-type="pmid">33015372</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B45">
        <label>45.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Ghassemi, M., Oakden-Rayner, L. and Beam, A.L. (2021) The False Hope of Current Approaches to Explainable Artificial Intelligence in Health Care. <italic>The</italic><italic>Lancet</italic><italic>Digital</italic><italic>Health</italic>, 3, e745-e750. https://doi.org/10.1016/s2589-7500(21)00208-9 <pub-id pub-id-type="doi">10.1016/s2589-7500(21)00208-9</pub-id><pub-id pub-id-type="pmid">34711379</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1016/s2589-7500(21)00208-9">https://doi.org/10.1016/s2589-7500(21)00208-9</ext-link></mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Ghassemi, M.</string-name>
              <string-name>Oakden-Rayner, L.</string-name>
              <string-name>Beam, A.L.</string-name>
            </person-group>
            <year>2021</year>
            <article-title>The False Hope of Current Approaches to Explainable Artificial Intelligence in Health Care</article-title>
            <source>The Lancet Digital Health</source>
            <volume>7500</volume>
            <issue>21</issue>
            <pub-id pub-id-type="doi">10.1016/s2589-7500(21)00208-9</pub-id>
            <pub-id pub-id-type="pmid">34711379</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
    </ref-list>
  </back>
</article>