<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.4 20241031//EN" "JATS-journalpublishing1-4.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article" dtd-version="1.4" xml:lang="en">
  <front>
    <journal-meta>
      <journal-id journal-id-type="publisher-id">Oalib</journal-id>
      <journal-title-group>
        <journal-title>Open Access Library Journal</journal-title>
      </journal-title-group>
      <issn pub-type="epub">2333-9721</issn>
      <issn pub-type="ppub">2333-9705</issn>
      <publisher>
        <publisher-name>Scientific Research Publishing</publisher-name>
      </publisher>
    </journal-meta>
    <article-meta>
      <article-id pub-id-type="doi">10.4236/oalib.1115677</article-id>
      <article-id pub-id-type="publisher-id">Oalib-153493</article-id>
      <article-categories>
        <subj-group>
          <subject>Article</subject>
        </subj-group>
        <subj-group>
          <subject>Biomedical</subject>
          <subject>Life Sciences</subject>
          <subject>Business</subject>
          <subject>Economics</subject>
          <subject>Chemistry</subject>
          <subject>Materials Science</subject>
          <subject>Computer Science</subject>
          <subject>Communications</subject>
          <subject>Earth</subject>
          <subject>Environmental Sciences</subject>
          <subject>Engineering</subject>
          <subject>Medicine</subject>
          <subject>Healthcare</subject>
          <subject>Physics</subject>
          <subject>Mathematics</subject>
          <subject>Social Sciences</subject>
          <subject>Humanities</subject>
        </subj-group>
      </article-categories>
      <title-group>
        <article-title>Reinforcement Learning for Personalised Antidepressant Sequencing in Treatment-Resistant Bipolar Depression: A Simulation-Based Policy Optimisation Study</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author" corresp="yes">
          <contrib-id contrib-id-type="orcid">0000-0001-9101-072X</contrib-id>
          <name name-style="western">
            <surname>Filippis</surname>
            <given-names>Rocco de</given-names>
          </name>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <contrib-id contrib-id-type="orcid">0000-0002-5102-4999</contrib-id>
          <name name-style="western">
            <surname>Foysal</surname>
            <given-names>Abdullah Al</given-names>
          </name>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
      </contrib-group>
      <aff id="aff1"><label>1</label> Department of Psychiatry, Istituto di Psicopatologia, Rome, Italy </aff>
      <aff id="aff2"><label>2</label> EPSM74, La Roche-sur-Foron, France </aff>
      <aff id="aff3"><label>3</label> Department of Informatics, Bioengineering, Robotics and Systems Engineering, University of Genoa, Genoa, Italy </aff>
      <author-notes>
        <fn fn-type="conflict" id="fn-conflict">
          <p>The authors declare no conflicts of interest.</p>
        </fn>
      </author-notes>
      <pub-date pub-type="epub">
        <day>03</day>
        <month>08</month>
        <year>2026</year>
      </pub-date>
      <pub-date pub-type="collection">
        <month>08</month>
        <year>2026</year>
      </pub-date>
      <volume>13</volume>
      <issue>08</issue>
      <fpage>1</fpage>
      <lpage>15</lpage>
      <history>
        <date date-type="received">
          <day>22</day>
          <month>06</month>
          <year>2026</year>
        </date>
        <date date-type="accepted">
          <day>24</day>
          <month>08</month>
          <year>2026</year>
        </date>
        <date date-type="published">
          <day>27</day>
          <month>08</month>
          <year>2026</year>
        </date>
      </history>
      <permissions>
        <copyright-statement>© 2026 by the authors and Scientific Research Publishing Inc.</copyright-statement>
        <copyright-year>2026</copyright-year>
        <license license-type="open-access">
          <license-p> This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license ( <ext-link ext-link-type="uri" xlink:href="https://creativecommons.org/licenses/by/4.0/">https://creativecommons.org/licenses/by/4.0/</ext-link> ). </license-p>
        </license>
      </permissions>
      <self-uri content-type="doi" xlink:href="https://doi.org/10.4236/oalib.1115677">https://doi.org/10.4236/oalib.1115677</self-uri>
      <abstract>
        <p>Treatment-resistant bipolar depression (TRD-BD), defined in this simulation as failure to achieve remission after at least two adequate pharmacological trials for the current bipolar depressive episode, including at least one antidepressant trial delivered with mood-stabilising treatment, is managed through a sequence of uncertain decisions: continue, switch antidepressant class, or augment with a mood stabiliser or atypical antipsychotic. Current practice relies on consensus bipolar-disorder guidelines and sequential-care principles; the STAR*D study is referenced only as a structural example of multistep treatment sequencing in unipolar major depression, not as a validated bipolar-depression protocol. The sequencing problem is therefore a natural candidate for reinforcement learning (RL), but no RL framework has yet been developed specifically for personalised antidepressant sequencing in bipolar depression. We formulated antidepressant sequencing in TRD-BD as a finite-horizon Markov Decision Process and developed an offline RL framework for policy optimisation. A clinically grounded simulator encoded a 15-dimensional policy state comprising MADRS, YMRS, prior treatment failures, weeks on current treatment, side-effect burden, current drug class, a three-dimensional responder-profile covariate, age, and BD subtype, together with an 8-action treatment space (continue; switch to SSRI, SNRI, bupropion, or MAOI; augment with lithium, quetiapine, or lamotrigine). Although the responder profile represents an underlying biological trait, it was supplied to the policy as an oracle covariate in this proof-of-concept simulation; in real practice, it would usually be unobserved, making the clinical problem partially observable. The reward integrated MADRS improvement, remission, side-effect burden, switching cost, and manic-switch risk. A Deep Q-Network (DQN) was trained on 6172 transitions from 800 simulated trajectories generated by a stochastic behavioural clinician policy. Trajectories terminated at remission, manic switch, intolerable adverse effects, or the eight-stage horizon, yielding a mean logged trajectory length of 7.72 stages. Exact behavioural-policy action probabilities were recorded by the simulator and used for importance-weighted off-policy estimators. DQN was compared with Fitted Q-Iteration, the behavioural clinician policy, a fixed sequential protocol, and a random policy. The proposed DQN policy achieved an estimated discounted return of 41.6 (95% CI: 39.0 - 44.0), substantially exceeding Fitted Q-Iteration (24.6), the clinician policy (12.8), and the fixed protocol (12.1). The DQN policy achieved a remission rate of 61.8% (95% CI: 57.0% - 66.0%), nearly quadrupling the clinician policy’s 16.5% and reaching remission in fewer treatment stages (mean 6.5 vs 7.7). Mean MADRS reduction was 24.3 points versus 15.6 for the clinician policy. Bayesian posterior analysis assigned the DQN policy a 100.0% probability of being optimal. Policy analysis revealed that the learned strategy strongly favoured lithium augmentation selected in 88% of states over the clinician’s switch-heavy, lower-augmentation pattern. Within this synthetic environment, reinforcement learning derived a personalised sequencing policy that outperformed the simulated behavioural and fixed-protocol comparators. The learned preference for early lithium augmentation reflects the simulator parameters and should be interpreted as a proof-of-concept result rather than a clinical treatment recommendation. Validation with real TRD-BD trajectories and conservative offline-RL methods is required before any clinical inference or deployment.</p>
      </abstract>
      <kwd-group kwd-group-type="author-generated" xml:lang="en">
        <kwd>Reinforcement Learning</kwd>
        <kwd>Deep Q-Network</kwd>
        <kwd>Treatment-Resistant Bipolar Depression</kwd>
        <kwd>Antidepressant Sequencing</kwd>
        <kwd>Markov Decision Process</kwd>
        <kwd>Offline RL</kwd>
        <kwd>Off-Policy Evaluation</kwd>
        <kwd>Personalised Medicine</kwd>
        <kwd>Policy Optimisation</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec1">
      <title>1. Introduction</title>
      <p>The pharmacological management of treatment-resistant bipolar depression is, in its essential structure, a sequential decision problem. Throughout this study, TRD-BD is defined consistently as failure to achieve remission after at least two adequate pharmacological trials for the current bipolar depressive episode, including at least one antidepressant trial administered with mood-stabilising treatment. At each review, the clinician must choose whether to continue, switch antidepressant class, or augment with a mood stabiliser or atypical antipsychotic under uncertainty about the patient’s response profile [<xref ref-type="bibr" rid="B1">1</xref>]. Each decision changes the clinical state and therefore constrains the next decision. This is the structure of a sequential decision process and motivates reinforcement learning [<xref ref-type="bibr" rid="B2">2</xref>].</p>
      <p>Current clinical practice uses consensus bipolar-disorder guidelines and stepwise treatment algorithms. STAR*D is cited here only as a well-known example of sequential treatment staging in unipolar major depressive disorder, not as a bipolar-depression trial or validated bipolar protocol [<xref ref-type="bibr" rid="B3">3</xref>]. Bipolar-specific decisions should instead be interpreted in the context of bipolar treatment guidance and evidence [<xref ref-type="bibr" rid="B4">4</xref>]. Such population-level algorithms remain limited because they cannot fully adapt to evolving symptoms, accumulated adverse effects, partial response, and individual vulnerability.</p>
      <p>Reinforcement learning offers a principled framework for learning personalised, state-adaptive treatment policies from data. In the offline RL setting, the clinically realistic scenario where a policy must be learned from historical logged treatment data rather than through live experimentation on patients, algorithms such as Fitted Q-Iteration [<xref ref-type="bibr" rid="B5">5</xref>] and Deep Q-Networks [<xref ref-type="bibr" rid="B6">6</xref>] can estimate the optimal action-value function Q*(s, a) and derive the greedy policy that maximises expected cumulative clinical benefit. RL has been applied to dynamic treatment regimes in sepsis management [<xref ref-type="bibr" rid="B7">7</xref>], mechanical ventilation weaning [<xref ref-type="bibr" rid="B8">8</xref>] and HIV therapy sequencing [<xref ref-type="bibr" rid="B9">9</xref>], but never to antidepressant sequencing in bipolar depression.</p>
      <p>We present five contributions: 1) the first formulation of antidepressant sequencing in TRD-BD as a Markov Decision Process with a clinically grounded state space, action space, and reward function; 2) a clinically realistic patient simulator with latent per-patient drug-response profiles, enabling the personalisation problem to be posed and solved; 3) an offline Deep Q-Network policy trained on 6172 logged transitions, achieving 41.6 estimated return versus 12.8 for the clinician policy; 4) off-policy evaluation via direct method, weighted importance sampling, and doubly robust estimators; and 5) interpretable policy analysis revealing the learned strategy’s preference for early lithium augmentation over repeated antidepressant switching.</p>
    </sec>
    <sec id="sec2">
      <title>2. Background and Related Work</title>
      <sec id="sec2dot1">
        <title>2.1. Sequential Treatment Decisions in TRD-BD</title>
        <p>In this paper, treatment-resistant bipolar depression denotes failure to achieve remission after at least two adequate pharmacological trials during the current bipolar depressive episode, including at least one antidepressant trial given with mood-stabilising treatment. This operational definition is used consistently for simulator eligibility, initial-state generation, and interpretation of results. TRD-BD is associated with substantial morbidity, functional impairment, and suicide risk, while the relative merits of switching, mood-stabiliser augmentation, and atypical-antipsychotic addition remain contested [<xref ref-type="bibr" rid="B10">10</xref>].</p>
      </sec>
      <sec id="sec2dot2">
        <title>2.2. Reinforcement Learning for Dynamic Treatment Regimes</title>
        <p>A dynamic treatment regime (DTR) is a sequence of decision rules that map a patient’s evolving clinical state to a recommended treatment [<xref ref-type="bibr" rid="B11">11</xref>]. RL provides the natural computational framework for learning optimal DTRs: the patient is the environment, the clinician (or learned policy) is the agent, treatments are actions, and clinical outcomes define the reward. The action-value function Q(s, a) quantifies the expected cumulative reward of taking action a in state s and following the optimal policy thereafter; the Bellman optimality equation Q*(s, a) = E[r + <italic>γ</italic> max<sub>a</sub><sub>′</sub> Q*(s′, a′)] defines the recursive structure that value-based RL algorithms exploit.</p>
      </sec>
      <sec id="sec2dot3">
        <title>2.3. Offline Reinforcement Learning</title>
        <p>In clinical applications, online RL learning through live trial-and-error on patients is ethically and practically impossible. Offline (batch) RL instead learns a policy from a fixed dataset of logged transitions collected under a behavioural policy (here, observed clinician practice) [<xref ref-type="bibr" rid="B12">12</xref>]. Fitted Q-Iteration iteratively fits a regression model to the Bellman target, while Deep Q-Networks parametrise the Q-function with a neural network trained by temporal-difference learning. A central challenge in offline RL is distributional shift the learned policy may select actions rarely observed in the logged data, where value estimates are unreliable motivating off-policy evaluation methods that quantify policy value with appropriate uncertainty.</p>
      </sec>
      <sec id="sec2dot4">
        <title>2.4. Off-Policy Evaluation</title>
        <p>Off-policy evaluation (OPE) estimates the value of a target policy using data collected under a different behavioural policy [<xref ref-type="bibr" rid="B13">13</xref>]. The direct method (DM) uses the learned Q-function to estimate value but inherits its model bias; importance sampling (IS) reweights observed returns by the policy ratio but suffers high variance; the doubly robust (DR) estimator combines both, achieving lower variance than IS and lower bias than DM [<xref ref-type="bibr" rid="B14">14</xref>]. Robust OPE is essential for clinical RL because it provides the value estimate with uncertainty that would inform any decision to deploy a learned policy.</p>
      </sec>
    </sec>
    <sec id="sec3">
      <title>3. Methods</title>
      <sec id="sec3dot1">
        <title>3.1. MDP Formulation</title>
        <p>We formulated antidepressant sequencing in TRD-BD as a finite-horizon discounted decision process (<italic>γ</italic> = 0.95; maximum horizon = 8 treatment stages). The policy state s ∈ ℝ<sup>15</sup> encoded normalised MADRS, YMRS, number of prior treatment failures, weeks on current treatment, side-effect burden, current drug class (one-hot over five classes), a three-dimensional responder-profile covariate, age, and BD subtype. The responder profile was generated as a latent biological trait but was deliberately exposed to the policy as an oracle covariate in this proof-of-concept experiment; this preserves a fully observed MDP in the reported analysis. In real clinical use, the profile would normally be unavailable or only imperfectly inferred, so the corresponding problem would be a partially observable MDP and performance would likely be lower. The action space comprised eight choices: continue; switch to SSRI, SNRI, bupropion, or MAOI; or augment with lithium, quetiapine, or lamotrigine. The reward r = ΔMADRS + 55∙𝟙[remission] − 3∙(side-effect burden × propensity) − 0.5·𝟙[switch] − 8·𝟙[manic switch] integrated symptom improvement, remission (MADRS &lt; 10 and YMRS &lt; 8), adverse-effect burden, switching cost, and manic-switch harm.</p>
      </sec>
      <sec id="sec3dot2">
        <title>3.2. Patient Simulator</title>
        <p>The simulator instantiated each patient with a three-dimensional responder profile that modulated treatment-specific MADRS change. Treatment-response magnitudes and remission tendencies were parameterised from bipolar-depression treatment evidence and guideline syntheses; action-specific side-effect burdens were informed by comparative pharmacological tolerability evidence summarised in those sources; and antidepressant-associated manic-switch risks were set according to bipolar antidepressant-safety evidence, with greater risk assigned to antidepressant switching than to mood-stabiliser augmentation. Patients began with severe depressive symptoms (MADRS approximately 34) and two to four prior adequate treatment failures. All transition parameters remained synthetic and were not fitted to patient-level clinical data; the simulator is therefore a proof-of-concept environment rather than a validated disease model.</p>
      </sec>
      <sec id="sec3dot3">
        <title>3.3. Logged Data Generation and Offline RL</title>
        <p>A stochastic behavioural clinician policy generated 6172 logged transitions from N = 800 simulated patients. It combined stepwise escalation, side-effect-driven switching, and exploratory variability (<italic>ε</italic> = 0.30); STAR*D informed only the generic concept of sequential escalation, whereas treatment choices and risks were bipolar-specific. Trajectories ended early at remission, manic switch, or intolerable side-effect burden, and otherwise stopped at the eight-stage horizon. The resulting mean trajectory length was 6172/800 = 7.72 stages (range 1 - 8), explaining why the transition count was below the theoretical maximum of 6400. The simulator stored the exact probability assigned by the behavioural policy to every selected action, so weighted importance sampling and doubly robust estimation used known propensities rather than estimated ones. Offline DQN was selected because the state-action relationship was nonlinear and the discrete action space was moderate, while target networks and replay-based optimisation offered a direct neural comparator to Fitted Q-Iteration. Because standard offline DQN can extrapolate to poorly supported actions under distributional shift, its use was explicitly exploratory and was accompanied by support-sensitive off-policy evaluation and comparison with a conservative offline-RL direction in the limitations. The DQN used two 128-unit ReLU hidden layers and temporal-difference learning with a target network; Fitted Q-Iteration used a gradient-boosting Q-function for 15 iterations. Both yielded greedy policies <italic>π</italic>(s) = argmaxₐ Q(s, a).</p>
      </sec>
      <sec id="sec3dot4">
        <title>3.4. Evaluation and Off-Policy Evaluation</title>
        <p>Policies were evaluated by Monte Carlo rollout on 400 fresh simulated patients not used for training, measuring discounted return, remission, MADRS reduction, and stages to termination. Off-policy evaluation on the logged dataset used the direct method, weighted importance sampling with clipped ratios, and the doubly robust estimator. The denominator probabilities in the importance ratios were the exact action probabilities emitted and logged by the simulator’s behavioural policy; target-policy probabilities were derived from the evaluated policy, with deterministic actions represented using the stated smoothing/clipping procedure. Bayesian policy comparison used bootstrap value distributions from 2000 resamples. These estimators assess internal consistency within the simulator but do not establish real-world clinical value.</p>
      </sec>
    </sec>
    <sec id="sec4">
      <title>4. Results</title>
      <sec id="sec4dot1">
        <title>4.1. Offline RL Training Convergence</title>
        <p><xref ref-type="fig" rid="fig1">Figure 1</xref><xref ref-type="fig" rid="fig1">Figure 1</xref> presents optimisation traces for both offline RL algorithms. Fitted Q-Iteration stabilised after approximately eight iterations, and the DQN training objective stabilised over approximately 40 epochs. These curves indicate numerical convergence on the logged synthetic dataset; they do not rule out offline-RL extrapolation error or distributional shift, which must be assessed separately through policy support and off-policy evaluation.</p>
        <fig id="fig1">
          <label>Figure 1</label>
          <graphic xlink:href="https://html.scirp.org/file/1115677-rId16.jpeg?20260827033829" />
        </fig>
        <p><xref ref-type="fig" rid="fig1">Figure 1</xref>. Offline RL training convergence. Fitted Q-Iteration (dark, lower x-axis) converges in ~8 iterations; Deep Q-Network (purple, upper x-axis) converges over ~40 epochs. Both show stable monotonic improvement in estimated initial-state value V(s<sub>0</sub>) without divergence.</p>
      </sec>
      <sec id="sec4dot2">
        <title>4.2. Policy Value Comparison</title>
        <p><bold>Table</bold><bold>1</bold> and <xref ref-type="fig" rid="fig2">Figure 2</xref> present the primary policy comparison. The proposed DQN policy achieved an estimated discounted return of 41.6 (95% CI: 39.0 - 44.0), substantially exceeding Fitted Q-Iteration (24.6), the clinician policy (12.8), the STAR*D-style fixed protocol (12.1), and the random policy (12.9). The 95% confidence intervals for the DQN policy do not overlap those of any comparator, establishing a statistically robust advantage. The clinician policy, notably, performed only marginally better than random in terms of discounted return, reflecting its switch-heavy, low-augmentation strategy that the reward structure penalises.</p>
        <p><bold>Table 1.</bold>Policy comparison—held-out simulated patients (n = 400).</p>
        <table-wrap id="tbl1">
          <label>Table 1</label>
          <table>
            <tbody>
              <tr>
                <td>
                  <bold>Policy</bold>
                </td>
                <td>
                  <bold>Est. Return</bold>
                </td>
                <td>
                  <bold>Remission %</bold>
                </td>
                <td>
                  <bold>MADRS</bold>
                  <bold>↓</bold>
                </td>
                <td>
                  <bold>Stages</bold>
                </td>
              </tr>
              <tr>
                <td>Random</td>
                <td>12.9</td>
                <td>21.5</td>
                <td>17.6</td>
                <td>7.59</td>
              </tr>
              <tr>
                <td>Fixed Protocol</td>
                <td>12.1</td>
                <td>26.0</td>
                <td>15.9</td>
                <td>7.38</td>
              </tr>
              <tr>
                <td>Clinician</td>
                <td>12.8</td>
                <td>16.5</td>
                <td>15.6</td>
                <td>7.72</td>
              </tr>
              <tr>
                <td>Fitted Q-Iteration</td>
                <td>24.6</td>
                <td>39.0</td>
                <td>19.6</td>
                <td>7.25</td>
              </tr>
              <tr>
                <td>
                  <bold>DQN (proposed)</bold>
                </td>
                <td>
                  <bold>41.6</bold>
                </td>
                <td>
                  <bold>61.8</bold>
                </td>
                <td>
                  <bold>24.3</bold>
                </td>
                <td>
                  <bold>6.55</bold>
                </td>
              </tr>
            </tbody>
          </table>
        </table-wrap>
        <p>Estimated discounted return (<italic>γ</italic> = 0.95), remission rate (MADRS &lt; 10, YMRS &lt; 8), mean MADRS reduction, and mean treatment stages to termination. DQN 95% CI for return: 39.0 - 44.0; non-overlapping with all comparators. Evaluated on 400 held-out simulated patients not seen in training.</p>
        <fig id="fig2">
          <label>Figure 2</label>
          <graphic xlink:href="https://html.scirp.org/file/1115677-rId17.jpeg?20260827033829" />
        </fig>
        <p><bold>Figure 2.</bold> Policy value comparison with 95% bootstrap CI. The DQN policy (dark) achieves estimated return 41.6, with CI non-overlapping all comparators. Fitted Q-Iteration is second; clinician, fixed protocol, and random policies cluster well below.</p>
      </sec>
      <sec id="sec4dot3">
        <title>
          4.3. MDP Formulation Diagram (See
          <xref ref-type="fig" rid="fig3">Figure 3</xref>
          )
        </title>
        <fig id="fig3">
          <label>Figure 3</label>
          <graphic xlink:href="https://html.scirp.org/file/1115677-rId18.jpeg?20260827033829" />
        </fig>
        <p><bold>Figure 3.</bold> Markov decision process for antidepressant sequencing in TRD-BD.</p>
        <p>State (clinical features) → RL agent policy → action (treatment choice) → patient environment transition → reward and next state, repeated over up to 8 treatment stages. Reward integrates MADRS improvement, remission bonus, and penalties for side effects, switching, and manic switching (See <xref ref-type="fig" rid="fig3">Figure 3</xref>).</p>
      </sec>
      <sec id="sec4dot4">
        <title>4.4. Clinical Outcomes: Remission and Symptom Reduction</title>
        <p><xref ref-type="fig" rid="fig4">Figure 4</xref><xref ref-type="fig" rid="fig4">Figure 4</xref> presents clinical outcomes. The DQN policy achieved a remission rate of 61.8% (95% CI: 57.0% - 66.0%), nearly quadrupling the clinician policy’s 16.5% and substantially exceeding Fitted Q-Iteration (39.0%) and the fixed protocol (26.0%). The DQN policy also achieved the largest mean MADRS reduction (24.3 points vs 15.6 for the clinician), while reaching termination in the fewest treatment stages (mean 6.55 vs 7.72), indicating that the learned policy not only achieves higher remission but does so faster—a clinically critical property given the morbidity and suicide risk of prolonged TRD-BD.</p>
        <fig id="fig4">
          <label>Figure 4</label>
          <graphic xlink:href="https://html.scirp.org/file/1115677-rId19.jpeg?20260827033829" />
        </fig>
        <p><bold>Figure 4.</bold> Clinical outcomes by policy. Left: remission rate (DQN 61.8% vs clinician 16.5%). Right: mean MADRS reduction (DQN 24.3 vs clinician 15.6). Error bars: 95% bootstrap CI. The DQN policy dominates on both clinical endpoints.</p>
      </sec>
      <sec id="sec4dot5">
        <title>4.5. Learned Treatment Transition Policy</title>
        <p><xref ref-type="fig" rid="fig5">Figure 5</xref><xref ref-type="fig" rid="fig5">Figure 5</xref> presents the learned treatment transition policy as a heatmap of action probability conditioned on current drug class. The most striking feature is the learned policy’s strong and consistent preference for lithium augmentation across nearly all states: from the no-drug initial state and from most antidepressant classes, the DQN policy preferentially selects lithium augmentation. This reflects the simulator’s embedded clinical reality lithium augmentation carries high response efficacy in bipolar depression with low manic-switch risk which the RL policy discovers and exploits. The policy retains some state-dependence: from states with high side-effect burden, the policy shifts toward better-tolerated options, demonstrating the personalisation the MDP formulation enables.</p>
        <fig id="fig5">
          <label>Figure 5</label>
          <graphic xlink:href="https://html.scirp.org/file/1115677-rId20.jpeg?20260827033829" />
        </fig>
        <p><bold>Figure 5.</bold> Learned treatment transition policy (DQN). Heatmap of P (next action | current drug class). The learned policy strongly favours lithium augmentation across most states, reflecting its high efficacy and low manic-switch risk in bipolar depression, while retaining state-dependent adaptation.</p>
      </sec>
      <sec id="sec4dot6">
        <title>4.6. Action Selection: Learned Policy vs Clinician</title>
        <p><xref ref-type="fig" rid="fig6">Figure 6</xref><xref ref-type="fig" rid="fig6">Figure 6</xref> contrasts the action-selection frequencies of the DQN policy and the behavioural clinician policy. The divergence is clinically informative. The clinician policy distributes actions broadly, with the single most frequent action being “continue current treatment” (41%) and a relatively even spread across switching and augmentation options. The DQN policy, in contrast, concentrates 88% of its action mass on lithium augmentation, with minimal use of antidepressant switching. This reflects the core learned insight: in TRD-BD, decisive early mood-stabiliser augmentation outperforms the clinician’s more cautious, switch-heavy, “wait-and-see” pattern a finding consistent with emerging clinical evidence favouring augmentation over switching in bipolar depression [<xref ref-type="bibr" rid="B15">15</xref>].</p>
        <fig id="fig6">
          <label>Figure 6</label>
          <graphic xlink:href="https://html.scirp.org/file/1115677-rId21.jpeg?20260827033829" />
        </fig>
        <p><bold>Figure 6.</bold> Action selection: DQN policy (dark) vs clinician (green). The clinician favours “continue” (41%) and distributes broadly; the DQN policy concentrates 88% on lithium augmentation, reflecting the learned preference for decisive augmentation over switching.</p>
      </sec>
      <sec id="sec4dot7">
        <title>4.7. DQN Policy Value Estimates</title>
        <p><bold>Table</bold><bold>2</bold> presents off-policy evaluation of the DQN policy on the logged data using three estimators. The direct method estimated the policy value at 19.5, the weighted importance sampling estimator at 0.2 (reflecting the high variance and downward bias characteristic of IS under substantial policy divergence and clipped weights), and the doubly robust estimator at 15.6. The DR estimate, which balances the model bias of DM against the variance of IS, provides the most reliable single value estimate and is consistent with the Monte Carlo rollout value. The substantial divergence between the DQN policy and the behavioural clinician policy evident in the action distributions of <xref ref-type="fig" rid="fig6">Figure 6</xref><xref ref-type="fig" rid="fig6">Figure 6</xref> explains the low IS estimate: the learned policy frequently selects actions (lithium augmentation) that the clinician selected comparatively rarely, producing extreme importance weights that, even after clipping, depress the IS estimate [<xref ref-type="bibr" rid="B16">16</xref>]-[<xref ref-type="bibr" rid="B18">18</xref>].</p>
        <p><bold>Table 2.</bold>Off-policy evaluation of the DQN policy (logged data).</p>
        <table-wrap id="tbl2">
          <label>Table 2</label>
          <table>
            <tbody>
              <tr>
                <td>
                  <bold>OPE Estimator</bold>
                </td>
                <td>
                  <bold>Value Estimate</bold>
                </td>
                <td>
                  <bold>Property</bold>
                </td>
              </tr>
              <tr>
                <td>Direct Method (DM)</td>
                <td>19.5</td>
                <td>Low variance, model bias</td>
              </tr>
              <tr>
                <td>Weighted Importance Sampling</td>
                <td>0.2</td>
                <td>Unbiased, high variance</td>
              </tr>
              <tr>
                <td>
                  <bold>Doubly Robust (DR)</bold>
                </td>
                <td>
                  <bold>15.6</bold>
                </td>
                <td>
                  <bold>Balanced bias-variance</bold>
                </td>
              </tr>
            </tbody>
          </table>
        </table-wrap>
        <p>Off-policy value estimates on the logged clinician data. The doubly robust estimate balances the model bias of DM against the variance of IS. Low IS reflects substantial policy divergence (the learned policy favours actions the clinician used rarely).</p>
      </sec>
      <sec id="sec4dot8">
        <title>4.8. Bayesian Policy Comparison</title>
        <fig id="fig7">
          <label>Figure 7</label>
          <graphic xlink:href="https://html.scirp.org/file/1115677-rId22.jpeg?20260827033829" />
        </fig>
        <p><bold>Figure 7.</bold> Bayesian policy comparison. Left: bootstrap value posterior distributions (2000 resamples); the DQN posterior dominates all comparators. Right: probability each policy is optimal—DQN 100.0%, all others negligible.</p>
        <p><xref ref-type="fig" rid="fig7">Figure 7</xref><xref ref-type="fig" rid="fig7">Figure 7</xref> presents the Bayesian policy comparison. The left panel shows the bootstrap value posterior distributions for all five policies; the DQN policy’s posterior is cleanly separated from and dominates all comparators. The right panel quantifies the probability that each policy is optimal: the DQN policy is optimal with 100.0% posterior probability, while no other policy achieves non-negligible probability of optimality. This decisive Bayesian separation confirms that the DQN policy’s advantage is robust to sampling uncertainty in the policy value estimates.</p>
      </sec>
      <sec id="sec4dot9">
        <title>4.9. Representative Patient Trajectories</title>
        <p><xref ref-type="fig" rid="fig8">Figure 8</xref><xref ref-type="fig" rid="fig8">Figure 8</xref> presents MADRS trajectories for two representative patients under the DQN policy versus the clinician policy. In both cases, the DQN policy reaches the remission threshold (MADRS &lt; 10) faster and more reliably than the clinician policy. The per-stage action labels reveal the mechanism: the DQN policy applies decisive augmentation early (typically lithium augmentation by the first or second stage), driving rapid MADRS reduction, whereas the clinician policy’s more cautious switching pattern produces slower, less complete symptom resolution. These trajectories illustrate at the individual-patient level the population-level advantage quantified in <bold>Table</bold><bold>1</bold>, <bold>Table 2</bold> and <xref ref-type="fig" rid="fig2">Figures 2-4</xref>.</p>
        <fig id="fig8">
          <label>Figure 8</label>
          <graphic xlink:href="https://html.scirp.org/file/1115677-rId23.jpeg?20260827033829" />
        </fig>
        <p><bold>Figure 8.</bold> Representative patient MADRS trajectories. DQN policy (dark) versus clinician policy (green) on two patients. The DQN policy reaches the remission threshold (red dotted line) faster through decisive early augmentation. Labels indicate the DQN-selected action at each treatment stage.</p>
      </sec>
    </sec>
    <sec id="sec5">
      <title>5. Discussion</title>
      <p>The central finding is that an offline reinforcement learning policy, trained on logged clinician treatment trajectories, can derive an antidepressant sequencing strategy for TRD-BD that substantially outperforms both the behavioural clinician policy and population-averaged fixed protocols achieving nearly four times the remission rate of the clinician policy in simulation. The mechanism of this advantage is interpretable and clinically meaningful: the learned policy favours decisive early mood-stabiliser augmentation (particularly lithium) over the clinician’s switch-heavy, cautious pattern. This learned preference aligns with a growing clinical evidence base suggesting that augmentation strategies outperform repeated antidepressant switching in bipolar depression, where antidepressants carry both limited efficacy and manic-switch risk [<xref ref-type="bibr" rid="B19">19</xref>]-[<xref ref-type="bibr" rid="B21">21</xref>].</p>
      <p>The personalisation capacity demonstrated here depends partly on the responder-profile covariate. In the reported proof-of-concept experiment, this simulator-generated biological profile was observable to the policy, functioning as an oracle feature. This design tests whether a policy can exploit patient heterogeneity when that heterogeneity is measured, but it is optimistic relative to clinical practice. If the responder profile is hidden, the task becomes partially observable and would require history-based state estimation, recurrent policies, Bayesian belief states, or measurable pharmacogenomic proxies before comparable personalisation claims could be evaluated [<xref ref-type="bibr" rid="B22">22</xref>].</p>
      <p>The off-policy evaluation results illustrate a fundamental tension in clinical RL. The large divergence between the learned policy and the behavioural clinician policy which drives the learned policy’s superior performance simultaneously makes off-policy value estimation difficult, because the logged data contain few examples of the actions the learned policy prefers [<xref ref-type="bibr" rid="B23">23</xref>]. This is the offline RL distributional-shift problem in its clinical form: the more a learned policy improves on current practice, the harder it is to validate that improvement from historical data alone. This argues for prospective, carefully monitored evaluation of any RL-derived treatment policy, rather than reliance on retrospective off-policy estimates.</p>
    </sec>
    <sec id="sec6">
      <title>6. Limitations</title>
      <p>The patient simulator is synthetic, and the learned policy’s performance reflects its reward and transition rules rather than observed clinical effectiveness. The strong preference for lithium augmentation is partly induced by the simulator’s efficacy, side-effect, and manic-switch parameters and must not be interpreted as a treatment recommendation. The responder profile was exposed to the policy as an oracle covariate, whereas in practice it would generally be latent, making the real problem partially observable. STAR*D contributed only the abstract idea of staged treatment sequencing and is a unipolar-depression study; it does not validate the bipolar treatment protocol used here. Standard offline DQN is vulnerable to distributional shift and overestimation for actions with weak behavioural-policy support. Although exact behavioural action probabilities were available in the simulator, the marked divergence between the DQN and clinician policies produced unstable importance-weighted estimates. The reward weights encode value judgements that require stakeholder elicitation, and the eight discrete actions omit dose, timing, combination, and psychotherapy decisions [<xref ref-type="bibr" rid="B24">24</xref>][<xref ref-type="bibr" rid="B25">25</xref>].</p>
    </sec>
    <sec id="sec7">
      <title>7. Conclusion</title>
      <p>This simulation-based proof-of-concept study formulates personalised sequencing for TRD-BD—defined as non-remission after at least two adequate pharmacological trials, including one antidepressant trial with mood-stabilising treatment—as a sequential decision problem. Within the synthetic environment, a DQN achieved higher simulated return and remission than Fitted Q-Iteration, the behavioural policy, a fixed sequential protocol, and random treatment. These results demonstrate methodological feasibility, not clinical superiority. The lithium-dominant policy is contingent on simulator assumptions, access to an oracle responder profile, the selected reward weights, and the logged-policy support. Priorities are therefore: calibration against real longitudinal bipolar-depression data; evaluation when responder heterogeneity is hidden or imperfectly measured; stakeholder-defined reward construction; and conservative, uncertainty-aware offline RL that penalises unsupported actions before any prospective clinical study.</p>
    </sec>
    <sec id="sec8">
      <title>Author Contributions</title>
      <p>Conceptualization, RDF and AAF; methodology, AAF and RDF; software, AAF; validation, RDF and AAF; formal analysis, AAF; investigation, RDF and AAF; resources, RDF; data curation, AAF; writing original draft preparation, AAF; writing review and editing, RDF and AAF; visualization, AAF; supervision, RDF; project administration, RDF. All authors have read and agreed to the published version of the manuscript.</p>
      <p><bold>Initials:</bold> RDF, Rocco de Filippis; AAF, Abdullah Al Foysal.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <title>References</title>
      <ref id="B1">
        <label>1.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Tundo, A., Filippis, R.d. and Proietti, L. (2015) Pharmacologic Approaches to Treatment Resistant Depression: Evidences and Personal Experience. <italic>World</italic><italic>Journal</italic><italic>of</italic><italic>Psychiatry</italic>, 5, 330-341. https://doi.org/10.5498/wjp.v5.i3.330 <pub-id pub-id-type="doi">10.5498/wjp.v5.i3.330</pub-id><pub-id pub-id-type="pmid">26425446</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.5498/wjp.v5.i3.330">https://doi.org/10.5498/wjp.v5.i3.330</ext-link></mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Tundo, A.</string-name>
              <string-name>Filippis, R.</string-name>
              <string-name>Proietti, L.</string-name>
            </person-group>
            <year>2015</year>
            <article-title>Pharmacologic Approaches to Treatment Resistant Depression: Evidences and Personal Experience</article-title>
            <source>World Journal of Psychiatry</source>
            <volume>5</volume>
            <pub-id pub-id-type="doi">10.5498/wjp.v5.i3.330</pub-id>
            <pub-id pub-id-type="pmid">26425446</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B2">
        <label>2.</label>
        <citation-alternatives>
          <mixed-citation publication-type="book">Sutton, R.S. and Barto, A.G. (2018) Reinforcement Learning: An Introduction. 2nd Edition, MIT Press.</mixed-citation>
          <element-citation publication-type="book">
            <person-group person-group-type="author">
              <string-name>Sutton, R.S.</string-name>
              <string-name>Barto, A.G.</string-name>
              <string-name>Edition, M</string-name>
            </person-group>
            <year>2018</year>
            <article-title>Reinforcement Learning: An Introduction</article-title>
            <source>2nd Edition</source>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B3">
        <label>3.</label>
        <citation-alternatives>
          <mixed-citation publication-type="report">Rush, A.J., Trivedi, M.H., Wisniewski, S.R., Nierenberg, A.A., Stewart, J.W., Warden, D., <italic>et al</italic>. (2006) Acute and Longer-Term Outcomes in Depressed Outpatients Requiring One or Several Treatment Steps: A STAR*D Report. <italic>American Journal of Psychiatry</italic>, 163, 1905-1917. https://doi.org/10.1176/ajp.2006.163.11.1905 <pub-id pub-id-type="doi">10.1176/ajp.2006.163.11.1905</pub-id><pub-id pub-id-type="pmid">17074942</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1176/ajp.2006.163.11.1905">https://doi.org/10.1176/ajp.2006.163.11.1905</ext-link></mixed-citation>
          <element-citation publication-type="report">
            <person-group person-group-type="author">
              <string-name>Rush, A.J.</string-name>
              <string-name>Trivedi, M.H.</string-name>
              <string-name>Wisniewski, S.R.</string-name>
              <string-name>Nierenberg, A.A.</string-name>
              <string-name>Stewart, J.W.</string-name>
              <string-name>Warden, D.</string-name>
            </person-group>
            <year>2006</year>
            <article-title>Acute and Longer-Term Outcomes in Depressed Outpatients Requiring One or Several Treatment Steps: A STAR*D Report</article-title>
            <source>American Journal of Psychiatry</source>
            <volume>163</volume>
            <pub-id pub-id-type="doi">10.1176/ajp.2006.163.11.1905</pub-id>
            <pub-id pub-id-type="pmid">17074942</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B4">
        <label>4.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Hui Poon, S., Sim, K. and J. Baldessarini, R. (2015) Pharmacological Approaches for Treatment-Resistant Bipolar Disorder. <italic>Current Neuropharmacology</italic>, 13, 592-604. https://doi.org/10.2174/1570159x13666150630171954 <pub-id pub-id-type="doi">10.2174/1570159x13666150630171954</pub-id><pub-id pub-id-type="pmid">26467409</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.2174/1570159x13666150630171954">https://doi.org/10.2174/1570159x13666150630171954</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Poon, S.</string-name>
              <string-name>Sim, K.</string-name>
              <string-name>Baldessarini, R.</string-name>
            </person-group>
            <year>2015</year>
            <article-title>Pharmacological Approaches for Treatment-Resistant Bipolar Disorder</article-title>
            <source>Current Neuropharmacology</source>
            <volume>13</volume>
            <pub-id pub-id-type="doi">10.2174/1570159x13666150630171954</pub-id>
            <pub-id pub-id-type="pmid">26467409</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B5">
        <label>5.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Ernst, D., Geurts, P. and Wehenkel, L. (2005) Tree-Based Batch Mode Reinforcement Learning. <italic>Journal of Machine Learning Research</italic>, 6, 503-556.</mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Ernst, D.</string-name>
              <string-name>Geurts, P.</string-name>
              <string-name>Wehenkel, L.</string-name>
            </person-group>
            <year>2005</year>
            <article-title>Tree-Based Batch Mode Reinforcement Learning</article-title>
            <source>Journal of Machine Learning Research</source>
            <volume>6</volume>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B6">
        <label>6.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A.A., Veness, J., Bellemare, M.G., <italic>et al</italic>. (2015) Human-Level Control through Deep Reinforcement Learning. <italic>Nature</italic>, 518, 529-533. https://doi.org/10.1038/nature14236 <pub-id pub-id-type="doi">10.1038/nature14236</pub-id><pub-id pub-id-type="pmid">25719670</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1038/nature14236">https://doi.org/10.1038/nature14236</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Mnih, V.</string-name>
              <string-name>Kavukcuoglu, K.</string-name>
              <string-name>Silver, D.</string-name>
              <string-name>Rusu, A.A.</string-name>
              <string-name>Veness, J.</string-name>
              <string-name>Bellemare, M.G.</string-name>
            </person-group>
            <year>2015</year>
            <article-title>Human-Level Control through Deep Reinforcement Learning</article-title>
            <source>Nature</source>
            <volume>518</volume>
            <pub-id pub-id-type="doi">10.1038/nature14236</pub-id>
            <pub-id pub-id-type="pmid">25719670</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B7">
        <label>7.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Komorowski, M., Celi, L.A., Badawi, O., Gordon, A.C. and Faisal, A.A. (2018) The Artificial Intelligence Clinician Learns Optimal Treatment Strategies for Sepsis in Intensive Care. <italic>Nature Medicine</italic>, 24, 1716-1720. https://doi.org/10.1038/s41591-018-0213-5 <pub-id pub-id-type="doi">10.1038/s41591-018-0213-5</pub-id><pub-id pub-id-type="pmid">30349085</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1038/s41591-018-0213-5">https://doi.org/10.1038/s41591-018-0213-5</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Komorowski, M.</string-name>
              <string-name>Celi, L.A.</string-name>
              <string-name>Badawi, O.</string-name>
              <string-name>Gordon, A.C.</string-name>
              <string-name>Faisal, A.A.</string-name>
            </person-group>
            <year>2018</year>
            <article-title>The Artificial Intelligence Clinician Learns Optimal Treatment Strategies for Sepsis in Intensive Care</article-title>
            <source>Nature Medicine</source>
            <volume>24</volume>
            <pub-id pub-id-type="doi">10.1038/s41591-018-0213-5</pub-id>
            <pub-id pub-id-type="pmid">30349085</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B8">
        <label>8.</label>
        <citation-alternatives>
          <mixed-citation publication-type="confproc">Prasad, N., Cheng, L.-F., Chivers, C., Draugelis, M. and Engelhardt, B.E. (2017) A Reinforcement Learning Approach to Weaning of Mechanical Ventilation in Intensive Care Units. <italic>Proceedings of the</italic>33 <italic>rd Conference on Uncertainty in Artificial Intelligence</italic> ( <italic>UAI</italic> 2017), Sydney, 11-15 August 2017.</mixed-citation>
          <element-citation publication-type="confproc">
            <person-group person-group-type="author">
              <string-name>Prasad, N.</string-name>
              <string-name>Cheng, L.</string-name>
              <string-name>Chivers, C.</string-name>
              <string-name>Draugelis, M.</string-name>
              <string-name>Engelhardt, B.E.</string-name>
            </person-group>
            <year>2017</year>
            <article-title>A Reinforcement Learning Approach to Weaning of Mechanical Ventilation in Intensive Care Units</article-title>
            <source>Proceedings of the 33rd Conference on Uncertainty in Artificial Intelligence (UAI 2017)</source>
            <volume>11</volume>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B9">
        <label>9.</label>
        <citation-alternatives>
          <mixed-citation publication-type="confproc">Parbhoo, S., Bogojeska, J., Zazzi, M., Roth, V. and Doshi-Velez, F. (2017) Combining Kernel and Model Based Learning for HIV Therapy Selection. <italic>AMIA Summits on Translational Science Proceedings</italic>, San Francisco, 27-30 March 2017, 239-248.</mixed-citation>
          <element-citation publication-type="confproc">
            <person-group person-group-type="author">
              <string-name>Parbhoo, S.</string-name>
              <string-name>Bogojeska, J.</string-name>
              <string-name>Zazzi, M.</string-name>
              <string-name>Roth, V.</string-name>
              <string-name>Doshi-Velez, F.</string-name>
              <string-name>Proceedings, S</string-name>
            </person-group>
            <year>2017</year>
            <article-title>Combining Kernel and Model Based Learning for HIV Therapy Selection</article-title>
            <source>AMIA Summits on Translational Science Proceedings</source>
            <volume>27</volume>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B10">
        <label>10.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Yatham, L.N., Kennedy, S.H., Parikh, S.V., <italic>et al</italic>. (2018) Canadian Network for Mood and Anxiety Treatments (CANMAT) and International Society for Bipolar Disorders (ISBD) 2018 Guidelines for the Management of Patients with Bipolar Disorder. <italic>Bipolar Disorders</italic>, 20, 97-170.</mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Yatham, L.N.</string-name>
              <string-name>Kennedy, S.H.</string-name>
              <string-name>Parikh, S.V.</string-name>
            </person-group>
            <year>2018</year>
            <article-title>Canadian Network for Mood and Anxiety Treatments (CANMAT) and International Society for Bipolar Disorders (ISBD) 2018 Guidelines for the Management of Patients with Bipolar Disorder</article-title>
            <source>Bipolar Disorders</source>
            <volume>20</volume>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B11">
        <label>11.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Chakraborty, B. and Murphy, S.A. (2014) Dynamic Treatment Regimes. <italic>Annual</italic><italic>Review</italic><italic>of</italic><italic>Statistics</italic><italic>and</italic><italic>Its</italic><italic>Application</italic>, 1, 447-464. https://doi.org/10.1146/annurev-statistics-022513-115553 <pub-id pub-id-type="doi">10.1146/annurev-statistics-022513-115553</pub-id><pub-id pub-id-type="pmid">25401119</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1146/annurev-statistics-022513-115553">https://doi.org/10.1146/annurev-statistics-022513-115553</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Chakraborty, B.</string-name>
              <string-name>Murphy, S.A.</string-name>
            </person-group>
            <year>2014</year>
            <article-title>Dynamic Treatment Regimes</article-title>
            <source>Annual Review of Statistics and Its Application</source>
            <volume>1</volume>
            <pub-id pub-id-type="doi">10.1146/annurev-statistics-022513-115553</pub-id>
            <pub-id pub-id-type="pmid">25401119</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B12">
        <label>12.</label>
        <citation-alternatives>
          <mixed-citation publication-type="web">Levine, S., Kumar, A., Tucker, G. and Fu, J. (2020) Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems. https://arxiv.org/abs/2005.01643</mixed-citation>
          <element-citation publication-type="web">
            <person-group person-group-type="author">
              <string-name>Levine, S.</string-name>
              <string-name>Kumar, A.</string-name>
              <string-name>Tucker, G.</string-name>
              <string-name>Fu, J.</string-name>
              <string-name>Tutorial, R</string-name>
            </person-group>
            <year>2020</year>
            <article-title>Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems</article-title>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B13">
        <label>13.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Gottesman, O., Johansson, F., Komorowski, M., Faisal, A., Sontag, D., Doshi-Velez, F., <italic>et al</italic>. (2019) Guidelines for Reinforcement Learning in Healthcare. <italic>Nature</italic><italic>Medicine</italic>, 25, 16-18. https://doi.org/10.1038/s41591-018-0310-5 <pub-id pub-id-type="doi">10.1038/s41591-018-0310-5</pub-id><pub-id pub-id-type="pmid">30617332</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1038/s41591-018-0310-5">https://doi.org/10.1038/s41591-018-0310-5</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Gottesman, O.</string-name>
              <string-name>Johansson, F.</string-name>
              <string-name>Komorowski, M.</string-name>
              <string-name>Faisal, A.</string-name>
              <string-name>Sontag, D.</string-name>
              <string-name>Doshi-Velez, F.</string-name>
            </person-group>
            <year>2019</year>
            <article-title>Guidelines for Reinforcement Learning in Healthcare</article-title>
            <source>Nature Medicine</source>
            <volume>25</volume>
            <pub-id pub-id-type="doi">10.1038/s41591-018-0310-5</pub-id>
            <pub-id pub-id-type="pmid">30617332</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B14">
        <label>14.</label>
        <citation-alternatives>
          <mixed-citation publication-type="confproc">Jiang, N. and Li, L. (2016) Doubly Robust Off-Policy Value Evaluation for Reinforcement Learning. <italic>Proceedings of the</italic>33 <italic>rd International Conference on Machine Learning</italic>, <italic>PMLR</italic>, Vol. 48, 652-661.</mixed-citation>
          <element-citation publication-type="confproc">
            <person-group person-group-type="author">
              <string-name>Jiang, N.</string-name>
              <string-name>Li, L.</string-name>
              <string-name>Learning, P</string-name>
              <string-name>MLR, V</string-name>
            </person-group>
            <year>2016</year>
            <article-title>Doubly Robust Off-Policy Value Evaluation for Reinforcement Learning</article-title>
            <source>Proceedings of the 33rd International Conference on Machine Learning</source>
            <volume>652</volume>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B15">
        <label>15.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Kumar, A., Zhou, A., Tucker, G. and Levine, S. (2020) Conservative Q-Learning for Offline Reinforcement Learning. <italic>Advances in Neural Information Processing Systems</italic> 33, 6-12 December 2020, 1179-1191.</mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Kumar, A.</string-name>
              <string-name>Zhou, A.</string-name>
              <string-name>Tucker, G.</string-name>
              <string-name>Levine, S.</string-name>
            </person-group>
            <year>2020</year>
            <article-title>Conservative Q-Learning for Offline Reinforcement Learning</article-title>
            <source>Advances in Neural Information Processing Systems 33</source>
            <volume>6</volume>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B16">
        <label>16.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Watkins, C.J.C.H. and Dayan, P. (1992) Technical Note: Q-Learning. <italic>Machine</italic><italic>Learning</italic>, 8, 279-292. https://doi.org/10.1023/a:1022676722315 <pub-id pub-id-type="doi">10.1023/a:1022676722315</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1023/a:1022676722315">https://doi.org/10.1023/a:1022676722315</ext-link></mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Watkins, C.J.C.H.</string-name>
              <string-name>Dayan, P.</string-name>
            </person-group>
            <year>1992</year>
            <article-title>Technical Note: Q-Learning</article-title>
            <source>Machine Learning</source>
            <volume>8</volume>
            <fpage>102267</fpage>
            <pub-id pub-id-type="doi">10.1023/a:1022676722315</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B17">
        <label>17.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Bellman, R. (1957) A Markovian Decision Process. <italic>Indiana</italic><italic>University</italic><italic>Mathematics</italic><italic>Journal</italic>, 6, 679-684. https://doi.org/10.1512/iumj.1957.6.56038 <pub-id pub-id-type="doi">10.1512/iumj.1957.6.56038</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1512/iumj.1957.6.56038">https://doi.org/10.1512/iumj.1957.6.56038</ext-link></mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Bellman, R.</string-name>
            </person-group>
            <year>1957</year>
            <article-title>A Markovian Decision Process</article-title>
            <source>Indiana University Mathematics Journal</source>
            <volume>6</volume>
            <pub-id pub-id-type="doi">10.1512/iumj.1957.6.56038</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B18">
        <label>18.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Sachs, G.S., Nierenberg, A.A., Calabrese, J.R., Marangell, L.B., Wisniewski, S.R., Gyulai, L., <italic>et al</italic>. (2007) Effectiveness of Adjunctive Antidepressant Treatment for Bipolar Depression. <italic>New</italic><italic>England</italic><italic>Journal</italic><italic>of</italic><italic>Medicine</italic>, 356, 1711-1722. https://doi.org/10.1056/nejmoa064135 <pub-id pub-id-type="doi">10.1056/nejmoa064135</pub-id><pub-id pub-id-type="pmid">17392295</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1056/nejmoa064135">https://doi.org/10.1056/nejmoa064135</ext-link></mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Sachs, G.S.</string-name>
              <string-name>Nierenberg, A.A.</string-name>
              <string-name>Calabrese, J.R.</string-name>
              <string-name>Marangell, L.B.</string-name>
              <string-name>Wisniewski, S.R.</string-name>
              <string-name>Gyulai, L.</string-name>
            </person-group>
            <year>2007</year>
            <article-title>Effectiveness of Adjunctive Antidepressant Treatment for Bipolar Depression</article-title>
            <source>New England Journal of Medicine</source>
            <volume>356</volume>
            <pub-id pub-id-type="doi">10.1056/nejmoa064135</pub-id>
            <pub-id pub-id-type="pmid">17392295</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B19">
        <label>19.</label>
        <citation-alternatives>
          <mixed-citation publication-type="report">Pacchiarotti, I., Bond, D. J., Baldessarini, R. J., <italic>et al</italic>. (2013) The International Society for Bipolar Disorders (ISBD) Task Force Report on Antidepressant Use in Bipolar Disorders. <italic>American Journal of Psychiatry</italic>, 170, 1249-1262.</mixed-citation>
          <element-citation publication-type="report">
            <person-group person-group-type="author">
              <string-name>Pacchiarotti, I.</string-name>
              <string-name>Bond, D.</string-name>
              <string-name>Baldessarini, R.</string-name>
            </person-group>
            <year>2013</year>
            <article-title>The International Society for Bipolar Disorders (ISBD) Task Force Report on Antidepressant Use in Bipolar Disorders</article-title>
            <source>American Journal of Psychiatry</source>
            <volume>170</volume>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B20">
        <label>20.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Nahum-Shani, I., Qian, M., Almirall, D., Pelham, W.E., Gnagy, B., Fabiano, G.A., <italic>et al</italic>. (2012) Experimental Design and Primary Data Analysis Methods for Comparing Adaptive Interventions. <italic>Psychological</italic><italic>Methods</italic>, 17, 457-477. https://doi.org/10.1037/a0029372 <pub-id pub-id-type="doi">10.1037/a0029372</pub-id><pub-id pub-id-type="pmid">23025433</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1037/a0029372">https://doi.org/10.1037/a0029372</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Nahum-Shani, I.</string-name>
              <string-name>Qian, M.</string-name>
              <string-name>Almirall, D.</string-name>
              <string-name>Pelham, W.E.</string-name>
              <string-name>Gnagy, B.</string-name>
              <string-name>Fabiano, G.A.</string-name>
            </person-group>
            <year>2012</year>
            <article-title>Experimental Design and Primary Data Analysis Methods for Comparing Adaptive Interventions</article-title>
            <source>Psychological Methods</source>
            <volume>17</volume>
            <pub-id pub-id-type="doi">10.1037/a0029372</pub-id>
            <pub-id pub-id-type="pmid">23025433</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B21">
        <label>21.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Geurts, P., Ernst, D. and Wehenkel, L. (2006) Extremely Randomized Trees. <italic>Machine</italic><italic>Learning</italic>, 63, 3-42. https://doi.org/10.1007/s10994-006-6226-1 <pub-id pub-id-type="doi">10.1007/s10994-006-6226-1</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1007/s10994-006-6226-1">https://doi.org/10.1007/s10994-006-6226-1</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Geurts, P.</string-name>
              <string-name>Ernst, D.</string-name>
              <string-name>Wehenkel, L.</string-name>
            </person-group>
            <year>2006</year>
            <article-title>Extremely Randomized Trees</article-title>
            <source>Machine Learning</source>
            <volume>63</volume>
            <pub-id pub-id-type="doi">10.1007/s10994-006-6226-1</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B22">
        <label>22.</label>
        <citation-alternatives>
          <mixed-citation publication-type="confproc">Thomas, P.S. and Brunskill, E. (2016) Data-Efficient Off-Policy Policy Evaluation for Reinforcement Learning. <italic>Proceedings of the</italic>33 <italic>rd International Conference on Machine Learning</italic>, <italic>PMLR</italic>, Vol. 48, 2139-2148.</mixed-citation>
          <element-citation publication-type="confproc">
            <person-group person-group-type="author">
              <string-name>Thomas, P.S.</string-name>
              <string-name>Brunskill, E.</string-name>
              <string-name>Learning, P</string-name>
              <string-name>MLR, V</string-name>
            </person-group>
            <year>2016</year>
            <article-title>Data-Efficient Off-Policy Policy Evaluation for Reinforcement Learning</article-title>
            <source>Proceedings of the 33rd International Conference on Machine Learning</source>
            <volume>2139</volume>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B23">
        <label>23.</label>
        <citation-alternatives>
          <mixed-citation publication-type="web">Raghu, A., Komorowski, M., Ahmed, I., Celi, L., Szolovits, P. and Ghassemi, M. (2017) Deep Reinforcement Learning for Sepsis Treatment. https://arxiv.org/abs/1711.09602</mixed-citation>
          <element-citation publication-type="web">
            <person-group person-group-type="author">
              <string-name>Raghu, A.</string-name>
              <string-name>Komorowski, M.</string-name>
              <string-name>Ahmed, I.</string-name>
              <string-name>Celi, L.</string-name>
              <string-name>Szolovits, P.</string-name>
              <string-name>Ghassemi, M.</string-name>
            </person-group>
            <year>2017</year>
            <article-title>Deep Reinforcement Learning for Sepsis Treatment</article-title>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B24">
        <label>24.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Montague, P.R., Dolan, R.J., Friston, K.J. and Dayan, P. (2012) Computational Psychiatry. <italic>Trends</italic><italic>in</italic><italic>Cognitive</italic><italic>Sciences</italic>, 16, 72-80. https://doi.org/10.1016/j.tics.2011.11.018 <pub-id pub-id-type="doi">10.1016/j.tics.2011.11.018</pub-id><pub-id pub-id-type="pmid">22177032</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1016/j.tics.2011.11.018">https://doi.org/10.1016/j.tics.2011.11.018</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Montague, P.R.</string-name>
              <string-name>Dolan, R.J.</string-name>
              <string-name>Friston, K.J.</string-name>
              <string-name>Dayan, P.</string-name>
            </person-group>
            <year>2012</year>
            <article-title>Computational Psychiatry</article-title>
            <source>Trends in Cognitive Sciences</source>
            <volume>16</volume>
            <pub-id pub-id-type="doi">10.1016/j.tics.2011.11.018</pub-id>
            <pub-id pub-id-type="pmid">22177032</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B25">
        <label>25.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Moazemi, S., Vahdati, S., Li, J., Kalkhoff, S., Castano, L.J.V., Dewitz, B., <italic>et al</italic>. (2023) Artificial Intelligence for Clinical Decision Support for Monitoring Patients in Cardiovascular ICUs: A Systematic Review. <italic>Frontiers in Medicine</italic>, 10, 18 p. https://doi.org/10.3389/fmed.2023.1109411 <pub-id pub-id-type="doi">10.3389/fmed.2023.1109411</pub-id><pub-id pub-id-type="pmid">37064042</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.3389/fmed.2023.1109411">https://doi.org/10.3389/fmed.2023.1109411</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Moazemi, S.</string-name>
              <string-name>Vahdati, S.</string-name>
              <string-name>Li, J.</string-name>
              <string-name>Kalkhoff, S.</string-name>
              <string-name>Castano, L.J.V.</string-name>
              <string-name>Dewitz, B.</string-name>
            </person-group>
            <year>2023</year>
            <article-title>Artificial Intelligence for Clinical Decision Support for Monitoring Patients in Cardiovascular ICUs: A Systematic Review</article-title>
            <source>Frontiers in Medicine</source>
            <volume>10</volume>
            <pub-id pub-id-type="doi">10.3389/fmed.2023.1109411</pub-id>
            <pub-id pub-id-type="pmid">37064042</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
    </ref-list>
  </back>
</article>