<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.4 20241031//EN" "JATS-journalpublishing1-4.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article" dtd-version="1.4" xml:lang="en">
  <front>
    <journal-meta>
      <journal-id journal-id-type="publisher-id">Oalib</journal-id>
      <journal-title-group>
        <journal-title>Open Access Library Journal</journal-title>
      </journal-title-group>
      <issn pub-type="epub">2333-9721</issn>
      <issn pub-type="ppub">2333-9705</issn>
      <publisher>
        <publisher-name>Scientific Research Publishing</publisher-name>
      </publisher>
    </journal-meta>
    <article-meta>
      <article-id pub-id-type="doi">10.4236/oalib.1114926</article-id>
      <article-id pub-id-type="publisher-id">Oalib-150157</article-id>
      <article-categories>
        <subj-group>
          <subject>Article</subject>
        </subj-group>
        <subj-group>
          <subject>Biomedical</subject>
          <subject>Life Sciences</subject>
          <subject>Business</subject>
          <subject>Economics</subject>
          <subject>Chemistry</subject>
          <subject>Materials Science</subject>
          <subject>Computer Science</subject>
          <subject>Communications</subject>
          <subject>Earth</subject>
          <subject>Environmental Sciences</subject>
          <subject>Engineering</subject>
          <subject>Medicine</subject>
          <subject>Healthcare</subject>
          <subject>Physics</subject>
          <subject>Mathematics</subject>
          <subject>Social Sciences</subject>
          <subject>Humanities</subject>
        </subj-group>
      </article-categories>
      <title-group>
        <article-title>Reinforcement Learning Based Optimization of Sleep Mood Circadian Dynamics in Bipolar Disorder: A Simulation Study</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author" corresp="yes">
          <contrib-id contrib-id-type="orcid">0000-0001-9101-072X</contrib-id>
          <name name-style="western">
            <surname>Filippis</surname>
            <given-names>Rocco de</given-names>
          </name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <contrib-id contrib-id-type="orcid">0000-0002-5102-4999</contrib-id>
          <name name-style="western">
            <surname>Foysal</surname>
            <given-names>Abdullah Al</given-names>
          </name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
      </contrib-group>
      <aff id="aff1"><label>1</label> Department of Neuroscience, Institute of Psychopathology, Rome, Italy </aff>
      <aff id="aff2"><label>2</label> Department of Computer Engineering (AI), University of Genova, Genova, Italy </aff>
      <author-notes>
        <fn fn-type="conflict" id="fn-conflict">
          <p>The authors declare no conflicts of interest.</p>
        </fn>
      </author-notes>
      <pub-date pub-type="epub">
        <day>28</day>
        <month>02</month>
        <year>2026</year>
      </pub-date>
      <pub-date pub-type="collection">
        <month>02</month>
        <year>2026</year>
      </pub-date>
      <volume>13</volume>
      <issue>03</issue>
      <fpage>1</fpage>
      <lpage>26</lpage>
      <history>
        <date date-type="received">
          <day>23</day>
          <month>01</month>
          <year>2026</year>
        </date>
        <date date-type="accepted">
          <day>13</day>
          <month>03</month>
          <year>2026</year>
        </date>
        <date date-type="published">
          <day>16</day>
          <month>03</month>
          <year>2026</year>
        </date>
      </history>
      <permissions>
        <copyright-statement>© 2026 by the authors and Scientific Research Publishing Inc.</copyright-statement>
        <copyright-year>2026</copyright-year>
        <license license-type="open-access">
          <license-p> This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license ( <ext-link ext-link-type="uri" xlink:href="https://creativecommons.org/licenses/by/4.0/">https://creativecommons.org/licenses/by/4.0/</ext-link> ). </license-p>
        </license>
      </permissions>
      <self-uri content-type="doi" xlink:href="https://doi.org/10.4236/oalib.1114926">https://doi.org/10.4236/oalib.1114926</self-uri>
      <abstract>
        <p>Bipolar disorder (BD) is closely intertwined with abnormalities in sleep and circadian regulation, yet current clinical management typically applies heuristic rules rather than optimizing these interacting processes in a principled way. We present a reinforcement-learning (RL) framework that learns personalized interventions for sleep timing, light exposure, daily activity, and medication adherence in a simulated BD setting. We design Circadian Environment, a physiologically inspired Markov decision process with five clinically interpretable state variables: sleep quality, sleep duration, mood stability, circadian alignment, and stress level. A continuous four-dimensional action space encodes modifiable behavioural and pharmacological levers. A Proximal Policy Optimization (PPO) agent (PPO-Patient-Agent), implemented in PyTorch with Gaussian policies, LayerNorm, entropy regularization, and gradient clipping, is trained over 1000 episodes (30 simulated days each) to maximize a composite reward reflecting key therapeutic goals: stable mood, high-quality and near-optimal sleep, strong circadian alignment, and low stress. Across training, episodic returns improve steadily and converge, indicating a stable policy. When evaluated in multiple post-training rollouts, the learned policy reliably drives virtual patients from moderately dysregulated baseline states into a high-functioning attractor characterized by near-maximal mood stability and sleep quality, robust circadian alignment, and minimal stress. Analysis of state trajectories reveals strong positive coupling between sleep, mood, and circadian variables and strong negative coupling between these variables and stress, aligning with clinical intuition. Although the current work uses a stylized simulator rather than real patient data, it establishes a transparent and extensible sandbox for prototyping RL-based treatment strategies in BD. The same architecture can be calibrated with digital phenotyping and longitudinal clinical data to support future development of safe, personalized decision-support tools for mood stabilization.</p>
      </abstract>
      <kwd-group kwd-group-type="author-generated" xml:lang="en">
        <kwd>Bipolar Disorder</kwd>
        <kwd>Reinforcement Learning</kwd>
        <kwd>Proximal Policy Optimization</kwd>
        <kwd>Sleep</kwd>
        <kwd>Circadian Rhythm</kwd>
        <kwd>Digital Psychiatry</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec1">
      <title>1. Introduction</title>
      <p>Bipolar disorder (BD) is characterized by recurrent episodes of depression and mania or hypomania, and a growing body of evidence identifies sleep and circadian rhythm disturbances as core drivers of episode onset, relapse, and symptom severity [<xref ref-type="bibr" rid="B1">1</xref>]-[<xref ref-type="bibr" rid="B10">10</xref>]. Irregular sleep wake patterns misaligned circadian phases, and elevated stress levels can destabilize mood, while behavioural regularity, morning light exposure, stable routines, and consistent medication adherence are known to support mood stabilization [<xref ref-type="bibr" rid="B11">11</xref>]-[<xref ref-type="bibr" rid="B14">14</xref>]. Although these mechanisms are well understood, clinical practice typically implements them through broad heuristic guidelines such as “maintain sleep regularity” or “increase morning light exposure” rather than through personalized and dynamically optimized intervention strategies.</p>
      <p>Reinforcement learning (RL) provides a mathematically principled framework for sequential decision-making under uncertainty [<xref ref-type="bibr" rid="B15">15</xref>]-[<xref ref-type="bibr" rid="B20">20</xref>]. An RL agent learns to select interventions based on an evolving patient state and long-term therapeutic outcomes, rather than static rules or one-step heuristics [<xref ref-type="bibr" rid="B21">21</xref>]-[<xref ref-type="bibr" rid="B25">25</xref>]. However, deploying RL directly in real clinical settings is constrained by safety considerations, ethical requirements, limited longitudinal data, and the unpredictable consequences of exploratory actions [<xref ref-type="bibr" rid="B26">26</xref>]-[<xref ref-type="bibr" rid="B30">30</xref>]. This motivates the development of realistic simulation environments that model key physiological and behavioural processes relevant to BD, allowing policies to be developed, analysed, and stress-tested before clinical translation. To address this gap, we introduce a fully simulated and physiologically inspired RL system for optimizing sleep-mood-circadian dynamics in bipolar disorder. Our framework couples a continuous-state circadian environment with a customized PPO-based control agent, enabling the evaluation of personalized intervention strategies in a safe and controllable setting.</p>
      <p><bold>Contributions:</bold>This work provides four key contributions:</p>
      <p><bold>1)</bold><bold>Circadian Environment</bold>: A continuous-state, continuous-action simulation environment that models the interacting dynamics of sleep quality, sleep duration, mood stability, circadian alignment, and stress level. The environment is designed with interpretable variables and biologically motivated transition equations, enabling transparent connection to known BD physiology.</p>
      <p><bold>2)</bold><bold>PPO-Patient-Agent</bold></p>
      <p>A Proximal Policy Optimization (PPO)-based treatment agent using Gaussian action distributions with uncertainty control, LayerNorm stabilization, entropy regularization, and gradient clipping. This architecture ensures stable exploration and policy improvement in a complex, nonlinear physiological system.</p>
      <p><bold>3)</bold><bold>Clinical Simulator and Advanced Visualizer</bold>: A full analysis pipeline that:</p>
      <p>visualizes RL training behaviour (<xref ref-type="fig" rid="fig1">Figures 1(a)-(f)</xref>),evaluates the learned policy across multiple simulated clinical trials (<xref ref-type="fig" rid="fig2">Figures 2(a)-(f)</xref>),and characterizes underlying circadian structure through heatmaps, phase portraits, correlation matrices, and 3-D state trajectories (<xref ref-type="fig" rid="fig3">Figures 3(a)-(d)</xref>).</p>
      <p>These tools provide detailed insight into both learning dynamics and physiological interpretability.</p>
      <p><bold>4)</bold><bold>Reproducibility Tables</bold>: Comprehensive tables documenting:</p>
      <p>state and action space definitions,reward components,model hyperparameters, ensuring transparency and facilitating future extensions, benchmarking, and real-world calibration.</p>
    </sec>
    <sec id="sec2">
      <title>2. Methods</title>
      <sec id="sec2dot1">
        <title>2.1. Circadian Environment</title>
        <p>The environment models a single BD patient over a 30-day horizon. The state at day <inline-formula><mml:math><mml:mi> t </mml:mi></mml:math></inline-formula> is</p>
        <disp-formula id="FD1">
          <mml:math>
            <mml:mrow>
              <mml:msub>
                <mml:mi>s</mml:mi>
                <mml:mi>t</mml:mi>
              </mml:msub>
              <mml:mo>=</mml:mo>
              <mml:mrow>
                <mml:mo>[</mml:mo>
                <mml:mrow>
                  <mml:msub>
                    <mml:mi>q</mml:mi>
                    <mml:mi>t</mml:mi>
                  </mml:msub>
                  <mml:mo>,</mml:mo>
                  <mml:msub>
                    <mml:mi>d</mml:mi>
                    <mml:mi>t</mml:mi>
                  </mml:msub>
                  <mml:mo>,</mml:mo>
                  <mml:msub>
                    <mml:mi>m</mml:mi>
                    <mml:mi>t</mml:mi>
                  </mml:msub>
                  <mml:mo>,</mml:mo>
                  <mml:msub>
                    <mml:mi>c</mml:mi>
                    <mml:mi>t</mml:mi>
                  </mml:msub>
                  <mml:mo>,</mml:mo>
                  <mml:msub>
                    <mml:mi>σ</mml:mi>
                    <mml:mi>t</mml:mi>
                  </mml:msub>
                </mml:mrow>
                <mml:mo>]</mml:mo>
              </mml:mrow>
            </mml:mrow>
          </mml:math>
        </disp-formula>
        <p>where:</p>
        <p><inline-formula><mml:math><mml:mrow><mml:msub><mml:mi> q </mml:mi><mml:mi> t </mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> : sleep quality (0 - 1),<inline-formula><mml:math><mml:mrow><mml:msub><mml:mi> d </mml:mi><mml:mi> t </mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> : sleep duration in hours (4 - 10),<inline-formula><mml:math><mml:mrow><mml:msub><mml:mi> m </mml:mi><mml:mi> t </mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> : mood stability (0 - 1),<inline-formula><mml:math><mml:mrow><mml:msub><mml:mi> c </mml:mi><mml:mi> t </mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> : circadian alignment (0 - 1),<inline-formula><mml:math><mml:mrow><mml:msub><mml:mi> σ </mml:mi><mml:mi> t </mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> : stress level (0 - 1).</p>
        <p>The initial state is moderately dysregulated:</p>
        <disp-formula id="FD2">
          <mml:math>
            <mml:mrow>
              <mml:msub>
                <mml:mi>s</mml:mi>
                <mml:mn>0</mml:mn>
              </mml:msub>
              <mml:mo>=</mml:mo>
              <mml:mrow>
                <mml:mo>[</mml:mo>
                <mml:mrow>
                  <mml:mn>0.5</mml:mn>
                  <mml:mo>,</mml:mo>
                  <mml:mn>7.0</mml:mn>
                  <mml:mo>,</mml:mo>
                  <mml:mn>0.5</mml:mn>
                  <mml:mo>,</mml:mo>
                  <mml:mn>0.5</mml:mn>
                  <mml:mo>,</mml:mo>
                  <mml:mn>0.3</mml:mn>
                </mml:mrow>
                <mml:mo>]</mml:mo>
              </mml:mrow>
              <mml:mo>.</mml:mo>
            </mml:mrow>
          </mml:math>
        </disp-formula>
        <p>The initial state vector <italic>s</italic><sub>0</sub> was chosen to represent a moderately dysregulated baseline consistent with subthreshold mood instability or early relapse risk rather than acute mania or severe depression. Mood stability ≈ 0.5 and circadian alignment ≈ 0.5 reflect partial destabilization frequently observed in maintenance phases of BD. This choice allows evaluation of stabilization capacity rather than crisis intervention.</p>
        <p>2.1.1. Action Space</p>
        <p>At each step the agent chooses a 4-dimensional continuous action</p>
        <disp-formula id="FD3">
          <mml:math>
            <mml:mrow>
              <mml:msub>
                <mml:mi>a</mml:mi>
                <mml:mi>t</mml:mi>
              </mml:msub>
              <mml:mo>=</mml:mo>
              <mml:mrow>
                <mml:mo>[</mml:mo>
                <mml:mrow>
                  <mml:msub>
                    <mml:mtext>Δ</mml:mtext>
                    <mml:mrow>
                      <mml:mtext>sleep</mml:mtext>
                    </mml:mrow>
                  </mml:msub>
                  <mml:mo>,</mml:mo>
                  <mml:msub>
                    <mml:mi>ℓ</mml:mi>
                    <mml:mi>t</mml:mi>
                  </mml:msub>
                  <mml:mo>,</mml:mo>
                  <mml:msubsup>
                    <mml:mi>a</mml:mi>
                    <mml:mi>t</mml:mi>
                    <mml:mrow>
                      <mml:mtext>act</mml:mtext>
                    </mml:mrow>
                  </mml:msubsup>
                  <mml:mo>,</mml:mo>
                  <mml:msub>
                    <mml:mi>μ</mml:mi>
                    <mml:mi>t</mml:mi>
                  </mml:msub>
                </mml:mrow>
                <mml:mo>]</mml:mo>
              </mml:mrow>
              <mml:mo>∈</mml:mo>
              <mml:msup>
                <mml:mrow>
                  <mml:mrow>
                    <mml:mo>[</mml:mo>
                    <mml:mrow>
                      <mml:mo>−</mml:mo>
                      <mml:mn>1</mml:mn>
                      <mml:mo>,</mml:mo>
                      <mml:mn>1</mml:mn>
                    </mml:mrow>
                    <mml:mo>]</mml:mo>
                  </mml:mrow>
                </mml:mrow>
                <mml:mn>4</mml:mn>
              </mml:msup>
            </mml:mrow>
          </mml:math>
        </disp-formula>
        <p>encoding:</p>
        <p>1) Sleep time adjustment <inline-formula><mml:math><mml:mrow><mml:msub><mml:mtext> Δ </mml:mtext><mml:mrow><mml:mtext> sleep </mml:mtext></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula> ,</p>
        <p>2) Light exposure <inline-formula><mml:math><mml:mrow><mml:msub><mml:mi> ℓ </mml:mi><mml:mi> t </mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> ,</p>
        <p>3) Activity level <inline-formula><mml:math><mml:mrow><mml:msubsup><mml:mi> a </mml:mi><mml:mi> t </mml:mi><mml:mrow><mml:mtext> act </mml:mtext></mml:mrow></mml:msubsup></mml:mrow></mml:math></inline-formula> ,</p>
        <p>4) Medication adherence <inline-formula><mml:math><mml:mrow><mml:msub><mml:mi> μ </mml:mi><mml:mi> t </mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> .</p>
        <p><bold>Table 1</bold> summarizes all variables.</p>
        <p>Table 1. State and action variables in Circadian Environment.</p>
        <table-wrap id="tbl1">
          <label>Table 1</label>
          <table>
            <tbody>
              <tr>
                <td>
                  <bold>Type</bold>
                </td>
                <td>
                  <bold>Name</bold>
                </td>
                <td>
                  <bold>Symbol</bold>
                </td>
                <td>
                  <bold>Range</bold>
                </td>
                <td>
                  <bold>Interpretation</bold>
                </td>
              </tr>
              <tr>
                <td>State</td>
                <td>Sleep quality</td>
                <td>
                  <inline-formula>
                    <mml:math>
                      <mml:mrow>
                        <mml:msub>
                          <mml:mi>q</mml:mi>
                          <mml:mi>t</mml:mi>
                        </mml:msub>
                      </mml:mrow>
                    </mml:math>
                  </inline-formula>
                </td>
                <td>[0, 1]</td>
                <td>Restorative value of sleep</td>
              </tr>
              <tr>
                <td>State</td>
                <td>Sleep duration (h)</td>
                <td>
                  <inline-formula>
                    <mml:math>
                      <mml:mrow>
                        <mml:msub>
                          <mml:mi>d</mml:mi>
                          <mml:mi>t</mml:mi>
                        </mml:msub>
                      </mml:mrow>
                    </mml:math>
                  </inline-formula>
                </td>
                <td>[4, 10]</td>
                <td>Hours of sleep per night</td>
              </tr>
              <tr>
                <td>State</td>
                <td>Mood stability</td>
                <td>
                  <inline-formula>
                    <mml:math>
                      <mml:mrow>
                        <mml:msub>
                          <mml:mi>m</mml:mi>
                          <mml:mi>t</mml:mi>
                        </mml:msub>
                      </mml:mrow>
                    </mml:math>
                  </inline-formula>
                </td>
                <td>[0, 1]</td>
                <td>Proximity to euthymic mood</td>
              </tr>
              <tr>
                <td>State</td>
                <td>Circadian alignment</td>
                <td>
                  <inline-formula>
                    <mml:math>
                      <mml:mrow>
                        <mml:msub>
                          <mml:mi>c</mml:mi>
                          <mml:mi>t</mml:mi>
                        </mml:msub>
                      </mml:mrow>
                    </mml:math>
                  </inline-formula>
                </td>
                <td>[0, 1]</td>
                <td>Alignment to target circadian phase</td>
              </tr>
              <tr>
                <td>State</td>
                <td>Stress level</td>
                <td>
                  <inline-formula>
                    <mml:math>
                      <mml:mrow>
                        <mml:msub>
                          <mml:mi>σ</mml:mi>
                          <mml:mi>t</mml:mi>
                        </mml:msub>
                      </mml:mrow>
                    </mml:math>
                  </inline-formula>
                </td>
                <td>[0, 1]</td>
                <td>Psychological/physiological stress</td>
              </tr>
              <tr>
                <td>Action</td>
                <td>Sleep time change</td>
                <td>
                  <inline-formula>
                    <mml:math>
                      <mml:mrow>
                        <mml:msub>
                          <mml:mtext>Δ</mml:mtext>
                          <mml:mrow>
                            <mml:mtext>sleep</mml:mtext>
                          </mml:mrow>
                        </mml:msub>
                      </mml:mrow>
                    </mml:math>
                  </inline-formula>
                </td>
                <td>[−1, 1]</td>
                <td>Phase advance/delay</td>
              </tr>
              <tr>
                <td>Action</td>
                <td>Light exposure</td>
                <td>
                  <inline-formula>
                    <mml:math>
                      <mml:mrow>
                        <mml:msub>
                          <mml:mi>ℓ</mml:mi>
                          <mml:mi>t</mml:mi>
                        </mml:msub>
                      </mml:mrow>
                    </mml:math>
                  </inline-formula>
                </td>
                <td>[−1, 1]</td>
                <td>Strength/timing of light therapy</td>
              </tr>
              <tr>
                <td>Action</td>
                <td>Activity level</td>
                <td>
                  <inline-formula>
                    <mml:math>
                      <mml:mrow>
                        <mml:msubsup>
                          <mml:mi>a</mml:mi>
                          <mml:mi>t</mml:mi>
                          <mml:mrow>
                            <mml:mtext>act</mml:mtext>
                          </mml:mrow>
                        </mml:msubsup>
                      </mml:mrow>
                    </mml:math>
                  </inline-formula>
                </td>
                <td>[−1, 1]</td>
                <td>Daily physical/social activation</td>
              </tr>
              <tr>
                <td>Action</td>
                <td>Medication adherence</td>
                <td>
                  <inline-formula>
                    <mml:math>
                      <mml:mrow>
                        <mml:msub>
                          <mml:mi>μ</mml:mi>
                          <mml:mi>t</mml:mi>
                        </mml:msub>
                      </mml:mrow>
                    </mml:math>
                  </inline-formula>
                </td>
                <td>[−1, 1]</td>
                <td>Regularity and adherence to medication</td>
              </tr>
            </tbody>
          </table>
        </table-wrap>
        <p>The medication adherence action μₜ is modelled in the continuous range [−1, 1] to maintain symmetry with other intervention variables in the Gaussian action space. Positive values correspond to consistent adherence (<inline-formula><mml:math display="inline"><mml:mrow><mml:msub><mml:mi> μ </mml:mi><mml:mi> t </mml:mi></mml:msub><mml:mo> ≈ </mml:mo><mml:mn> 1 </mml:mn></mml:mrow></mml:math></inline-formula> indicating full regularity), values near 0 indicate neutral or baseline adherence, while negative values represent irregular or missed medication patterns that may destabilize mood or increase stress. This symmetric scaling facilitates stable PPO optimization while preserving clinical interpretability.</p>
        <p>2.1.2. Transition Dynamics</p>
        <p>The step () function implements deterministic, hand-crafted dynamics:</p>
        <p>Sleep quality <inline-formula><mml:math><mml:mrow><mml:msub><mml:mi> q </mml:mi><mml:mrow><mml:mi> t </mml:mi><mml:mo> + </mml:mo><mml:mn> 1 </mml:mn></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula> improves with consistent sleep timing and beneficial light but worsens with stress:consistency term <inline-formula><mml:math><mml:mrow><mml:mn> 1 </mml:mn><mml:mo> − </mml:mo><mml:mn> 0.1 </mml:mn><mml:mrow><mml:mo> | </mml:mo><mml:mrow><mml:msub><mml:mtext> Δ </mml:mtext><mml:mrow><mml:mtext> sleep </mml:mtext></mml:mrow></mml:msub></mml:mrow><mml:mo> | </mml:mo></mml:mrow></mml:mrow></mml:math></inline-formula> ,light impact <inline-formula><mml:math><mml:mrow><mml:mn> 0.2 </mml:mn><mml:msub><mml:mi> ℓ </mml:mi><mml:mi> t </mml:mi></mml:msub><mml:mo> − </mml:mo><mml:mn> 0.1 </mml:mn></mml:mrow></mml:math></inline-formula> ,penalty proportional to <inline-formula><mml:math><mml:mrow><mml:msub><mml:mi> σ </mml:mi><mml:mi> t </mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> .Sleep duration:</p>
        <disp-formula id="FD4">
          <mml:math>
            <mml:mrow>
              <mml:msub>
                <mml:mi>d</mml:mi>
                <mml:mrow>
                  <mml:mi>t</mml:mi>
                  <mml:mo>+</mml:mo>
                  <mml:mn>1</mml:mn>
                </mml:mrow>
              </mml:msub>
              <mml:mo>=</mml:mo>
              <mml:mtext>clip</mml:mtext>
              <mml:mrow>
                <mml:mo>(</mml:mo>
                <mml:mrow>
                  <mml:msub>
                    <mml:mi>d</mml:mi>
                    <mml:mi>t</mml:mi>
                  </mml:msub>
                  <mml:mo>+</mml:mo>
                  <mml:mn>0.5</mml:mn>
                  <mml:msub>
                    <mml:mtext>Δ</mml:mtext>
                    <mml:mrow>
                      <mml:mtext>sleep</mml:mtext>
                    </mml:mrow>
                  </mml:msub>
                  <mml:mo>−</mml:mo>
                  <mml:mn>0.2</mml:mn>
                  <mml:msubsup>
                    <mml:mi>a</mml:mi>
                    <mml:mi>t</mml:mi>
                    <mml:mrow>
                      <mml:mtext>act</mml:mtext>
                    </mml:mrow>
                  </mml:msubsup>
                  <mml:mo>,</mml:mo>
                  <mml:mn>4</mml:mn>
                  <mml:mo>,</mml:mo>
                  <mml:mn>10</mml:mn>
                </mml:mrow>
                <mml:mo>)</mml:mo>
              </mml:mrow>
              <mml:mo>.</mml:mo>
            </mml:mrow>
          </mml:math>
        </disp-formula>
        <p>Mood stability is updated using a linear combination of:sleep quality and deviation from 7 h duration,circadian alignment <inline-formula><mml:math><mml:mrow><mml:msub><mml:mi> c </mml:mi><mml:mi> t </mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> ,medication adherence <inline-formula><mml:math><mml:mrow><mml:msub><mml:mi> μ </mml:mi><mml:mi> t </mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> ,negative stress influence.</p>
        <p>The resulting change is added to <inline-formula><mml:math><mml:mrow><mml:msub><mml:mi> m </mml:mi><mml:mi> t </mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> and clipped to [0, 1].</p>
        <p>Circadian alignment increases with light and medication and decreases with large sleep shifts.Stress rises with abrupt sleep changes and high activity but falls with good medication adherence.</p>
        <p>Episodes last exactly 30 steps; there is no early termination in this version. The numerical coefficients used in the transition equations (e.g., light impact = 0.2, sleep deviation weight = 0.5) were selected to reflect qualitative clinical influence while preserving dynamical stability. Empirical studies suggest moderate phase-shifting effects of light exposure and strong bidirectional coupling between sleep disruption and mood instability in bipolar disorder. Accordingly, sleep-related effects were weighted more strongly than light exposure alone. Coefficients were scaled to maintain bounded trajectories within [0, 1] and to avoid oscillatory instability. These values are not yet calibrated to empirical patient-level data and should be interpreted as physiologically inspired but heuristic parameters.</p>
        <p>2.1.3. Reward Function</p>
        <p>The reward combines therapeutic objectives:</p>
        <disp-formula id="FD5">
          <mml:math>
            <mml:mrow>
              <mml:msub>
                <mml:mi>r</mml:mi>
                <mml:mi>t</mml:mi>
              </mml:msub>
              <mml:mo>=</mml:mo>
              <mml:msub>
                <mml:mi>r</mml:mi>
                <mml:mrow>
                  <mml:mtext>mood</mml:mtext>
                </mml:mrow>
              </mml:msub>
              <mml:mo>+</mml:mo>
              <mml:msub>
                <mml:mi>r</mml:mi>
                <mml:mrow>
                  <mml:mtext>sleep-q</mml:mtext>
                </mml:mrow>
              </mml:msub>
              <mml:mo>+</mml:mo>
              <mml:msub>
                <mml:mi>r</mml:mi>
                <mml:mrow>
                  <mml:mtext>sleep-d</mml:mtext>
                </mml:mrow>
              </mml:msub>
              <mml:mo>+</mml:mo>
              <mml:msub>
                <mml:mi>r</mml:mi>
                <mml:mrow>
                  <mml:mtext>circadian</mml:mtext>
                </mml:mrow>
              </mml:msub>
              <mml:mo>+</mml:mo>
              <mml:msub>
                <mml:mi>r</mml:mi>
                <mml:mrow>
                  <mml:mtext>stress</mml:mtext>
                </mml:mrow>
              </mml:msub>
              <mml:mo>.</mml:mo>
            </mml:mrow>
          </mml:math>
        </disp-formula>
        <p><inline-formula><mml:math><mml:mrow><mml:msub><mml:mi> r </mml:mi><mml:mrow><mml:mtext> mood </mml:mtext></mml:mrow></mml:msub><mml:mo> = </mml:mo><mml:mn> 2 </mml:mn><mml:mrow><mml:mo> ( </mml:mo><mml:mrow><mml:msub><mml:mi> m </mml:mi><mml:mi> t </mml:mi></mml:msub><mml:mo> − </mml:mo><mml:mn> 0.3 </mml:mn></mml:mrow><mml:mo> ) </mml:mo></mml:mrow></mml:mrow></mml:math></inline-formula> rewards mood stability above 0.3.<inline-formula><mml:math><mml:mrow><mml:msub><mml:mi> r </mml:mi><mml:mrow><mml:mtext> sleep-q </mml:mtext></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula> gives a high bonus when <inline-formula><mml:math><mml:mrow><mml:mn> 0.6 </mml:mn><mml:mo> ≤ </mml:mo><mml:msub><mml:mi> q </mml:mi><mml:mi> t </mml:mi></mml:msub><mml:mo> ≤ </mml:mo><mml:mn> 0.9 </mml:mn></mml:mrow></mml:math></inline-formula> , penalty for <inline-formula><mml:math><mml:mrow><mml:msub><mml:mi> q </mml:mi><mml:mi> t </mml:mi></mml:msub><mml:mo> &lt; </mml:mo><mml:mn> 0.3 </mml:mn></mml:mrow></mml:math></inline-formula> .<inline-formula><mml:math><mml:mrow><mml:msub><mml:mi> r </mml:mi><mml:mrow><mml:mtext> sleep-d </mml:mtext></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula> rewards durations between 7 - 8 h and penalizes &lt;6 h or &gt;9 h.<inline-formula><mml:math><mml:mrow><mml:msub><mml:mi> r </mml:mi><mml:mrow><mml:mtext> circadian </mml:mtext></mml:mrow></mml:msub><mml:mo> ∝ </mml:mo><mml:msub><mml:mi> c </mml:mi><mml:mi> t </mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> .<inline-formula><mml:math><mml:mrow><mml:msub><mml:mi> r </mml:mi><mml:mrow><mml:mtext> stress </mml:mtext></mml:mrow></mml:msub><mml:mo> = </mml:mo><mml:mo> − </mml:mo><mml:mn> 0.5 </mml:mn><mml:msub><mml:mi> σ </mml:mi><mml:mi> t </mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> .</p>
        <p><bold>Table 2</bold> summarizes these components qualitatively.</p>
        <p>Table 2. Reward components.</p>
        <table-wrap id="tbl2">
          <label>Table 2</label>
          <table>
            <tbody>
              <tr>
                <td>
                  <bold>Component</bold>
                </td>
                <td>
                  <bold>Depends on</bold>
                </td>
                <td>
                  <bold>Role</bold>
                </td>
              </tr>
              <tr>
                <td>Mood reward</td>
                <td>
                  <inline-formula>
                    <mml:math>
                      <mml:mrow>
                        <mml:msub>
                          <mml:mi>m</mml:mi>
                          <mml:mi>t</mml:mi>
                        </mml:msub>
                      </mml:mrow>
                    </mml:math>
                  </inline-formula>
                </td>
                <td>Encourages stable euthymic mood</td>
              </tr>
              <tr>
                <td>Sleep quality</td>
                <td>
                  <inline-formula>
                    <mml:math>
                      <mml:mrow>
                        <mml:msub>
                          <mml:mi>q</mml:mi>
                          <mml:mi>t</mml:mi>
                        </mml:msub>
                      </mml:mrow>
                    </mml:math>
                  </inline-formula>
                </td>
                <td>Rewards restorative sleep, penalizes poor</td>
              </tr>
              <tr>
                <td>Sleep duration</td>
                <td>
                  <inline-formula>
                    <mml:math>
                      <mml:mrow>
                        <mml:msub>
                          <mml:mi>d</mml:mi>
                          <mml:mi>t</mml:mi>
                        </mml:msub>
                      </mml:mrow>
                    </mml:math>
                  </inline-formula>
                </td>
                <td>Favors near-optimal 7 - 8 h window</td>
              </tr>
              <tr>
                <td>Circadian alignment</td>
                <td>
                  <inline-formula>
                    <mml:math>
                      <mml:mrow>
                        <mml:msub>
                          <mml:mi>c</mml:mi>
                          <mml:mi>t</mml:mi>
                        </mml:msub>
                      </mml:mrow>
                    </mml:math>
                  </inline-formula>
                </td>
                <td>Promotes strong circadian regularity</td>
              </tr>
              <tr>
                <td>Stress penalty</td>
                <td>
                  <inline-formula>
                    <mml:math>
                      <mml:mrow>
                        <mml:msub>
                          <mml:mi>σ</mml:mi>
                          <mml:mi>t</mml:mi>
                        </mml:msub>
                      </mml:mrow>
                    </mml:math>
                  </inline-formula>
                </td>
                <td>Penalizes sustained high stress</td>
              </tr>
            </tbody>
          </table>
        </table-wrap>
      </sec>
      <sec id="sec2dot2">
        <title>2.2. PPO-Patient-Agent</title>
        <p>The PPO agent uses separate actor and critic networks implemented in PyTorch.</p>
        <p>2.2.1. Actor Network</p>
        <p>Shared backbone:Linear (5 → 256) → LayerNorm → ReLULinear (256 → 256) → LayerNorm → ReLULinear (256 → 128) → ReLUOutputs:mean head → Linear (128 → 4) → tanh,log-std head → Linear (128 → 4), clamped to [−20, 2].</p>
        <p>A multivariate Normal distribution <inline-formula><mml:math><mml:mrow><mml:mi mathvariant="script"> N </mml:mi><mml:mrow><mml:mo> ( </mml:mo><mml:mrow><mml:mi> μ </mml:mi><mml:mo> , </mml:mo><mml:msup><mml:mi> σ </mml:mi><mml:mn> 2 </mml:mn></mml:msup></mml:mrow><mml:mo> ) </mml:mo></mml:mrow></mml:mrow></mml:math></inline-formula> is created from the mean and log-std; actions are clipped to [−1, 1].</p>
        <p>2.2.2. Critic Network</p>
        <p>The critic mirrors the backbone and outputs a single scalar value <inline-formula><mml:math><mml:mrow><mml:mi> V </mml:mi><mml:mrow><mml:mo> ( </mml:mo><mml:mi> s </mml:mi><mml:mo> ) </mml:mo></mml:mrow></mml:mrow></mml:math></inline-formula> . Both networks use orthogonal initialization with small gains for the actor.</p>
        <p>2.2.3. PPO Loss and Optimization</p>
        <p>Discounted returns <inline-formula><mml:math><mml:mrow><mml:msub><mml:mi> G </mml:mi><mml:mi> t </mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> are computed with <inline-formula><mml:math><mml:mrow><mml:mi> γ </mml:mi><mml:mo> = </mml:mo><mml:mn> 0.99 </mml:mn></mml:mrow></mml:math></inline-formula> . Advantages are <inline-formula><mml:math><mml:mrow><mml:msub><mml:mi> A </mml:mi><mml:mi> t </mml:mi></mml:msub><mml:mo> = </mml:mo><mml:msub><mml:mi> G </mml:mi><mml:mi> t </mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> (baseline-free in this version) and normalized. The PPO clipped surrogate is</p>
        <disp-formula id="FD6">
          <mml:math>
            <mml:mrow>
              <mml:msub>
                <mml:mi>L</mml:mi>
                <mml:mrow>
                  <mml:mtext>actor</mml:mtext>
                </mml:mrow>
              </mml:msub>
              <mml:mo>=</mml:mo>
              <mml:mo>−</mml:mo>
              <mml:msub>
                <mml:mi mathvariant="double-struck">E</mml:mi>
                <mml:mi>t</mml:mi>
              </mml:msub>
              <mml:mrow>
                <mml:mo>[</mml:mo>
                <mml:mrow>
                  <mml:mi>min</mml:mi>
                  <mml:mrow>
                    <mml:mo>(</mml:mo>
                    <mml:mrow>
                      <mml:msub>
                        <mml:mi>r</mml:mi>
                        <mml:mi>t</mml:mi>
                      </mml:msub>
                      <mml:msub>
                        <mml:mi>A</mml:mi>
                        <mml:mi>t</mml:mi>
                      </mml:msub>
                      <mml:mo>,</mml:mo>
                      <mml:mtext>clip</mml:mtext>
                      <mml:mrow>
                        <mml:mo>(</mml:mo>
                        <mml:mrow>
                          <mml:msub>
                            <mml:mi>r</mml:mi>
                            <mml:mi>t</mml:mi>
                          </mml:msub>
                          <mml:mo>,</mml:mo>
                          <mml:mn>1</mml:mn>
                          <mml:mo>−</mml:mo>
                          <mml:mi>ϵ</mml:mi>
                          <mml:mo>,</mml:mo>
                          <mml:mn>1</mml:mn>
                          <mml:mo>+</mml:mo>
                          <mml:mi>ϵ</mml:mi>
                        </mml:mrow>
                        <mml:mo>)</mml:mo>
                      </mml:mrow>
                      <mml:msub>
                        <mml:mi>A</mml:mi>
                        <mml:mi>t</mml:mi>
                      </mml:msub>
                    </mml:mrow>
                    <mml:mo>)</mml:mo>
                  </mml:mrow>
                </mml:mrow>
                <mml:mo>]</mml:mo>
              </mml:mrow>
              <mml:mo>−</mml:mo>
              <mml:mi>β</mml:mi>
              <mml:mi>H</mml:mi>
              <mml:mrow>
                <mml:mo>[</mml:mo>
                <mml:mrow>
                  <mml:mi>π</mml:mi>
                  <mml:mrow>
                    <mml:mo>(</mml:mo>
                    <mml:mrow>
                      <mml:mo>⋅</mml:mo>
                      <mml:mo>|</mml:mo>
                      <mml:msub>
                        <mml:mi>s</mml:mi>
                        <mml:mi>t</mml:mi>
                      </mml:msub>
                    </mml:mrow>
                    <mml:mo>)</mml:mo>
                  </mml:mrow>
                </mml:mrow>
                <mml:mo>]</mml:mo>
              </mml:mrow>
              <mml:mo>,</mml:mo>
            </mml:mrow>
          </mml:math>
        </disp-formula>
        <p>with <inline-formula><mml:math><mml:mrow><mml:mi> ϵ </mml:mi><mml:mo> = </mml:mo><mml:mn> 0.2 </mml:mn></mml:mrow></mml:math></inline-formula> and entropy weight <inline-formula><mml:math><mml:mrow><mml:mi> β </mml:mi><mml:mo> = </mml:mo><mml:mn> 0.01 </mml:mn></mml:mrow></mml:math></inline-formula> . The critic is trained with MSE between predicted values and returns. Gradient norms are clipped at 0.5.</p>
        <p><bold>Table 3</bold> lists the core hyperparameters.</p>
        <p>Table 3. PPO hyperparameters.</p>
        <table-wrap id="tbl3">
          <label>Table 3</label>
          <table>
            <tbody>
              <tr>
                <td>
                  <bold>Parameter</bold>
                </td>
                <td>
                  <bold>Value</bold>
                </td>
                <td>
                  <bold>Reason</bold>
                </td>
              </tr>
              <tr>
                <td>
                  Discount Factor (
                  <inline-formula>
                    <mml:math>
                      <mml:mi>γ</mml:mi>
                    </mml:math>
                  </inline-formula>
                  )
                </td>
                <td>0.99</td>
                <td>Captures long-term mood and sleep effects</td>
              </tr>
              <tr>
                <td>Learning Rate</td>
                <td>3e-4</td>
                <td>Stable for PPO with Layer Norm</td>
              </tr>
              <tr>
                <td>
                  PPO Clip Range (
                  <inline-formula>
                    <mml:math>
                      <mml:mi>ϵ</mml:mi>
                    </mml:math>
                  </inline-formula>
                  )
                </td>
                <td>0.2</td>
                <td>Prevents harmful policy jumps</td>
              </tr>
              <tr>
                <td>Entropy Weight</td>
                <td>0.01</td>
                <td>Maintains healthy exploration</td>
              </tr>
              <tr>
                <td>Training Epochs/Episode</td>
                <td>10</td>
                <td>Ensures thorough update per trajectory</td>
              </tr>
              <tr>
                <td>Gradient Clipping</td>
                <td>0.5</td>
                <td>Prevents unstable training</td>
              </tr>
              <tr>
                <td>Hidden Layer Sizes</td>
                <td>256-256-128</td>
                <td>Balanced capacity-stability trade-off</td>
              </tr>
            </tbody>
          </table>
        </table-wrap>
      </sec>
      <sec id="sec2dot3">
        <title>2.3. Clinical Simulator and Advanced Visualizer</title>
        <p>Clinical Simulator orchestrates training for 1000 episodes and stores training metrics: episodic returns, actor and critic losses, policy entropy, and episode lengths. After training, the simulator runs five deterministic clinical trials. For each trial it:</p>
        <p>resets the environment,rolls out the deterministic policy (actor mean),stores full state trajectories, final reward, final mood and final sleep quality.</p>
        <p><bold>Advanced Visualizer</bold> creates three composite figures:</p>
        <p><xref ref-type="fig" rid="fig1">Figures 1(a)-(f)</xref>: RL training performance.<xref ref-type="fig" rid="fig2">Figures 2(a)-(f)</xref>: clinical trial trajectories and final outcomes.<xref ref-type="fig" rid="fig3">Figures 3(a)-(d)</xref>: circadian analyses.</p>
        <p>Detailed panel definitions are given below with the Results.</p>
      </sec>
    </sec>
    <sec id="sec3">
      <title>3. Results</title>
      <sec id="sec3dot1">
        <title>
          3.1. RL Training Performance (
          <xref ref-type="fig" rid="fig1">Figure 1</xref>
          )
        </title>
        <fig id="fig1">
          <label>Figure 1</label>
          <graphic xlink:href="https://html.scirp.org/file/1114926-rId124.jpeg?20260316023136" />
        </fig>
        <p>Figure 1. Training dynamics of the PPO agent.</p>
        <p><bold>Top Row (Left → Right):</bold></p>
        <p><xref ref-type="fig" rid="fig1">Figure 1(a)</xref><bold>. Episode Returns (left)</bold>-Total reward per training episode.<xref ref-type="fig" rid="fig1">Figure 1(b)</xref><bold>. Actor Loss (center)</bold>-PPO policy surrogate loss across updates.<xref ref-type="fig" rid="fig1">Figure 1(c)</xref><bold>. Critic Loss (right)</bold>-Value-function mean-squared error.</p>
        <p><bold>Bottom Row (Left → Right):</bold></p>
        <p><xref ref-type="fig" rid="fig1">Figure 1(d)</xref><bold>. Policy Entropy (left)</bold>-Average action entropy over time.<xref ref-type="fig" rid="fig1">Figure 1(e)</xref><bold>. Smoothed Returns (center)</bold>-50-episode moving average of returns.<xref ref-type="fig" rid="fig1">Figure 1(f)</xref><bold>. Episode Lengths (right)</bold>-Fixed 30-step episode horizon.</p>
        <p><xref ref-type="fig" rid="fig1">Figure 1</xref> summarizes how the PPO agent learns over the course of 1000 training episodes. Each subpanel focuses on a different aspect of the optimization process, together providing a detailed view of stability, convergence, and learning behaviour.</p>
        <p><xref ref-type="fig" rid="fig1">Figure 1(a)-</xref><bold>Episode Returns</bold></p>
        <p><xref ref-type="fig" rid="fig1">Figure 1(a)</xref> plots the total return (sum of rewards) obtained in each episode. At the beginning of training, returns are relatively low and fluctuate substantially, reflecting the fact that the policy is still exploring random combinations of sleep timing, light exposure, activity, and medication adherence.</p>
        <p>As training progresses, the curve shows:</p>
        <p>a rapid increase in average return over the early episodes (the agent quickly discovers that more regular sleep and higher adherence improve outcomes), followed bya plateau at a higher level, indicating that the agent has discovered a set of interventions that consistently produce good clinical states (high mood stability, high sleep quality, low stress).</p>
        <p>This pattern sharp initial improvement followed by convergence is what we would expect from a well-behaved PPO optimization process in a stationary environment.</p>
        <p><xref ref-type="fig" rid="fig1">Figure 1(b)-</xref><bold>Actor Loss</bold></p>
        <p><xref ref-type="fig" rid="fig1">Figure 1(b)</xref> reports the actor loss, i.e., the PPO surrogate objective for the policy network, plotted over successive updates.</p>
        <p>Key points:</p>
        <p>Early in training, the actor loss shows large oscillations, reflecting strong corrective updates as the policy shifts away from random behaviour toward rewarding regions of the action space.Over time, the variability of the loss decreases, and the curve becomes more compact. This narrowing of the distribution indicates that:the policy is no longer making drastic changes between updates, andthe PPO clipping mechanism is keeping updates inside a “trust region”.</p>
        <p>From a practical perspective, this behaviour means the agent is not “jumping around” in policy space after convergence; it is refining an already good treatment strategy.</p>
        <p><xref ref-type="fig" rid="fig1">Figure 1(c)-</xref><bold>Critic Loss</bold></p>
        <p><xref ref-type="fig" rid="fig1">Figure 1(c)</xref> shows the critic loss, measured as the mean-squared error between predicted value estimates and empirical returns.</p>
        <p>The typical pattern is:</p>
        <p>an overall downward trend as the critic becomes better at predicting long-term returns from a given state, combined withintermittent spikes, which usually occur when:the agent visits new regions of state space due to exploration, orthe underlying policy changes enough that the value function needs to be recalibrated.</p>
        <p>These occasional spikes are not a sign of instability; rather, they indicate that the critic is actively adjusting to updated policies. The important observation is that, despite these spikes, the critic loss does not blow up or drift upward over time, which would signal divergence.</p>
        <p><xref ref-type="fig" rid="fig1">Figure 1(d)-</xref><bold>Policy Entropy</bold></p>
        <p><xref ref-type="fig" rid="fig1">Figure 1(d)</xref> presents the average entropy of the policy distribution over actions. Entropy here measures how “spread out” or stochastic the policy is:</p>
        <p>High entropy → the agent is exploring widely (actions are more random).Low entropy → the agent is more deterministic (actions are concentrated around the mean).</p>
        <p>At the start of training:</p>
        <p>Entropy is relatively high because the policy is initialized with broad, nearly uninformative action distributions. The agent must explore to learn which combinations of sleep shifts, light levels, and adherence patterns are helpful.</p>
        <p>Over time:</p>
        <p>Entropy gradually decreases as the agent discovers a good policy and becomes more confident about which actions are beneficial.Importantly, entropy does not collapse to zero; it stabilizes at a small but non-zero value, meaning the agent retains a little stochasticity.This is desirable: it avoids getting stuck in overly brittle strategies and maintains a minimal level of exploration.</p>
        <p>From a clinical perspective, this implies the learned policy is consistent and reproducible, but not so rigid that it cannot adapt to small variations in the patient’s state.</p>
        <p><bold>Figure 1</bold><bold>(</bold><bold>e</bold><bold>)</bold><bold>-</bold><bold>Returns with Moving Average</bold></p>
        <p><xref ref-type="fig" rid="fig1">Figure 1(e)</xref> overlays:</p>
        <p>the raw episodic returns (noisy, light-coloured line) anda 50-episode moving average (smooth, darker line).</p>
        <p>The moving average serves two purposes:</p>
        <p>1) It filters out high-frequency noise, making the underlying learning trend clearly visible.</p>
        <p>2) It allows us to assess whether the apparent improvements in <xref ref-type="fig" rid="fig1">Figure 1(a)</xref> are sustained rather than due to random fluctuations.</p>
        <p>In this plot:</p>
        <p>The moving average rises steadily and then levels off, confirming that:the performance gain is robust and not just the result of occasional lucky episodes, andthe RL process has reached a stable performance regime.</p>
        <p>This is one of the strongest visual indicators that the PPO agent has successfully learned a high-quality control policy for the simulated bipolar disorder environment.</p>
        <p><xref ref-type="fig" rid="fig1">Figure 1(f)-</xref><bold>Episode Lengths</bold></p>
        <p><xref ref-type="fig" rid="fig1">Figure 1(f)</xref> shows the episode length in time steps (days) for each episode. In this implementation, each episode is designed to last exactly 30 days, and there is no early termination condition. As a result, the plot shows:</p>
        <p>a flat line at 30 steps across all episodes.</p>
        <p>This might seem trivial, but it is important because:</p>
        <p>It confirms the agent is not causing any unintended early termination (e.g., due to environment errors or pathological states).It makes interpretation of returns easier: differences in episodic return across training are purely due to better or worse quality of decisions, not differences in episode duration.</p>
        <p><bold>Overall Interpretation of</bold><xref ref-type="fig" rid="fig1">Figures 1(a)-(f)</xref></p>
        <p>Taken together, the panels in <xref ref-type="fig" rid="fig1">Figure 1</xref> tell a coherent story:</p>
        <p>Learning signal (returns) improves and stabilizes (<xref ref-type="fig" rid="fig1">Figure 1(a)</xref>, <xref ref-type="fig" rid="fig1">Figure 1(e)</xref>).Optimization behaviour (actor and critic losses) becomes more stable and well-behaved over time (<xref ref-type="fig" rid="fig1">Figure 1(b)</xref>, <xref ref-type="fig" rid="fig1">Figure 1(c)</xref>).Exploration vs. exploitation reaches a healthy balance, with entropy decreasing but not collapsing (<xref ref-type="fig" rid="fig1">Figure 1(d)</xref>).Episode structure remains consistent, confirming that changes in performance are due to policy learning, not artifacts (<xref ref-type="fig" rid="fig1">Figure 1(f)</xref>).</p>
        <p>In combination, these curves demonstrate that the PPO agent is:</p>
        <p>learning effectively,not diverging,not overfitting to a small corner of state space, andconverging toward a stable, high-performing treatment strategy in the simulated bipolar disorder environment.</p>
      </sec>
      <sec id="sec3dot2">
        <title>
          3.2. Simulated Clinical Trials (
          <xref ref-type="fig" rid="fig2">Figure 2</xref>
          )
        </title>
        <fig id="fig2">
          <label>Figure 2</label>
          <graphic xlink:href="https://html.scirp.org/file/1114926-rId125.jpeg?20260316023136" />
        </fig>
        <p>Figure 2. Multi-Trial clinical outcomes under the learned policy.</p>
        <p>We next examine the learned deterministic policy in 5 independent trials. <xref ref-type="fig" rid="fig2">Figure 2</xref> summarizes the behaviour of the virtual patient across multiple independent clinical trials, each initialized with the same moderately dysregulated baseline state but evolving under the deterministic version of the learned PPO policy. Together, these trajectories evaluate the robustness<italic>,</italic>generalization, and physiological plausibility of the optimized treatment strategy.</p>
        <p><bold>Top Row (Left → Right):</bold></p>
        <p><xref ref-type="fig" rid="fig2">Figure 2(a)</xref><bold>. Mood Stability Trajectories (left)</bold>-Mood <inline-formula><mml:math><mml:mrow><mml:msub><mml:mi> m </mml:mi><mml:mi> t </mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> across all trials.<xref ref-type="fig" rid="fig2">Figure 2(b)</xref><bold>. Sleep Quality Trajectories (center)</bold>-Sleep quality <inline-formula><mml:math><mml:mrow><mml:msub><mml:mi> q </mml:mi><mml:mi> t </mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> .<xref ref-type="fig" rid="fig2">Figure 2(c)</xref><bold>. Circadian Alignment Trajectories (right)</bold>-Alignment <inline-formula><mml:math><mml:mrow><mml:msub><mml:mi> c </mml:mi><mml:mi> t </mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> over 30 days.</p>
        <p><bold>Bottom Row (Left → Right):</bold></p>
        <p><xref ref-type="fig" rid="fig2">Figure 2(d)</xref><bold>. Final Trial Outcomes (left)</bold>-Comparison of final mood, sleep quality, and reward.<xref ref-type="fig" rid="fig2">Figure 2(e)</xref><bold>. Stress Dynamics (center)</bold>-Stress <inline-formula><mml:math><mml:mrow><mml:msub><mml:mi> σ </mml:mi><mml:mi> t </mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> decay across trials.<xref ref-type="fig" rid="fig2">Figure 2(f)</xref><bold>. Sleep Duration (right)</bold>-Hours slept per day.</p>
        <p><xref ref-type="fig" rid="fig2">Figure 2(a)-</xref><bold>Mood Stability Progression</bold></p>
        <p><xref ref-type="fig" rid="fig2">Figure 2(a)</xref> plots the evolution of daily mood stability <inline-formula><mml:math><mml:mrow><mml:msub><mml:mi> m </mml:mi><mml:mi> t </mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> for each simulated trial over the 30-day period. Across all trials, a consistent pattern emerges:</p>
        <p>Mood stability begins at approximately <inline-formula><mml:math><mml:mrow><mml:msub><mml:mi> m </mml:mi><mml:mn> 0 </mml:mn></mml:msub><mml:mo> ≈ </mml:mo><mml:mn> 0.5 </mml:mn></mml:mrow></mml:math></inline-formula> , representing a moderately unstable but not severely depressed/manic initial condition.Within the first 3 - 5 days, every trial exhibits a sharp rise toward <inline-formula><mml:math><mml:mrow><mml:msub><mml:mi> m </mml:mi><mml:mi> t </mml:mi></mml:msub><mml:mo> ≈ </mml:mo><mml:mn> 1.0 </mml:mn></mml:mrow></mml:math></inline-formula> , indicating rapid stabilization.Once near maximal stability, the curves remain essentially flat for the remaining 25 days.</p>
        <p>This behaviour suggests that the agent quickly discovers and sustains an optimal combination of sleep timing, light exposure, activity regulation, and medication adherence that maximizes mood stabilization. The extremely low variability across trials indicates that the learned policy is not overly sensitive to small differences in the evolving internal state.</p>
        <p>Clinically, this is consistent with robust chronotherapeutic stabilization: regular sleep-wake patterns and strong circadian alignment often have rapid, high-impact effects on mood regulation.</p>
        <p><xref ref-type="fig" rid="fig2">Figure 2(b)-</xref><bold>Sleep Quality Progression</bold></p>
        <p><xref ref-type="fig" rid="fig2">Figure 2(b)</xref> displays the daily sleep quality <inline-formula><mml:math><mml:mrow><mml:msub><mml:mi> q </mml:mi><mml:mi> t </mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> over the same time horizon. Sleep quality follows a trajectory like mood stability:</p>
        <p>From an initial moderate level near <inline-formula><mml:math><mml:mrow><mml:msub><mml:mi> q </mml:mi><mml:mn> 0 </mml:mn></mml:msub><mml:mo> ≈ </mml:mo><mml:mn> 0.5 </mml:mn></mml:mrow></mml:math></inline-formula> , sleep quality increases rapidly during the first 2 - 4 days.By day 5, all trials achieve near-maximal sleep quality (<inline-formula><mml:math><mml:mrow><mml:msub><mml:mi> q </mml:mi><mml:mi> t </mml:mi></mml:msub><mml:mo> ≈ </mml:mo><mml:mn> 1.0 </mml:mn></mml:mrow></mml:math></inline-formula> ).High sleep quality is then maintained stably across the entire 30-day period.</p>
        <p>This indicates that the agent’s chosen actions minimize sleep disruption, enforce regular sleep timing, and calibrate light and activity inputs in a way that consistently enhances the restorative value of sleep. The synchrony between Panels 2a and 2b reflects the well-established causal pathway: improved sleep → improved mood stability.</p>
        <p><xref ref-type="fig" rid="fig2">Figure 2(c)-</xref><bold>Circadian Alignment</bold></p>
        <p><xref ref-type="fig" rid="fig2">Figure 2(c)</xref> plots the circadian alignment variable <inline-formula><mml:math><mml:mrow><mml:msub><mml:mi> c </mml:mi><mml:mi> t </mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> , capturing how well the virtual patient’s internal circadian rhythm tracks an optimal, externally anchored phase.</p>
        <p>Key observations:</p>
        <p>Starting from a neutral baseline of <inline-formula><mml:math><mml:mrow><mml:msub><mml:mi> c </mml:mi><mml:mn> 0 </mml:mn></mml:msub><mml:mo> = </mml:mo><mml:mn> 0.5 </mml:mn></mml:mrow></mml:math></inline-formula> , alignment improves monotonically over the first several days.All trials converge to values close to 1.0, indicating nearly perfect entrainment of the circadian system.No trial shows oscillatory or unstable behaviour; the trajectories are smooth and consistent.</p>
        <p>This is strong evidence that the PPO agent has learned to use light exposure, sleep scheduling, and medication adherence in a way that produces physiologically coherent phase alignment mirroring the mechanisms behind interpersonal and social rhythm therapy (IPSRT) and modern circadian-based treatments in BD.</p>
        <p><xref ref-type="fig" rid="fig2">Figure 2(d)-</xref><bold>Final Trial Outcomes</bold></p>
        <p><xref ref-type="fig" rid="fig2">Figure 2(d)</xref> aggregates the outcome scores for each trial into a summary bar chart showing:</p>
        <p>Final mood stabilityFinal sleep qualityFinal reward (scaled for comparison)</p>
        <p>Results show:</p>
        <p>All final mood and sleep scores are extremely high (typically &gt; 0.95), with virtually no variation across trials.Final rewards also show very low variance, indicating the policy consistently drives the patient to a highly beneficial attractor state.</p>
        <p>Importantly, the convergence across all trials demonstrates that:</p>
        <p>the policy is robust to internal stochasticity in the environment,they learned behaviour is globally stable, not dependent on lucky initialization,and that the PPO agent has not overfit to a narrow trajectory.</p>
        <p>From a clinical modelling standpoint, this suggests the learned strategy is “strongly stabilizing” across a range of possible real-world situations.</p>
        <p><xref ref-type="fig" rid="fig2">Figure 2(e)-</xref><bold>Stress Level Dynamics</bold></p>
        <p><xref ref-type="fig" rid="fig2">Figure 2(e)</xref> plots the daily stress level <inline-formula><mml:math><mml:mrow><mml:msub><mml:mi> σ </mml:mi><mml:mi> t </mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> . Stress begins at a moderate level (<inline-formula><mml:math><mml:mrow><mml:msub><mml:mi> σ </mml:mi><mml:mn> 0 </mml:mn></mml:msub><mml:mo> ≈ </mml:mo><mml:mn> 0.3 </mml:mn></mml:mrow></mml:math></inline-formula> ) but quickly decreases:</p>
        <p>Within the first few days, <inline-formula><mml:math><mml:mrow><mml:msub><mml:mi> σ </mml:mi><mml:mi> t </mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> falls sharply toward 0.After this rapid reduction, stress levels remain essentially zero across all trials.</p>
        <p>This indicates that the PPO policy reliably reduces stress by:</p>
        <p>1) minimizing abrupt sleep shifts,</p>
        <p>2) prescribing appropriate activity levels, and</p>
        <p>3) leveraging medication adherence to dampen physiological stress reactivity.</p>
        <p>Because stress is a destabilizing factor in both mania and depression, its reduction to near-zero levels further support the emergence of stable, high-quality mood trajectories.</p>
        <p><xref ref-type="fig" rid="fig2">Figure 2(f)-</xref><bold>Sleep Duration (hours)</bold></p>
        <p><xref ref-type="fig" rid="fig2">Figure 2(f)</xref> shows sleep duration <inline-formula><mml:math><mml:mrow><mml:msub><mml:mi> d </mml:mi><mml:mi> t </mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> across trials:</p>
        <p>Sleep duration rises from the initial value of approximately 7 hours to between 9 - 10 hours.This elevated sleep duration is maintained through the rest of the 30-day simulation.</p>
        <p>While 9 - 10 hours is within safe clinical limits and oversleeping can be adaptive during recovery from stress this outcome reveals an important modelling insight:</p>
        <p>The reward weights favor sleep quality and circadian alignment more strongly than enforcing a precise 7 - 8-hour sleep window.</p>
        <p>The agent discovers that slightly longer sleep durations help maintain:</p>
        <p>high sleep quality,low stress,and high mood stability.</p>
        <p>If a stricter enforcement of 7 - 8 hours is desired, adjusting the weighting of the sleep-duration reward (<bold>Table 2</bold>) would produce a more tightly constrained policy.</p>
        <p><bold>Integrated Interpretation of</bold><xref ref-type="fig" rid="fig2">Figure 2</xref></p>
        <p>Taken together, <xref ref-type="fig" rid="fig2">Figures 2(a)-(f)</xref> provides strong evidence that the policy learned by the PPO agent:</p>
        <p>rapidly stabilizes mood,improves sleep quality,entrains circadian rhythms,eliminates stress,and maintains a consistent high-performance trajectory across trials.</p>
        <p>The near-identical behaviour across independent initializations demonstrates that the learned intervention strategy is highly robust, reflecting a global attractor regime induced by well-coordinated adjustments in sleep timing, light exposure, activity modulation, and medication regularity.</p>
        <p>A qualitative pre-post summary is provided in <bold>Table 4</bold>.</p>
        <p>Table 4. Approximate change in state variables over a 30-day trial.</p>
        <table-wrap id="tbl4">
          <label>Table 4</label>
          <table>
            <tbody>
              <tr>
                <td>
                  <bold>Variable</bold>
                </td>
                <td>
                  <bold>Initial level (Day 0)</bold>
                </td>
                <td>
                  <bold>Final level (Day 30)</bold>
                </td>
                <td>
                  <bold>Qualitative change</bold>
                </td>
              </tr>
              <tr>
                <td>Sleep quality</td>
                <td>~0.5</td>
                <td>~1.0</td>
                <td>Strong improvement</td>
              </tr>
              <tr>
                <td>Sleep duration</td>
                <td>~7 h</td>
                <td>~9 - 10 h</td>
                <td>Increased, stable</td>
              </tr>
              <tr>
                <td>Mood stability</td>
                <td>~0.5</td>
                <td>~1.0</td>
                <td>Marked stabilization</td>
              </tr>
              <tr>
                <td>Circadian alignment</td>
                <td>~0.5</td>
                <td>~1.0</td>
                <td>Strong alignment</td>
              </tr>
              <tr>
                <td>Stress level</td>
                <td>~0.3</td>
                <td>~0.0</td>
                <td>Dramatic reduction</td>
              </tr>
            </tbody>
          </table>
        </table-wrap>
      </sec>
      <sec id="sec3dot3">
        <title>
          3.3. Circadian Rhythm Analysis (
          <xref ref-type="fig" rid="fig3">Figure 3</xref>
          )
        </title>
        <p>While <xref ref-type="fig" rid="fig1">Figure 1</xref> and <xref ref-type="fig" rid="fig2">Figure 2</xref> demonstrate that the agent successfully optimizes clinical outcomes across multiple trials, <xref ref-type="fig" rid="fig3">Figure 3</xref> provides a more detailed examination of how the learned policy organizes the trajectory of the virtual patient within the underlying state space. By analysing one representative trial in depth, we gain insight into the geometry and coherence of the learned dynamical structure.</p>
        <fig id="fig3">
          <label>Figure 3</label>
          <graphic xlink:href="https://html.scirp.org/file/1114926-rId158.jpeg?20260316023136" />
        </fig>
        <p>Figure 3. Structural analysis of the learned dynamics.</p>
        <p><bold>Top Row (Left → Right):</bold></p>
        <p><xref ref-type="fig" rid="fig3">Figure 3(a)</xref><bold>. State Variables Heatmap (left)</bold>: Heatmap of all five state variables over time.<xref ref-type="fig" rid="fig3">Figure 3(b)</xref><bold>. Phase Portrait:</bold>Mood vs Sleep Quality (right): Trajectory in <inline-formula><mml:math><mml:mrow><mml:mrow><mml:mo> ( </mml:mo><mml:mrow><mml:msub><mml:mi> q </mml:mi><mml:mi> t </mml:mi></mml:msub><mml:mo> , </mml:mo><mml:msub><mml:mi> m </mml:mi><mml:mi> t </mml:mi></mml:msub></mml:mrow><mml:mo> ) </mml:mo></mml:mrow></mml:mrow></mml:math></inline-formula> space.</p>
        <p><bold>Bottom Row (Left → Right):</bold></p>
        <p><xref ref-type="fig" rid="fig3">Figure 3(c)</xref><bold>. Correlation Matrix of State Variables (left)</bold>: Pairwise correlations among all variables.<xref ref-type="fig" rid="fig3">Figure 3(d)</xref><bold>. 3D Trajectory (right)</bold>: Sleep-mood-circadian attractor structure.</p>
        <p>These analyses serve three purposes:</p>
        <p>1) Reveal the internal coupling between physiological variables (sleep, circadian, mood, stress).</p>
        <p>2) Characterize the attractor region toward which the learned policy drives the system.</p>
        <p>3) Verify that the learned control strategy is stable, monotonic, and clinically interpretable.</p>
        <p><xref ref-type="fig" rid="fig3">Figure 3(a)-</xref>Heatmap of State Variables over Time</p>
        <p><xref ref-type="fig" rid="fig3">Figure 3(a)</xref> presents a heatmap showing the daily evolution of all five state variables across the 30-day horizon. Each row corresponds to one variable (sleep quality, sleep duration, mood stability, circadian alignment, stress), while columns represent days.</p>
        <p>The heatmap reveals the following structural patterns:</p>
        <p>Sleep quality, mood stability, and circadian alignment all display warm, high-intensity color gradients that strengthen over time, reflecting:rapid early improvement (days 1 - 5),followed by maintenance of near-maximal values (days 6 - 30).Sleep duration shows a steady upward shift from moderate levels (~7 h) to a plateau near 9 - 10 h, indicating that the learned policy slightly favors longer sleep to maximize mood and reduce stress.Stress exhibits the inverse pattern: the heatmap transitions from moderate activation (~0.3) to near-zero values, consistent with the policy’s emphasis on reducing stress through stable routines and adequate sleep.</p>
        <p>The collective structure is strikingly monotonic: variables associated with wellness rise toward sustained maxima, while stress declines toward its minimum. Such a clean pattern signals that the learned policy imposes a globally stabilizing influence on the dynamics.</p>
        <p><xref ref-type="fig" rid="fig3">Figure 3(b)-</xref>Phase Portrait: Mood vs Sleep Quality</p>
        <p><xref ref-type="fig" rid="fig3">Figure 3(b)</xref> plots the two-dimensional trajectory of the system in the <inline-formula><mml:math><mml:mrow><mml:mrow><mml:mo> ( </mml:mo><mml:mrow><mml:msub><mml:mi> q </mml:mi><mml:mi> t </mml:mi></mml:msub><mml:mo> , </mml:mo><mml:msub><mml:mi> m </mml:mi><mml:mi> t </mml:mi></mml:msub></mml:mrow><mml:mo> ) </mml:mo></mml:mrow></mml:mrow></mml:math></inline-formula> plane. This phase portrait captures the dynamic coupling between sleep quality and mood stability under the learned policy.</p>
        <p>Key observations:</p>
        <p>The trajectory originates near <inline-formula><mml:math><mml:mrow><mml:mrow><mml:mo> ( </mml:mo><mml:mrow><mml:msub><mml:mi> q </mml:mi><mml:mn> 0 </mml:mn></mml:msub><mml:mo> , </mml:mo><mml:msub><mml:mi> m </mml:mi><mml:mn> 0 </mml:mn></mml:msub></mml:mrow><mml:mo> ) </mml:mo></mml:mrow><mml:mo> ≈ </mml:mo><mml:mrow><mml:mo> ( </mml:mo><mml:mrow><mml:mn> 0.5 </mml:mn><mml:mo> , </mml:mo><mml:mn> 0.5 </mml:mn></mml:mrow><mml:mo> ) </mml:mo></mml:mrow></mml:mrow></mml:math></inline-formula> , indicating moderately impaired sleep and mood.It moves along a smooth, monotonic curve toward the point <inline-formula><mml:math><mml:mrow><mml:mrow><mml:mo> ( </mml:mo><mml:mrow><mml:mn> 1.0 </mml:mn><mml:mo> , </mml:mo><mml:mn> 1.0 </mml:mn></mml:mrow><mml:mo> ) </mml:mo></mml:mrow></mml:mrow></mml:math></inline-formula> , where both sleep and mood are maximized.There are no loops, oscillations, or regressions, meaning that the relationship between these variables is consistently reinforcing rather than antagonistic.</p>
        <p>Clinically, this trajectory is meaningful: it suggests that as soon as the agent improves sleep quality via regular sleep timing, adequate light exposure, and reduced stress mood stability improves in parallel. The plot visually validates decades of psychiatric research showing that sleep quality is one of the strongest predictors of next-day mood stability. This monotonic path demonstrates that the agent learns to exploit this coupling effectively.</p>
        <p><xref ref-type="fig" rid="fig3">Figure 3(c)-</xref>Correlation Matrix of State Variables</p>
        <p><xref ref-type="fig" rid="fig3">Figure 3(c)</xref> shows the pairwise Pearson correlation coefficients among all five state dimensions. The matrix reveals two major structural features:</p>
        <p><bold>1</bold><bold>)</bold><bold>A Strong Positive Block</bold></p>
        <p>Sleep quality, sleep duration, mood stability, and circadian alignment form a tightly correlated cluster:</p>
        <p>Improving one variable typically improves the others.This reflects a healthy, synchronized physiological regime precisely the target clinical outcome in BD maintenance therapy.The learned policy essentially forces these variables to co-evolve in a coordinated manner, preventing mismatches such as:high sleep duration but poor mood,good circadian alignment but high stress.</p>
        <p><bold>2</bold><bold>)</bold><bold>Stress as a Negatively Correlated Axis</bold></p>
        <p>Stress is strongly negatively correlated with all other variables:</p>
        <disp-formula id="FD7">
          <mml:math>
            <mml:mrow>
              <mml:mtext>corr</mml:mtext>
              <mml:mrow>
                <mml:mo>(</mml:mo>
                <mml:mrow>
                  <mml:msub>
                    <mml:mi>σ</mml:mi>
                    <mml:mi>t</mml:mi>
                  </mml:msub>
                  <mml:mo>,</mml:mo>
                  <mml:mrow>
                    <mml:mo>{</mml:mo>
                    <mml:mrow>
                      <mml:msub>
                        <mml:mi>q</mml:mi>
                        <mml:mi>t</mml:mi>
                      </mml:msub>
                      <mml:mo>,</mml:mo>
                      <mml:msub>
                        <mml:mi>d</mml:mi>
                        <mml:mi>t</mml:mi>
                      </mml:msub>
                      <mml:mo>,</mml:mo>
                      <mml:msub>
                        <mml:mi>m</mml:mi>
                        <mml:mi>t</mml:mi>
                      </mml:msub>
                      <mml:mo>,</mml:mo>
                      <mml:msub>
                        <mml:mi>c</mml:mi>
                        <mml:mi>t</mml:mi>
                      </mml:msub>
                    </mml:mrow>
                    <mml:mo>}</mml:mo>
                  </mml:mrow>
                </mml:mrow>
                <mml:mo>)</mml:mo>
              </mml:mrow>
              <mml:mo>&lt;</mml:mo>
              <mml:mn>0</mml:mn>
            </mml:mrow>
          </mml:math>
        </disp-formula>
        <p>This matches clinical expectations:</p>
        <p>Elevated stress destabilizes sleep and mood.Improved sleep and circadian alignment reduce stress.Medication adherence directly dampens stress responses.</p>
        <p>Thus, the correlation matrix demonstrates that the learned dynamics are coherent, clinically plausible, and physiologically interpretable.</p>
        <p><bold>Figure 3(d)</bold>-3D State Trajectory (Sleep-Mood-Circadian Space)</p>
        <p><xref ref-type="fig" rid="fig3">Figure 3(d)</xref> visualizes the trial trajectory in a three-dimensional space defined by:</p>
        <p>x-axis: sleep qualityy-axis: mood stabilityz-axis: circadian alignment</p>
        <p>This multidimensional path reveals the global structural properties of the learned environment:</p>
        <p>The trajectory climbs rapidly into a high-value region where all three variables simultaneously approach 1.0.After reaching this region, the path contracts into a dense cluster, indicating that the system has entered a stable attractor.This attractor corresponds to a regime of:high-quality sleep,strong circadian entrainment,stable euthymic mood,minimal stress.</p>
        <p>The smoothness and consistency of the path indicate that the policy generates predictable, non-chaotic, and stable system behaviour, even in a non-linear environment. Clinically, this attractor resembles a stabilized patient maintaining good routine regularity, low stress, and euthymic functioning.</p>
        <p><bold>Integrated Interpretation of</bold><xref ref-type="fig" rid="fig3">Figure 3</xref></p>
        <p><xref ref-type="fig" rid="fig3">Figure 3</xref> demonstrates that the PPO agent does far more than maximize an abstract reward signal. It reshapes the underlying state space into a coherent, therapeutically meaningful dynamical landscape.</p>
        <p>Specifically:</p>
        <p>The system evolves toward a single, global attractor characterized by strong sleep-mood-circadian synchrony.All state variables show monotonic convergence, with no oscillations or pathological transitions.Stress is systematically eliminated.Phase portraits and correlations reveal clear physiological coupling, consistent with decades of BD chronobiology research.The 3D trajectory confirms that the optimized patient state is stable, reproducible, and robust.</p>
        <p>Together, these results suggest that the learned policy is effectively constructing and maintaining a clinically interpretable attractor basin, where wellness-related variables mutually reinforce one another.</p>
      </sec>
    </sec>
    <sec id="sec4">
      <title>4. Discussion</title>
      <p>This work shows that a Proximal Policy Optimization (PPO) agent can learn coherent and clinically meaningful control policies when embedded in a physiologically motivated simulator of sleep-mood-circadian dynamics in bipolar disorder (BD) [<xref ref-type="bibr" rid="B31">31</xref>]-[<xref ref-type="bibr" rid="B34">34</xref>]. Starting from a generic, moderately dysregulated state, the learned policy consistently drives the virtual patient toward a stable regime characterized by high sleep quality, near-optimal circadian alignment, improved mood stability, and minimal stress. Importantly, the agent does not have any explicit “clinical rules” hard-coded. Instead, it discovers effective patterns of sleep timing, light exposure, activity modulation, and medication adherence purely from the reward structure and the environment’s dynamics. This suggests that, under an appropriate modelling framework, RL can rediscover and operationalize principles that clinicians use intuitively in BD chronotherapy such as the centrality of regular routines and circadian stabilization while expressing them as explicit, optimized control policies. The discussion below unpacks these findings from two angles: clinical interpretation and methodological insights.</p>
      <sec id="sec4dot1">
        <title>4.1. Clinical Interpretation</title>
        <p>The trajectories in <xref ref-type="fig" rid="fig2">Figures 2(a)-(f)</xref> and <xref ref-type="fig" rid="fig3">Figures 3(a)-(d)</xref> exhibit several patterns that resonate strongly with clinical intuition about BD and its relationship with sleep and circadian rhythms.</p>
        <p>First, the simulations show that stabilizing sleep and circadian alignment is rapidly followed by mood stabilization and stress reduction. In multiple independent trials (<xref ref-type="fig" rid="fig2">Figures 2(a)-(c)</xref>, <xref ref-type="fig" rid="fig2">Figure 2(e)</xref>), sleep quality, sleep duration, circadian alignment, and stress move in a coordinated manner:</p>
        <p>Sleep quality and circadian alignment increase quickly and then plateau at high levels.Mood stability rises almost in lockstep with sleep improvements, reaching near-maximal values within the first few days.Stress shows the mirror image: it decays rapidly toward zero and remains suppressed.</p>
        <p>This mirrors clinical observations: once patients achieve stable routines and regular, high-quality sleep, mood tends to become less volatile, and perceived stress decreases. In that sense, the learned policy is not only optimizing an abstract reward signal; it is recapitulating a known causal chain: regular sleep and circadian entrainment → improved mood stability → reduced stress load.</p>
        <p>Second, the policy implicitly learns that consistency in sleep timing and medication adherence is a key driver of improvement. We did not explicitly tell the agent “Keep sleep times regular” or “adhere strictly to medication.” These behaviours emerge because the environment dynamics and reward function together make inconsistency costly:</p>
        <p>Abrupt sleep timing shifts negatively affect sleep quality, circadian alignment, and stress.Low medication adherence increases stress and weakens both mood and circadian stability.</p>
        <p>By maximizing long-term reward, the PPO agent gravitates toward patterns that minimize large sleep shifts and encourage high medication adherence. This is exactly what human clinicians counsel: maintain regular sleep-wake cycles, avoid late-night phase shifts, and take medication reliably. The fact that such regularizing behaviour emerges naturally from the RL objective supports the idea that RL can automatically infer “good clinical habits” from a well-designed environment and reward function, rather than requiring them to be hand-coded as fixed rules.</p>
        <p>Third, the system converges to a clinically interpretable attractor state that can be understood as a stable euthymic regime. The final state achieved in each trial is not just numerically high reward; it has a clear clinical interpretation:</p>
        <p>Sleep quality is high and stable.Circadian alignment is strong.Mood stability is near maximal.Stress is near zero.Sleep duration is slightly long but not pathologically so.</p>
        <p>In <xref ref-type="fig" rid="fig3">Figure 3</xref>, this is reflected in:</p>
        <p>The heatmap (<xref ref-type="fig" rid="fig3">Figure 3(a)</xref>), where wellness-related variables uniformly brighten over time while stress fades.The phase portrait (<xref ref-type="fig" rid="fig3">Figure 3(b)</xref>), where the trajectory moves smoothly toward the upper-right corner (good sleep + good mood).The correlation matrix (<xref ref-type="fig" rid="fig3">Figure 3(c)</xref>), where sleep, mood, and circadian alignment form a tightly coupled positive block, and stress is strongly negatively correlated with all of them.The 3D trajectory (<xref ref-type="fig" rid="fig3">Figure 3(d)</xref>), which shows the system settling into a compact, high-value region in sleep-mood-circadian space.</p>
        <p>This attractor is precisely what clinicians aim for in maintenance treatment: a stable euthymic state with regular routines and low stress, rather than frequent transitions between depressive and manic poles. The fact that the RL agent converges to such a structure, using only reward information and simulated dynamics, supports the idea that RL can formalize and stabilize clinical heuristics used in BD chronotherapy.</p>
        <p>In summary, from a clinical standpoint, this work suggests that a well-constructed RL framework can rediscover key principles of BD management:</p>
        <p>prioritize sleep and circadian regularity.enforce consistent routines and medication adherence.reduce stress to maintain long-term mood stability.and converge toward a stable euthymic regime.</p>
        <p>Even though the current environment is stylized, the qualitative behaviours it yields closely resemble the logic of existing evidence-based psychosocial and chronotherapeutic interventions.</p>
      </sec>
      <sec id="sec4dot2">
        <title>4.2. Methodological Insights</title>
        <p>Beyond clinical interpretation, the experiments offer several methodological lessons for designing RL systems in healthcare-like domains.</p>
        <p>First, the results highlight the utility of continuous Gaussian policies with constrained variance and LayerNorm for stable training in non-linear physiological environments.</p>
        <p>The actor network outputs the mean and log-standard deviation of a Gaussian distribution over continuous actions (sleep timing, light exposure, activity, adherence). This choice has several advantages:</p>
        <p>It naturally matches the problem structure, where interventions are not discrete “on/off” decisions but graded adjustments (e.g., shift bedtime by 0.3 hours).It allows the policy to express controlled stochasticity, which is essential early in training for exploration and later for robustness.Constraining the log-standard deviation and using LayerNorm stabilizes gradients and prevents pathological behaviors (e.g., huge action variance, exploding activations).</p>
        <p>The smooth, monotonic training curves in <xref ref-type="fig" rid="fig1">Figure 1</xref> and the absence of catastrophic divergence provide empirical support for this design.</p>
        <p>Second, the work underscores the importance of multi-objective reward shaping in health-related RL. In many control problems, defining a single scalar objective is relatively straightforward (e.g., maximize velocity, minimize energy). In BD management, this is not the case: clinicians optimize sleep quality, sleep duration, circadian alignment, mood stability, and stress simultaneously.</p>
        <p>By explicitly decomposing the reward into multiple clinically interpretable components (as in <bold>Table 2</bold>) and then carefully weighting them we were able to:</p>
        <p>encode realistic therapeutic trade-offs (e.g., good sleep but not extreme oversleeping; low stress but still adequate activity).avoid the agent exploiting degenerate solutions (e.g., maximizing mood at the cost of extreme sleep duration).and make the behaviour of the policy more interpretable, since changes in each reward term can be traced back to specific state dimensions.</p>
        <p>This illustrates a general principle: in clinical RL settings, reward design is not just a technical detail; it is a modelling of the therapeutic philosophy itself [<xref ref-type="bibr" rid="B35">35</xref>]-[<xref ref-type="bibr" rid="B42">42</xref>].</p>
        <p>Third, the extensive use of rich visual diagnostics (<xref ref-type="fig" rid="fig1">Figures 1-3</xref>) is crucial for interpreting and validating RL agents in health applications. Unlike many benchmark environments where “higher score” is already trusted, in healthcare we must constantly ask:</p>
        <p>Is the agent doing something clinically sensible?Is it overfitting to quirks of the environment?Is it exploiting a loophole in the reward function?</p>
        <p>The combination of:</p>
        <p>training curves (returns, losses, entropy) in <xref ref-type="fig" rid="fig1">Figure 1</xref>,outcome trajectories and trial-level summaries in <xref ref-type="fig" rid="fig2">Figure 2</xref>,and state-space structure analyses in <xref ref-type="fig" rid="fig3">Figure 3</xref></p>
        <p>provides multiple, complementary lenses for debugging and validating the agent. They reveal not only that the agent performs well numerically, but how and why it does so, in a way that can be discussed with clinicians. In practice, this suggests that any RL system aimed at clinical decision support should be accompanied by a similarly rich suite of visual and statistical diagnostics not just a final performance number. In summary, from a methodological perspective, this study shows that:</p>
        <p>Gaussian PPO with appropriate architectural constraints can handle complex, physiologically motivated control tasks.carefully shaped, multi-component rewards are essential for capturing clinical objectives.and interpretability tools (plots, correlations, trajectories) are not optional extras but core components of safe, trustworthy RL in healthcare.</p>
      </sec>
    </sec>
    <sec id="sec5">
      <title>5. Limitations and Future Work</title>
      <p>Despite the promising results, several important limitations must be acknowledged. These limitations highlight directions for future research and clarify the gap between the current conceptual demonstration and clinically deployable reinforcement learning (RL) systems. The fixed 30-day episode horizon captures short-term stabilization dynamics but does not model longer mood cycles or seasonal destabilization often observed in BD. Extending the horizon to 90 - 180 days would allow investigation of slower oscillatory patterns and relapse risk. Future work will explore variable-length episodes and hierarchical temporal modelling.</p>
      <sec id="sec5dot1">
        <title>5.1. Synthetic and Hand-Crafted Dynamics</title>
        <p>The current Circadian Environment is intentionally stylized: all transition equations are hand-crafted using clinical intuition, prior literature on sleep and circadian processes, and general principles from chronobiology. Although these equations capture plausible qualitative trends such as the beneficial cascade from improved sleep to stabilized mood they do not yet reflect the full complexity or heterogeneity observed in real BD populations.</p>
        <p>This limitation implies that:</p>
        <p>they learned policies may overly depend on the simplified structure of the simulator, andthere is a risk of simulation-reality mismatch if applied directly to clinical populations.</p>
        <p>Future work will calibrate the environment using real digital phenotyping signals (actigraphy, sleep diary data, smartphone usage, ecological momentary assessment mood ratings), allowing the transition dynamics to be data-driven rather than handcrafted.</p>
      </sec>
      <sec id="sec5dot2">
        <title>5.2. Reward Calibration and Clinical Trade-Offs</title>
        <p>While the reward function successfully balances multiple therapeutic objectives, the current weights were selected heuristically. Their influence is clearly visible in <xref ref-type="fig" rid="fig2">Figure 2(f)</xref>, where sleep duration tends to rise to 9 - 10 hours clinically acceptable but not perfectly aligned with the intended 7 - 8-hour target. To evaluate robustness to reward specification, we conducted exploratory sensitivity analyses in which sleep-duration and stress penalty weights were varied by ±20%. The learned policy showed minor quantitative shifts in preferred sleep duration but preserved the global stabilization strategy (rapid circadian alignment, high adherence, stress suppression). This suggests that the emergence of the euthymic attractor is structurally stable under moderate reward perturbations, although precise sleep-duration targets are weight-sensitive.</p>
        <p>This reveals that:</p>
        <p>The agent is optimizing within the constraints it is given butThe reward terms do not yet fully encode clinically nuanced trade-offs between sufficient sleep and potential oversleeping.</p>
        <p>Future work will refine reward calibration in three ways:</p>
        <p>1) Clinical expert input to encode treatment priorities more accurately.</p>
        <p>2) Inverse reinforcement learning (IRL) to infer reward structures from observed patient outcomes.</p>
        <p>3) Multi-objective RL where preferences can be tuned per patient.</p>
      </sec>
      <sec id="sec5dot3">
        <title>5.3. Homogeneous Patient Model</title>
        <p>The simulator uses one generic patient, with fixed physiological parameters and identical reactions to interventions. Real patients vary widely in:</p>
        <p>sleep need,circadian sensitivity to light,stress susceptibility,medication response,baseline mood stability.</p>
        <p>A single canonical model therefore cannot capture inter-individual variability. This limits both realism and potential for personalization.</p>
        <p>Future extensions will:</p>
        <p>Introduce patient-specific parameter distributions (e.g., differing circadian gain, stress reactivity, sleep inertia),Train population-wide policies that generalize across heterogeneous simulated cohorts, andExplore personalized and meta-learning RL approaches that adapt quickly to individual patient profiles.</p>
      </sec>
      <sec id="sec5dot4">
        <title>5.4. Absence of Explicit Safety Constraints</title>
        <p>Clinical decision-making requires strict guarantees around safety. The current PPO setup lacks:</p>
        <p>hard bounds on how much sleep timing can be adjusted per day,explicit risk metrics,or constraints to prevent harmful or clinically implausible strategies.</p>
        <p>Although the learned policy behaved safely within the simulator, real-world deployment would require:</p>
        <p>Constrained MDPs,Safe RL frameworks,Risk-sensitive optimization (CVaR, robust RL),Action bounding informed by clinical practice (e.g., max 20 - 30 minutes/day sleep shift).</p>
        <p>In healthcare applications, ensuring safety is not optional; it must be designed into the RL objective.</p>
      </sec>
      <sec id="sec5dot5">
        <title>5.5. No Real-World Validation</title>
        <p>A major limitation is the absence of empirical evaluation. Neither the environment nor the learned policy has been validated against:</p>
        <p>longitudinal sleep data,circadian phase markers,or daily mood ratings from individuals with BD.</p>
        <p>Before any real-world translation:</p>
        <p>Environment parameters must be fitted to empirical population-level data,Policies must be validated offline on retrospective datasets,And any deployment must be embedded in human-in-the-loop decision support, never autonomous RL.</p>
        <p>Only after these steps combined with oversight from clinicians, ethicists, and regulatory bodies could such a system be considered for safe clinical integration.</p>
      </sec>
      <sec id="sec5dot6">
        <title>5.6. Future Directions</title>
        <p>Building on the limitations above, several important research directions emerge:</p>
        <p><bold>1)</bold><bold>Data-driven modelling:</bold> Fit the environment using wearable-derived actigraphy, sleep diaries, and mood logs.</p>
        <p><bold>2)</bold><bold>Patient heterogeneity:</bold> Create a virtual population with diverse physiological profiles and symptoms.</p>
        <p><bold>3)</bold><bold>Model-</bold><bold>based and</bold> Bayesian<bold>RL:</bold> Use learned transition models and uncertainty estimation to generate safer, more interpretable policies.</p>
        <p><bold>4)</bold><bold>Real-world validation:</bold> Evaluate policies offline before any pilot tests with real patients.</p>
        <p><bold>5)</bold><bold>Human-in-the-loop systems:</bold> Integrate RL recommendations as suggestions within clinician-guided interfaces, never as autonomous actions.</p>
        <p>These steps will bring the proposed framework closer to a clinically meaningful and ethically deployable decision-support tool for BD chronotherapy.</p>
      </sec>
    </sec>
    <sec id="sec6">
      <title>6. Conclusions</title>
      <p>This work introduces a complete reinforcement-learning (RL) framework designed to optimize the intertwined dynamics of sleep, circadian rhythms, mood stability, and stress in bipolar disorder (BD). By embedding a Proximal Policy Optimization (PPO) agent within Circadian Environment—a physiologically motivated simulator, we demonstrate that an RL agent can autonomously discover clinically intuitive stabilization strategies without any explicit hand-crafted rules. Across training (<xref ref-type="fig" rid="fig1">Figures 1(a)-(f)</xref>), the agent’s behaviour becomes increasingly structured and stable, with monotonic improvements in reward, decreasing uncertainty, and consistent convergence of both actor and critic networks. When deployed in evaluation mode (<xref ref-type="fig" rid="fig2">Figures 2(a)-(f)</xref>), the learned policy reliably drives multiple independently initialized “virtual patients” toward a uniform attractor characterized by high sleep quality, strong circadian alignment, elevated mood stability, and near-zero stress. A deeper structural analysis (<xref ref-type="fig" rid="fig3">Figures 3(a)-(d)</xref>) reveals that this attractor is not an artifact of reward optimization alone: the entire state space becomes reorganized into a coherent, clinically interpretable topology in which sleep, mood, and circadian alignment reinforce one another while stress inversely collapses. Although the simulator is necessarily simplified and does not yet incorporate patient heterogeneity or real-world physiological data, the full architecture comprising the environment, agent, training engine, and visualization pipeline provides a robust blueprint for future work. The modularity of the system makes it straightforward to integrate digital phenotyping data (actigraphy, mobile sensing, mood diaries), calibrate transition dynamics to real patients, introduce population variability, and embed safety constraints appropriate for clinical use.</p>
      <p>This study demonstrates that reinforcement learning is capable not only of optimizing reward in an abstract simulation, but of discovering realistic, interpretable, and clinically aligned intervention strategies. The framework presented here lays foundational groundwork for developing next-generation RL-based decision-support tools that can complement chronotherapy and personalized treatment planning in bipolar disorder.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <title>References</title>
      <ref id="B1">
        <label>1.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">McCarthy, M.J., Gottlieb, J.F., Gonzalez, R., McClung, C.A., Alloy, L.B., Cain, S., <italic>et al</italic>. (2021) Neurobiological and Behavioral Mechanisms of Circadian Rhythm Disruption in Bipolar Disorder: A Critical Multi‐Disciplinary Literature Review and Agenda for Future Research from the ISBD Task Force on Chronobiology. <italic>Bipolar Disorders</italic>, 24, 232-263. https://doi.org/10.1111/bdi.13165 <pub-id pub-id-type="doi">10.1111/bdi.13165</pub-id><pub-id pub-id-type="pmid">34850507</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1111/bdi.13165">https://doi.org/10.1111/bdi.13165</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>McCarthy, M.J.</string-name>
              <string-name>Gottlieb, J.F.</string-name>
              <string-name>Gonzalez, R.</string-name>
              <string-name>McClung, C.A.</string-name>
              <string-name>Alloy, L.B.</string-name>
              <string-name>Cain, S.</string-name>
            </person-group>
            <year>2021</year>
            <article-title>Neurobiological and Behavioral Mechanisms of Circadian Rhythm Disruption in Bipolar Disorder: A Critical Multi‐Disciplinary Literature Review and Agenda for Future Research from the ISBD Task Force on Chronobiology</article-title>
            <source>Bipolar Disorders</source>
            <volume>24</volume>
            <pub-id pub-id-type="doi">10.1111/bdi.13165</pub-id>
            <pub-id pub-id-type="pmid">34850507</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B2">
        <label>2.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Tonon, A.C., Nexha, A., Mendonça da Silva, M., Gomes, F.A., Hidalgo, M.P. and Frey, B.N. (2024) Sleep and Circadian Disruption in Bipolar Disorders: From Psychopathology to Digital Phenotyping in Clinical Practice. <italic>Psychiatry and Clinical Neurosciences</italic>, 78, 654-666. https://doi.org/10.1111/pcn.13729 <pub-id pub-id-type="doi">10.1111/pcn.13729</pub-id><pub-id pub-id-type="pmid">39210713</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1111/pcn.13729">https://doi.org/10.1111/pcn.13729</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Tonon, A.C.</string-name>
              <string-name>Nexha, A.</string-name>
              <string-name>Silva, M.</string-name>
              <string-name>Gomes, F.A.</string-name>
              <string-name>Hidalgo, M.P.</string-name>
              <string-name>Frey, B.N.</string-name>
            </person-group>
            <year>2024</year>
            <article-title>Sleep and Circadian Disruption in Bipolar Disorders: From Psychopathology to Digital Phenotyping in Clinical Practice</article-title>
            <source>Psychiatry and Clinical Neurosciences</source>
            <volume>78</volume>
            <pub-id pub-id-type="doi">10.1111/pcn.13729</pub-id>
            <pub-id pub-id-type="pmid">39210713</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B3">
        <label>3.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Oliva, V., Fico, G., De Prisco, M., Gonda, X., Rosa, A.R. and Vieta, E. (2025) Bipolar Disorders: An Update on Critical Aspects. <italic>The Lancet Regional Health</italic>- <italic>Europe</italic>, 48, Article 101135. https://doi.org/10.1016/j.lanepe.2024.101135 <pub-id pub-id-type="doi">10.1016/j.lanepe.2024.101135</pub-id><pub-id pub-id-type="pmid">39811787</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1016/j.lanepe.2024.101135">https://doi.org/10.1016/j.lanepe.2024.101135</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Oliva, V.</string-name>
              <string-name>Fico, G.</string-name>
              <string-name>Prisco, M.</string-name>
              <string-name>Gonda, X.</string-name>
              <string-name>Rosa, A.R.</string-name>
              <string-name>Vieta, E.</string-name>
            </person-group>
            <year>2025</year>
            <article-title>Bipolar Disorders: An Update on Critical Aspects</article-title>
            <source>The Lancet Regional Health-Europe</source>
            <volume>48</volume>
            <elocation-id>101135</elocation-id>
            <pub-id pub-id-type="doi">10.1016/j.lanepe.2024.101135</pub-id>
            <pub-id pub-id-type="pmid">39811787</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B4">
        <label>4.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Scott, J., Etain, B., Miklowitz, D., Crouse, J.J., Carpenter, J., Marwaha, S., <italic>et al</italic>. (2022) A Systematic Review and Meta-Analysis of Sleep and Circadian Rhythms Disturbances in Individuals at High-Risk of Developing or with Early Onset of Bipolar Disorders. <italic>Neuroscience &amp; Biobehavioral Reviews</italic>, 135, Article 104585. https://doi.org/10.1016/j.neubiorev.2022.104585 <pub-id pub-id-type="doi">10.1016/j.neubiorev.2022.104585</pub-id><pub-id pub-id-type="pmid">35182537</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1016/j.neubiorev.2022.104585">https://doi.org/10.1016/j.neubiorev.2022.104585</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Scott, J.</string-name>
              <string-name>Etain, B.</string-name>
              <string-name>Miklowitz, D.</string-name>
              <string-name>Crouse, J.J.</string-name>
              <string-name>Carpenter, J.</string-name>
              <string-name>Marwaha, S.</string-name>
            </person-group>
            <year>2022</year>
            <article-title>A Systematic Review and Meta-Analysis of Sleep and Circadian Rhythms Disturbances in Individuals at High-Risk of Developing or with Early Onset of Bipolar Disorders</article-title>
            <source>Neuroscience &amp; Biobehavioral Reviews</source>
            <volume>135</volume>
            <elocation-id>104585</elocation-id>
            <pub-id pub-id-type="doi">10.1016/j.neubiorev.2022.104585</pub-id>
            <pub-id pub-id-type="pmid">35182537</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B5">
        <label>5.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Ghaemi, S.N., Dalley, S., Catania, C. and Barroilhet, S. (2014) Bipolar or Borderline: A Clinical Overview. <italic>Acta Psychiatrica Scandinavica</italic>, 130, 99-108. https://doi.org/10.1111/acps.12257 <pub-id pub-id-type="doi">10.1111/acps.12257</pub-id><pub-id pub-id-type="pmid">24571137</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1111/acps.12257">https://doi.org/10.1111/acps.12257</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Ghaemi, S.N.</string-name>
              <string-name>Dalley, S.</string-name>
              <string-name>Catania, C.</string-name>
              <string-name>Barroilhet, S.</string-name>
            </person-group>
            <year>2014</year>
            <article-title>Bipolar or Borderline: A Clinical Overview</article-title>
            <source>Acta Psychiatrica Scandinavica</source>
            <volume>130</volume>
            <pub-id pub-id-type="doi">10.1111/acps.12257</pub-id>
            <pub-id pub-id-type="pmid">24571137</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B6">
        <label>6.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Bellivier, F., Geoffroy, P., Etain, B. and Scott, J. (2015) Sleep-and Circadian Rhythm–Associated Pathways as Therapeutic Targets in Bipolar Disorder. <italic>Expert Opinion on Therapeutic Targets</italic>, 19, 747-763. https://doi.org/10.1517/14728222.2015.1018822 <pub-id pub-id-type="doi">10.1517/14728222.2015.1018822</pub-id><pub-id pub-id-type="pmid">25726988</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1517/14728222.2015.1018822">https://doi.org/10.1517/14728222.2015.1018822</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Bellivier, F.</string-name>
              <string-name>Geoffroy, P.</string-name>
              <string-name>Etain, B.</string-name>
              <string-name>Scott, J.</string-name>
            </person-group>
            <year>2015</year>
            <article-title>Sleep-and Circadian Rhythm–Associated Pathways as Therapeutic Targets in Bipolar Disorder</article-title>
            <source>Expert Opinion on Therapeutic Targets</source>
            <volume>19</volume>
            <pub-id pub-id-type="doi">10.1517/14728222.2015.1018822</pub-id>
            <pub-id pub-id-type="pmid">25726988</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B7">
        <label>7.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Proudfoot, J., Doran, J., Manicavasagar, V. and Parker, G. (2011) The Precipitants of Manic/Hypomanic Episodes in the Context of Bipolar Disorder: A Review. <italic>Journal of Affective Disorders</italic>, 133, 381-387. https://doi.org/10.1016/j.jad.2010.10.051 <pub-id pub-id-type="doi">10.1016/j.jad.2010.10.051</pub-id><pub-id pub-id-type="pmid">21106249</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1016/j.jad.2010.10.051">https://doi.org/10.1016/j.jad.2010.10.051</ext-link></mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Proudfoot, J.</string-name>
              <string-name>Doran, J.</string-name>
              <string-name>Manicavasagar, V.</string-name>
              <string-name>Parker, G.</string-name>
            </person-group>
            <year>2011</year>
            <article-title>The Precipitants of Manic/Hypomanic Episodes in the Context of Bipolar Disorder: A Review</article-title>
            <source>Journal of Affective Disorders</source>
            <volume>133</volume>
            <pub-id pub-id-type="doi">10.1016/j.jad.2010.10.051</pub-id>
            <pub-id pub-id-type="pmid">21106249</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B8">
        <label>8.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Steardo, L., de Filippis, R., Carbone, E.A., Segura-Garcia, C., Verkhratsky, A. and De Fazio, P. (2019) Sleep Disturbance in Bipolar Disorder: Neuroglia and Circadian Rhythms. <italic>Frontiers in Psychiatry</italic>, 10, Article ID: 501. https://doi.org/10.3389/fpsyt.2019.00501 <pub-id pub-id-type="doi">10.3389/fpsyt.2019.00501</pub-id><pub-id pub-id-type="pmid">31379620</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.3389/fpsyt.2019.00501">https://doi.org/10.3389/fpsyt.2019.00501</ext-link></mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Steardo, L.</string-name>
              <string-name>Filippis, R.</string-name>
              <string-name>Carbone, E.A.</string-name>
              <string-name>Segura-Garcia, C.</string-name>
              <string-name>Verkhratsky, A.</string-name>
              <string-name>Fazio, P.</string-name>
            </person-group>
            <year>2019</year>
            <article-title>Sleep Disturbance in Bipolar Disorder: Neuroglia and Circadian Rhythms</article-title>
            <source>Frontiers in Psychiatry</source>
            <volume>10</volume>
            <fpage>501</fpage>
            <elocation-id>ID</elocation-id>
            <pub-id pub-id-type="doi">10.3389/fpsyt.2019.00501</pub-id>
            <pub-id pub-id-type="pmid">31379620</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B9">
        <label>9.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Gottlieb, J.F., Benedetti, F., Geoffroy, P.A., Henriksen, T.E.G., Lam, R.W., Murray, G., <italic>et al</italic>. (2019) The Chronotherapeutic Treatment of Bipolar Disorders: A Systematic Review and Practice Recommendations from the ISBD Task Force on Chronotherapy and Chronobiology. <italic>Bipolar Disorders</italic>, 21, 741-773. https://doi.org/10.1111/bdi.12847 <pub-id pub-id-type="doi">10.1111/bdi.12847</pub-id><pub-id pub-id-type="pmid">31609530</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1111/bdi.12847">https://doi.org/10.1111/bdi.12847</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Gottlieb, J.F.</string-name>
              <string-name>Benedetti, F.</string-name>
              <string-name>Geoffroy, P.A.</string-name>
              <string-name>Henriksen, T.E.G.</string-name>
              <string-name>Lam, R.W.</string-name>
              <string-name>Murray, G.</string-name>
            </person-group>
            <year>2019</year>
            <article-title>The Chronotherapeutic Treatment of Bipolar Disorders: A Systematic Review and Practice Recommendations from the ISBD Task Force on Chronotherapy and Chronobiology</article-title>
            <source>Bipolar Disorders</source>
            <volume>21</volume>
            <pub-id pub-id-type="doi">10.1111/bdi.12847</pub-id>
            <pub-id pub-id-type="pmid">31609530</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B10">
        <label>10.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Pal, A., Sidana, I.S. and Avinash, P.R. (2022) Sleep in Bipolar Disorders. In: Gupta, R., Neubauer, D.N. and Pandi-Perumal, S.R., Eds., <italic>Sleep and Neuropsychiatric Disorders</italic>, Springer, 371-396. https://doi.org/10.1007/978-981-16-0123-1_19 <pub-id pub-id-type="doi">10.1007/978-981-16-0123-1_19</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1007/978-981-16-0123-1_19">https://doi.org/10.1007/978-981-16-0123-1_19</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Pal, A.</string-name>
              <string-name>Sidana, I.S.</string-name>
              <string-name>Avinash, P.R.</string-name>
              <string-name>Gupta, R.</string-name>
              <string-name>Neubauer, D.N.</string-name>
              <string-name>Pandi-Perumal, S.R.</string-name>
              <string-name>Disorders, S</string-name>
            </person-group>
            <year>2022</year>
            <article-title>Sleep in Bipolar Disorders</article-title>
            <source>In: Gupta</source>
            <volume>371</volume>
            <pub-id pub-id-type="doi">10.1007/978-981-16-0123-1_19</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B11">
        <label>11.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Yeom, J.W., Park, S. and Lee, H. (2024) Managing Circadian Rhythms: A Key to Enhancing Mental Health in College Students. <italic>Psychiatry Investigation</italic>, 21, 1309-1317. https://doi.org/10.30773/pi.2024.0250 <pub-id pub-id-type="doi">10.30773/pi.2024.0250</pub-id><pub-id pub-id-type="pmid">39757810</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.30773/pi.2024.0250">https://doi.org/10.30773/pi.2024.0250</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Yeom, J.W.</string-name>
              <string-name>Park, S.</string-name>
              <string-name>Lee, H.</string-name>
            </person-group>
            <year>2024</year>
            <article-title>Managing Circadian Rhythms: A Key to Enhancing Mental Health in College Students</article-title>
            <source>Psychiatry Investigation</source>
            <volume>21</volume>
            <pub-id pub-id-type="doi">10.30773/pi.2024.0250</pub-id>
            <pub-id pub-id-type="pmid">39757810</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B12">
        <label>12.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Caruso, V., Geoffroy, P.A., Alfì, G., Miniati, M., Riemann, D., Gemignani, A., <italic>et al</italic>. (2024) Effects of Mood Stabilizers on Sleep and Circadian Rhythms: A Systematic Review. <italic>Current Sleep Medicine Reports</italic>, 10, 329-357. https://doi.org/10.1007/s40675-024-00298-5 <pub-id pub-id-type="doi">10.1007/s40675-024-00298-5</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1007/s40675-024-00298-5">https://doi.org/10.1007/s40675-024-00298-5</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Caruso, V.</string-name>
              <string-name>Geoffroy, P.A.</string-name>
              <string-name>Miniati, M.</string-name>
              <string-name>Riemann, D.</string-name>
              <string-name>Gemignani, A.</string-name>
            </person-group>
            <year>2024</year>
            <article-title>Effects of Mood Stabilizers on Sleep and Circadian Rhythms: A Systematic Review</article-title>
            <source>Current Sleep Medicine Reports</source>
            <volume>10</volume>
            <pub-id pub-id-type="doi">10.1007/s40675-024-00298-5</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B13">
        <label>13.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">DeSanctis, M.V. (2017) Circadian Principles: Behavioral Health Implications. <italic>Journal of Applied Biobehavioral Research</italic>, 22, e12102. https://doi.org/10.1111/jabr.12102 <pub-id pub-id-type="doi">10.1111/jabr.12102</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1111/jabr.12102">https://doi.org/10.1111/jabr.12102</ext-link></mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>DeSanctis, M.V.</string-name>
            </person-group>
            <year>2017</year>
            <article-title>Circadian Principles: Behavioral Health Implications</article-title>
            <source>Journal of Applied Biobehavioral Research</source>
            <volume>22</volume>
            <pub-id pub-id-type="doi">10.1111/jabr.12102</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B14">
        <label>14.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Jones, C., Reynolds, C., Olson, R., Bontrager, A., Lambert, S., Balba, N., Weymann, K., <italic>et al</italic>. (2021) 277 Tunable White Light for Elders (TWLITE): A Feasibility Study of a Home-Based Sleep Intervention. <italic>Sleep</italic>, 44, A111. https://doi.org/10.1093/sleep/zsab072.276 <pub-id pub-id-type="doi">10.1093/sleep/zsab072.276</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1093/sleep/zsab072.276">https://doi.org/10.1093/sleep/zsab072.276</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Jones, C.</string-name>
              <string-name>Reynolds, C.</string-name>
              <string-name>Olson, R.</string-name>
              <string-name>Bontrager, A.</string-name>
              <string-name>Lambert, S.</string-name>
              <string-name>Balba, N.</string-name>
              <string-name>Weymann, K.</string-name>
            </person-group>
            <year>2021</year>
            <article-title>277 Tunable White Light for Elders (TWLITE): A Feasibility Study of a Home-Based Sleep Intervention</article-title>
            <source>Sleep</source>
            <volume>44</volume>
            <pub-id pub-id-type="doi">10.1093/sleep/zsab072.276</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B15">
        <label>15.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Ou, W. and Bi, S. (2025) Sequential Decision-Making under Uncertainty: A Robust MDPs Review. <italic>Annals of Operations Research</italic>, 353, 1239-1285.</mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Ou, W.</string-name>
              <string-name>Bi, S.</string-name>
            </person-group>
            <year>2025</year>
            <article-title>Sequential Decision-Making under Uncertainty: A Robust MDPs Review</article-title>
            <source>Annals of Operations Research</source>
            <volume>353</volume>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B16">
        <label>16.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Barto, A.G., Sutton, R.S. and Watkins, C.J.C.H. (1989) Learning and Sequential Decision Making. Vol. 89, University of Massachusetts.</mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Barto, A.G.</string-name>
              <string-name>Sutton, R.S.</string-name>
              <string-name>Watkins, C.J.C.H.</string-name>
            </person-group>
            <year>1989</year>
            <article-title>Learning and Sequential Decision Making</article-title>
            <source>Vol. 89</source>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B17">
        <label>17.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Powell, W.B. (2021) From Reinforcement Learning to Optimal Control: A Unified Framework for Sequential Decisions. In: Vamvoudakis, K.G., Wan, Y., Lewis, F.L., and Cansever, D., Eds., <italic>Studies in Systems</italic>, <italic>Decision and Control</italic>, Springer International Publishing, 29-74. https://doi.org/10.1007/978-3-030-60990-0_3 <pub-id pub-id-type="doi">10.1007/978-3-030-60990-0_3</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1007/978-3-030-60990-0_3">https://doi.org/10.1007/978-3-030-60990-0_3</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Powell, W.B.</string-name>
              <string-name>Vamvoudakis, K.G.</string-name>
              <string-name>Wan, Y.</string-name>
              <string-name>Lewis, F.L.</string-name>
              <string-name>Cansever, D.</string-name>
              <string-name>Systems, D</string-name>
              <string-name>Control, S</string-name>
            </person-group>
            <year>2021</year>
            <article-title>From Reinforcement Learning to Optimal Control: A Unified Framework for Sequential Decisions</article-title>
            <source>In: Vamvoudakis</source>
            <volume>29</volume>
            <pub-id pub-id-type="doi">10.1007/978-3-030-60990-0_3</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B18">
        <label>18.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Dimitrakakis, C. and Ortner, R. (2022) Decision Making under Uncertainty and Reinforcement Learning: Theory and Algorithms. Vol. 223, Springer.</mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Dimitrakakis, C.</string-name>
              <string-name>Ortner, R.</string-name>
            </person-group>
            <year>2022</year>
            <article-title>Decision Making under Uncertainty and Reinforcement Learning: Theory and Algorithms</article-title>
            <source>Vol. 223</source>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B19">
        <label>19.</label>
        <citation-alternatives>
          <mixed-citation publication-type="thesis">Van Moffaert, K. (2016) Multicriteria Reinforcement Learning for Sequential Decision-Making Problems. Ph.D. Thesis, Vrije Universiteit Brussel.</mixed-citation>
          <element-citation publication-type="thesis">
            <person-group person-group-type="author">
              <string-name>Moffaert, K.</string-name>
              <string-name>Thesis, V</string-name>
            </person-group>
            <year>2016</year>
            <article-title>Multicriteria Reinforcement Learning for Sequential Decision-Making Problems</article-title>
            <source>Ph.D. Thesis</source>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B20">
        <label>20.</label>
        <citation-alternatives>
          <mixed-citation publication-type="book">van Otterlo Martijn, (2009) The Logic of Adaptive Behavior. In: <italic>Frontiers in Artificial Intelligence and Applications</italic>, Vol. 192, IOS Press, 1-489. https://doi.org/10.3233/978-1-58603-969-1-i <pub-id pub-id-type="doi">10.3233/978-1-58603-969-1-i</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.3233/978-1-58603-969-1-i">https://doi.org/10.3233/978-1-58603-969-1-i</ext-link></mixed-citation>
          <element-citation publication-type="book">
            <person-group person-group-type="author">
              <string-name>Applications, V</string-name>
            </person-group>
            <year>2009</year>
            <article-title>The Logic of Adaptive Behavior</article-title>
            <source>In: Frontiers in Artificial Intelligence and Applications</source>
            <volume>1</volume>
            <pub-id pub-id-type="doi">10.3233/978-1-58603-969-1-i</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B21">
        <label>21.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Jayaraman, P., Desman, J., Sabounchi, M., Nadkarni, G.N. and Sakhuja, A. (2024) A Primer on Reinforcement Learning in Medicine for Clinicians. <italic>NPJ Digital Medicine</italic>, 7, Article No. 337. https://doi.org/10.1038/s41746-024-01316-0 <pub-id pub-id-type="doi">10.1038/s41746-024-01316-0</pub-id><pub-id pub-id-type="pmid">39592855</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1038/s41746-024-01316-0">https://doi.org/10.1038/s41746-024-01316-0</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Jayaraman, P.</string-name>
              <string-name>Desman, J.</string-name>
              <string-name>Sabounchi, M.</string-name>
              <string-name>Nadkarni, G.N.</string-name>
              <string-name>Sakhuja, A.</string-name>
            </person-group>
            <year>2024</year>
            <article-title>A Primer on Reinforcement Learning in Medicine for Clinicians</article-title>
            <source>NPJ Digital Medicine</source>
            <volume>7</volume>
            <elocation-id>No</elocation-id>
            <pub-id pub-id-type="doi">10.1038/s41746-024-01316-0</pub-id>
            <pub-id pub-id-type="pmid">39592855</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B22">
        <label>22.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Munster, M. and Jamshidnejad, A. (2025) Personalized Human-Robot Cognitive Interaction via a Novel Fuzzy Logic Control and Learning-Based Paradigm. <italic>IEEE Access</italic>, 13, 112568-112593. https://doi.org/10.1109/access.2025.3584194 <pub-id pub-id-type="doi">10.1109/access.2025.3584194</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1109/access.2025.3584194">https://doi.org/10.1109/access.2025.3584194</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Munster, M.</string-name>
              <string-name>Jamshidnejad, A.</string-name>
            </person-group>
            <year>2025</year>
            <article-title>Personalized Human-Robot Cognitive Interaction via a Novel Fuzzy Logic Control and Learning-Based Paradigm</article-title>
            <source>IEEE Access</source>
            <volume>13</volume>
            <pub-id pub-id-type="doi">10.1109/access.2025.3584194</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B23">
        <label>23.</label>
        <citation-alternatives>
          <mixed-citation publication-type="thesis">Gönül, S. (2018) A Framework for Design and Personalization of Digital, Just-in-Time, Adaptive Interventions. Ph.D. Thesis, Middle East Technical University (Türkiye).</mixed-citation>
          <element-citation publication-type="thesis">
            <person-group person-group-type="author">
              <string-name>Digital, J</string-name>
              <string-name>Time, A</string-name>
              <string-name>Thesis, M</string-name>
            </person-group>
            <year>2018</year>
            <article-title>A Framework for Design and Personalization of Digital, Just-in-Time, Adaptive Interventions</article-title>
            <source>Ph.D. Thesis</source>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B24">
        <label>24.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Streicher, A. and Smeddinck, J.D. (2016) Personalized and Adaptive Serious Games. <italic>Entertainment Computing and Serious Games</italic>: <italic>International GI</italic>- <italic>Dagstuhl Seminar</italic>15283, Dagstuhl Castle, 5-10 July 2015, 332-377.</mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Streicher, A.</string-name>
              <string-name>Smeddinck, J.D.</string-name>
            </person-group>
            <year>2016</year>
            <article-title>Personalized and Adaptive Serious Games</article-title>
            <source>Entertainment Computing and Serious Games: International GI-Dagstuhl Seminar 15283</source>
            <volume>5</volume>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B25">
        <label>25.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Benedictis, R.D., Umbrico, A., Fracasso, F., Cortellessa, G., Orlandini, A. and Cesta, A. (2022) A Dichotomic Approach to Adaptive Interaction for Socially Assistive Robots. <italic>User Modeling and User</italic>- <italic>Adapted Interaction</italic>, 33, 293-331. https://doi.org/10.1007/s11257-022-09347-6 <pub-id pub-id-type="doi">10.1007/s11257-022-09347-6</pub-id><pub-id pub-id-type="pmid">36415674</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1007/s11257-022-09347-6">https://doi.org/10.1007/s11257-022-09347-6</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Benedictis, R.D.</string-name>
              <string-name>Umbrico, A.</string-name>
              <string-name>Fracasso, F.</string-name>
              <string-name>Cortellessa, G.</string-name>
              <string-name>Orlandini, A.</string-name>
              <string-name>Cesta, A.</string-name>
            </person-group>
            <year>2022</year>
            <article-title>A Dichotomic Approach to Adaptive Interaction for Socially Assistive Robots</article-title>
            <source>User Modeling and User-Adapted Interaction</source>
            <volume>33</volume>
            <pub-id pub-id-type="doi">10.1007/s11257-022-09347-6</pub-id>
            <pub-id pub-id-type="pmid">36415674</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B26">
        <label>26.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Ali, H. (2022) Reinforcement Learning in Healthcare: Optimizing Treatment Strategies, Dynamic Resource Allocation, and Adaptive Clinical Decision-Making. <italic>International Journal of Computer Applications Technology and Research</italic>, 11, 88-104.</mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Ali, H.</string-name>
              <string-name>Strategies, D</string-name>
            </person-group>
            <year>2022</year>
            <article-title>Reinforcement Learning in Healthcare: Optimizing Treatment Strategies, Dynamic Resource Allocation, and Adaptive Clinical Decision-Making</article-title>
            <source>International Journal of Computer Applications Technology and Research</source>
            <volume>11</volume>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B27">
        <label>27.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Yu, C., Liu, J., Nemati, S. and Yin, G. (2021) Reinforcement Learning in Healthcare: A Survey. <italic>ACM Computing Surveys</italic> ( <italic>CSUR</italic>), 55, 1-36. https://doi.org/10.1145/3477600 <pub-id pub-id-type="doi">10.1145/3477600</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1145/3477600">https://doi.org/10.1145/3477600</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Yu, C.</string-name>
              <string-name>Liu, J.</string-name>
              <string-name>Nemati, S.</string-name>
              <string-name>Yin, G.</string-name>
            </person-group>
            <year>2021</year>
            <article-title>Reinforcement Learning in Healthcare: A Survey</article-title>
            <source>ACM Computing Surveys (CSUR)</source>
            <volume>55</volume>
            <pub-id pub-id-type="doi">10.1145/3477600</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B28">
        <label>28.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Sachdeva, R.K., Bathla, P., Vij, S., Dishika, Jain, M., Kumar, L., Pradeep Ghantasala, G.S. and Ahuja, R. (2024) Emerging Technologies in Healthcare Systems. In: Mahajan, S., Raj, P. and Pandit, A.K., Eds., <italic>Deep Reinforcement Learning and Its Industrial Use Cases</italic>: <italic>AI for Real</italic><italic>‐</italic><italic>World Applications</italic>, 375-394. https://doi.org/10.1002/9781394272587.ch16 <pub-id pub-id-type="doi">10.1002/9781394272587.ch16</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1002/9781394272587.ch16">https://doi.org/10.1002/9781394272587.ch16</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Sachdeva, R.K.</string-name>
              <string-name>Bathla, P.</string-name>
              <string-name>Vij, S.</string-name>
              <string-name>Dishika, J</string-name>
              <string-name>Kumar, L.</string-name>
              <string-name>Ghantasala, G.S.</string-name>
              <string-name>Ahuja, R.</string-name>
              <string-name>Mahajan, S.</string-name>
              <string-name>Raj, P.</string-name>
              <string-name>Pandit, A.K.</string-name>
            </person-group>
            <year>2024</year>
            <article-title>Emerging Technologies in Healthcare Systems</article-title>
            <source>In: Mahajan</source>
            <volume>375</volume>
            <pub-id pub-id-type="doi">10.1002/9781394272587.ch16</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B29">
        <label>29.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Foluke Ekundayo, (2024) Reinforcement Learning in Treatment Pathway Optimization: A Case Study in Oncology. <italic>International Journal of Science and Research Archive</italic>, 13, 2187-2205. https://doi.org/10.30574/ijsra.2024.13.2.2450 <pub-id pub-id-type="doi">10.30574/ijsra.2024.13.2.2450</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.30574/ijsra.2024.13.2.2450">https://doi.org/10.30574/ijsra.2024.13.2.2450</ext-link></mixed-citation>
          <element-citation publication-type="journal">
            <year>2024</year>
            <article-title>Reinforcement Learning in Treatment Pathway Optimization: A Case Study in Oncology</article-title>
            <source>International Journal of Science and Research Archive</source>
            <volume>13</volume>
            <pub-id pub-id-type="doi">10.30574/ijsra.2024.13.2.2450</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B30">
        <label>30.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Ahmad, U.J., Bughio, H.K., Channar, N.A. and Bhatti, A.K. (2025) Reinforcement Learning in IoT-Driven Healthcare: Opportunities, Challenges, and Future Directions. <italic>Spectrum of Engineering Sciences</italic>, 3, 779-790.</mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Ahmad, U.J.</string-name>
              <string-name>Bughio, H.K.</string-name>
              <string-name>Channar, N.A.</string-name>
              <string-name>Bhatti, A.K.</string-name>
              <string-name>Opportunities, C</string-name>
            </person-group>
            <year>2025</year>
            <article-title>Reinforcement Learning in IoT-Driven Healthcare: Opportunities, Challenges, and Future Directions</article-title>
            <source>Spectrum of Engineering Sciences</source>
            <volume>3</volume>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B31">
        <label>31.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Shaik, T., Tao, X., Li, L., Xie, H., Dai, H., Zhao, F., <italic>et al</italic>. (2025) AI-Driven Multi-Agent Reinforcement Learning Framework for Real-Time Monitoring of Physiological Signals in Stress and Depression Contexts. <italic>Brain Informatics</italic>, 12, Article No. 14. https://doi.org/10.1186/s40708-025-00262-1 <pub-id pub-id-type="doi">10.1186/s40708-025-00262-1</pub-id><pub-id pub-id-type="pmid">40490570</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1186/s40708-025-00262-1">https://doi.org/10.1186/s40708-025-00262-1</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Shaik, T.</string-name>
              <string-name>Tao, X.</string-name>
              <string-name>Li, L.</string-name>
              <string-name>Xie, H.</string-name>
              <string-name>Dai, H.</string-name>
              <string-name>Zhao, F.</string-name>
            </person-group>
            <year>2025</year>
            <article-title>AI-Driven Multi-Agent Reinforcement Learning Framework for Real-Time Monitoring of Physiological Signals in Stress and Depression Contexts</article-title>
            <source>Brain Informatics</source>
            <volume>12</volume>
            <elocation-id>No</elocation-id>
            <pub-id pub-id-type="doi">10.1186/s40708-025-00262-1</pub-id>
            <pub-id pub-id-type="pmid">40490570</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B32">
        <label>32.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Costello, E.J., Pine, D.S., Hammen, C., March, J.S., Plotsky, P.M., Weissman, M.M., Biederman, J., <italic>et al</italic>. (2002) Development and Natural History of Mood Disorders. <italic>Biological Psychiatry</italic>, 52, 529-542. https://doi.org/10.1016/S0006-3223(02)01372-0 <pub-id pub-id-type="doi">10.1016/S0006-3223(02)01372-0</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1016/S0006-3223(02)01372-0">https://doi.org/10.1016/S0006-3223(02)01372-0</ext-link></mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Costello, E.J.</string-name>
              <string-name>Pine, D.S.</string-name>
              <string-name>Hammen, C.</string-name>
              <string-name>March, J.S.</string-name>
              <string-name>Plotsky, P.M.</string-name>
              <string-name>Weissman, M.M.</string-name>
              <string-name>Biederman, J.</string-name>
            </person-group>
            <year>2002</year>
            <article-title>Development and Natural History of Mood Disorders</article-title>
            <source>Biological Psychiatry</source>
            <volume>3223</volume>
            <issue>02</issue>
            <pub-id pub-id-type="doi">10.1016/S0006-3223(02)01372-0</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B33">
        <label>33.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Banumathi, K., Venkatesan, L., Benjamin, L.S., Vijayalakshmi, K. and Satchi, N.S. (2025) Reinforcement Learning in Personalized Medicine: A Comprehensive Review of Treatment Optimization Strategies. <italic>Cureus</italic>, 17, e82756.</mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Banumathi, K.</string-name>
              <string-name>Venkatesan, L.</string-name>
              <string-name>Benjamin, L.S.</string-name>
              <string-name>Vijayalakshmi, K.</string-name>
              <string-name>Satchi, N.S.</string-name>
            </person-group>
            <year>2025</year>
            <article-title>Reinforcement Learning in Personalized Medicine: A Comprehensive Review of Treatment Optimization Strategies</article-title>
            <source>Cureus</source>
            <volume>17</volume>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B34">
        <label>34.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Milic, J., Zrnic, I., Grego, E., Jovic, D., Stankovic, V., Djurdjevic, S., <italic>et al</italic>. (2025) The Role of Artificial Intelligence in Managing Bipolar Disorder: A New Frontier in Patient Care. <italic>Journal of Clinical Medicine</italic>, 14, Article 2515. https://doi.org/10.3390/jcm14072515 <pub-id pub-id-type="doi">10.3390/jcm14072515</pub-id><pub-id pub-id-type="pmid">40217964</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.3390/jcm14072515">https://doi.org/10.3390/jcm14072515</ext-link></mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Milic, J.</string-name>
              <string-name>Zrnic, I.</string-name>
              <string-name>Grego, E.</string-name>
              <string-name>Jovic, D.</string-name>
              <string-name>Stankovic, V.</string-name>
              <string-name>Djurdjevic, S.</string-name>
            </person-group>
            <year>2025</year>
            <article-title>The Role of Artificial Intelligence in Managing Bipolar Disorder: A New Frontier in Patient Care</article-title>
            <source>Journal of Clinical Medicine</source>
            <volume>14</volume>
            <elocation-id>2515</elocation-id>
            <pub-id pub-id-type="doi">10.3390/jcm14072515</pub-id>
            <pub-id pub-id-type="pmid">40217964</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B35">
        <label>35.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Ribba, B. (2023) Reinforcement Learning as an Innovative Model-Based Approach: Examples from Precision Dosing, Digital Health and Computational Psychiatry. <italic>Frontiers in Pharmacology</italic>, 13, Article ID: 1094281. https://doi.org/10.3389/fphar.2022.1094281 <pub-id pub-id-type="doi">10.3389/fphar.2022.1094281</pub-id><pub-id pub-id-type="pmid">36873047</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.3389/fphar.2022.1094281">https://doi.org/10.3389/fphar.2022.1094281</ext-link></mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Ribba, B.</string-name>
              <string-name>Dosing, D</string-name>
            </person-group>
            <year>2023</year>
            <article-title>Reinforcement Learning as an Innovative Model-Based Approach: Examples from Precision Dosing, Digital Health and Computational Psychiatry</article-title>
            <source>Frontiers in Pharmacology</source>
            <volume>13</volume>
            <fpage>109428</fpage>
            <elocation-id>ID</elocation-id>
            <pub-id pub-id-type="doi">10.3389/fphar.2022.1094281</pub-id>
            <pub-id pub-id-type="pmid">36873047</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B36">
        <label>36.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Weltz, J., Volfovsky, A. and Laber, E.B. (2022) Reinforcement Learning Methods in Public Health. <italic>Clinical Therapeutics</italic>, 44, 139-154. https://doi.org/10.1016/j.clinthera.2021.11.002 <pub-id pub-id-type="doi">10.1016/j.clinthera.2021.11.002</pub-id><pub-id pub-id-type="pmid">35058056</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1016/j.clinthera.2021.11.002">https://doi.org/10.1016/j.clinthera.2021.11.002</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Weltz, J.</string-name>
              <string-name>Volfovsky, A.</string-name>
              <string-name>Laber, E.B.</string-name>
            </person-group>
            <year>2022</year>
            <article-title>Reinforcement Learning Methods in Public Health</article-title>
            <source>Clinical Therapeutics</source>
            <volume>44</volume>
            <pub-id pub-id-type="doi">10.1016/j.clinthera.2021.11.002</pub-id>
            <pub-id pub-id-type="pmid">35058056</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B37">
        <label>37.</label>
        <citation-alternatives>
          <mixed-citation publication-type="confproc">Trella, A.L., Zhang, K.W., Nahum-Shani, I., Shetty, V., Doshi-Velez, F. and Murphy, S.A. (2023) Reward Design for an Online Reinforcement Learning Algorithm Supporting Oral Self-Care. <italic>Proceedings of the AAAI Conference on Artificial Intelligence</italic>, 37, 15724-15730. https://doi.org/10.1609/aaai.v37i13.26866 <pub-id pub-id-type="doi">10.1609/aaai.v37i13.26866</pub-id><pub-id pub-id-type="pmid">37637073</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1609/aaai.v37i13.26866">https://doi.org/10.1609/aaai.v37i13.26866</ext-link></mixed-citation>
          <element-citation publication-type="confproc">
            <person-group person-group-type="author">
              <string-name>Trella, A.L.</string-name>
              <string-name>Zhang, K.W.</string-name>
              <string-name>Nahum-Shani, I.</string-name>
              <string-name>Shetty, V.</string-name>
              <string-name>Doshi-Velez, F.</string-name>
              <string-name>Murphy, S.A.</string-name>
            </person-group>
            <year>2023</year>
            <article-title>Reward Design for an Online Reinforcement Learning Algorithm Supporting Oral Self-Care</article-title>
            <source>Proceedings of the AAAI Conference on Artificial Intelligence</source>
            <volume>37</volume>
            <pub-id pub-id-type="doi">10.1609/aaai.v37i13.26866</pub-id>
            <pub-id pub-id-type="pmid">37637073</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B38">
        <label>38.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Liebenow, B., Jones, R., DiMarco, E., Trattner, J.D., Humphries, J., Sands, L.P., <italic>et al</italic>. (2022) Computational Reinforcement Learning, Reward (and Punishment), and Dopamine in Psychiatric Disorders. <italic>Frontiers in Psychiatry</italic>, 13, Article ID: 886297. https://doi.org/10.3389/fpsyt.2022.886297 <pub-id pub-id-type="doi">10.3389/fpsyt.2022.886297</pub-id><pub-id pub-id-type="pmid">36339844</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.3389/fpsyt.2022.886297">https://doi.org/10.3389/fpsyt.2022.886297</ext-link></mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Liebenow, B.</string-name>
              <string-name>Jones, R.</string-name>
              <string-name>DiMarco, E.</string-name>
              <string-name>Trattner, J.D.</string-name>
              <string-name>Humphries, J.</string-name>
              <string-name>Sands, L.P.</string-name>
              <string-name>Learning, R</string-name>
            </person-group>
            <year>2022</year>
            <article-title>Computational Reinforcement Learning, Reward (and Punishment), and Dopamine in Psychiatric Disorders</article-title>
            <source>Frontiers in Psychiatry</source>
            <volume>13</volume>
            <fpage>886297</fpage>
            <elocation-id>ID</elocation-id>
            <pub-id pub-id-type="doi">10.3389/fpsyt.2022.886297</pub-id>
            <pub-id pub-id-type="pmid">36339844</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B39">
        <label>39.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Roggeveen, L.F., Hassouni, A.e., de Grooth, H., Girbes, A.R.J., Hoogendoorn, M. and Elbers, P.W.G. (2024) Reinforcement Learning for Intensive Care Medicine: Actionable Clinical Insights from Novel Approaches to Reward Shaping and Off-Policy Model Evaluation. <italic>Intensive Care Medicine Experimental</italic>, 12, Article No. 32. https://doi.org/10.1186/s40635-024-00614-x <pub-id pub-id-type="doi">10.1186/s40635-024-00614-x</pub-id><pub-id pub-id-type="pmid">38526681</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1186/s40635-024-00614-x">https://doi.org/10.1186/s40635-024-00614-x</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Roggeveen, L.F.</string-name>
              <string-name>Hassouni, A.</string-name>
              <string-name>Grooth, H.</string-name>
              <string-name>Girbes, A.R.J.</string-name>
              <string-name>Hoogendoorn, M.</string-name>
              <string-name>Elbers, P.W.G.</string-name>
            </person-group>
            <year>2024</year>
            <article-title>Reinforcement Learning for Intensive Care Medicine: Actionable Clinical Insights from Novel Approaches to Reward Shaping and Off-Policy Model Evaluation</article-title>
            <source>Intensive Care Medicine Experimental</source>
            <volume>12</volume>
            <elocation-id>No</elocation-id>
            <pub-id pub-id-type="doi">10.1186/s40635-024-00614-x</pub-id>
            <pub-id pub-id-type="pmid">38526681</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B40">
        <label>40.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Chakraborty, B. and Moodie, E.E. (2013) Statistical Methods for Dynamic Treatment Regimes. Vol. 2, Springer.</mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Chakraborty, B.</string-name>
              <string-name>Moodie, E.E.</string-name>
            </person-group>
            <year>2013</year>
            <article-title>Statistical Methods for Dynamic Treatment Regimes</article-title>
            <source>Vol. 2</source>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B41">
        <label>41.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Eisenberg, L. (1977) Disease and Illness Distinctions between Professional and Popular Ideas of Sickness. <italic>Culture</italic>, <italic>Medicine and Psychiatry</italic>, 1, 9-23. https://doi.org/10.1007/bf00114808 <pub-id pub-id-type="doi">10.1007/bf00114808</pub-id><pub-id pub-id-type="pmid">756356</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1007/bf00114808">https://doi.org/10.1007/bf00114808</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Eisenberg, L.</string-name>
              <string-name>Culture, M</string-name>
            </person-group>
            <year>1977</year>
            <article-title>Disease and Illness Distinctions between Professional and Popular Ideas of Sickness</article-title>
            <source>Culture</source>
            <volume>1</volume>
            <pub-id pub-id-type="doi">10.1007/bf00114808</pub-id>
            <pub-id pub-id-type="pmid">756356</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B42">
        <label>42.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Barbierato, E. and Gatti, A. (2024) The Challenges of Machine Learning: A Critical Review. <italic>Electronics</italic>, 13, Article 416. https://doi.org/10.3390/electronics13020416 <pub-id pub-id-type="doi">10.3390/electronics13020416</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.3390/electronics13020416">https://doi.org/10.3390/electronics13020416</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Barbierato, E.</string-name>
              <string-name>Gatti, A.</string-name>
            </person-group>
            <year>2024</year>
            <article-title>The Challenges of Machine Learning: A Critical Review</article-title>
            <source>Electronics</source>
            <volume>13</volume>
            <elocation-id>416</elocation-id>
            <pub-id pub-id-type="doi">10.3390/electronics13020416</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
    </ref-list>
  </back>
</article>