TITLE:
Reinforcement Learning Based Optimization of Sleep Mood Circadian Dynamics in Bipolar Disorder: A Simulation Study
AUTHORS:
Rocco de Filippis, Abdullah Al Foysal
KEYWORDS:
Bipolar Disorder, Reinforcement Learning, Proximal Policy Optimization, Sleep, Circadian Rhythm, Digital Psychiatry
JOURNAL NAME:
Open Access Library Journal,
Vol.13 No.3,
March
16,
2026
ABSTRACT: Bipolar disorder (BD) is closely intertwined with abnormalities in sleep and circadian regulation, yet current clinical management typically applies heuristic rules rather than optimizing these interacting processes in a principled way. We present a reinforcement-learning (RL) framework that learns personalized interventions for sleep timing, light exposure, daily activity, and medication adherence in a simulated BD setting. We design Circadian Environment, a physiologically inspired Markov decision process with five clinically interpretable state variables: sleep quality, sleep duration, mood stability, circadian alignment, and stress level. A continuous four-dimensional action space encodes modifiable behavioural and pharmacological levers. A Proximal Policy Optimization (PPO) agent (PPO-Patient-Agent), implemented in PyTorch with Gaussian policies, LayerNorm, entropy regularization, and gradient clipping, is trained over 1000 episodes (30 simulated days each) to maximize a composite reward reflecting key therapeutic goals: stable mood, high-quality and near-optimal sleep, strong circadian alignment, and low stress. Across training, episodic returns improve steadily and converge, indicating a stable policy. When evaluated in multiple post-training rollouts, the learned policy reliably drives virtual patients from moderately dysregulated baseline states into a high-functioning attractor characterized by near-maximal mood stability and sleep quality, robust circadian alignment, and minimal stress. Analysis of state trajectories reveals strong positive coupling between sleep, mood, and circadian variables and strong negative coupling between these variables and stress, aligning with clinical intuition. Although the current work uses a stylized simulator rather than real patient data, it establishes a transparent and extensible sandbox for prototyping RL-based treatment strategies in BD. The same architecture can be calibrated with digital phenotyping and longitudinal clinical data to support future development of safe, personalized decision-support tools for mood stabilization.