Interpretable Machine Learning Framework for Predicting Treatment Resistance in Psychiatric Disorders Using Synthetic Pharmacogenomic and Clinical Data

Abstract

This study presents a comprehensive and interpretable machine learning pipeline for predicting treatment resistance in psychiatric disorders using synthetically generated, multimodal data. The simulated dataset integrates gene expression profiles and clinically relevant features, such as age of onset, BMI, comorbidities, treatment response scores, and disease progression markers, thereby mimicking real-world pharmacogenomic and clinical complexities. The primary objective is to assess the feasibility of using artificial intelligence for individualized treatment stratification in psychiatry, with a strong emphasis on transparency and clinical relevance. Three classifiers—Random Forest, Gradient Boosting, and Calibrated Support Vector Machine (SVM)—were trained and evaluated using a balanced dataset of 2000 synthetic patients. Rigorous model validation was performed using key metrics, including ROC-AUC, F1-score, and balanced accuracy. Random Forest achieved the highest ROC-AUC (0.80) and balanced accuracy (0.72), followed closely by Gradient Boosting. Calibrated SVM exhibited lower performance but added methodological diversity. Feature importance was assessed using both traditional methods and permutation-based analysis, consistently highlighting GENE_1, GENE_6, Rapid_Cycling, and Lithium_Response as dominant predictors. Local model explainability was further enhanced using LIME, which provided individualized insights into each prediction. A full visual clinical report was generated for a sample patient, including gene expression summaries, predicted drug responses, and actionable recommendations. The pipeline demonstrates a reproducible and explainable framework for AI-driven clinical decision support in psychiatry. By bridging genetic markers, treatment outcomes, and machine learning interpretability, this study offers a template for precision psychiatry using realistic simulation models. The findings reinforce the value of integrating interpretability into ML models to promote trust and applicability in clinical practice.

Share and Cite:

de Filippis, R. and Al Foysal, A. (2025) Interpretable Machine Learning Framework for Predicting Treatment Resistance in Psychiatric Disorders Using Synthetic Pharmacogenomic and Clinical Data . Open Access Library Journal, 12, 1-26. doi: 10.4236/oalib.1113870.

1. Introduction

Treatment-resistant psychiatric disorders, such as major depressive disorder and bipolar disorder, represent a major burden for patients and healthcare systems alike [1]-[3]. These cases often fail to respond to multiple lines of standard pharmacological interventions, resulting in prolonged suffering, increased hospitalization rates, and higher healthcare costs [4]-[6]. Identifying patients at risk of treatment resistance early in the clinical pathway remains a critical challenge in psychiatry [7]-[10]. Recent advances in pharmacogenomics and Machine Learning (ML) have created promising avenues for personalized medicine, offering the ability to predict individual treatment response based on molecular and clinical profiles [11]-[13]. However, building such predictive models requires access to large, high-quality datasets that combine gene expression data, treatment outcomes, and diverse clinical features [14] [15]. In real-world settings, such multimodal datasets are often fragmented, limited in size, and burdened with missing or biased information [16] [17]. These constraints significantly limit the generalizability, robustness, and explainability of many ML-based healthcare solutions [18]-[20]. To address these limitations, this study proposes a novel and fully synthetic dataset that simulates realistic pharmacogenomic and clinical interactions observed in psychiatric populations. The dataset is used to train and evaluate a range of supervised machine learning classifiers—Random Forest, Gradient Boosting, and Calibrated Support Vector Machine—within a structured pipeline that includes preprocessing, feature selection, model evaluation, and interpretability. Techniques such as permutation importance and LIME (Local Interpretable Model-Agnostic Explanations) are used to enhance transparency, while individualized clinical reports demonstrate the model’s utility in real-world decision-making contexts [21]-[23]. By combining synthetic data generation, model performance evaluation, and patient-level interpretability, this study aims to validate a scalable framework for predicting treatment resistance in psychiatry. The overarching goal is to bridge the gap between algorithmic performance and clinical relevance, contributing to the future of explainable, AI-powered precision psychiatry.

2. Literature Review

2.1. Understanding Treatment Resistance in Psychiatry

Treatment resistance is a persistent and challenging issue in the clinical management of psychiatric disorders [24]-[26]. A significant portion of patients fail to respond adequately to standard pharmacological treatments, even after multiple therapeutic attempts [27] [28]. This lack of response is influenced by a complex interplay of genetic, biological, and psychosocial factors that vary widely across individuals. Traditional diagnostic approaches and treatment guidelines are often unable to accommodate this variability, resulting in suboptimal outcomes for many patients [29]-[31]. This has led to increasing interest in predictive models that can anticipate treatment resistance before clinical failure occurs.

2.2. Role of Genetic and Clinical Predictors

Both genetic markers and clinical features hold considerable potential in forecasting treatment outcomes [32] [33]. Specific genetic variations are known to influence drug metabolism, neurochemical pathways, and receptor sensitivity. At the same time, clinical variables such as age at onset, illness duration, number of prior episodes, comorbidities, and early treatment response indicators offer valuable insights into the trajectory of psychiatric disorders. When used in isolation, these features may have limited predictive power; however, when combined, they form a robust foundation for personalized predictions [34] [35]. This convergence of genetic and clinical signals is increasingly recognized as essential for precision psychiatry.

2.3. The Emergence of Machine Learning in Psychiatry

Machine learning has rapidly become a transformative tool in mental health research, enabling the development of predictive models that can uncover hidden patterns within complex, high-dimensional data [36] [37]. Models such as Random Forests, Support Vector Machines, and Gradient Boosting classifiers are particularly well-suited to the integration of heterogeneous data sources. These models can learn intricate, nonlinear relationships between features and outcomes, making them ideal candidates for forecasting treatment resistance. However, the adoption of machine learning in psychiatry remains limited by concerns related to model transparency, interpretability, and clinical applicability.

2.4. The Need for Explainable AI in Clinical Decision-Making

As machine learning models begin to inform clinical decisions, the need for interpretability becomes critical. Clinicians must understand not only what a model predicts but why it arrives at a particular conclusion. Explainable AI techniques such as feature importance rankings, local instance explanations, and intuitive visualizations serve to bridge this gap [38] [39]. By translating model outputs into human-understandable reasoning, these tools foster trust and accountability in real-world applications. In psychiatry, where decisions have long-term consequences and ethical weight, transparent AI is not a luxury but a necessity.

2.5. Synthetic Data as a Catalyst for Innovation

The availability of high-quality, labelled clinical data often poses a major limitation to the development of reliable machine learning models [40] [41]. Privacy concerns, data access restrictions, and small sample sizes can hinder model training and evaluation. In response, synthetic data generation has emerged as a powerful alternative. By simulating realistic patient cohorts with configurable clinical and genetic features, researchers can design controlled experiments, test model robustness, and explore feature interactions without ethical or logistical constraints. Importantly, synthetic data allows for the introduction of biologically meaningful signal patterns, enhancing the relevance of downstream analyses [42] [43].

2.6. Toward a Unified and Interpretable Framework

Despite progress in machine learning applications and synthetic data modelling, a fully integrated approach—combining biologically grounded synthetic data, predictive modelling, explainability, and clinical reporting—remains underdeveloped. Many existing studies focus on isolated components: model performance, gene associations, or feature selection [44]-[47]. There is a clear need for end-to-end frameworks that not only predict treatment outcomes but also provide interpretable, actionable feedback tailored to individual patients.

This research addresses that need by building a comprehensive pipeline that begins with realistic synthetic data and culminates in individual clinical insights. By aligning interpretability with predictive power, the proposed framework aims to advance the use of AI in psychiatric care and contribute meaningfully to the field of precision mental health.

3. Methodology

This section outlines the complete machine learning pipeline for predicting treatment resistance using synthetically simulated pharmacogenomic and clinical data. The process is divided into eight interconnected stages, visualized in Figure 1 below, which maps the sequential logic and integration points across data generation, modelling, evaluation, and explainability.

3.1. Synthetic Data Generation

To overcome limitations of real-world patient data availability, we generated a synthetic dataset comprising 2000 samples, each with 50 gene expression features and 15 clinical variables. Gene features were initialized using standard normal distributions. To inject domain-relevant signals:

  • GENE_1 was amplified in lithium responders.

  • GENE_6 was suppressed in valproate-sensitive individuals.

  • GENE_10 was boosted in lamotrigine responders.

Figure 1. Pipeline roadmap for treatment resistance prediction.

These controlled distortions simulate biological variability aligned with known pharmacogenetic interactions.

3.2. Clinical Feature Simulation

The clinical component includes demographics (e.g., age, BMI), illness characteristics (e.g., depression/mania episodes, age of onset), comorbidities (e.g., anxiety, substance use), and baseline severity scores (MADRS, YMRS). Notably:

  • Treatment response scores for lithium, valproate, and lamotrigine were modelled using beta distributions to generate bimodal separability between responders and non-responders.

  • A binary rapid cycling indicator was calculated from cumulative episode history using a probability threshold.

3.3. Target Definition: Treatment Resistance

The binary target variable Treatment Resistant was constructed using a weighted probabilistic function. Factors increasing the resistance probability included:

  • Low response scores to any of the three mood stabilizers.

  • Genetic risk signals in GENE_1, GENE_6, or GENE_10.

  • The presence of rapid cycling.

The output of this function was then sampled using a Bernoulli distribution, resulting in a resistance prevalence of ~46.5%.

3.4. Data Preprocessing

Features were visually inspected through distribution plots. Genes with extremely low variance were identified for removal; in this case, none met the exclusion threshold. The dataset was split into training (80%) and test (20%) sets using stratified sampling to preserve label balance. In addition to the 80/20 stratified split, stratified 5-fold cross-validation was conducted for robust performance estimation. This ensured balanced label distribution across folds and reduced variance in metric reporting.

3.5. Model Training

Three classification models were trained:

  • Random Forest: 300 trees, depth of 10, class_weight = balanced.

  • Gradient Boosting: 200 estimators, learning rate 0.05, depth of 7.

  • Calibrated SVM: RBF kernel with probability calibration (sigmoid) and data standardization via StandardScaler.

Training was conducted on the stratified training set. Each model was evaluated independently on the held-out test set. Hyperparameters for each model were selected via randomized grid search within a 5-fold cross-validation setup on the training set. The optimal values—such as 300 estimators and depth = 10 for Random Forest—were chosen based on ROC-AUC performance stability.

3.6. Evaluation Metrics

Model performance was quantified using:

  • Accuracy, Precision, Recall, F1-score.

  • Balanced Accuracy.

  • ROC-AUC Score.

Visual diagnostics included:

  • Confusion matrices to reveal class-wise error patterns.

  • ROC curves to inspect sensitivity-specificity trade-offs.

  • Yellow brick-generated classification heatmaps.

All performance metrics were averaged across the 5 validation folds, with final model testing conducted on the holdout 20% set. This two-level evaluation ensured both internal consistency and fair generalization assessment.

3.7. Explainability: Feature & Local Importance

To ensure interpretability:

  • Permutation Importance (n = 10 repeats) ranked top contributors based on model score degradation.

  • Standard Feature Importances were extracted from tree-based models for visual comparison.

  • LIME (Local Interpretable Model-agnostic Explanations) was used to explain specific test samples. LIME decomposes model predictions into a weighted sum of interpretable features, which is useful for clinical transparency.

3.8. Clinical Report Generator

A custom reporting tool was implemented for individual-level explainability, producing:

  • Prediction and confidence score.

  • Visual summaries of top 10 features and gene markers.

  • Treatment response bars.

  • Clinical flagging for actionable features (e.g., low lithium response + GENE_1 signal).

  • A set of human-readable clinical recommendations.

This report simulates the usability of the model as a clinical decision support system.

4. Results and Visualization

This section presents the outcomes of the machine learning framework applied to the synthetically generated dataset. It encompasses visual distribution analysis, model performance metrics, feature importance evaluation, and interpretability-based clinical insights. The five figures included illustrate critical components of the pipeline: gene distributions, classifier performance, and interpretability metrics.

4.1. Feature Distribution Analysis

The distribution of the first 15 gene expression features is presented in Figure 2. Most features approximate Gaussian curves, yet a few exhibit distinctly non-normal or multimodal patterns. Notably:

Figure 2. Distribution histograms for the first 15 gene features, revealing Gaussian, skewed, and multimodal patterns.

  • GENE_1 displays a skewed bimodal distribution, reflecting the artificial signal injection for lithium response simulation.

  • GENE_6 and GENE_11 also show skewness and peaks that correspond to treatment-related transformations in the data generation stage.

Such patterns confirm that the synthetic data generation effectively introduces biological heterogeneity representative of real pharmacogenomic data.

4.2. Model Performance Evaluation

Model accuracy and diagnostic metrics were evaluated across Random Forest, Gradient Boosting, and Calibrated SVM classifiers [48] [49]. Random Forest was selected for initial visualization due to its interpretability and robust performance [50] [51].

The confusion matrix for the Random Forest classifier is shown in Figure 3, indicating a balanced prediction capability:

  • True Positives (Resistant correctly identified): 128.

  • True Negatives (Responsive correctly identified): 162.

  • False Positives: 52.

  • False Negatives: 58.

The ROC curve is presented in Figure 4, reflecting the model’s ability to discriminate between classes across thresholds. An area under the curve (AUC) of 0.80 confirms strong discriminative performance.

Additionally, a detailed classification report heatmap is illustrated in Figure 5. It reveals:

  • F1-score of 0.699 for the resistant class (label 1).

  • F1-score of 0.747 for the responsive class (label 0).

  • Balanced overall precision and recall across both categories.

Figure 3. Confusion matrix for Random Forest classifier showing model accuracy across treatment classes.

Figure 4. ROC curve for Random Forest showing an AUC of 0.80.

Figure 5. Classification report heatmap displaying per-class precision, recall, and F1-score for the Random Forest classifier.

4.3. Summary of Classifier Metrics

Model

AUC

F1 Score

Balanced Accuracy

Random Forest

0.80

0.699

0.72

Gradient Boosting

0.80

0.708

0.72

Calibrated SVM

0.75

0.626

0.66

These results confirm that both tree-based ensemble models outperformed SVM, offering stronger generalization to unseen patient samples.

4.4. Feature Importance Analysis: Permutation Scores (Random Forest)

A permutation importance analysis was conducted to interpret the underlying drivers of the Random Forest classifier. This method assesses the decrease in model performance when the values of a given feature are randomly permuted, thereby estimating the true dependency of the model on that feature. As shown in Figure 6, the most influential features include both genetic markers and clinical indicators. Specifically:

  • GENE_1 shows the highest importance with a substantial gap compared to all other features.

  • GENE_6, Rapid_Cycling, and treatment response scores such as Valproate_ Response and Lithium_Response also rank among the most impactful.

  • Clinical attributes like Age_Onset, Lamotrigine_Response, and BMI add moderate but meaningful contributions.

The results indicate that treatment resistance prediction is strongly driven by a blend of molecular and phenotypic factors. The boxplots represent the distribution of importance values across multiple permutation runs, capturing both the central tendency and variance of each feature’s influence. Features with narrow interquartile ranges and high mean importance, such as GENE_1, are considered the most stable and reliable predictors.

Figure 6. Permutation importance (Top 15 Features)—Random Forest Model. Box plots display the change in prediction accuracy when each feature is permuted. The most important features include GENE_1, GENE_6, Rapid_Cycling, and core treatment response scores. Higher values and tighter confidence intervals denote stronger and more stable feature influence.

4.5. Feature Importance Analysis: Tree-Based Scores (Random Forest)

In addition to permutation-based evaluation, we extracted the built-in feature importance scores from the Random Forest model, based on the frequency and quality of feature splits in the decision trees.

As shown in Figure 7, the feature importance ranking largely overlaps with the permutation results. The most critical features remain:

  • GENE_1 (most decisive contributor).

  • GENE_6, Rapid_Cycling, and Lithium_Response.

  • Clinical markers such as Valproate_Response, Lamotrigine_Response, and Age_Onset.

This consistency across analytical methods strengthens confidence in the model’s biological grounding and highlights that both genetic and treatment-derived clinical features contribute significantly to model behavior.

Figure 7. Random forest feature importance (Top 15).

The bar chart in Figure 7 illustrates feature importance scores calculated from decision node contributions. Results align closely with permutation rankings, confirming the dominant role of GENE_1 and mood stabilizer response indicators.

4.6. Performance Evaluation: Gradient Boosting Classifier

To benchmark Random Forest against another ensemble method, we trained and evaluated a Gradient Boosting Classifier. The confusion matrix, shown in Figure 8, reveals a well-balanced classification outcome:

  • 152 responsive and 136 resistant patients were correctly identified.

  • 112 instances were misclassified, like the Random Forest model.

The matrix shows true positive and negative rates across treatment response classes. The model maintains consistent classification across both categories. The model’s ROC curve, presented in Figure 9, demonstrates strong discriminatory capacity, achieving an AUC of 0.80, nearly identical to that of Random Forest.

The model achieves an AUC of 0.80, indicating a high sensitivity and specificity trade-off across thresholds.

The classification report in Figure 10 confirms that the model performs comparably across both classes:

  • F1-score: 0.708 (Resistant), 0.731 (Responsive).

  • Balanced precision and recall.

Gradient Boosting shows consistent metrics across classes, slightly improving on resistant class recall compared to Random Forest.

Figure 8. Confusion matrix—gradient boosting model.

Figure 9. ROC curve—gradient boosting classifier.

Figure 10. Classification metrics heatmap—gradient boosting.

4.7. Feature Importance Analysis: Permutation and Built-in Scores (Gradient Boosting)

As with Random Forest, we examined both permutation importance and built-in feature scores for the Gradient Boosting model.

Figure 11. Permutation importance—gradient boosting (Top 15 Features).

Figure 11 displays the permutation-based results. Like prior findings, GENE_1, GENE_6, Rapid_Cycling, and the three drug response scores dominate the ranking. New entries, such as Illness_Duration, GENE_32, and GENE_24, suggest the Gradient Boosting model captures subtle temporal and genomic effects.

This plot highlights features that most affect model performance when permuted. Both genetic and temporal clinical variables were influential.

Figure 12 shows the standard Gradient Boosting feature importance scores. These reinforce previous findings, again led by GENE_1, GENE_6, and mood stabilizer response indicators.

Figure 12. Gradient boosting feature importance scores.

Top features remain consistent, with strong emphasis on GENE_1, GENE_6, Rapid_Cycling, and medication response metrics.

4.8. Calibrated SVM Performance and Local Explainability

To complement ensemble-based classifiers, a Calibrated Support Vector Machine (SVM) model was introduced [52]-[54]. The SVM was paired with Platt scaling to enable probabilistic outputs, which are essential for ROC curve evaluation and explainability integration such as LIME.

The confusion matrix in Figure 13 shows that the model correctly identified 150 responsive and 114 resistant patients. However, it also produced 64 false positives and 72 false negatives—more errors compared to the Random Forest and Gradient Boosting models.

Figure 13. Confusion matrix—calibrated SVM.

Figure 14. ROC Curve—Calibrated SVM.

The model shows slightly reduced performance in classifying treatment-resistant individuals, reflecting its linear margin limitations in high-dimensional biomedical data. The ROC curve shown in Figure 14 presents an AUC of 0.75, which, while acceptable, is noticeably lower than the AUC of 0.80 observed in tree-based models. The curve shape indicates moderate discrimination capacity, especially in mid-sensitivity ranges.

The AUC value of 0.75 indicates fair classification capability but suggests that the SVM model underperforms when distinguishing resistant vs. responsive cases in complex feature space.

The performance metrics in Figure 15 confirm this trend:

  • F1-score for resistant class: 0.626.

  • F1-score for responsive class: 0.688.

  • Overall accuracy: 66%.

  • Balanced accuracy: 65.69%.

Figure 15. Classification report – calibrated SVM.

Lower recall and F1-score in the resistant group confirm limited generalization power compared to ensemble methods. However, the calibrated output makes the model usable in probabilistic frameworks. To explore model decisions at the patient level, LIME (Local Interpretable Model-Agnostic Explanations) was used [55] [56]. Figure 16 presents a bar chart of feature contributions for patient #0, who was classified as treatment responsive by the Random Forest model with approximately 80% confidence.

The chart in Figure 16 shows how individual features contributed to reducing (red) or increasing (green) the likelihood of treatment resistance. Key variables include Rapid_Cycling, GENE_1, GENE_6, and Lithium_Response.

Insights from LIME:

  • Decreased resistance risk:

  • Rapid_Cycling ≤ 0.0

  • GENE_1 ≤ −0.70, GENE_6 > 0.66

  • High scores in Lithium_Response and Valproate_Response

  • Mild risk contributors:

  • GENE_11 > 0.64

  • Low BMI, presence of GENE_23, and elevated Illness_Duration

Figure 16. LIME explanation for sample 0.

These interpretations provide personalized transparency, reinforcing the clinical plausibility of the model and increasing its trustworthiness in psychiatric decision support.

4.9. Individual-Level Clinical Report and Actionable Insights

To translate predictive results into clinically actionable intelligence, we generated a full visual and interpretive report for patient 0, predicted by the Random Forest classifier to be treatment responsive. This integrated visualization, presented in Figure 17, combines feature-level insights, genetic expression data, treatment response estimations, and model-derived clinical recommendations. For interpretive clarity, we describe its five subpanels as Figures 17(a)-(e).

The Random Forest classifier predicted this patient to be treatment responsive, with a resistance probability of 19.80%, significantly below the clinical threshold. The prediction was correct, aligning with the true label. This affirms the model’s interpretability and real-world applicability in patient stratification.

Figure 17(a)—Top Predictive Featuresl: This subpanel displays the top 10 features by global importance score derived from the Random Forest model. It highlights GENE_1, GENE_6, and Rapid_Cycling as the most influential predictors. Notably, pharmacological response markers—Lithium_Response, Valproate_Response, and Lamotrigine_Response—also appear prominently, emphasizing the model’s balanced use of genetic and clinical features.

Figure 17. Clinical interpretation report for patient 0. (a) Top predictive features ranked by global importance. (b) Patient-level values for Top 10 features. (c) Predicted treatment response scores with resistance threshold. (d) Top genetic marker expression levels. (e) Model-based clinical recommendations.

Figure 17(b)—Patient Feature Values: This bar chart shows the actual values of patient 0 for the same top 10 features. The high values in GENE_6 and GENE_10 align with treatment responsiveness, while GENE_1’s low expression contributes negatively to resistance risk. The clinical indicator Rapid_Cycling is zero, a protective factor. This subpanel supports direct interpretability of the input vector that drove the model’s decision.

Figure 17(c)—Treatment Response Scores: This plot illustrates predicted scores for three core mood stabilizers: Lithium (0.86), Valproate (0.77), and Lamotrigine (0.77). All values lie well above the resistance threshold (0.3), represented by a red dashed line. This provides a strong therapeutic signal indicating that all three medications are likely to be effective in this individual case.

Figure 17(d)—Genetic Marker Expression: The genetic panel focuses on the expression levels of the top 5 genes. GENE_3 and GENE_4 show high expression and are typically correlated with positive outcomes. Conversely, GENE_1 and GENE_5 are under-expressed, both of which were associated with treatment resistance in broader model training. This provides a molecular justification for the model’s response prediction.

Figure 17(e)—Clinical Recommendations: Based on model inference and feature analysis, four clear, model-backed interventions were proposed:

1) Begin standard mood stabilizer regimen prioritizing lithium and lamotrigine.

2) Monitor early treatment progress to validate predicted efficacy.

3) Plan for maintenance phase due to promising response indicators.

4) Schedule structured psychiatric follow-ups to ensure continued adherence and outcome tracking.

This personalized set of recommendations demonstrates how machine learning can go beyond risk scores to directly support clinical decisions in psychiatric practice.

5. Discussion

This study demonstrates the feasibility and effectiveness of using interpretable machine learning models trained on synthetically generated, biologically informed datasets to predict treatment resistance in psychiatric conditions. The integration of pharmacogenomic and clinical features in a simulated but realistic data environment allowed us to capture key patterns observed in real-world psychiatric populations while overcoming common limitations such as missing values, small sample sizes, and privacy restrictions. Our results revealed that ensemble-based models—particularly Random Forest and Gradient Boosting—outperformed the Calibrated SVM in both overall classification performance and robustness across evaluation metrics. The Random Forest classifier achieved a balanced accuracy of 0.72 and an AUC of 0.80, indicating a strong ability to generalize across both treatment-resistant and responsive patient profiles. These findings highlight the suitability of tree-based models in high-dimensional biomedical settings, where variable interactions and non-linear relationships are common. The identification of key predictors such as GENE_1, GENE_6, Rapid_Cycling, and Lithium_Response across both traditional and permutation-based importance metrics reinforces the biological credibility of the synthetic dataset. These features are consistent with known pharmacogenomic and clinical determinants of mood stabilizer efficacy, lending face validity to the simulation approach. Moreover, the LIME explanations provided highly granular insights into individual predictions, effectively bridging the gap between model output and clinical reasoning. Such interpretability tools are critical for building clinician trust and facilitating the integration of AI systems into psychiatric decision-making workflows [57]-[59]. Importantly, the individualized clinical report developed in this study goes beyond abstract model metrics to demonstrate how machine learning predictions can be translated into patient-specific recommendations. This form of output is crucial for real-world applicability, particularly in complex domains like psychiatry where diagnostic ambiguity and treatment heterogeneity are common [60]-[62]. Nevertheless, while synthetic data provides a controlled environment for model development, future research must validate these findings using real-world patient cohorts. Incorporating longitudinal data, such as treatment timelines and symptom evolution, could further improve model precision and provide deeper insight into temporal treatment dynamics. Additionally, extending the framework to include multimodal inputs—such as neuroimaging, electronic health records, or wearable data—may enhance generalizability and clinical relevance. This work presents a replicable, interpretable, and clinically grounded machine learning framework for precision psychiatry, setting the stage for future translational AI research in mental health care.

6. Conclusion

This study presents a fully integrated, explainable machine learning framework for predicting treatment resistance in psychiatric disorders using a synthetically generated, multimodal dataset. By simulating realistic interactions between genetic markers, clinical variables, and pharmacological response indicators, we addressed a fundamental limitation in psychiatric AI research—the lack of large, high-quality, and privacy-compliant datasets [63] [64]. The proposed synthetic pipeline serves as a testbed for developing and validating predictive models that are both statistically rigorous and clinically interpretable. Among the models tested, Random Forest and Gradient Boosting consistently achieved strong performance, with ROC-AUC values reaching 0.80 and balanced accuracies exceeding 0.72. These results highlight the potential of ensemble learning to handle complex, non-linear relationships inherent in psychiatric data. Additionally, feature importance and LIME-based explanations ensured that the model’s decisions were transparent, reinforcing trust and enabling real-time clinical interpretation [65] [66]. The inclusion of a comprehensive, individualized clinical report further demonstrates how machine learning outputs can be translated into actionable insights at the patient level. This bridges the gap between data science and frontline psychiatry, facilitating personalized treatment planning informed by biological and clinical indicators [67] [68]. In conclusion, our framework offers a reproducible and interpretable approach to AI-driven precision psychiatry. It serves as a foundation for future research that incorporates real patient data, longitudinal trajectories, and multimodal signals to advance the clinical application of machine learning in mental health care.

Conflicts of Interest

The authors declare no conflicts of interest.

Conflicts of Interest

The authors declare no conflicts of interest.

References

[1] Zhdanava, M., Pilon, D., Ghelerter, I., Chow, W., Joshi, K., Lefebvre, P., et al. (2021) The Prevalence and National Burden of Treatment-Resistant Depression and Major Depressive Disorder in the United States. The Journal of Clinical Psychiatry, 82, 20m13699.[CrossRef] [PubMed]
[2] Baig-Ward, K.M., Jha, M.K. and Trivedi, M.H. (2023) The Individual and Societal Burden of Treatment-Resistant Depression: An Overview. Psychiatric Clinics of North America, 46, 211-226.[CrossRef] [PubMed]
[3] Sousa, R.D., Gouveia, M., Nunes da Silva, C., Rodrigues, A.M., Cardoso, G., Antunes, A.F., et al. (2022) Treatment-resistant Depression and Major Depression with Suicide Risk—The Cost of Illness and Burden of Disease. Frontiers in Public Health, 10, Article 898491.[CrossRef] [PubMed]
[4] Benjamin, D.M. (2003) Reducing Medication Errors and Increasing Patient Safety: Case Studies in Clinical Pharmacology. Journal of Clinical Pharmacology, 43, 768-783.[CrossRef] [PubMed]
[5] Al Hamid, A., Ghaleb, M., Aljadhey, H. and Aslanpour, Z. (2014) A Systematic Review of Hospitalization Resulting from Medicine‐related Problems in Adult Patients. British Journal of Clinical Pharmacology, 78, 202-217.[CrossRef] [PubMed]
[6] Schneiderman, L.J. and Jecker, N.S. (2011) Wrong Medicine: Doctors, Patients, and Futile Treatment. JHU Press.
[7] Howes, O.D., Thase, M.E. and Pillinger, T. (2021) Treatment Resistance in Psychiatry: State of the Art and New Directions. Molecular Psychiatry, 27, 58-72.[CrossRef] [PubMed]
[8] Kane, J.M., Agid, O., Baldwin, M.L., Howes, O., Lindenmayer, J., Marder, S., et al. (2019) Clinical Guidance on the Identification and Management of Treatment-Resistant Schizophrenia. The Journal of Clinical Psychiatry, 80, 18com12123.[CrossRef] [PubMed]
[9] Perlis, R.H. (2013) A Clinical Risk Stratification Tool for Predicting Treatment Resistance in Major Depressive Disorder. Biological Psychiatry, 74, 7-14.[CrossRef] [PubMed]
[10] Dodd, S., Bauer, M., Carvalho, A.F., Eyre, H., Fava, M., Kasper, S., et al. (2020) A Clinical Approach to Treatment Resistance in Depressed Patients: What to Do When the Usual Treatments Don’t Work Well Enough? The World Journal of Biological Psychiatry, 22, 483-494.[CrossRef] [PubMed]
[11] Vadapalli, S., Abdelhalim, H., Zeeshan, S. and Ahmed, Z. (2022) Artificial Intelligence and Machine Learning Approaches Using Gene Expression and Variant Data for Personalized Medicine. Briefings in Bioinformatics, 23, bbac191.[CrossRef] [PubMed]
[12] Taherdoost, H. and Ghofrani, A. (2024) AI’s Role in Revolutionizing Personalized Medicine by Reshaping Pharmacogenomics and Drug Therapy. Intelligent Pharmacy, 2, 643-650.
[13] Feng, F., Shen, B., Mou, X., Li, Y. and Li, H. (2021) Large-Scale Pharmacogenomic Studies and Drug Response Prediction for Personalized Cancer Medicine. Journal of Genetics and Genomics, 48, 540-551.[CrossRef] [PubMed]
[14] Martínez-García, M. and Hernández-Lemus, E. (2022) Data Integration Challenges for Machine Learning in Precision Medicine. Frontiers in Medicine, 8, Article 784455.[CrossRef] [PubMed]
[15] Zitnik, M., Nguyen, F., Wang, B., Leskovec, J., Goldenberg, A. and Hoffman, M.M. (2019) Machine Learning for Integrating Data in Biology and Medicine: Principles, Practice, and Opportunities. Information Fusion, 50, 71-91.[CrossRef] [PubMed]
[16] Muller, H. and Unay, D. (2017) Retrieval from and Understanding of Large-Scale Multi-Modal Medical Datasets: A Review. IEEE Transactions on Multimedia, 19, 2093-2104.[CrossRef]
[17] Santos, D. and Barbeiro, R.B. (2024) Improving the Robustness of Multimodal AI with Asynchronous and Missing Inputs. Ph.D. Thesis, NOVA University Lisbon.
[18] Kavyashree, N., Surekha, R., Arya, A., Phadke, M.M., Khan, M.A. and Khetre, R.D. (2024) ML and AI Challenges and Applications in Healthcare. African Journal of Biological Sciences, 6, 509-527.
[19] Ennab, M. and Mcheick, H. (2024) Enhancing Interpretability and Accuracy of AI Models in Healthcare: A Comprehensive Review on Challenges and Future Directions. Frontiers in Robotics and AI, 11, Article 1444763.[CrossRef] [PubMed]
[20] Petersen, E., Potdevin, Y., Mohammadi, E., Zidowitz, S., Breyer, S., Nowotka, D., et al. (2022) Responsible and Regulatory Conform Machine Learning for Medicine: A Survey of Challenges and Solutions. IEEE Access, 10, 58375-58418.[CrossRef]
[21] Mienye, I.D., Obaido, G., Jere, N., Mienye, E., Aruleba, K., Emmanuel, I.D., et al. (2024) A Survey of Explainable Artificial Intelligence in Healthcare: Concepts, Applications, and Challenges. Informatics in Medicine Unlocked, 51, Article 101587.[CrossRef]
[22] Tursunalieva, A., Alexander, D.L.J., Dunne, R., Li, J., Riera, L. and Zhao, Y. (2024) Making Sense of Machine Learning: A Review of Interpretation Techniques and Their Applications. Applied Sciences, 14, Article 496.[CrossRef]
[23] Valente, F., Paredes, S., Henriques, J., Rocha, T., de Carvalho, P. and Morais, J. (2022) Interpretability, Personalization and Reliability of a Machine Learning Based Clinical Decision Support System. Data Mining and Knowledge Discovery, 36, 1140-1173.[CrossRef]
[24] Velligan, D.I., Weiden, P.J., Sajatovic, M., Scott, J., Carpenter, D., Ross, R., et al. (2010) Strategies for Addressing Adherence Problems in Patients with Serious and Persistent Mental Illness: Recommendations from the Expert Consensus Guidelines. Journal of Psychiatric Practice, 16, 306-324.[CrossRef] [PubMed]
[25] Pompili, M. and Fiorillo, A. (2017) Editorial: Unmet Needs in Modern Psychiatry. CNS & Neurological Disorders-Drug Targets, 16, 857-857.[CrossRef] [PubMed]
[26] Voineskos, D., Daskalakis, Z.J. and Blumberger, D.M. (2020) Management of Treatment-Resistant Depression: Challenges and Strategies. Neuropsychiatric Disease and Treatment, 16, 221-234.[CrossRef] [PubMed]
[27] Huang, J. and Hunt, R.H. (1999) Treatment after Failure: The Problem of “Non-Responders”. Gut, 45, i40-i44.[CrossRef] [PubMed]
[28] Eichler, H., Abadie, E., Breckenridge, A., Flamion, B., Gustafsson, L.L., Leufkens, H., et al. (2011) Bridging the Efficacy-Effectiveness Gap: A Regulator’s Perspective on Addressing Variability of Drug Response. Nature Reviews Drug Discovery, 10, 495-506.[CrossRef] [PubMed]
[29] Croft, P., Altman, D.G., Deeks, J.J., Dunn, K.M., Hay, A.D., Hemingway, H., et al. (2015) The Science of Clinical Practice: Disease Diagnosis or Patient Prognosis? Evidence about “What Is Likely to Happen” Should Shape Clinical Practice. BMC Medicine, 13, Article No. 20.[CrossRef] [PubMed]
[30] Kent, D.M., Steyerberg, E. and van Klaveren, D. (2018) Personalized Evidence Based Medicine: Predictive Approaches to Heterogeneous Treatment Effects. BMJ, 363, k4245.[CrossRef] [PubMed]
[31] Stein, D.J., Shoptaw, S.J., Vigo, D.V., Lund, C., Cuijpers, P., Bantjes, J., et al. (2022) Psychiatric Diagnosis and Treatment in the 21st Century: Paradigm Shifts versus Incremental Integration. World Psychiatry, 21, 393-414.[CrossRef] [PubMed]
[32] Roukos, D.H., Murray, S. and Briasoulis, E. (2007) Molecular Genetic Tools Shape a Roadmap Towards a More Accurate Prognostic Prediction and Personalized Management of Cancer. Cancer Biology & Therapy, 6, 308-312.[CrossRef] [PubMed]
[33] Nevins, J.R., Huang, E.S., Dressman, H., Pittman, J., Huang, A.T. and West, M. (2003) Towards Integrated Clinico-Genomic Models for Personalized Medicine: Combining Gene Expression Signatures and Clinical Factors in Breast Cancer Outcomes Prediction. Human Molecular Genetics, 12, R153-R157.[CrossRef] [PubMed]
[34] Wottschel, V., Alexander, D.C., Kwok, P.P., Chard, D.T., Stromillo, M.L., De Stefano, N., et al. (2015) Predicting Outcome in Clinically Isolated Syndrome Using Machine Learning. NeuroImage: Clinical, 7, 281-287.[CrossRef] [PubMed]
[35] Lambin, P., Leijenaar, R.T.H., Deist, T.M., Peerlings, J., de Jong, E.E.C., van Timmeren, J., et al. (2017) Radiomics: The Bridge between Medical Imaging and Personalized Medicine. Nature Reviews Clinical Oncology, 14, 749-762.[CrossRef] [PubMed]
[36] Miotto, R., Wang, F., Wang, S., Jiang, X. and Dudley, J.T. (2017) Deep Learning for Healthcare: Review, Opportunities and Challenges. Briefings in Bioinformatics, 19, 1236-1246.[CrossRef] [PubMed]
[37] Thieme, A., Belgrave, D. and Doherty, G. (2020) Machine Learning in Mental Health: A Systematic Review of the HCI Literature to Support the Development of Effective and Implementable ML Systems. ACM Transactions on Computer-Human Interaction, 27, 1-53.[CrossRef]
[38] Alicioglu, G. and Sun, B. (2022) A Survey of Visual Analytics for Explainable Artificial Intelligence Methods. Computers & Graphics, 102, 502-520.[CrossRef]
[39] Minh, D., Wang, H.X., Li, Y.F. and Nguyen, T.N. (2022) Explainable Artificial Intelligence: A Comprehensive Review. Artificial Intelligence Review, 55, 3503-3568.
[40] Ghassemi, M., Naumann, T., Schulam, P., Beam, A.L., Chen, I.Y. and Ranganath, R. (2020) A Review of Challenges and Opportunities in Machine Learning for Health. AMIA Summits on Translational Science Proceedings, 2020, 191-200.
[41] Ching, T., Himmelstein, D.S., Beaulieu-Jones, B.K., Kalinin, A.A., Do, B.T., Way, G.P., et al. (2018) Opportunities and Obstacles for Deep Learning in Biology and Medicine. Journal of the Royal Society Interface, 15, Article ID: 20170387.[CrossRef] [PubMed]
[42] Achuthan, S., Chatterjee, R., Kotnala, S., Mohanty, A., Bhattacharya, S., Salgia, R., et al. (2022) Leveraging Deep Learning Algorithms for Synthetic Data Generation to Design and Analyze Biological Networks. Journal of Biosciences, 47, Article No. 43.[CrossRef] [PubMed]
[43] van Breugel, B., Liu, T., Oglic, D. and van der Schaar, M. (2024) Synthetic Data in Biomedicine via Generative Artificial Intelligence. Nature Reviews Bioengineering, 2, 991-1004.[CrossRef]
[44] Ang, J.C., Mirzal, A., Haron, H. and Hamed, H.N.A. (2016) Supervised, Unsupervised, and Semi-Supervised Feature Selection: A Review on Gene Selection. IEEE/ACM Transactions on Computational Biology and Bioinformatics, 13, 971-989.[CrossRef] [PubMed]
[45] Alelyani, S., Tang, J. and Liu, H. (2018) Feature Selection for Clustering: A Review. In: Aggarwal, C.C. and Reddy, C.K., Eds., Data Clustering, Chapman and Hall/CRC, 29-60.[CrossRef]
[46] Hall, M.A. (1999) Correlation-Based Feature Selection for Machine Learning. Ph.D. Thesis, The University of Waikato.
[47] Liu, H. and Motoda, H. (2012) Feature Selection for Knowledge Discovery and Data Mining. Springer.
[48] Ozcift, A. and Gulten, A. (2011) Classifier Ensemble Construction with Rotation Forest to Improve Medical Diagnosis Performance of Machine Learning Algorithms. Computer Methods and Programs in Biomedicine, 104, 443-451.[CrossRef] [PubMed]
[49] Mohammadagha, M. (2025) Hyperparameter Optimization Strategies for Tree-Based Machine Learning Models Prediction: A Comparative Study of Ada-Boost, Decision Trees, and Random Forest. Decision Trees, and Random Forest.
[50] Neto, M.P. and Paulovich, F.V. (2021) Explainable Matrix—Visualization for Global and Local Interpretability of Random Forest Classification Ensembles. IEEE Transactions on Visualization and Computer Graphics, 27, 1427-1437.[CrossRef] [PubMed]
[51] Marchese Robinson, R.L., Palczewska, A., Palczewski, J. and Kidley, N. (2017) Comparison of the Predictive Performance and Interpretability of Random Forest and Linear Models on Benchmark Data Sets. Journal of Chemical Information and Modeling, 57, 1773-1792.[CrossRef] [PubMed]
[52] Kim, H., Pang, S., Je, H., Kim, D. and Yang Bang, S. (2003) Constructing Support Vector Machine Ensemble. Pattern Recognition, 36, 2757-2767.[CrossRef]
[53] Wang, S., Mathew, A., Chen, Y., Xi, L., Ma, L. and Lee, J. (2009) Empirical Analysis of Support Vector Machine Ensemble Classifiers. Expert Systems with Applications, 36, 6466-6476.[CrossRef]
[54] Mehmood, Z. and Asghar, S. (2021) Customizing SVM as a Base Learner with AdaBoost Ensemble to Learn from Multi-Class Problems: A Hybrid Approach AdaBoost-MSVM. Knowledge-Based Systems, 217, Article ID: 106845.[CrossRef]
[55] Barr Kumarakulasinghe, N., Blomberg, T., Liu, J., Saraiva Leao, A. and Papapetrou, P. (2020) Evaluating Local Interpretable Model-Agnostic Explanations on Clinical Machine Learning Classification Models. 2020 IEEE 33rd International Symposium on Computer-Based Medical Systems (CBMS), Rochester, 28-30 July 2020, 7-12.[CrossRef]
[56] Zafar, M.R. and Khan, N.M. (2019) DLIME: A Deterministic Local Interpretable Model-Agnostic Explanations Approach for Computer-Aided Diagnosis Systems. arXiv: 1906.10263.
[57] Shamszare, H. and Choudhury, A. (2023) Clinicians’ Perceptions of Artificial Intelligence: Focus on Workload, Risk, Trust, Clinical Decision Making, and Clinical Integration. Healthcare, 11, Article 2308.[CrossRef] [PubMed]
[58] Golden, G., Popescu, C., Israel, S., Perlman, K., Armstrong, C., Fratila, R., et al. (2024) Applying Artificial Intelligence to Clinical Decision Support in Mental Health: What Have We Learned? Health Policy and Technology, 13, Article ID: 100844.[CrossRef]
[59] Datta Burton, S., Mahfoud, T., Aicardi, C. and Rose, N. (2021) Clinical Translation of Computational Brain Models: Understanding the Salience of Trust in Clinician-Researcher Relationships. Interdisciplinary Science Reviews, 46, 138-157.[CrossRef]
[60] Koch, E., Pardiñas, A.F., O’Connell, K.S., Selvaggi, P., Camacho Collados, J., Babic, A., et al. (2024) How Real-World Data Can Facilitate the Development of Precision Medicine Treatment in Psychiatry. Biological Psychiatry, 96, 543-551.[CrossRef] [PubMed]
[61] Clark, L.A., Cuthbert, B., Lewis-Fernández, R., Narrow, W.E. and Reed, G.M. (2017) Three Approaches to Understanding and Classifying Mental Disorder: ICD-11, DSM-5, and the National Institute of Mental Health’s Research Domain Criteria (RDOC). Psychological Science in the Public Interest, 18, 72-145. [Google Scholar] [CrossRef] [PubMed]
[62] McGorry, P.D., Hickie, I.B., Yung, A.R., Pantelis, C. and Jackson, H.J. (2006) Clinical Staging of Psychiatric Disorders: A Heuristic Framework for Choosing Earlier, Safer and More Effective Interventions. Australian and New Zealand Journal of Psychiatry, 40, 616-622.[CrossRef] [PubMed]
[63] Innocent, E.K. (2024) Enhancing Data Security in Healthcare with Synthetic Data Generation: An Autoencoder and Variational Autoencoder Approach. Master’s Thesis, Oslo Metropolitan University.
[64] Basri, M.A. (2024) Evaluating the Usefulness of Synthetic Data in Healthcare: Applications in Predictive Modeling and Privacy Protection. Master’s Thesis, University of Waterloo.
[65] Alharthi, A., Alqurashi, A., Alharbi, T., Alammar, M., Aldosari, N., Bouchekara, H., Shaaban, Y., Shahriar, M.S. and Al Ayidh, A. (2024) The Role of Explainable AI in Revolutionizing Human Health Monitoring. arXiv: 2409.07347.
[66] Hassija, V., Chamola, V., Mahapatra, A., Singal, A., Goel, D., Huang, K., et al. (2023) Interpreting Black-Box Models: A Review on Explainable Artificial Intelligence. Cognitive Computation, 16, 45-74.[CrossRef]
[67] Barron, D.S., Baker, J.T., Budde, K.S., Bzdok, D., Eickhoff, S.B., Friston, K.J., et al. (2021) Decision Models and Technology Can Help Psychiatry Develop Biomarkers. Frontiers in Psychiatry, 12, Article 706655.[CrossRef] [PubMed]
[68] Prasad, A. (2025) Predictive Analytics in Clinical Psychology: Role of Machine Learning in Future Mental Health Care. In: Bansal, R., Maqableh, T., Shuklaa, G., Rabby, F. and Lathabhavan, R., Eds., Transforming Neuropsychology and Cognitive Psychology with AI and Machine Learning, IGI Global, 313-332.[CrossRef]

Copyright © 2026 by authors and Scientific Research Publishing Inc.

Creative Commons License

This work and the related PDF file are licensed under a Creative Commons Attribution 4.0 International License.