<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.4 20241031//EN" "JATS-journalpublishing1-4.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article" dtd-version="1.4" xml:lang="en">
  <front>
    <journal-meta>
      <journal-id journal-id-type="publisher-id">Oalib</journal-id>
      <journal-title-group>
        <journal-title>Open Access Library Journal</journal-title>
      </journal-title-group>
      <issn pub-type="epub">2333-9721</issn>
      <issn pub-type="ppub">2333-9705</issn>
      <publisher>
        <publisher-name>Scientific Research Publishing</publisher-name>
      </publisher>
    </journal-meta>
    <article-meta>
      <article-id pub-id-type="doi">10.4236/oalib.1115348</article-id>
      <article-id pub-id-type="publisher-id">Oalib-151676</article-id>
      <article-categories>
        <subj-group>
          <subject>Article</subject>
        </subj-group>
        <subj-group>
          <subject>Biomedical</subject>
          <subject>Life Sciences</subject>
          <subject>Business</subject>
          <subject>Economics</subject>
          <subject>Chemistry</subject>
          <subject>Materials Science</subject>
          <subject>Computer Science</subject>
          <subject>Communications</subject>
          <subject>Earth</subject>
          <subject>Environmental Sciences</subject>
          <subject>Engineering</subject>
          <subject>Medicine</subject>
          <subject>Healthcare</subject>
          <subject>Physics</subject>
          <subject>Mathematics</subject>
          <subject>Social Sciences</subject>
          <subject>Humanities</subject>
        </subj-group>
      </article-categories>
      <title-group>
        <article-title>A Counterfactual Explainability Framework for Transparent, Actionable, and Clinician Validated Psychiatric Treatment Decision Support</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author" corresp="yes">
          <contrib-id contrib-id-type="orcid">0000-0001-9101-072X</contrib-id>
          <name name-style="western">
            <surname>Filippis</surname>
            <given-names>Rocco de</given-names>
          </name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <contrib-id contrib-id-type="orcid">0000-0002-5102-4999</contrib-id>
          <name name-style="western">
            <surname>Foysal</surname>
            <given-names>Abdullah Al</given-names>
          </name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
      </contrib-group>
      <aff id="aff1"><label>1</label> Department of Neuroscience, Institute of Psychopathology, Rome, Italy </aff>
      <aff id="aff2"><label>2</label> Department of Computer Engineering (AI), University of Genova, Genova, Italy </aff>
      <author-notes>
        <fn fn-type="conflict" id="fn-conflict">
          <p>The authors declare no conflicts of interest.</p>
        </fn>
      </author-notes>
      <pub-date pub-type="epub">
        <day>06</day>
        <month>05</month>
        <year>2026</year>
      </pub-date>
      <pub-date pub-type="collection">
        <month>05</month>
        <year>2026</year>
      </pub-date>
      <volume>13</volume>
      <issue>05</issue>
      <fpage>1</fpage>
      <lpage>25</lpage>
      <history>
        <date date-type="received">
          <day>14</day>
          <month>04</month>
          <year>2026</year>
        </date>
        <date date-type="accepted">
          <day>26</day>
          <month>05</month>
          <year>2026</year>
        </date>
        <date date-type="published">
          <day>29</day>
          <month>05</month>
          <year>2026</year>
        </date>
      </history>
      <permissions>
        <copyright-statement>© 2026 by the authors and Scientific Research Publishing Inc.</copyright-statement>
        <copyright-year>2026</copyright-year>
        <license license-type="open-access">
          <license-p> This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license ( <ext-link ext-link-type="uri" xlink:href="https://creativecommons.org/licenses/by/4.0/">https://creativecommons.org/licenses/by/4.0/</ext-link> ). </license-p>
        </license>
      </permissions>
      <self-uri content-type="doi" xlink:href="https://doi.org/10.4236/oalib.1115348">https://doi.org/10.4236/oalib.1115348</self-uri>
      <abstract>
        <p>Psychiatric treatment decisions are among the most consequential and least transparent clinical choices a physician makes. A machine learning model that predicts treatment non-response is only clinically useful if it can also answer the question every psychiatrist immediately asks: what would need to change for this patient to respond? Standard black-box models cannot answer this question. Counterfactual explanation methods propose to fill this gap, but existing approaches generate scenarios that are mathematically optimal yet clinically implausible changing features that cannot be acted upon, violating causal constraints between clinical variables, or ignoring the patient’s circumstances and preferences. We introduce CounterPsych, an end-to-end counterfactual explainability framework specifically designed for psychiatric treatment decision support. CounterPsych combines a Bayesian outcome predictor a ten-member deep ensemble with calibrated uncertainty (ECE = 0.021) with a constrained counterfactual generator that produces what-if treatment scenarios satisfying four simultaneous validity criteria: clinical plausibility, medical actionability, causal consistency, and patient-preference alignment. The counterfactual generator is built on a novel proximity-constrained gradient search with clinical validity filtering and diversity regularization, producing sparse, realistic recourse plans with a mean of 2.3 feature changes per counterfactual. Trained and validated on a retrospective-prospective cohort of 2,480 psychiatric outpatients across five diagnostic categories and four clinical sites, CounterPsych achieves treatment outcome prediction accuracy of 94.1%, AUC-ROC of 0.977, and macro-F1 of 0.919. In a prospective clinician evaluation with 24 consultant psychiatrists, CounterPsych counterfactuals received mean ratings of 4.42/5 for clinical plausibility and 4.51/5 for trustworthiness substantially outperforming the best prior counterfactual method (DiCE: 3.21/5 and 3.14/5). CounterPsych is the first counterfactual explanation framework validated for psychiatric treatment decisions through direct clinician evaluation, establishing a new standard for clinically meaningful machine learning explainability in psychiatry.</p>
      </abstract>
      <kwd-group kwd-group-type="author-generated" xml:lang="en">
        <kwd>Counterfactual Explanations</kwd>
        <kwd>Explainable AI</kwd>
        <kwd>Psychiatric Treatment</kwd>
        <kwd>Clinical Decision Support</kwd>
        <kwd>Algorithmic Recourse</kwd>
        <kwd>Treatment Response Prediction</kwd>
        <kwd>Bayesian Deep Learning</kwd>
        <kwd>Interpretable Machine Learning</kwd>
        <kwd>XAI</kwd>
        <kwd>What-If Scenarios</kwd>
        <kwd>Clinician Evaluation</kwd>
        <kwd>Causal Consistency</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec1">
      <title>1. Introduction</title>
      <p>A psychiatric clinical decision is not a lookup table. When a consultant decides to switch a patient from sertraline to venlafaxine, reduce the quetiapine dose, and add structured sleep hygiene, they are reasoning across a high-dimensional space of interacting clinical variables under genuine uncertainty about the outcome. Machine learning models have grown increasingly capable of predicting whether a given treatment will work for a given patient, but a prediction is not an explanation, and an explanation is not a recommendation. What clinicians need is not a probability score. They need to understand why the model predicts non-response, and crucially, what could change to turn a predicted failure into a predicted success. This is the counterfactual question, and it is the question that most clinical AI systems cannot answer [<xref ref-type="bibr" rid="B1">1</xref>][<xref ref-type="bibr" rid="B2">2</xref>].</p>
      <p>The concept of counterfactual explanation originates in causal reasoning: a counterfactual is a statement of the form ‘if X had been different, Y would have been different.’ Applied to machine learning models, a counterfactual explanation identifies the minimal change to a patient’s input features that would flip the model’s prediction from an unfavourable outcome to a favourable one [<xref ref-type="bibr" rid="B3">3</xref>]. The clinical translation is direct: if the model predicts non-response, the counterfactual answers ‘what would this patient’s situation need to look like for the model to predict response instead?’ Done correctly, this provides clinicians with actionable, patient-specific recourse not generic treatment guidelines but personalized, data-driven recommendations grounded in the model’s learned representations of this patient’s particular combination of clinical factors [<xref ref-type="bibr" rid="B4">4</xref>].</p>
      <p>Psychiatric medicine is precisely where counterfactual explainability matters most and where it has been least developed. Treatment decisions in psychiatry are high-stakes, individually variable, evidence-sparse, and ethically complex [<xref ref-type="bibr" rid="B5">5</xref>]. The same medication at the same dose produces full remission in one patient and adverse effects in another. Clinical guidelines provide population-level recommendations, not patient-level predictions. The treating psychiatrist is expected to integrate published evidence, clinical experience, patient preference, and biological context into a decision for a person standing in front of them and to document and justify that decision [<xref ref-type="bibr" rid="B6">6</xref>]. A counterfactual explanation system that can show the clinician ‘this patient’s model-predicted outcome improves from non-response to response if we increase therapy frequency from 0 to 4 sessions per month and reduce the quetiapine dose by 150 mg’ provides exactly the kind of actionable, case-specific reasoning support that psychiatric practice needs and currently lacks.</p>
      <p>The regulatory pressure is building. The EU AI Act (Regulation EU 2024/1689) classifies AI systems supporting clinical diagnosis and treatment decisions as high-risk, requiring that their outputs be interpretable and that affected individuals have the right to a meaningful explanation of any automated decision [<xref ref-type="bibr" rid="B7">7</xref>]. GDPR Article 22 establishes the right to explanation for automated decisions affecting individuals in legally or similarly significant ways [<xref ref-type="bibr" rid="B8">8</xref>]. Standard black-box prediction models gradient boosting classifiers, and deep neural networks satisfy neither requirement. Counterfactual explanation methods satisfy both: they provide not just a prediction but an account of what would need to change for the prediction to differ, which is operationally the most clinically useful form of explanation available [<xref ref-type="bibr" rid="B9">9</xref>].</p>
      <p>The problem is that existing counterfactual methods were not designed for psychiatric clinical data. DiCE and its variants optimize for diversity and proximity in feature space without enforcing clinical plausibility generating counterfactuals that change features no clinician can act on (genetic markers, age at first episode), violate causal relationships between clinical variables (suggesting lower depression scores without any intervention that would produce them), or ignore the patient’s real-world constraints and preferences. CounterPsych is designed to solve these problems: to produce counterfactual explanations that are not just mathematically valid but clinically meaningful, medically actionable, causally consistent, and directly useful in a psychiatric consultation.</p>
      <sec id="sec1dot1">
        <title>Summary of Contributions</title>
        <p><bold>CounterPsych framework:</bold>The first end-to-end counterfactual explainability framework designed and validated specifically for psychiatric treatment decisions, combining a Bayesian outcome predictor with a clinically constrained counterfactual generator and a multi-criteria validity filter.<bold>Clinical validity constraints:</bold>A formal specification of four psychiatric-domain counterfactual validity criteria, clinical plausibility, medical actionability, causal consistency, and patient-preference alignment implemented as hard constraints in the optimization objective, ensuring generated counterfactuals are clinically meaningful by construction.<bold>Proximity-constrained gradient search:</bold>A novel counterfactual generation algorithm combining gradient-based search in the outcome predictor’s latent space with proximity regularization and diversity promotion, producing sparse counterfactuals with a mean of 2.3 feature changes and clinician plausibility ratings of 4.42/5.<bold>Bayesian outcome predictor:</bold>A ten-member deep ensemble with MC-Dropout achieving treatment outcome prediction accuracy of 94.1%, AUC-ROC of 0.977, macro-F1 of 0.919, and ECE of 0.021, providing the calibrated probabilistic foundation on which counterfactual generation depends.<bold>Prospective clinician evaluation:</bold>A structured evaluation with 24 consultant psychiatrists rating CounterPsych counterfactuals on five clinical quality dimensions, achieving the highest published clinician plausibility and trustworthiness ratings for any psychiatric AI explanation method.<bold>Open clinical cohort:</bold>A retrospective-prospective dataset of 2,480 psychiatric outpatients across five diagnostic categories and four clinical sites, with treatment outcome labels adjudicated by consensus at 12-week follow-up, constituting the largest clinically labelled dataset for psychiatric treatment response prediction with counterfactual annotations.</p>
      </sec>
    </sec>
    <sec id="sec2">
      <title>2. Background and Related Work</title>
      <sec id="sec2dot1">
        <title>2.1. Counterfactual Explanations: Foundations and Formal Definition</title>
        <p>The counterfactual explanation framework was formalized in the context of algorithmic accountability by Wachter and colleagues, who proposed that individuals affected by automated decisions have a right to be told what minimal change to their situation would have produced a different outcome. Formally, given a classifier f: X → Y and a factual instance x with predicted class y = f(x), a counterfactual explanation x’ satisfies: 1) f(x’) = y’ ≠ y (prediction flip); 2) ‖x’ − x‖_p is minimized (proximity fewest, smallest changes); and 3) x’ ∈ D (data manifold the counterfactual should look like a plausible real patient). The foundational DiCE framework [<xref ref-type="bibr" rid="B10">10</xref>] extended this to generate multiple diverse counterfactuals simultaneously, improving clinical utility by showing several alternative recourse paths rather than one.</p>
        <p>The concept of algorithmic recourse extends counterfactual explanations to the specific question of actionability: not just ‘what would need to be different’ but ‘what can realistically be done to produce a different outcome’ [<xref ref-type="bibr" rid="B11">11</xref>]. This distinction is critical in clinical settings where many features in a patient’s record age, genetic background, illness history cannot be changed, and where the clinical value of a counterfactual depends entirely on whether the suggested changes are within the clinician’s and patient’s power to implement. A comprehensive review of counterfactual explanation methods [<xref ref-type="bibr" rid="B12">12</xref>] identifies five dimensions of counterfactual quality: correctness (the prediction flips), proximity (minimum feature change), sparsity (few features changed), plausibility (the counterfactual resembles real instances), and actionability (the changes are implementable).</p>
        <p>Prior methods have addressed these dimensions partially and in isolation. Actionable recourse frameworks [<xref ref-type="bibr" rid="B13">13</xref>] enforce immutability constraints (features that cannot change) but do not address causal consistency between mutable features. Prototype-based counterfactual generation [<xref ref-type="bibr" rid="B14">14</xref>] improves plausibility by anchoring counterfactuals to real training instances but sacrifices proximity and actionability in feature-rich clinical datasets. FACE [<xref ref-type="bibr" rid="B15">15</xref>] generates feasible counterfactuals by constraining search to high-density regions of the data manifold, but its density estimation scales poorly to the mixed continuous-categorical feature spaces typical of electronic health record data.</p>
        <p>The causal consistency requirement is particularly challenging and particularly important in psychiatry. Clinical variables do not change independently: reducing a depression score presupposes an intervention that produces that reduction; reducing a medication dose changes biomarker levels and side effect profiles; increasing therapy frequency affects engagement metrics. Counterfactuals that ignore these causal dependencies are not just implausible they are incoherent. The causally constrained counterfactual framework [<xref ref-type="bibr" rid="B16">16</xref>] provides the formal machinery for this constraint, and CounterPsych implements a psychiatric-domain instantiation of this framework that has not previously been attempted.</p>
      </sec>
      <sec id="sec2dot2">
        <title>2.2. Explainable AI in Psychiatric Clinical Practice</title>
        <p>Machine learning has been applied to psychiatric clinical prediction with increasing sophistication over the past decade. Dwyer and colleagues [<xref ref-type="bibr" rid="B17">17</xref>] provided an influential overview of ML applications across diagnostic classification, treatment response prediction, and relapse risk estimation, identifying the transition from research to clinical deployment as the major unresolved challenge. A systematic scoping review of ML in mental health [<xref ref-type="bibr" rid="B18">18</xref>] catalogued 28 prediction tasks across 15 psychiatric conditions, finding that treatment response prediction and medication selection are the two tasks where ML accuracy most clearly exceeds clinical heuristics and precisely where interpretability is most urgently needed because the decisions are both high-stakes and individually variable.</p>
        <p>The relationship between predictive accuracy and interpretability has been the subject of sustained debate. A systematic review of ML versus logistic regression for clinical prediction [<xref ref-type="bibr" rid="B19">19</xref>] found that complex models rarely outperform logistic regression on clinical datasets of modest size, suggesting that interpretability may be achievable without sacrificing predictive accuracy in many psychiatric applications. The dominant current approaches to ML interpretability in clinical settings SHAP (SHapley Additive exPlanations) [<xref ref-type="bibr" rid="B20">20</xref>] and LIME (Local Interpretable Model-agnostic Explanations) [<xref ref-type="bibr" rid="B21">21</xref>] provide feature importance scores that answer the question ‘which features drove this prediction?’ but do not answer the clinically more useful question ‘what would need to change for the prediction to be different?’ This is precisely the gap that counterfactual explanation addresses.</p>
      </sec>
      <sec id="sec2dot3">
        <title>2.3. Treatment Outcome Prediction in Psychiatry</title>
        <p>Treatment outcome in psychiatry is typically operationalized through validated rating scales: the Hamilton Depression Rating Scale (HDRS-17) [<xref ref-type="bibr" rid="B22">22</xref>] for depressive outcomes, the Young Mania Rating Scale (YMRS) [<xref ref-type="bibr" rid="B23">23</xref>] for manic outcomes, and the Global Assessment of Functioning (GAF) [<xref ref-type="bibr" rid="B24">24</xref>] for broad functional outcomes. Response is conventionally defined as a ≥ 50% reduction in the primary symptom scale score from baseline to endpoint. Remission requires score reduction to below a clinical threshold (HDRS-17 ≤ 7 for remission in depression). The International Society for Bipolar Disorders has published standardized nomenclature for course and outcome that CounterPsych adopts as its labelling framework [<xref ref-type="bibr" rid="B25">25</xref>].</p>
        <p>Prior ML systems for treatment outcome prediction have demonstrated that baseline symptom severity, illness duration, number of prior episodes, and medication adherence are reliably predictive of treatment response across psychiatric conditions. However, these models have been deployed as black-box score generators without any mechanism for translating their predictions into clinical guidance. CounterPsych is the first system to close this gap providing not just a response probability but a set of clinically constrained, actionable interventions that the model predicts would change the outcome.</p>
      </sec>
    </sec>
    <sec id="sec3">
      <title>3. Dataset and Cohort Design</title>
      <sec id="sec3dot1">
        <title>3.1. Study Population</title>
        <p>The CounterPsych dataset was assembled through a retrospective-prospective observational design across four outpatient psychiatric centers in Italy, Germany, Switzerland, and the United Kingdom (Primary ethics approval: IRB Ref. UNIGE-2024-CPX-01; GDPR Article 9 compliance). Retrospective records spanned January 2016 to December 2022; prospective enrolment ran from January 2023 to December 2024. Inclusion criteria: adults aged 18 - 70 with confirmed DSM-5 diagnosis in one of five categories (major depressive disorder, bipolar disorder, schizophrenia spectrum, obsessive-compulsive disorder, generalised anxiety disorder); initiation or modification of a pharmacological treatment regimen at the index visit; minimum 12-week follow-up with at least one post-treatment clinical assessment; and capacity to provide informed consent. Exclusion criteria: active substance use disorder with ongoing intoxication; neurological comorbidity affecting cognition; and concurrent enrolment in a conflicting clinical trial. The demographic, diagnostic, and treatment-related characteristics of the CounterPsych cohort are summarized in <bold>Table 1</bold>.</p>
      </sec>
      <sec id="sec3dot2">
        <title>3.2. Cohort Characteristics</title>
        <p><bold>Table 1.</bold> CounterPsych cohort demographic, diagnostic, and treatment characteristics.</p>
        <table-wrap id="tbl1">
          <label>Table 1</label>
          <table>
            <tbody>
              <tr>
                <td>
                  <bold>Characteristic</bold>
                </td>
                <td colspan="2">
                  <bold>Value/Distribution</bold>
                </td>
                <td>
                  <bold>Notes</bold>
                </td>
              </tr>
              <tr>
                <td>Total patients</td>
                <td colspan="2">
                  <bold>2480</bold>
                </td>
                <td>After exclusion criteria</td>
              </tr>
              <tr>
                <td>Mean age (years ± SD)</td>
                <td colspan="2">41.2 ± 14.8</td>
                <td>Range: 18 - 70</td>
              </tr>
              <tr>
                <td>Female (%)</td>
                <td colspan="2">54.7%</td>
                <td>
                </td>
              </tr>
              <tr>
                <td colspan="2">Diagnosis</td>
                <td>MDD 38%/BD 26%/SCZ 18%/OCD 11%/GAD 7%</td>
                <td>DSM-5</td>
              </tr>
              <tr>
                <td colspan="2">Mean illness duration (years ± SD)</td>
                <td>9.7 ± 8.2</td>
                <td>Since first diagnosis</td>
              </tr>
              <tr>
                <td colspan="2">Mean HDRS-17 at baseline</td>
                <td>19.4 ± 5.8</td>
                <td>Moderate-severe range</td>
              </tr>
              <tr>
                <td colspan="2">Mean GAF at baseline</td>
                <td>52.3 ± 12.1</td>
                <td>Moderate impairment</td>
              </tr>
              <tr>
                <td colspan="2">Mean concurrent medications</td>
                <td>2.4 ± 1.2</td>
                <td>Range: 1 - 7</td>
              </tr>
              <tr>
                <td colspan="2">Treatment outcome distribution</td>
                <td>Response 42%/Partial 34%/Non-response 24%</td>
                <td>12-week follow-up</td>
              </tr>
              <tr>
                <td colspan="2">
                  Inter-rater reliability (
                  <italic>κ</italic>
                  )
                </td>
                <td>
                  <italic>
                    <bold>κ</bold>
                  </italic>
                  <bold>= 0.86 (95% CI: 0.83</bold>
                  <bold>-</bold>
                  <bold>0.89)</bold>
                </td>
                <td>Outcome adjudication</td>
              </tr>
              <tr>
                <td colspan="2">Train/Val/Test split</td>
                <td>70/15/15%</td>
                <td>Patient-stratified</td>
              </tr>
            </tbody>
          </table>
        </table-wrap>
      </sec>
      <sec id="sec3dot3">
        <title>3.3. Feature Engineering</title>
        <p>The 112-dimensional feature vector for each patient was constructed from five domains. Clinical state (32 features): HDRS-17 total and subscale scores [<xref ref-type="bibr" rid="B26">26</xref>], YMRS total score, GAF global score, Clinical Global Impression (CGI) severity and improvement, self-reported energy, sleep, and appetite ratings, and functional impairment across occupational, social, and self-care domains. Treatment profile (24 features): current medications with dose equivalents, medication adherence rate, number of prior medication trials, duration of current regimen, psychotherapy type and frequency, and cumulative anticholinergic burden. Patient history (20 features): illness duration, number of prior episodes, number of prior hospitalizations, prior treatment response history (coded as a binary response vector over previous medication trials), and prior ECT or psychotherapy history. Biomarkers (16 features): serum drug levels where available, thyroid function, inflammatory markers (CRP, IL-6), EEG alpha asymmetry, and genetic CYP2D6 and CYP3A4 metabolizer status. Demographics (20 features): age, gender, education level, employment status, social support index, living situation, and comorbid physical health conditions.</p>
        <p>Data was extracted from clinical records using a standardized OMOP Common Data Model harmonization pipeline applied to each site’s EHR system [<xref ref-type="bibr" rid="B27">27</xref>]. Missing biomarker data (present in 31.4% of patients) were imputed using multivariate imputation by chained equations (MICE), with missingness indicators included as auxiliary features. Medication doses were standardized to chlorpromazine equivalents for antipsychotics and diazepam equivalents for benzodiazepines using published conversion tables from the Maudsley Prescribing Guidelines [<xref ref-type="bibr" rid="B28">28</xref>]. All the features were confirmed to precede the index treatment initiation visit. CGI-Improvement was derived from the most recent prior clinical assessment, not the 12-week follow-up, reflecting the patient’s trajectory entering the index visit. Medication adherence was computed over the 4-week window prior to the index visit using prescription refill records and clinician-documented adherence ratings. Serum drug levels and inflammatory markers (CRP, IL-6, thyroid) were extracted from the most recent laboratory record predating the index visit by no more than 8 weeks. CYP2D6/CYP3A4 metabolizer status was extracted from any prior genotyping record. No post-treatment information entered the predictor at any stage; outcome labels at 12-week follow-up were held strictly separate from the feature construction pipeline.</p>
      </sec>
      <sec id="sec3dot4">
        <title>3.4. Outcome Labelling and Inter-Rater Reliability</title>
        <p>Treatment outcomes were assessed at 12-week follow-up by a trained clinical rater using all available clinical information: repeat ratings on the HDRS-17 and GAF, clinical notes from interim visits, patient self-report, and collateral from treating clinicians. Response was defined as ≥50% reduction in HDRS-17 from baseline; remission required HDRS-17 ≤ 7 and GAF ≥ 61. Partial response was defined as 25% - 49% HDRS-17 reduction. Non-response was defined as &lt;25% reduction or clinical worsening. For non-depressive presentations, equivalent operationalizations were applied using YMRS (mania) and CGI (schizophrenia, OCD, GAD). All borderline classifications were reviewed by a second independent rater, achieving inter-rater reliability of <italic>κ</italic> = 0.86 (95% CI: 0.83 - 0.89). Cross-Diagnostic Outcome Harmonization. Outcome operationalization was adapted by primary diagnosis while maintaining the three-class label structure (Response/Partial Response/Non-Response) across all five diagnostic categories. For MDD (n = 942): HDRS-17 ≥ 50% reduction = Response; 25% - 49% = Partial Response; &lt;25% = non-response. For BD (n = 645): YMRS ≥50% reduction from mania baseline, or HDRS-17 ≥ 50% reduction from depressive baseline, depending on index episode polarity = Response; 25% - 49% = Partial Response; &lt;25% = non-response. For Schizophrenia spectrum (n = 446): CGI-Improvement score ≤ 2 (much/very much improved) = Response; CGI-I = 3 (minimally improved) = Partial Response; CGI-I ≥ 4 = non-response. For OCD (n = 273): Y-BOCS ≥ 35% reduction = Response; 25% - 34% = Partial Response; &lt;25% = non-response. For GAD (n = 174): GAD-7 ≥ 50% reduction = Response; 25% - 49% = Partial Response; &lt;25% = non-response. Scale availability by diagnosis is reported in Supplementary Table S3. The three-class label structure was applied uniformly across disorders; all threshold definitions were pre-specified in the study protocol prior to data analysis.</p>
      </sec>
    </sec>
    <sec id="sec4">
      <title>4. Architecture and Technical Specification</title>
      <p>CounterPsych has four integrated modules: 1) a multimodal outcome predictor generating calibrated treatment response probabilities; 2) a proximity-constrained counterfactual generator producing candidate what-if scenarios; 3) a clinical validity filter enforcing psychiatric domain constraints; and 4) a diversity regularization mechanism ensuring that the returned counterfactual set covers multiple actionable recourse paths. <xref ref-type="fig" rid="fig1">Figure 1</xref><xref ref-type="fig" rid="fig1">Figure 1</xref> presents the complete architecture.</p>
      <fig id="fig1">
        <label>Figure 1</label>
        <graphic xlink:href="https://html.scirp.org/file/1115348-rId16.jpeg?20260529040049" />
      </fig>
      <p><bold>Figure 1.</bold> CounterPsych End-to-End Architecture. Five clinical input modalities are encoded by domain-specific modules and fused into a unified patient representation. The Bayesian outcome predictor (M = 10 ensemble, ECE = 0.021) generates treatment response probabilities with calibrated uncertainty. The Counterfactual Generator performs proximity-constrained gradient search in the predictor’s latent space. The Validity Filter enforces four clinical constraint categories. A recourse feedback loop connects the filtered counterfactuals back to the predictor to verify prediction flip. The final output includes the prediction, multiple diverse counterfactual scenarios, and per-feature attribution scores.</p>
      <sec id="sec4dot1">
        <title>4.1. Module 1: Multimodal Outcome Predictor</title>
        <p>4.1.1. Encoder Architecture</p>
        <p>Each of the five feature domains is encoded by a domain-specific module. Clinical state and treatment profile features are encoded by a 4-layer Clinical Transformer [<xref ref-type="bibr" rid="B29">29</xref>] with 8 attention heads and <italic>d</italic><sub>model</sub> = 256, operating on the feature vector with learned positional encodings that encode feature-type rather than sequence position. Patient history features are encoded by a 3-layer Temporal Convolutional Network [<xref ref-type="bibr" rid="B30">30</xref>] with dilations <italic>d</italic><italic><sub>l</sub></italic> ∈ {1, 2, 4} operating on the patient’s longitudinal treatment history vector each prior treatment trial is a timestep, and the TCN learns to extract trajectory patterns (e.g., a sequence of partial responses that predicts eventual non-response). Biomarker and demographic features are each encoded by 3-layer MLPs with residual connections. All encoders produce 128-dimensional embeddings that are concatenated and projected to a unified 512-dimensional patient state representation <inline-formula><mml:math display="inline"><mml:mrow><mml:msub><mml:mi> h </mml:mi><mml:mrow><mml:mtext> patient </mml:mtext></mml:mrow></mml:msub><mml:mo> ∈ </mml:mo><mml:msup><mml:mi> ℝ </mml:mi><mml:mrow><mml:mn> 512 </mml:mn></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula> .</p>
        <p>4.1.2. Bayesian Ensemble Output</p>
        <p>The outcome predictor deploys a Bayesian deep ensemble of <italic>M</italic> = 10 independently initialized network instances, following the deep ensemble methodology [<xref ref-type="bibr" rid="B31">31</xref>]. At inference time, Monte Carlo Dropout [<xref ref-type="bibr" rid="B32">32</xref>] with <italic>p</italic> = 0.3 is maintained active across all ensemble members, generating <italic>T</italic> = 50 stochastic forward passes per member. The treatment outcome probability vector and its uncertainty are computed as:</p>
        <disp-formula id="FD1">
          <label>(1)</label>
          <mml:math display="inline">
            <mml:mrow>
              <mml:mover accent="true">
                <mml:mi>p</mml:mi>
                <mml:mo>¯</mml:mo>
              </mml:mover>
              <mml:mrow>
                <mml:mo>(</mml:mo>
                <mml:mrow>
                  <mml:mi>y</mml:mi>
                  <mml:mo>|</mml:mo>
                  <mml:mi>x</mml:mi>
                </mml:mrow>
                <mml:mo>)</mml:mo>
              </mml:mrow>
              <mml:mo>=</mml:mo>
              <mml:mrow>
                <mml:mo>(</mml:mo>
                <mml:mrow>
                  <mml:mrow>
                    <mml:mn>1</mml:mn>
                    <mml:mo>/</mml:mo>
                    <mml:mi>M</mml:mi>
                  </mml:mrow>
                  <mml:mo>⋅</mml:mo>
                  <mml:mi>T</mml:mi>
                </mml:mrow>
                <mml:mo>)</mml:mo>
              </mml:mrow>
              <mml:mstyle displaystyle="true">
                <mml:msubsup>
                  <mml:mo>∑</mml:mo>
                  <mml:mrow>
                    <mml:mi>m</mml:mi>
                    <mml:mo>=</mml:mo>
                    <mml:mn>1</mml:mn>
                  </mml:mrow>
                  <mml:mi>M</mml:mi>
                </mml:msubsup>
                <mml:mrow>
                  <mml:mstyle displaystyle="true">
                    <mml:msubsup>
                      <mml:mo>∑</mml:mo>
                      <mml:mrow>
                        <mml:mi>t</mml:mi>
                        <mml:mo>=</mml:mo>
                        <mml:mn>1</mml:mn>
                      </mml:mrow>
                      <mml:mi>T</mml:mi>
                    </mml:msubsup>
                    <mml:mi>p</mml:mi>
                  </mml:mstyle>
                </mml:mrow>
              </mml:mstyle>
              <mml:mrow>
                <mml:mo>(</mml:mo>
                <mml:mrow>
                  <mml:mi>y</mml:mi>
                  <mml:mo>|</mml:mo>
                  <mml:mi>x</mml:mi>
                  <mml:mo>,</mml:mo>
                  <mml:msub>
                    <mml:mi>θ</mml:mi>
                    <mml:mi>m</mml:mi>
                  </mml:msub>
                  <mml:mo>,</mml:mo>
                  <mml:msub>
                    <mml:mi>ε</mml:mi>
                    <mml:mi>t</mml:mi>
                  </mml:msub>
                </mml:mrow>
                <mml:mo>)</mml:mo>
              </mml:mrow>
            </mml:mrow>
          </mml:math>
        </disp-formula>
        <disp-formula id="FD2">
          <label>(2)</label>
          <mml:math display="inline">
            <mml:mrow>
              <mml:msub>
                <mml:mi>σ</mml:mi>
                <mml:mi>i</mml:mi>
              </mml:msub>
              <mml:mo>=</mml:mo>
              <mml:mi>s</mml:mi>
              <mml:mi>t</mml:mi>
              <mml:mi>d</mml:mi>
              <mml:mrow>
                <mml:mo>(</mml:mo>
                <mml:mrow>
                  <mml:mrow>
                    <mml:mo>{</mml:mo>
                    <mml:mrow>
                      <mml:msubsup>
                        <mml:mi>p</mml:mi>
                        <mml:mi>i</mml:mi>
                        <mml:mrow>
                          <mml:mi>m</mml:mi>
                          <mml:mo>,</mml:mo>
                          <mml:mi>t</mml:mi>
                        </mml:mrow>
                      </mml:msubsup>
                      <mml:mo>:</mml:mo>
                      <mml:mi>m</mml:mi>
                      <mml:mo>=</mml:mo>
                      <mml:mn>1</mml:mn>
                      <mml:mo>,</mml:mo>
                      <mml:mo>⋯</mml:mo>
                      <mml:mo>,</mml:mo>
                      <mml:mi>M</mml:mi>
                      <mml:mo>;</mml:mo>
                      <mml:mi>t</mml:mi>
                      <mml:mo>=</mml:mo>
                      <mml:mn>1</mml:mn>
                      <mml:mo>,</mml:mo>
                      <mml:mo>⋯</mml:mo>
                      <mml:mo>,</mml:mo>
                      <mml:mi>T</mml:mi>
                    </mml:mrow>
                    <mml:mo>}</mml:mo>
                  </mml:mrow>
                </mml:mrow>
                <mml:mo>)</mml:mo>
              </mml:mrow>
            </mml:mrow>
          </mml:math>
        </disp-formula>
        <p>When predictive entropy <inline-formula><mml:math display="inline"><mml:mrow><mml:mi> H </mml:mi><mml:mo> = </mml:mo><mml:mo> − </mml:mo><mml:mstyle displaystyle="true"><mml:msub><mml:mo> ∑ </mml:mo><mml:mi> c </mml:mi></mml:msub><mml:mrow><mml:msub><mml:mover accent="true"><mml:mi> p </mml:mi><mml:mo> ¯ </mml:mo></mml:mover><mml:mi> c </mml:mi></mml:msub><mml:mi> log </mml:mi><mml:mrow><mml:mo> ( </mml:mo><mml:mrow><mml:msub><mml:mover accent="true"><mml:mi> p </mml:mi><mml:mo> ¯ </mml:mo></mml:mover><mml:mi> c </mml:mi></mml:msub></mml:mrow><mml:mo> ) </mml:mo></mml:mrow></mml:mrow></mml:mstyle></mml:mrow></mml:math></inline-formula> exceeds threshold <italic>τ</italic> = 0.44 nats calibrated on the validation set via temperature scaling [<xref ref-type="bibr" rid="B33">33</xref>] the model flags the prediction as high-uncertainty and routes it for mandatory clinician review rather than automated counterfactual generation. The abstention rate on the test set is 8.7%; within non-abstained predictions, accuracy rises to 96.2%.</p>
      </sec>
      <sec id="sec4dot2">
        <title>4.2. Module 2: Proximity-Constrained Counterfactual Generator</title>
        <p>4.2.1. Formal Problem Statement</p>
        <p>Given a factual instance <italic>x</italic> with predicted class <italic>y</italic> = <italic>f</italic>(<italic>x</italic>) = ‘Non-Response’, we seek a counterfactual <italic>x</italic><italic>’</italic> satisfying:</p>
        <disp-formula id="FD3">
          <mml:math display="inline">
            <mml:mrow>
              <mml:msup>
                <mml:mi>x</mml:mi>
                <mml:mo>′</mml:mo>
              </mml:msup>
              <mml:mo>=</mml:mo>
              <mml:mi>arg</mml:mi>
              <mml:msub>
                <mml:mrow>
                  <mml:mi>min</mml:mi>
                </mml:mrow>
                <mml:msup>
                  <mml:mi>x</mml:mi>
                  <mml:mo>′</mml:mo>
                </mml:msup>
              </mml:msub>
              <mml:msub>
                <mml:mi>λ</mml:mi>
                <mml:mn>1</mml:mn>
              </mml:msub>
              <mml:mo>⋅</mml:mo>
              <mml:msub>
                <mml:mi>L</mml:mi>
                <mml:mrow>
                  <mml:mtext>proximity</mml:mtext>
                </mml:mrow>
              </mml:msub>
              <mml:mrow>
                <mml:mo>(</mml:mo>
                <mml:mrow>
                  <mml:mi>x</mml:mi>
                  <mml:mo>,</mml:mo>
                  <mml:msup>
                    <mml:mi>x</mml:mi>
                    <mml:mo>′</mml:mo>
                  </mml:msup>
                </mml:mrow>
                <mml:mo>)</mml:mo>
              </mml:mrow>
              <mml:mo>+</mml:mo>
              <mml:msub>
                <mml:mi>λ</mml:mi>
                <mml:mn>2</mml:mn>
              </mml:msub>
              <mml:mo>⋅</mml:mo>
              <mml:msub>
                <mml:mi>L</mml:mi>
                <mml:mrow>
                  <mml:mtext>validity</mml:mtext>
                </mml:mrow>
              </mml:msub>
              <mml:mrow>
                <mml:mo>(</mml:mo>
                <mml:msup>
                  <mml:mi>x</mml:mi>
                  <mml:mo>′</mml:mo>
                </mml:msup>
                <mml:mo>)</mml:mo>
              </mml:mrow>
              <mml:mo>+</mml:mo>
              <mml:msub>
                <mml:mi>λ</mml:mi>
                <mml:mn>3</mml:mn>
              </mml:msub>
              <mml:mo>⋅</mml:mo>
              <mml:msub>
                <mml:mi>L</mml:mi>
                <mml:mrow>
                  <mml:mtext>diversity</mml:mtext>
                </mml:mrow>
              </mml:msub>
              <mml:mrow>
                <mml:mo>(</mml:mo>
                <mml:mrow>
                  <mml:msup>
                    <mml:mi>x</mml:mi>
                    <mml:mo>′</mml:mo>
                  </mml:msup>
                  <mml:mo>,</mml:mo>
                  <mml:msub>
                    <mml:msup>
                      <mml:mi>X</mml:mi>
                      <mml:mo>′</mml:mo>
                    </mml:msup>
                    <mml:mrow>
                      <mml:mtext>prev</mml:mtext>
                    </mml:mrow>
                  </mml:msub>
                </mml:mrow>
                <mml:mo>)</mml:mo>
              </mml:mrow>
            </mml:mrow>
          </mml:math>
        </disp-formula>
        <disp-formula id="FD4">
          <label>(3)</label>
          <mml:math display="inline">
            <mml:mrow>
              <mml:mi>f</mml:mi>
              <mml:mrow>
                <mml:mo>(</mml:mo>
                <mml:msup>
                  <mml:mi>x</mml:mi>
                  <mml:mo>′</mml:mo>
                </mml:msup>
                <mml:mo>)</mml:mo>
              </mml:mrow>
              <mml:mo>∈</mml:mo>
              <mml:msub>
                <mml:mi>Y</mml:mi>
                <mml:mrow>
                  <mml:mtext>target</mml:mtext>
                </mml:mrow>
              </mml:msub>
              <mml:mrow>
                <mml:mo>(</mml:mo>
                <mml:mi>x</mml:mi>
                <mml:mo>)</mml:mo>
              </mml:mrow>
              <mml:mo>,</mml:mo>
              <mml:msup>
                <mml:mi>x</mml:mi>
                <mml:mo>′</mml:mo>
              </mml:msup>
              <mml:mo>∈</mml:mo>
              <mml:mi>C</mml:mi>
              <mml:mrow>
                <mml:mo>(</mml:mo>
                <mml:mi>x</mml:mi>
                <mml:mo>)</mml:mo>
              </mml:mrow>
              <mml:mo>,</mml:mo>
              <mml:msub>
                <mml:mrow>
                  <mml:mrow>
                    <mml:mo>‖</mml:mo>
                    <mml:mrow>
                      <mml:msup>
                        <mml:mi>x</mml:mi>
                        <mml:mo>′</mml:mo>
                      </mml:msup>
                      <mml:mo>−</mml:mo>
                      <mml:mi>x</mml:mi>
                    </mml:mrow>
                    <mml:mo>‖</mml:mo>
                  </mml:mrow>
                </mml:mrow>
                <mml:mn>0</mml:mn>
              </mml:msub>
              <mml:mo>≤</mml:mo>
              <mml:msub>
                <mml:mi>k</mml:mi>
                <mml:mrow>
                  <mml:mi>max</mml:mi>
                </mml:mrow>
              </mml:msub>
            </mml:mrow>
          </mml:math>
        </disp-formula>
        <p>where <italic>L</italic><sub>proximity</sub> is the weighted L1 distance between <italic>x</italic><italic>’</italic> and <italic>x</italic> (weighted by feature mutability scores assigned by clinical domain experts); <italic>L</italic><sub>validity</sub> is a differentiable penalty for violating clinical plausibility and causal consistency constraints; <italic>L</italic><sub>diversity</sub> promotes dissimilarity between <italic>x</italic><italic>’</italic> and previously generated counterfactuals <inline-formula><mml:math display="inline"><mml:mrow><mml:msub><mml:msup><mml:mi> X </mml:mi><mml:mo> ′ </mml:mo></mml:msup><mml:mrow><mml:mtext> prev </mml:mtext></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula> ; <italic>C</italic>(<italic>x</italic>) is the space of clinically actionable modifications to <italic>x</italic>; and <italic>k</italic><sub>max</sub> = 5 is the maximum allowed number of feature changes (in practice, the mean observed is 2.3).</p>
        <p><italic>Y</italic><sub>target</sub>(<italic>x</italic>) is the class-conditional recourse target, defined as follows. For patients predicted as Non-Response, <italic>Y</italic><sub>target</sub> = {‘Response’, ‘Partial Response’} the generator first seeks ‘Response’ and returns ‘Partial Response’ counterfactuals if no valid ‘Response’ counterfactual is found within 500 gradient iterations. For patients predicted as Partial Response, <italic>Y</italic><sub>target</sub> = {‘Response’} the recourse target is always full response. For patients predicted as Response, no counterfactual is generated. This formulation ensures clinically meaningful recourse is defined for all three prediction classes, not only non-response cases.</p>
        <p>4.2.2. Feature Mutability and Causal Constraint Graph</p>
        <p>A key innovation of CounterPsych is the explicit encoding of two domain-specific constraint types. First, each of the 112 features is assigned a mutability score <italic>m</italic><italic><sub>j</sub></italic> ∈ {0, 0.5, 1} by a panel of five consultant psychiatrists: immutable features (<italic>m</italic><italic><sub>j</sub></italic> = 0) include age, genetic markers, age at first episode, and number of prior episodes; partially mutable features (<italic>m</italic><italic><sub>j</sub></italic> = 0.5) include illness duration and number of prior hospitalizations; fully mutable features (<italic>m</italic><italic><sub>j</sub></italic> = 1) include medication dose, therapy frequency, sleep duration, medication adherence, and social support measures. Second, a causal constraint graph G<sub>C</sub> encodes the directed causal dependencies between clinical features: reducing the HDRS-17 score requires an intervention that produces that reduction (it cannot be changed independently); increasing therapy frequency is causally downstream of an actionable clinical decision; medication dose changes propagate to serum level estimates and side effect profiles.</p>
        <p>The causal constraint graph G<sub>C</sub> was constructed from published psychiatric treatment literature and validated by the same clinician panel. During counterfactual search, any proposed feature change that violates a causal constraint in G<sub>C</sub> is projected back onto the constraint-satisfying manifold through a differentiable constraint projection layer, ensuring that all generated counterfactuals respect the causal structure of the clinical domain.</p>
        <p><bold>Constraint Construction, Agreement, and Usage Statistics:</bold> The mutability map was constructed by a panel of five consultant psychiatrists who independently scored all 112 features as immutable (<italic>m</italic><italic><sub>j</sub></italic> = 0), partially mutable (<italic>m</italic><italic><sub>j</sub></italic> = 0.5), or fully mutable (<italic>m</italic><italic><sub>j</sub></italic> = 1). Inter-rater agreement for mutability classification was <italic>κ</italic> = 0.83 (95% CI: 0.79 - 0.87), indicating strong expert consensus. Disagreements were resolved by majority vote; two features (illness duration, number of prior hospitalizations) required structured panel discussion before consensus. The causal constraint graph G_C contains 47 directed edges encoding causal dependencies validated against CANMAT 2023, BAP 2019, and NICE Clinical Guidelines NG222. Patient-preference data (e.g., stated aversion to specific drug classes, appointment frequency constraints) was extractable from clinical records for 61.4% of patients (n = 1523/2480); for the remaining 38.6%, the patient-preference constraint was inactive. Constraint impact statistics across the full test set: the mutability filter excluded at least one proposed feature change in 34.7% of first-pass counterfactuals; the causal constraint projection modified at least one feature change in 28.3% of counterfactuals; the patient-preference filter rejected at least one candidate counterfactual entirely in 19.1% of cases where preference data was available, triggering a replacement search. These rejection rates confirm that all three constraint layers are operationally active and materially shape the counterfactual output.</p>
        <p>4.2.3. Gradient Search Algorithm</p>
        <p>The counterfactual search is implemented as a constrained gradient descent in the continuous feature space, starting from the factual instance x and following the gradient of the outcome predictor’s log-probability surface toward the target class. Unlike prototype-based methods [<xref ref-type="bibr" rid="B34">34</xref>], which anchor counterfactuals to training instances, our gradient search explores the full feature space subject to the mutability and causal constraints. The search terminates when the predictor’s probability for the target class exceeds 0.65 and all constraint violations are below a tolerance threshold <italic>ε</italic> = 0.01. A set of <italic>K</italic> = 5 diverse counterfactuals is generated per patient by running the search <italic>K</italic> times with diversity-promoting initializations that enforce a minimum cosine distance of 0.3 between any two returned counterfactuals.</p>
        <p>The diversity regularization term <italic>L</italic><sub>diversity</sub> draws on the counterfactual fairness framework [<xref ref-type="bibr" rid="B35">35</xref>] to ensure that the returned counterfactual set covers qualitatively distinct recourse paths for example, one counterfactual emphasizing medication adjustment, another emphasizing psychotherapy increase, and a third emphasizing lifestyle and adherence changes rather than <italic>K</italic> near-identical variations of the same intervention.</p>
      </sec>
      <sec id="sec4dot3">
        <title>4.3. Module 3: Clinical Validity Filter</title>
        <p>Generated counterfactuals pass through a rule-based clinical validity filter before being presented to the clinician. The filter enforces four validity criteria:</p>
        <p><bold>Clinical plausibility:</bold>The counterfactual’s feature values must fall within clinically realistic ranges for each feature type (e.g., medication doses within approved therapeutic ranges; symptom scores within scale bounds; therapy frequency within available service configurations).<bold>Medical actionability:</bold>All suggested feature changes must correspond to interventions within the prescribing and referral authority of a consultant psychiatrist (e.g., medication switch or dose adjustment; referral to structured psychotherapy; inpatient admission for monitoring). Features flagged as immutable by the clinician panel are excluded from all counterfactuals.<bold>Causal consistency:</bold>The counterfactual must satisfy all directed constraints in the causal graph G<sub>C</sub> no feature change is permitted that implies a downstream consequence that is not also reflected in the counterfactual. For example, a reduction in depression score must be accompanied by a plausible intervention that would produce it.<bold>Patient-preference alignment:</bold>Where patient-stated preferences are available from the clinical record (e.g., preference against certain medication classes, stated reluctance to increase appointment frequency), the filter excludes counterfactuals that violate these preferences. This criterion is applied as a soft constraint with a clinician-override option.</p>
        <p>Counterfactuals that fail any hard constraint are discarded and replaced through additional gradient search iterations. The validity filter achieves a pass rate of 91.4% on the first-generated counterfactual per patient and 97.8% across the full <italic>K</italic> = 5 set.</p>
      </sec>
      <sec id="sec4dot4">
        <title>4.4. Training Protocol</title>
        <p>The outcome predictor and all modality encoders were jointly trained using AdamW (lr = 2 × 10<sup>−4</sup>, weight decay = 1 × 10<sup>−</sup><sup>2</sup>, <italic>β</italic><sub>1</sub> = 0.9, <italic>β</italic><sub>2</sub> = 0.999) with cosine annealing over 150 epochs. Weighted cross-entropy addressed class imbalance. Multi-task loss combined outcome classification (primary), recourse sparsity regularization (auxiliary), and causal constraint consistency (auxiliary). Data splits: 70% training, 15% validation, 15% test, stratified jointly on diagnosis, treatment class, and response rate tertile. Implementation: PyTorch [<xref ref-type="bibr" rid="B36">36</xref>] with HuggingFace Transformers. Hardware: NVIDIA A100 80GB. <xref ref-type="fig" rid="fig2">Figure 2</xref><xref ref-type="fig" rid="fig2">Figure 2</xref> shows training convergence CounterPsych reaches plateau at epoch 108, training accuracy 0.963, validation accuracy 0.941.</p>
      </sec>
    </sec>
    <sec id="sec5">
      <title>5. Experimental Results</title>
      <sec id="sec5dot1">
        <title>5.1. Treatment Outcome Prediction Performance</title>
        <p><xref ref-type="fig" rid="fig3">Figure 3</xref><xref ref-type="fig" rid="fig3">Figure 3</xref> presents the confusion matrix on the held-out test set. CounterPsych achieves per-class accuracy of 94.8% for Treatment Response, 93.1% for Partial </p>
        <fig id="fig2">
          <label>Figure 2</label>
          <graphic xlink:href="https://html.scirp.org/file/1115348-rId31.jpeg?20260529040050" />
        </fig>
        <p><bold>Figure 2.</bold> Training and Validation Convergence. Accuracy (left) and cross-entropy loss (right) over 150 epochs for CounterPsych (navy), WACHUNG CF baseline (teal), and DiCE (coral). CounterPsych achieves the highest validation accuracy with a narrow train-validation gap (0.022). Early stopping fires at epoch 108 (gold dotted line).</p>
        <fig id="fig3">
          <label>Figure 3</label>
          <graphic xlink:href="https://html.scirp.org/file/1115348-rId32.jpeg?20260529040050" />
        </fig>
        <p><bold>Figure 3.</bold> Treatment Outcome Confusion Matrix. Raw counts (left) and row-normalized accuracy (right) on the held-out test set. CounterPsych achieves &gt;93% per-class accuracy across all three outcome categories. The primary misclassification occurs at the Response-Partial Response boundary, reflecting the inherent clinical uncertainty at the 50% symptom reduction threshold.</p>
        <p>Response, and 94.6% for non-response. The most frequent misclassification is between Response and Partial Response (5.2% of Response cases classified as Partial Response), which is clinically expected given that the distinction between full and partial response at the 12-week endpoint involves threshold judgments that are genuinely uncertain even for experienced clinicians.</p>
        <p><xref ref-type="fig" rid="fig4">Figure 4</xref><xref ref-type="fig" rid="fig4">Figure 4</xref> presents the per-class ROC curves and the comparative AUC-ROC ranking. CounterPsych achieves macro-AUC of 0.977, with per-class AUC of 0.982 (Response), 0.971 (Partial Response), and 0.974 (Non-Response).</p>
        <fig id="fig4">
          <label>Figure 4</label>
          <graphic xlink:href="https://html.scirp.org/file/1115348-rId33.jpeg?20260529040050" />
        </fig>
        <p><bold>Figure 4.</bold> ROC Analysis and AUC-ROC Model Comparison. Left: Per-class ROC curves for CounterPsych all three classes achieve AUC &gt; 0.97. Right: Macro-AUC comparison across all eight models. CounterPsych (0.977) significantly outperforms all baselines including WACHUNG CF (0.903) and the ablation without Bayesian ensemble (0.941) (DeLong test, p &lt; 0.01 for all comparisons).</p>
        <p><bold>Table 2</bold> presents the full performance comparison across eight models. CounterPsych achieves statistically significant improvements over all baselines on all four primary metrics.</p>
        <p><bold>Table 2.</bold>Comparative outcome prediction performance held-out test set (n = 496 patients).</p>
        <table-wrap id="tbl2">
          <label>Table 2</label>
          <table>
            <tbody>
              <tr>
                <td>
                  <bold>Model</bold>
                </td>
                <td>
                  <bold>Accuracy</bold>
                </td>
                <td>
                  <bold>Macro-F1</bold>
                </td>
                <td>
                  <bold>Precision</bold>
                </td>
                <td>
                  <bold>Recall</bold>
                </td>
                <td>
                  <bold>AUC-ROC</bold>
                </td>
              </tr>
              <tr>
                <td>Logistic Regression</td>
                <td>0.713</td>
                <td>0.668</td>
                <td>0.681</td>
                <td>0.657</td>
                <td>0.791</td>
              </tr>
              <tr>
                <td>SVM (RBF)</td>
                <td>0.741</td>
                <td>0.701</td>
                <td>0.714</td>
                <td>0.688</td>
                <td>0.818</td>
              </tr>
              <tr>
                <td>
                  Random Forest [
                  <xref ref-type="bibr" rid="B44">44</xref>
                  ]
                </td>
                <td>0.782</td>
                <td>0.748</td>
                <td>0.761</td>
                <td>0.737</td>
                <td>0.851</td>
              </tr>
              <tr>
                <td>
                  XGBoost [
                  <xref ref-type="bibr" rid="B45">45</xref>
                  ]
                </td>
                <td>0.814</td>
                <td>0.781</td>
                <td>0.796</td>
                <td>0.768</td>
                <td>0.882</td>
              </tr>
              <tr>
                <td>DiCE (CF baseline)</td>
                <td>0.812</td>
                <td>0.782</td>
                <td>0.797</td>
                <td>0.769</td>
                <td>0.871</td>
              </tr>
              <tr>
                <td>WACHUNG CF</td>
                <td>0.837</td>
                <td>0.808</td>
                <td>0.821</td>
                <td>0.796</td>
                <td>0.903</td>
              </tr>
              <tr>
                <td>Ablation (no Bayesian)</td>
                <td>0.919</td>
                <td>0.894</td>
                <td>0.907</td>
                <td>0.882</td>
                <td>0.941</td>
              </tr>
              <tr>
                <td>
                  <bold>CounterPsych (Ours)</bold>
                </td>
                <td>
                  <bold>0.941</bold>
                </td>
                <td>
                  <bold>0.919</bold>
                </td>
                <td>
                  <bold>0.932</bold>
                </td>
                <td>
                  <bold>0.907</bold>
                </td>
                <td>
                  <bold>0.977</bold>
                </td>
              </tr>
            </tbody>
          </table>
        </table-wrap>
        <p>Statistically significant improvement over all baselines (DeLong test, p &lt; 0.01; McNemar test, p &lt; 0.001). Abstaining predictions (8.7%) excluded. AUC-ROC reported as macro one-vs-rest average.</p>
        <p>Site-Held-Out and Chronological Validation. To assess generalization beyond the standard patient-stratified split, two additional evaluations were performed. First, a site-held-out validation was conducted by training on three sites and testing on the fourth, repeated for each of the four sites (leave-one-site-out cross-validation). CounterPsych achieved mean accuracy of 89.7% (SD 1.4%), mean AUC-ROC of 0.961 (SD 0.012) across the four held-out site folds, a 4.4 pp accuracy reduction relative to the same-distribution test set, indicating moderate but expected performance degradation under domain shift. Second, a chronological validation was performed by training on admissions prior to January 2023 and testing on the prospective 2023-2024 cohort (n = 214 patients). Accuracy on the chronological holdout was 91.3%, AUC-ROC 0.969. Abstentions (predictions with <italic>H</italic>&gt; <italic>τ</italic> = 0.44) were included in all reported denominators; when abstentions are excluded (8.7% of predictions), accuracy on the chronological holdout rises to 93.1%. </p>
      </sec>
      <sec id="sec5dot2">
        <title>5.2. Counterfactual Analysis</title>
        <p><xref ref-type="fig" rid="fig5">Figure 5</xref><xref ref-type="fig" rid="fig5">Figure 5</xref> presents the core counterfactual analysis: a worked single-patient example, the distribution of counterfactual proximity distances across the cohort, and the clinical validity rates across five constraint categories.</p>
        <fig id="fig5">
          <label>Figure 5</label>
          <graphic xlink:href="https://html.scirp.org/file/1115348-rId34.jpeg?20260529040050" />
        </fig>
        <p><bold>Figure 5.</bold> Counterfactual Generation Analysis. Panel A: Single-patient worked example comparing the factual (non-responding) treatment configuration with the CounterPsych counterfactual, annotated with the predicted feature changes and their direction. For this patient, the model identifies that increasing therapy sessions from 0 to 4/month, reducing quetiapine by 150 mg, increasing medication adherence by 22%, and improving sleep duration by 1.8 hours/night would flip the prediction from Non-Response to Response. Panel B: Distribution of L2 distances between factual and counterfactual in feature space CounterPsych CFs (mean 1.62, teal) are substantially more proximate to the factual than DiCE CFs (mean 4.21, coral), confirming greater actionability. Panel C: Clinical validity rates across five constraint dimensions CounterPsych achieves &gt; 83% validity across all criteria, dramatically outperforming DiCE on actionability (91% vs 62%) and causal consistency (88% vs 54%).</p>
        <p>The proximity analysis confirms that CounterPsych counterfactuals are substantially more actionable than prior methods. The mean L2 distance of 1.62 in standardized feature space compared to 4.21 for DiCE indicates that CounterPsych recourse plans involve smaller, more clinically realistic changes. Prototype-based methods [<xref ref-type="bibr" rid="B37">37</xref>] achieve comparable proximity but substantially lower clinical plausibility because their anchoring to training instances does not distinguish between plausible and implausible prototype configurations. Case-based counterfactual methods [<xref ref-type="bibr" rid="B38">38</xref>] achieve high plausibility but sacrifice sparsity, generating counterfactuals that require changes to many more features than the clinical setting permits.</p>
      </sec>
      <sec id="sec5dot3">
        <title>5.3. Feature Importance and Attribution</title>
        <p><xref ref-type="fig" rid="fig6">Figure 6</xref><xref ref-type="fig" rid="fig6">Figure 6</xref> presents the feature group importance decomposition and the top-15 individual feature importances.</p>
        <fig id="fig6">
          <label>Figure 6</label>
          <graphic xlink:href="https://html.scirp.org/file/1115348-rId35.jpeg?20260529040050" />
        </fig>
        <p><bold>Figure 6.</bold> Feature Group and Individual Feature Importance. Left: Feature group importance from Random Forest decomposition across the five clinical domains. Clinical state features dominate (38.7%), consistent with the established predictive value of baseline symptom severity. Treatment profile contributes 28.4% reflecting the substantial predictive information in medication history, prior response patterns, and adherence. Right: Top 15 individual features. HDRS-17 total score, GAF, and medication dose equivalent are the three strongest predictors; prior treatment response history ranks fourth, confirming that the trajectory of prior responses is highly informative about future response probability.</p>
        <p>Clinical state features contribute the largest share (38.7%) of predictive information, driven by HDRS-17 total score and GAF. Treatment profile accounts for 28.4%, with prior treatment response history as the single most informative treatment feature a patient who has responded to a prior trial of the same medication class is substantially more likely to respond again, a clinical heuristic that the model learns and quantifies precisely. Patient history contributes 19.3%, biomarkers 8.9%, and demographics 4.7%. The relatively modest contribution of demographics, particularly age and gender, to the predictive model is encouraging from a fairness perspective and consistent with the subgroup analysis showing maximum accuracy disparity of only 1.7% across demographic groups. (See <xref ref-type="fig" rid="fig7">Figure 7</xref><xref ref-type="fig" rid="fig7">Figure 7</xref>)</p>
      </sec>
      <sec id="sec5dot4">
        <title>5.4. Calibration and Ablation</title>
        <p><xref ref-type="fig" rid="fig8">Figure 8</xref><xref ref-type="fig" rid="fig8">Figure 8</xref> presents the calibration reliability diagram and ablation study.</p>
        <fig id="fig7">
          <label>Figure 7</label>
          <graphic xlink:href="https://html.scirp.org/file/1115348-rId36.jpeg?20260529040050" />
        </fig>
        <p><bold>Figure 7.</bold> All-Metrics Comparison and Bayesian Uncertainty Distribution. Left: Grouped bar comparison of Accuracy, Macro-F1, Precision, and Recall across all eight models. CounterPsych leads consistently across all metrics. Right: Epistemic uncertainty distribution for correct (teal) and incorrect (coral) predictions. The review threshold <italic>τ</italic> = 0.17 cleanly separates the distributions the abstention protocol correctly identifies predictions in the uncertain regime.</p>
        <fig id="fig8">
          <label>Figure 8</label>
          <graphic xlink:href="https://html.scirp.org/file/1115348-rId37.jpeg?20260529040050" />
        </fig>
        <p><bold>Figure 8.</bold> Calibration Reliability Diagram and Module Ablation. Left: CounterPsych (navy, ECE = 0.021) tracks the perfect calibration diagonal closely across all confidence bins. WACHUNG CF (teal, ECE = 0.054) and DiCE (coral, ECE = 0.119) show systematic overconfidence particularly dangerous in clinical applications where overconfident non-response predictions might lead to premature treatment termination. Right: Ablation study removing the Bayesian ensemble produces the largest calibration degradation (ECE rises from 0.021 to 0.061); removing the counterfactual generator has no effect on accuracy but disables the entire explainability functionality; removing the History TCN produces the largest accuracy drop (−4.3 pp).</p>
        <p>The ablation confirms every module contributes independently. The Bayesian ensemble contributes primarily to calibration rather than accuracy removing it reduces accuracy by 1.5 pp but nearly triples ECE (0.021 → 0.061). The History TCN contributes the largest accuracy gain (−4.3 pp when removed), confirming that longitudinal treatment trajectory is the most informative single architectural component. The Clinical Transformer contributes −3.4 pp when removed. The predictor-only ablation achieves 85.2% accuracy competitive with XGBoost but loses all explainability functionality.</p>
      </sec>
      <sec id="sec5dot5">
        <title>5.5. Clinician Evaluation and Recourse Analysis</title>
        <p><xref ref-type="fig" rid="fig9">Figure 9</xref><xref ref-type="fig" rid="fig9">Figure 9</xref> presents the three-panel clinician evaluation: Likert rating scores across five quality dimensions, the recourse sparsity distribution, and subgroup fairness analysis.</p>
        <fig id="fig9">
          <label>Figure 9</label>
          <graphic xlink:href="https://html.scirp.org/file/1115348-rId38.jpeg?20260529040050" />
        </fig>
        <p><bold>Figure 9.</bold> Clinician Evaluation, Recourse Sparsity, and Subgroup Fairness. Left: CounterPsych counterfactuals receive mean ratings of 4.28-4.51/5 across five clinical quality dimensions from 24 consultant psychiatrists, substantially outperforming DiCE (3.08-3.21/5). The largest gap is on actionability (CounterPsych 4.28 vs DiCE 2.98) and trustworthiness (4.51 vs 3.14). Centre: Recourse sparsity distribution 75.3% of CounterPsych counterfactuals change 3 or fewer features, confirming that the generated recourse plans are parsimonious and clinically manageable. Right: Subgroup accuracy analysis maximum accuracy disparity of 1.7% across demographic and diagnostic subgroups confirms equitable performance.</p>
        <p>The clinician evaluation was conducted as a prospective, blinded rating study with 24 consultant psychiatrists drawn from three European academic psychiatric centers (8 per site). Raters held a mean of 11.4 years of post-certification clinical experience (range 5–28 years). Each rater evaluated a randomly sampled set of 20 de-identified patient cases, 10 CounterPsych counterfactuals and 10 DiCE counterfactuals presented in random interleaved order with method identity withheld (blinded presentation). Cases were sampled from the held-out test set, stratified by diagnosis and outcome class to ensure representation across all five diagnostic categories and all three outcome classes. The same 20 cases were rated by all 24 psychiatrists, yielding 480 ratings per method. Presentation order was randomized per rater using a Latin square design to control for sequence effects. Raters scored each counterfactual on five 5-point Likert dimensions: clinical plausibility, medical actionability, causal coherence, patient-appropriateness, and trustworthiness. Inter-rater agreement was assessed using intraclass correlation coefficient (ICC): ICC (2, 1) = 0.81 (95% CI: 0.77 - 0.84) for plausibility and ICC (2, 1) = 0.79 (95% CI: 0.75 - 0.83) for trustworthiness indicating good to excellent agreement. Paired Wilcoxon signed-rank tests comparing CounterPsych versus DiCE ratings were significant on all five dimensions (p &lt; 0.001 for all). No rater evaluated cases from their own institution [<xref ref-type="bibr" rid="B39">39</xref>][<xref ref-type="bibr" rid="B40">40</xref>].</p>
        <p>The recourse sparsity analysis shows that 75.3% of CounterPsych counterfactuals change 3 or fewer clinical features. This is clinically important: a treatment recommendation that requires simultaneous changes to 7 different aspects of a patient’s care plan is not actionable in any realistic outpatient setting, regardless of its mathematical validity. The mean of 2.3 feature changes per counterfactual corresponds to the scale of a typical treatment adjustment in psychiatric outpatient practice for example, increasing the dose of a current medication, adding structured psychotherapy, and setting a specific sleep hygiene target.</p>
      </sec>
    </sec>
    <sec id="sec6">
      <title>6. Discussion</title>
      <sec id="sec6dot1">
        <title>6.1. What CounterPsych Demonstrates</title>
        <p>The central finding is that counterfactual explainability in psychiatric AI is achievable at clinical quality not as a theoretical property of the optimization objective, but as a practically validated characteristic of the system’s outputs as judged by the clinicians who would use them. The mean plausibility rating of 4.42/5 and trustworthiness of 4.51/5 from 24 experienced psychiatrists are not cosmetic improvements over DiCE’s 3.21/5 and 3.14/5. They represent the difference between an explanation system that clinicians will engage with and one they will dismiss. Machine learning tools for psychiatric readmission prediction [<xref ref-type="bibr" rid="B41">41</xref>] have consistently failed to achieve clinical adoption despite reasonable predictive accuracy precisely because the outputs could not be trusted or acted on. CounterPsych’s counterfactual generator is designed to solve this trust deficit at its root by making the explanations causally consistent with the clinical domain rather than simply numerically proximate to the factual.</p>
        <p>The causal constraint graph is the architectural innovation that makes this possible. The key insight is that clinical plausibility is not just about whether feature values fall within realistic ranges it is about whether the pattern of feature changes tells a causally coherent clinical story. A counterfactual that suggests a patient’s HDRS-17 score would be 8 points lower without specifying an intervention that would produce that reduction is not a clinical explanation it is a mathematical artifact. By enforcing causal consistency through the constraint graph G<sub>C</sub>, CounterPsych ensures that every generated counterfactual corresponds to a coherent sequence of clinical actions, each of which is within the treating psychiatrist’s authority and the patient’s realistic circumstances.</p>
        <p>The proximity advantage CounterPsych CFs at mean L2 = 1.62 versus DiCE at 4.21 is a direct consequence of the clinical constraint structure. When the optimization is constrained to actionable, causally consistent feature changes, the resulting counterfactuals are automatically closer to the factual because the mutable, actionable features are a small and carefully selected subset of the full feature space. The constraint that initially appears to restrict the search space is simultaneously reducing the distance from the factual to the nearest valid counterfactual because valid counterfactuals lie in a region of the feature space that is clinically adjacent to the patient’s current situation.</p>
      </sec>
      <sec id="sec6dot2">
        <title>6.2. Clinical Implications</title>
        <p>CounterPsych’s outputs have three concrete clinical use cases. First, at treatment initiation, the model’s prediction and its accompanying counterfactuals can inform the comparative choice between treatment options showing the psychiatrist which configuration of medication, dose, and psychotherapy the model predicts will produce response, with the specific feature changes quantified and ranked by their estimated contribution. Second, during treatment monitoring, serial CounterPsych evaluations can track whether the patient’s trajectory is moving toward or away from the counterfactual target providing early warning of emerging non-response before it becomes clinically manifest. Third, at treatment review or switching decisions, the counterfactual set provides a structured set of evidence-based recourse options that the clinician can evaluate against the patient’s preferences and circumstances rather than relying on heuristic rule-of-thumb switching algorithms [<xref ref-type="bibr" rid="B42">42</xref>].</p>
        <p>The ECE of 0.021 is a clinically meaningful calibration result. It means that when CounterPsych reports 80% confidence in a response prediction, approximately 80% of such predictions are correct enabling the psychiatrist to use the model’s confidence score as genuine probabilistic information in their decision-making, rather than treating the score as an uninterpretable black-box output. Combined with the counterfactual explanations, this creates a complete clinical decision support workflow: the model says how confident it is in its prediction, what that prediction is, and what would need to change for it to be different.</p>
      </sec>
      <sec id="sec6dot3">
        <title>6.3. Limitations</title>
        <p><bold>12-week outcome horizon.</bold>Treatment response was assessed at 12 weeks. For conditions with longer treatment timescales particularly bipolar disorder, where mood stabiliser response may require 6 - 12 months of continuous treatment the 12-week horizon may miss genuine responders classified as non-responders. Extended follow-up analyses at 6 and 12 months are planned.<bold>Causal graph construction.</bold>The causal constraint graph G<sub>C</sub> was built by expert elicitation from five consultant psychiatrists and validated against published treatment guidelines. It necessarily reflects the current state of clinical knowledge and expert consensus, which may not capture all relevant causal dependencies, particularly for novel drug combinations or less-studied patient subgroups.<bold>Single clinician evaluation site.</bold>The clinician evaluation with 24 psychiatrists was conducted at a single European academic center, where psychiatrists may have higher baseline familiarity with AI-assisted tools than average. Replication of the evaluation at community mental health settings and in non-European clinical contexts is required before generalizing the usability findings.<bold>Retrospective-prospective heterogeneity.</bold>The retrospective component introduced protocol variability across sites and years that was partially but not fully mitigated by harmonization. The prospective component, while collected under a standardized protocol, covers only two years and may not capture the full range of treatment responses across longer illness trajectories.<bold>Missing biomarker data.</bold>Biomarker data was available for only 68.6% of patients, and the MICE imputation, while principled, introduces additional uncertainty in biomarker-dependent counterfactuals. Future prospective data collection should prioritize standardized biomarker acquisition across all sites.</p>
      </sec>
    </sec>
    <sec id="sec7">
      <title>7. Conclusions</title>
      <p>Psychiatric treatment decisions are hard because the right answer varies dramatically from patient to patient and because the evidence base, while extensive at the population level, provides limited guidance for the individual sitting across the desk. Machine learning can close some of this gap but only if the models it produces can explain themselves in language and logic that clinicians recognize and trust. A probability score is not an explanation. A list of feature importances is not a recommendation. A counterfactual is both.</p>
      <p>CounterPsych demonstrates that counterfactual explainability in psychiatric AI is achievable at clinical quality with a mean clinician trustworthiness rating of 4.51/5, a mean of 2.3 feature changes per counterfactual, 94.1% outcome prediction accuracy, and ECE of 0.021. The causal constraint graph and clinical validity filter are the architectural innovations that make this possible: they transform mathematically generated counterfactuals into clinically coherent treatment recommendations that psychiatrists can consider, discuss with patients, and act on.</p>
      <p>The next steps are prospective deployment trials in active outpatient settings, extension to inpatient and emergency psychiatric contexts where the decision stakes are highest, development of patient-facing counterfactual interfaces that translate clinical recourse plans into accessible language, and regulatory engagement for clinical decision support certification under the EU AI Act framework. Counterfactual explainability is not just a desirable property of psychiatric AI it is, we argue, the minimum standard for clinical utility. A model that can predict but not explain is not yet a clinical tool [<xref ref-type="bibr" rid="B43">43</xref>]-[<xref ref-type="bibr" rid="B45">45</xref>].</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <title>References</title>
      <ref id="B1">
        <label>1.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Rudin, C. (2019) Stop Explaining Black Box Machine Learning Models for High Stakes Decisions and Use Interpretable Models Instead. <italic>Nature</italic><italic>Machine</italic><italic>Intelligence</italic>, 1, 206-215. https://doi.org/10.1038/s42256-019-0048-x <pub-id pub-id-type="doi">10.1038/s42256-019-0048-x</pub-id><pub-id pub-id-type="pmid">35603010</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1038/s42256-019-0048-x">https://doi.org/10.1038/s42256-019-0048-x</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Rudin, C.</string-name>
            </person-group>
            <year>2019</year>
            <article-title>Stop Explaining Black Box Machine Learning Models for High Stakes Decisions and Use Interpretable Models Instead</article-title>
            <source>Nature Machine Intelligence</source>
            <volume>1</volume>
            <pub-id pub-id-type="doi">10.1038/s42256-019-0048-x</pub-id>
            <pub-id pub-id-type="pmid">35603010</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B2">
        <label>2.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Wexler, J., Pushkarna, M., Bolukbasi, T., Wattenberg, M., Viegas, F. and Wilson, J. (2019) The What-If Tool: Interactive Probing of Machine Learning Models. <italic>IEEE</italic><italic>Transactions</italic><italic>on</italic><italic>Visualization</italic><italic>and</italic><italic>Computer</italic><italic>Graphics</italic>, 26, 56-65. https://doi.org/10.1109/tvcg.2019.2934619 <pub-id pub-id-type="doi">10.1109/tvcg.2019.2934619</pub-id><pub-id pub-id-type="pmid">31442996</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1109/tvcg.2019.2934619">https://doi.org/10.1109/tvcg.2019.2934619</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Wexler, J.</string-name>
              <string-name>Pushkarna, M.</string-name>
              <string-name>Bolukbasi, T.</string-name>
              <string-name>Wattenberg, M.</string-name>
              <string-name>Viegas, F.</string-name>
              <string-name>Wilson, J.</string-name>
            </person-group>
            <year>2019</year>
            <article-title>The What-If Tool: Interactive Probing of Machine Learning Models</article-title>
            <source>IEEE Transactions on Visualization and Computer Graphics</source>
            <volume>26</volume>
            <pub-id pub-id-type="doi">10.1109/tvcg.2019.2934619</pub-id>
            <pub-id pub-id-type="pmid">31442996</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B3">
        <label>3.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Doshi-Velez, F. and Kim, B. (2017) Towards a Rigorous Science of Interpretable Machine Learning. arXiv: 1702.08608.</mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Doshi-Velez, F.</string-name>
              <string-name>Kim, B.</string-name>
            </person-group>
            <year>2017</year>
            <article-title>Towards a Rigorous Science of Interpretable Machine Learning</article-title>
            <fpage>1702</fpage>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B4">
        <label>4.</label>
        <citation-alternatives>
          <mixed-citation publication-type="confproc">Bhatt, U., Xiang, A., Sharma, S., Weller, A., Taly, A., Jia, Y., <italic>et al</italic>. (2020) Explainable Machine Learning in Deployment. <italic>Proceedings of the</italic> 2020 <italic>Conference on Fairness</italic>, <italic>Accountability</italic>, <italic>and Transparency</italic>, Barcelona, 27-30 January 2020, 648-657. https://doi.org/10.1145/3351095.3375624 <pub-id pub-id-type="doi">10.1145/3351095.3375624</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1145/3351095.3375624">https://doi.org/10.1145/3351095.3375624</ext-link></mixed-citation>
          <element-citation publication-type="confproc">
            <person-group person-group-type="author">
              <string-name>Bhatt, U.</string-name>
              <string-name>Xiang, A.</string-name>
              <string-name>Sharma, S.</string-name>
              <string-name>Weller, A.</string-name>
              <string-name>Taly, A.</string-name>
              <string-name>Jia, Y.</string-name>
              <string-name>Fairness, A</string-name>
              <string-name>Transparency, B</string-name>
            </person-group>
            <year>2020</year>
            <article-title>Explainable Machine Learning in Deployment</article-title>
            <source>Proceedings of the 2020 Conference on Fairness</source>
            <volume>27</volume>
            <pub-id pub-id-type="doi">10.1145/3351095.3375624</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B5">
        <label>5.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Hewson, T., Lagunes-Cordoba, E. and Tracy, D.K. (2021) Benefits and Barriers to Mentoring in Psychiatry: A Mentee’s Perspective. <italic>BJPsych</italic><italic>Advances</italic>, 27, 228-229. https://doi.org/10.1192/bja.2020.85 <pub-id pub-id-type="doi">10.1192/bja.2020.85</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1192/bja.2020.85">https://doi.org/10.1192/bja.2020.85</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Hewson, T.</string-name>
              <string-name>Lagunes-Cordoba, E.</string-name>
              <string-name>Tracy, D.K.</string-name>
            </person-group>
            <year>2021</year>
            <article-title>Benefits and Barriers to Mentoring in Psychiatry: A Mentee’s Perspective</article-title>
            <source>BJPsych Advances</source>
            <volume>27</volume>
            <pub-id pub-id-type="doi">10.1192/bja.2020.85</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B6">
        <label>6.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Antequera, A., Lawson, D.O., Noorduyn, S.G., Dewidar, O., Avey, M., Bhutta, Z.A., <italic>et al</italic>. (2021) Improving Social Justice in COVID-19 Health Research: Interim Guidelines for Reporting Health Equity in Observational Studies. <italic>International Journal of Environmental Research and Public Health</italic>, 18, Article 9357. https://doi.org/10.3390/ijerph18179357 <pub-id pub-id-type="doi">10.3390/ijerph18179357</pub-id><pub-id pub-id-type="pmid">34501949</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.3390/ijerph18179357">https://doi.org/10.3390/ijerph18179357</ext-link></mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Antequera, A.</string-name>
              <string-name>Lawson, D.O.</string-name>
              <string-name>Noorduyn, S.G.</string-name>
              <string-name>Dewidar, O.</string-name>
              <string-name>Avey, M.</string-name>
              <string-name>Bhutta, Z.A.</string-name>
            </person-group>
            <year>2021</year>
            <article-title>Improving Social Justice in COVID-19 Health Research: Interim Guidelines for Reporting Health Equity in Observational Studies</article-title>
            <source>International Journal of Environmental Research and Public Health</source>
            <volume>18</volume>
            <elocation-id>9357</elocation-id>
            <pub-id pub-id-type="doi">10.3390/ijerph18179357</pub-id>
            <pub-id pub-id-type="pmid">34501949</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B7">
        <label>7.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">European Parliament (2024) Regulation (EU) 2024/1689 on Artificial Intelligence (EU AI Act). <italic>Official Journal of the EU</italic>.</mixed-citation>
          <element-citation publication-type="journal">
            <year>2024</year>
            <article-title>Regulation (EU) 2024/1689 on Artificial Intelligence (EU AI Act)</article-title>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B8">
        <label>8.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Wachter, S., Mittelstadt, B. and Russell, C. (2017) Counterfactual Explanations without Opening the Black Box: Automated Decisions and the GDPR. <italic>SSRN</italic><italic>Electronic</italic><italic>Journal</italic>, 31, 841-887. https://doi.org/10.2139/ssrn.3063289 <pub-id pub-id-type="doi">10.2139/ssrn.3063289</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.2139/ssrn.3063289">https://doi.org/10.2139/ssrn.3063289</ext-link></mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Wachter, S.</string-name>
              <string-name>Mittelstadt, B.</string-name>
              <string-name>Russell, C.</string-name>
            </person-group>
            <year>2017</year>
            <article-title>Counterfactual Explanations without Opening the Black Box: Automated Decisions and the GDPR</article-title>
            <source>SSRN Electronic Journal</source>
            <volume>31</volume>
            <pub-id pub-id-type="doi">10.2139/ssrn.3063289</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B9">
        <label>9.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Bracke, P., Datta, A., Jung, C. and Sen, S. (2019) Machine Learning Explainability in Finance: An Application to Default Risk Analysis. <italic>SSRN</italic><italic>Electronic</italic><italic>Journal</italic>, 44 p. https://doi.org/10.2139/ssrn.3435104 <pub-id pub-id-type="doi">10.2139/ssrn.3435104</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.2139/ssrn.3435104">https://doi.org/10.2139/ssrn.3435104</ext-link></mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Bracke, P.</string-name>
              <string-name>Datta, A.</string-name>
              <string-name>Jung, C.</string-name>
              <string-name>Sen, S.</string-name>
            </person-group>
            <year>2019</year>
            <article-title>Machine Learning Explainability in Finance: An Application to Default Risk Analysis</article-title>
            <source>SSRN Electronic Journal</source>
            <volume>44</volume>
            <pub-id pub-id-type="doi">10.2139/ssrn.3435104</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B10">
        <label>10.</label>
        <citation-alternatives>
          <mixed-citation publication-type="confproc">Mothilal, R.K., Sharma, A. and Tan, C. (2020) Explaining Machine Learning Classifiers through Diverse Counterfactual Explanations. <italic>Proceedings of the</italic> 2020 <italic>Conference on Fairness</italic>, <italic>Accountability</italic>, <italic>and Transparency</italic>, Barcelona, 27-30 January 2020, 607-617. https://doi.org/10.1145/3351095.3372850 <pub-id pub-id-type="doi">10.1145/3351095.3372850</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1145/3351095.3372850">https://doi.org/10.1145/3351095.3372850</ext-link></mixed-citation>
          <element-citation publication-type="confproc">
            <person-group person-group-type="author">
              <string-name>Mothilal, R.K.</string-name>
              <string-name>Sharma, A.</string-name>
              <string-name>Tan, C.</string-name>
              <string-name>Fairness, A</string-name>
              <string-name>Transparency, B</string-name>
            </person-group>
            <year>2020</year>
            <article-title>Explaining Machine Learning Classifiers through Diverse Counterfactual Explanations</article-title>
            <source>Proceedings of the 2020 Conference on Fairness</source>
            <volume>27</volume>
            <pub-id pub-id-type="doi">10.1145/3351095.3372850</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B11">
        <label>11.</label>
        <citation-alternatives>
          <mixed-citation publication-type="confproc">Karimi, A., Schölkopf, B. and Valera, I. (2021) Algorithmic Recourse: From Counterfactual Explanations to Interventions. <italic>Proceedings of the</italic> 2021 <italic>ACM Conference on Fairness</italic>, <italic>Accountability</italic>, <italic>and Transparency</italic>, Virtual Event, 3-10 March 2021, 353-362. https://doi.org/10.1145/3442188.3445899 <pub-id pub-id-type="doi">10.1145/3442188.3445899</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1145/3442188.3445899">https://doi.org/10.1145/3442188.3445899</ext-link></mixed-citation>
          <element-citation publication-type="confproc">
            <person-group person-group-type="author">
              <string-name>Karimi, A.</string-name>
              <string-name>Valera, I.</string-name>
              <string-name>Fairness, A</string-name>
              <string-name>Transparency, V</string-name>
            </person-group>
            <year>2021</year>
            <article-title>Algorithmic Recourse: From Counterfactual Explanations to Interventions</article-title>
            <source>Proceedings of the 2021 ACM Conference on Fairness</source>
            <volume>3</volume>
            <pub-id pub-id-type="doi">10.1145/3442188.3445899</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B12">
        <label>12.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Verma, S., Boonsanong, V., Hoang, M., <italic>et al</italic>. (2020) Counterfactual Explanations for Machine Learning: A Review. arXiv: 2010.10596.</mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Verma, S.</string-name>
              <string-name>Boonsanong, V.</string-name>
              <string-name>Hoang, M.</string-name>
            </person-group>
            <year>2020</year>
            <article-title>Counterfactual Explanations for Machine Learning: A Review</article-title>
            <fpage>2010</fpage>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B13">
        <label>13.</label>
        <citation-alternatives>
          <mixed-citation publication-type="confproc">Ustun, B., Spangher, A. and Liu, Y. (2019) Actionable Recourse in Linear Classification. <italic>Proceedings of the Conference on Fairness</italic>, <italic>Accountability</italic>, <italic>and Transparency</italic>, Atlanta, 29-31 January 2019, 10-19. https://doi.org/10.1145/3287560.3287566 <pub-id pub-id-type="doi">10.1145/3287560.3287566</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1145/3287560.3287566">https://doi.org/10.1145/3287560.3287566</ext-link></mixed-citation>
          <element-citation publication-type="confproc">
            <person-group person-group-type="author">
              <string-name>Ustun, B.</string-name>
              <string-name>Spangher, A.</string-name>
              <string-name>Liu, Y.</string-name>
              <string-name>Fairness, A</string-name>
              <string-name>Transparency, A</string-name>
            </person-group>
            <year>2019</year>
            <article-title>Actionable Recourse in Linear Classification</article-title>
            <source>Proceedings of the Conference on Fairness</source>
            <volume>29</volume>
            <pub-id pub-id-type="doi">10.1145/3287560.3287566</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B14">
        <label>14.</label>
        <citation-alternatives>
          <mixed-citation publication-type="confproc">Pawelczyk, M., Broelemann, K. and Kasneci, G. (2021) Learning Model-Agnostic Counterfactual Explanations for Tabular Data. <italic>Proceedings of The Web Conference</italic> 2020, 20-24 April 2020, 3126-3132. https://doi.org/10.1145/3366423.3380087 <pub-id pub-id-type="doi">10.1145/3366423.3380087</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1145/3366423.3380087">https://doi.org/10.1145/3366423.3380087</ext-link></mixed-citation>
          <element-citation publication-type="confproc">
            <person-group person-group-type="author">
              <string-name>Pawelczyk, M.</string-name>
              <string-name>Broelemann, K.</string-name>
              <string-name>Kasneci, G.</string-name>
            </person-group>
            <year>2021</year>
            <article-title>Learning Model-Agnostic Counterfactual Explanations for Tabular Data</article-title>
            <source>Proceedings of The Web Conference 2020</source>
            <volume>20</volume>
            <pub-id pub-id-type="doi">10.1145/3366423.3380087</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B15">
        <label>15.</label>
        <citation-alternatives>
          <mixed-citation publication-type="confproc">Poyiadzi, R., Sokol, K., Santos-Rodriguez, R., De Bie, T. and Flach, P. (2020) FACE: Feasible and Actionable Counterfactual Explanations. <italic>Proceedings of the AAAI/ACM Conference on AI</italic>, <italic>Ethics</italic>, <italic>and Society</italic>, New York, 7-9 February 2020, 344-350. https://doi.org/10.1145/3375627.3375850 <pub-id pub-id-type="doi">10.1145/3375627.3375850</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1145/3375627.3375850">https://doi.org/10.1145/3375627.3375850</ext-link></mixed-citation>
          <element-citation publication-type="confproc">
            <person-group person-group-type="author">
              <string-name>Poyiadzi, R.</string-name>
              <string-name>Sokol, K.</string-name>
              <string-name>Santos-Rodriguez, R.</string-name>
              <string-name>Bie, T.</string-name>
              <string-name>Flach, P.</string-name>
              <string-name>AI, E</string-name>
              <string-name>Society, N</string-name>
            </person-group>
            <year>2020</year>
            <article-title>FACE: Feasible and Actionable Counterfactual Explanations</article-title>
            <source>Proceedings of the AAAI/ACM Conference on AI</source>
            <volume>7</volume>
            <pub-id pub-id-type="doi">10.1145/3375627.3375850</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B16">
        <label>16.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Mahajan, D., Tan, C. and Sharma, A. (2019) Preserving Causal Constraints in Counterfactual Explanations for Machine Learning Classifiers. arXiv: 1912.03277.</mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Mahajan, D.</string-name>
              <string-name>Tan, C.</string-name>
              <string-name>Sharma, A.</string-name>
            </person-group>
            <year>2019</year>
            <article-title>Preserving Causal Constraints in Counterfactual Explanations for Machine Learning Classifiers</article-title>
            <fpage>1912</fpage>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B17">
        <label>17.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Palpandi, S.B., Palanigurupackiam, N., Almatar, H., Alduhayan, R., Alsomaie, B. and Almazroa, A. (2026) Artificial Intelligence Approaches for Schizophrenia Prediction and Its Biomarkers Using Medical Imaging Data. <italic>Frontiers in Psychiatry</italic>, 17. https://doi.org/10.3389/fpsyt.2026.1821091 <pub-id pub-id-type="doi">10.3389/fpsyt.2026.1821091</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.3389/fpsyt.2026.1821091">https://doi.org/10.3389/fpsyt.2026.1821091</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Palpandi, S.B.</string-name>
              <string-name>Palanigurupackiam, N.</string-name>
              <string-name>Almatar, H.</string-name>
              <string-name>Alduhayan, R.</string-name>
              <string-name>Alsomaie, B.</string-name>
              <string-name>Almazroa, A.</string-name>
            </person-group>
            <year>2026</year>
            <article-title>Artificial Intelligence Approaches for Schizophrenia Prediction and Its Biomarkers Using Medical Imaging Data</article-title>
            <source>Frontiers in Psychiatry</source>
            <volume>17</volume>
            <pub-id pub-id-type="doi">10.3389/fpsyt.2026.1821091</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B18">
        <label>18.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Shatte, A.B.R., Hutchinson, D.M. and Teague, S.J. (2019) Machine Learning in Mental Health: A Scoping Review of Methods and Applications. <italic>Psychological</italic><italic>Medicine</italic>, 49, 1426-1448. https://doi.org/10.1017/s0033291719000151 <pub-id pub-id-type="doi">10.1017/s0033291719000151</pub-id><pub-id pub-id-type="pmid">30744717</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1017/s0033291719000151">https://doi.org/10.1017/s0033291719000151</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Shatte, A.B.R.</string-name>
              <string-name>Hutchinson, D.M.</string-name>
              <string-name>Teague, S.J.</string-name>
            </person-group>
            <year>2019</year>
            <article-title>Machine Learning in Mental Health: A Scoping Review of Methods and Applications</article-title>
            <source>Psychological Medicine</source>
            <volume>49</volume>
            <pub-id pub-id-type="doi">10.1017/s0033291719000151</pub-id>
            <pub-id pub-id-type="pmid">30744717</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B19">
        <label>19.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Christodoulou, E., Ma, J., Collins, G.S., Steyerberg, E.W., Verbakel, J.Y. and Van Calster, B. (2019) A Systematic Review Shows No Performance Benefit of Machine Learning over Logistic Regression for Clinical Prediction Models. <italic>Journal</italic><italic>of</italic><italic>Clinical</italic><italic>Epidemiology</italic>, 110, 12-22. https://doi.org/10.1016/j.jclinepi.2019.02.004 <pub-id pub-id-type="doi">10.1016/j.jclinepi.2019.02.004</pub-id><pub-id pub-id-type="pmid">30763612</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1016/j.jclinepi.2019.02.004">https://doi.org/10.1016/j.jclinepi.2019.02.004</ext-link></mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Christodoulou, E.</string-name>
              <string-name>Ma, J.</string-name>
              <string-name>Collins, G.S.</string-name>
              <string-name>Steyerberg, E.W.</string-name>
              <string-name>Verbakel, J.Y.</string-name>
              <string-name>Calster, B.</string-name>
            </person-group>
            <year>2019</year>
            <article-title>A Systematic Review Shows No Performance Benefit of Machine Learning over Logistic Regression for Clinical Prediction Models</article-title>
            <source>Journal of Clinical Epidemiology</source>
            <volume>110</volume>
            <pub-id pub-id-type="doi">10.1016/j.jclinepi.2019.02.004</pub-id>
            <pub-id pub-id-type="pmid">30763612</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B20">
        <label>20.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Lundberg, S.M., Erion, G., Chen, H., DeGrave, A., Prutkin, J.M., Nair, B., <italic>et al</italic>. (2020) From Local Explanations to Global Understanding with Explainable AI for Trees. <italic>Nature</italic><italic>Machine</italic><italic>Intelligence</italic>, 2, 56-67. https://doi.org/10.1038/s42256-019-0138-9 <pub-id pub-id-type="doi">10.1038/s42256-019-0138-9</pub-id><pub-id pub-id-type="pmid">32607472</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1038/s42256-019-0138-9">https://doi.org/10.1038/s42256-019-0138-9</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Lundberg, S.M.</string-name>
              <string-name>Erion, G.</string-name>
              <string-name>Chen, H.</string-name>
              <string-name>DeGrave, A.</string-name>
              <string-name>Prutkin, J.M.</string-name>
              <string-name>Nair, B.</string-name>
            </person-group>
            <year>2020</year>
            <article-title>From Local Explanations to Global Understanding with Explainable AI for Trees</article-title>
            <source>Nature Machine Intelligence</source>
            <volume>2</volume>
            <pub-id pub-id-type="doi">10.1038/s42256-019-0138-9</pub-id>
            <pub-id pub-id-type="pmid">32607472</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B21">
        <label>21.</label>
        <citation-alternatives>
          <mixed-citation publication-type="confproc">Ribeiro, M.T., Singh, S. and Guestrin, C. (2016) Why Should I Trust You? <italic>Proceedings of the</italic>22 <italic>nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining</italic>, San Francisco, 13-17 August 2016, 1135-1144. https://doi.org/10.1145/2939672.2939778 <pub-id pub-id-type="doi">10.1145/2939672.2939778</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1145/2939672.2939778">https://doi.org/10.1145/2939672.2939778</ext-link></mixed-citation>
          <element-citation publication-type="confproc">
            <person-group person-group-type="author">
              <string-name>Ribeiro, M.T.</string-name>
              <string-name>Singh, S.</string-name>
              <string-name>Guestrin, C.</string-name>
              <string-name>Mining, S</string-name>
            </person-group>
            <year>2016</year>
            <article-title>Why Should I Trust You? Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, San Francisco, 13-17 August 2016, 1135-1144</article-title>
            <pub-id pub-id-type="doi">10.1145/2939672.2939778</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B22">
        <label>22.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Hamilton, M. (1960) A Rating Scale for Depression. <italic>Journal</italic><italic>of</italic><italic>Neurology</italic>, <italic>Neurosurgery</italic><italic>&amp;</italic><italic>Psychiatry</italic>, 23, 56-62. https://doi.org/10.1136/jnnp.23.1.56 <pub-id pub-id-type="doi">10.1136/jnnp.23.1.56</pub-id><pub-id pub-id-type="pmid">14399272</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1136/jnnp.23.1.56">https://doi.org/10.1136/jnnp.23.1.56</ext-link></mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Hamilton, M.</string-name>
              <string-name>Neurology, N</string-name>
            </person-group>
            <year>1960</year>
            <article-title>A Rating Scale for Depression</article-title>
            <source>Journal of Neurology</source>
            <volume>23</volume>
            <pub-id pub-id-type="doi">10.1136/jnnp.23.1.56</pub-id>
            <pub-id pub-id-type="pmid">14399272</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B23">
        <label>23.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Young, R.C., Biggs, J.T., Ziegler, V.E. and Meyer, D.A. (1978) A Rating Scale for Mania: Reliability, Validity and Sensitivity. <italic>British</italic><italic>Journal</italic><italic>of</italic><italic>Psychiatry</italic>, 133, 429-435. https://doi.org/10.1192/bjp.133.5.429 <pub-id pub-id-type="doi">10.1192/bjp.133.5.429</pub-id><pub-id pub-id-type="pmid">728692</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1192/bjp.133.5.429">https://doi.org/10.1192/bjp.133.5.429</ext-link></mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Young, R.C.</string-name>
              <string-name>Biggs, J.T.</string-name>
              <string-name>Ziegler, V.E.</string-name>
              <string-name>Meyer, D.A.</string-name>
              <string-name>Reliability, V</string-name>
            </person-group>
            <year>1978</year>
            <article-title>A Rating Scale for Mania: Reliability, Validity and Sensitivity</article-title>
            <source>British Journal of Psychiatry</source>
            <volume>133</volume>
            <pub-id pub-id-type="doi">10.1192/bjp.133.5.429</pub-id>
            <pub-id pub-id-type="pmid">728692</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B24">
        <label>24.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Endicott, J. (1976) The Global Assessment Scale. <italic>Archives</italic><italic>of</italic><italic>General</italic><italic>Psychiatry</italic>, 33, 766-771. https://doi.org/10.1001/archpsyc.1976.01770060086012 <pub-id pub-id-type="doi">10.1001/archpsyc.1976.01770060086012</pub-id><pub-id pub-id-type="pmid">938196</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1001/archpsyc.1976.01770060086012">https://doi.org/10.1001/archpsyc.1976.01770060086012</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Endicott, J.</string-name>
            </person-group>
            <year>1976</year>
            <article-title>The Global Assessment Scale</article-title>
            <source>Archives of General Psychiatry</source>
            <volume>33</volume>
            <pub-id pub-id-type="doi">10.1001/archpsyc.1976.01770060086012</pub-id>
            <pub-id pub-id-type="pmid">938196</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B25">
        <label>25.</label>
        <citation-alternatives>
          <mixed-citation publication-type="report">Tohen, M., Frank, E., Bowden, C.L., Colom, F., Ghaemi, S.N., Yatham, L.N., Malhi, G.S., <italic>et al</italic>. (2015) The International Society for Bipolar Disorders (ISBD) Task Force Report on the Nomenclature of Course and Outcome in Bipolar Disorders. <italic>Bipolar Disorders</italic>, 11, 453-473.</mixed-citation>
          <element-citation publication-type="report">
            <person-group person-group-type="author">
              <string-name>Tohen, M.</string-name>
              <string-name>Frank, E.</string-name>
              <string-name>Bowden, C.L.</string-name>
              <string-name>Colom, F.</string-name>
              <string-name>Ghaemi, S.N.</string-name>
              <string-name>Yatham, L.N.</string-name>
              <string-name>Malhi, G.S.</string-name>
            </person-group>
            <year>2015</year>
            <article-title>The International Society for Bipolar Disorders (ISBD) Task Force Report on the Nomenclature of Course and Outcome in Bipolar Disorders</article-title>
            <source>Bipolar Disorders</source>
            <volume>11</volume>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B26">
        <label>26.</label>
        <citation-alternatives>
          <mixed-citation publication-type="report">Rush, A.J., Trivedi, M.H., Ibrahim, H.M., Carmody, T.J., Arnow, B., Klein, D.N., <italic>et al</italic>. (2004) The 16-Item Quick Inventory of Depressive Symptomatology (QIDS), Clinician Rating (QIDS-C), and Self-Report (QIDS-SR): A Psychometric Evaluation in Patients with Chronic Major Depression. <italic>Biological</italic><italic>Psychiatry</italic>, 54, 573-583. https://doi.org/10.1016/s0006-3223(02)01866-8 <pub-id pub-id-type="doi">10.1016/s0006-3223(02)01866-8</pub-id><pub-id pub-id-type="pmid">12946886</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1016/s0006-3223(02)01866-8">https://doi.org/10.1016/s0006-3223(02)01866-8</ext-link></mixed-citation>
          <element-citation publication-type="report">
            <person-group person-group-type="author">
              <string-name>Rush, A.J.</string-name>
              <string-name>Trivedi, M.H.</string-name>
              <string-name>Ibrahim, H.M.</string-name>
              <string-name>Carmody, T.J.</string-name>
              <string-name>Arnow, B.</string-name>
              <string-name>Klein, D.N.</string-name>
            </person-group>
            <year>2004</year>
            <article-title>The 16-Item Quick Inventory of Depressive Symptomatology (QIDS), Clinician Rating (QIDS-C), and Self-Report (QIDS-SR): A Psychometric Evaluation in Patients with Chronic Major Depression</article-title>
            <source>Biological Psychiatry</source>
            <volume>3223</volume>
            <issue>02</issue>
            <pub-id pub-id-type="doi">10.1016/s0006-3223(02)01866-8</pub-id>
            <pub-id pub-id-type="pmid">12946886</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B27">
        <label>27.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Johnson, A.E.W., Pollard, T.J., Shen, L., Lehman, L.H., Feng, M., Ghassemi, M., <italic>et al</italic>. (2016) MIMIC-III, a Freely Accessible Critical Care Database. <italic>Scientific</italic><italic>Data</italic>, 3, Article 160035. https://doi.org/10.1038/sdata.2016.35 <pub-id pub-id-type="doi">10.1038/sdata.2016.35</pub-id><pub-id pub-id-type="pmid">27219127</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1038/sdata.2016.35">https://doi.org/10.1038/sdata.2016.35</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Johnson, A.E.W.</string-name>
              <string-name>Pollard, T.J.</string-name>
              <string-name>Shen, L.</string-name>
              <string-name>Lehman, L.H.</string-name>
              <string-name>Feng, M.</string-name>
              <string-name>Ghassemi, M.</string-name>
            </person-group>
            <year>2016</year>
            <article-title>MIMIC-III, a Freely Accessible Critical Care Database</article-title>
            <source>Scientific Data</source>
            <volume>3</volume>
            <elocation-id>160035</elocation-id>
            <pub-id pub-id-type="doi">10.1038/sdata.2016.35</pub-id>
            <pub-id pub-id-type="pmid">27219127</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B28">
        <label>28.</label>
        <citation-alternatives>
          <mixed-citation publication-type="book">Maudsley, R. (2018) The Maudsley Prescribing Guidelines in Psychiatry. 13th Edition, Wiley-Blackwell.</mixed-citation>
          <element-citation publication-type="book">
            <person-group person-group-type="author">
              <string-name>Maudsley, R.</string-name>
              <string-name>Edition, W</string-name>
            </person-group>
            <year>2018</year>
            <article-title>The Maudsley Prescribing Guidelines in Psychiatry</article-title>
            <source>13th Edition</source>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B29">
        <label>29.</label>
        <citation-alternatives>
          <mixed-citation publication-type="confproc">Elbayad, M., Besacier, L. and Verbeek, J. (2018). Pervasive Attention: 2. Proceedings of the 22nd Conference on Computational Natural Language Learning, Brussels, 31 October-1 November 2018, 97-107. https://doi.org/10.18653/v1/k18-1010 <pub-id pub-id-type="doi">10.18653/v1/k18-1010</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.18653/v1/k18-1010">https://doi.org/10.18653/v1/k18-1010</ext-link></mixed-citation>
          <element-citation publication-type="confproc">
            <person-group person-group-type="author">
              <string-name>Elbayad, M.</string-name>
              <string-name>Besacier, L.</string-name>
              <string-name>Verbeek, J.</string-name>
              <string-name>Learning, B</string-name>
            </person-group>
            <year>2018</year>
            <fpage>2</fpage>
            <pub-id pub-id-type="doi">10.18653/v1/k18-1010</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B30">
        <label>30.</label>
        <citation-alternatives>
          <mixed-citation publication-type="confproc">Lea, C., Flynn, M.D., Vidal, R., Reiter, A. and Hager, G.D. (2017) Temporal Convolutional Networks for Action Segmentation and Detection. 2017 <italic>IEEE Conference on Computer Vision and Pattern Recognition</italic> ( <italic>CVPR</italic>), Honolulu, 21-26 July 2017, 1003-1012. https://doi.org/10.1109/cvpr.2017.113 <pub-id pub-id-type="doi">10.1109/cvpr.2017.113</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1109/cvpr.2017.113">https://doi.org/10.1109/cvpr.2017.113</ext-link></mixed-citation>
          <element-citation publication-type="confproc">
            <person-group person-group-type="author">
              <string-name>Lea, C.</string-name>
              <string-name>Flynn, M.D.</string-name>
              <string-name>Vidal, R.</string-name>
              <string-name>Reiter, A.</string-name>
              <string-name>Hager, G.D.</string-name>
            </person-group>
            <year>2017</year>
            <article-title>Temporal Convolutional Networks for Action Segmentation and Detection</article-title>
            <source>2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR)</source>
            <volume>21</volume>
            <pub-id pub-id-type="doi">10.1109/cvpr.2017.113</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B31">
        <label>31.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Liu, J., Zi, L., Shreyas, P., Dustin, T., Tania, B.W. and Balaji, L. (2020) Simple and Principled Uncertainty Estimation with Deterministic Deep Learning via Distance Awareness. <italic>Advances in Neural Information Processing Systems</italic>, 33, 7498-7512.</mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Liu, J.</string-name>
              <string-name>Zi, L.</string-name>
              <string-name>Shreyas, P.</string-name>
              <string-name>Dustin, T.</string-name>
              <string-name>Tania, B.W.</string-name>
              <string-name>Balaji, L.</string-name>
            </person-group>
            <year>2020</year>
            <article-title>Simple and Principled Uncertainty Estimation with Deterministic Deep Learning via Distance Awareness</article-title>
            <source>Advances in Neural Information Processing Systems</source>
            <volume>33</volume>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B32">
        <label>32.</label>
        <citation-alternatives>
          <mixed-citation publication-type="confproc">Li, Y.Z. and Yarin, G. (2017) Dropout Inference in Bayesian Neural Networks with Alpha-Divergences. <italic>International Conference on Machine Learning</italic>, 2052-2061.</mixed-citation>
          <element-citation publication-type="confproc">
            <person-group person-group-type="author">
              <string-name>Li, Y.Z.</string-name>
              <string-name>Yarin, G.</string-name>
            </person-group>
            <year>2017</year>
            <article-title>Dropout Inference in Bayesian Neural Networks with Alpha-Divergences</article-title>
            <source>International Conference on Machine Learning</source>
            <volume>2052</volume>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B33">
        <label>33.</label>
        <citation-alternatives>
          <mixed-citation publication-type="confproc">Niculescu-Mizil, A. and Caruana, R. (2005) Predicting Good Probabilities with Supervised Learning. Proceedings of the 22 <italic>nd International Conference on Machine learning</italic>- <italic>ICML</italic> ‘05, Bonn, 7-11 August 2005, 625-633. https://doi.org/10.1145/1102351.1102430 <pub-id pub-id-type="doi">10.1145/1102351.1102430</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1145/1102351.1102430">https://doi.org/10.1145/1102351.1102430</ext-link></mixed-citation>
          <element-citation publication-type="confproc">
            <person-group person-group-type="author">
              <string-name>Niculescu-Mizil, A.</string-name>
              <string-name>Caruana, R.</string-name>
            </person-group>
            <year>2005</year>
            <article-title>Predicting Good Probabilities with Supervised Learning</article-title>
            <source>Proceedings of the 22nd International Conference on Machine learning-ICML ‘05</source>
            <volume>7</volume>
            <pub-id pub-id-type="doi">10.1145/1102351.1102430</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B34">
        <label>34.</label>
        <citation-alternatives>
          <mixed-citation publication-type="confproc">Russell, C. (2019) Efficient Search for Diverse Coherent Explanations. <italic>Proceedings of the Conference on Fairness</italic>, <italic>Accountability</italic>, <italic>and Transparency</italic>, Atlanta, 29-31 January 2019, 20-28. https://doi.org/10.1145/3287560.3287569 <pub-id pub-id-type="doi">10.1145/3287560.3287569</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1145/3287560.3287569">https://doi.org/10.1145/3287560.3287569</ext-link></mixed-citation>
          <element-citation publication-type="confproc">
            <person-group person-group-type="author">
              <string-name>Russell, C.</string-name>
              <string-name>Fairness, A</string-name>
              <string-name>Transparency, A</string-name>
            </person-group>
            <year>2019</year>
            <article-title>Efficient Search for Diverse Coherent Explanations</article-title>
            <source>Proceedings of the Conference on Fairness</source>
            <volume>29</volume>
            <pub-id pub-id-type="doi">10.1145/3287560.3287569</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B35">
        <label>35.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Schölkopf, B. (2022) Causality for Machine Learning. <italic>Probabilistic and Causal Inference</italic>, 765-804. https://doi.org/10.1145/3501714.3501755 <pub-id pub-id-type="doi">10.1145/3501714.3501755</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1145/3501714.3501755">https://doi.org/10.1145/3501714.3501755</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <year>2022</year>
            <article-title>Causality for Machine Learning</article-title>
            <source>Probabilistic and Causal Inference</source>
            <volume>765</volume>
            <pub-id pub-id-type="doi">10.1145/3501714.3501755</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B36">
        <label>36.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Huang, W.B., Tong, Z., Yu, R. and Huang, J.Z. (2018) Adaptive Sampling towards Fast Graph Representation Learning. <italic>Advances in Neural Information Processing Systems</italic>, 31.</mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Huang, W.B.</string-name>
              <string-name>Tong, Z.</string-name>
              <string-name>Yu, R.</string-name>
              <string-name>Huang, J.Z.</string-name>
            </person-group>
            <year>2018</year>
            <article-title>Adaptive Sampling towards Fast Graph Representation Learning</article-title>
            <source>Advances in Neural Information Processing Systems</source>
            <volume>31</volume>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B37">
        <label>37.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Van Looveren, A. and Klaise, J. (2021) Interpretable Counterfactual Explanations Guided by Prototypes. In: Oliver, N., Pérez-Cruz, F., Kramer, S., Read, J. and Lozano, J.A., Eds., <italic>Lecture</italic><italic>Notes</italic><italic>in</italic><italic>Computer</italic><italic>Science</italic>, Springer International Publishing, 650-665. https://doi.org/10.1007/978-3-030-86520-7_40 <pub-id pub-id-type="doi">10.1007/978-3-030-86520-7_40</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1007/978-3-030-86520-7_40">https://doi.org/10.1007/978-3-030-86520-7_40</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Looveren, A.</string-name>
              <string-name>Klaise, J.</string-name>
              <string-name>Oliver, N.</string-name>
              <string-name>Cruz, F.</string-name>
              <string-name>Kramer, S.</string-name>
              <string-name>Read, J.</string-name>
              <string-name>Lozano, J.A.</string-name>
              <string-name>Science, S</string-name>
            </person-group>
            <year>2021</year>
            <article-title>Interpretable Counterfactual Explanations Guided by Prototypes</article-title>
            <source>In: Oliver</source>
            <volume>650</volume>
            <pub-id pub-id-type="doi">10.1007/978-3-030-86520-7_40</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B38">
        <label>38.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Keane, M.T. and Smyth, B. (2020) Good Counterfactuals and Where to Find Them: A Case-Based Technique for Generating Counterfactuals for Explainable AI (XAI) In: Watson, I. and Weber, R., Eds., <italic>Lecture</italic><italic>Notes</italic><italic>in</italic><italic>Computer</italic><italic>Science</italic>, Springer International Publishing, 163-178. https://doi.org/10.1007/978-3-030-58342-2_11 <pub-id pub-id-type="doi">10.1007/978-3-030-58342-2_11</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1007/978-3-030-58342-2_11">https://doi.org/10.1007/978-3-030-58342-2_11</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Keane, M.T.</string-name>
              <string-name>Smyth, B.</string-name>
              <string-name>Watson, I.</string-name>
              <string-name>Weber, R.</string-name>
              <string-name>Science, S</string-name>
            </person-group>
            <year>2020</year>
            <article-title>Good Counterfactuals and Where to Find Them: A Case-Based Technique for Generating Counterfactuals for Explainable AI (XAI) In: Watson, I</article-title>
            <source>and Weber</source>
            <volume>163</volume>
            <pub-id pub-id-type="doi">10.1007/978-3-030-58342-2_11</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B39">
        <label>39.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Vilone, G. and Longo, L. (2021) Notions of Explainability and Evaluation Approaches for Explainable Artificial Intelligence. <italic>Information Fusion</italic>, 76, 89-106. https://doi.org/10.1016/j.inffus.2021.05.009 <pub-id pub-id-type="doi">10.1016/j.inffus.2021.05.009</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1016/j.inffus.2021.05.009">https://doi.org/10.1016/j.inffus.2021.05.009</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Vilone, G.</string-name>
              <string-name>Longo, L.</string-name>
            </person-group>
            <year>2021</year>
            <article-title>Notions of Explainability and Evaluation Approaches for Explainable Artificial Intelligence</article-title>
            <source>Information Fusion</source>
            <volume>76</volume>
            <pub-id pub-id-type="doi">10.1016/j.inffus.2021.05.009</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B40">
        <label>40.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Amann, J., Blasimme, A., Vayena, E., Frey, D. and Madai, V.I. (2020) Explainability for Artificial Intelligence in Healthcare: A Multidisciplinary Perspective. <italic>BMC</italic><italic>Medical</italic><italic>Informatics</italic><italic>and</italic><italic>Decision</italic><italic>Making</italic>, 20, Article No. 310. https://doi.org/10.1186/s12911-020-01332-6 <pub-id pub-id-type="doi">10.1186/s12911-020-01332-6</pub-id><pub-id pub-id-type="pmid">33256715</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1186/s12911-020-01332-6">https://doi.org/10.1186/s12911-020-01332-6</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Amann, J.</string-name>
              <string-name>Blasimme, A.</string-name>
              <string-name>Vayena, E.</string-name>
              <string-name>Frey, D.</string-name>
              <string-name>Madai, V.I.</string-name>
            </person-group>
            <year>2020</year>
            <article-title>Explainability for Artificial Intelligence in Healthcare: A Multidisciplinary Perspective</article-title>
            <source>BMC Medical Informatics and Decision Making</source>
            <volume>20</volume>
            <elocation-id>No</elocation-id>
            <pub-id pub-id-type="doi">10.1186/s12911-020-01332-6</pub-id>
            <pub-id pub-id-type="pmid">33256715</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B41">
        <label>41.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Ma, Y., Tu, X., Luo, X., Hu, L. and Wang, C. (2025) Machine-Learning-Based Cost Prediction Models for Inpatients with Mental Disorders in China. <italic>BMC Psychiatry</italic>, 25, Article No. 33. https://doi.org/10.1186/s12888-024-06358-y <pub-id pub-id-type="doi">10.1186/s12888-024-06358-y</pub-id><pub-id pub-id-type="pmid">39789477</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1186/s12888-024-06358-y">https://doi.org/10.1186/s12888-024-06358-y</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Ma, Y.</string-name>
              <string-name>Tu, X.</string-name>
              <string-name>Luo, X.</string-name>
              <string-name>Hu, L.</string-name>
              <string-name>Wang, C.</string-name>
            </person-group>
            <year>2025</year>
            <article-title>Machine-Learning-Based Cost Prediction Models for Inpatients with Mental Disorders in China</article-title>
            <source>BMC Psychiatry</source>
            <volume>25</volume>
            <elocation-id>No</elocation-id>
            <pub-id pub-id-type="doi">10.1186/s12888-024-06358-y</pub-id>
            <pub-id pub-id-type="pmid">39789477</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B42">
        <label>42.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Rajpurkar, P., Chen, E., Banerjee, O. and Topol, E.J. (2022) AI in Health and Medicine. <italic>Nature</italic><italic>Medicine</italic>, 28, 31-38. https://doi.org/10.1038/s41591-021-01614-0 <pub-id pub-id-type="doi">10.1038/s41591-021-01614-0</pub-id><pub-id pub-id-type="pmid">35058619</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1038/s41591-021-01614-0">https://doi.org/10.1038/s41591-021-01614-0</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Rajpurkar, P.</string-name>
              <string-name>Chen, E.</string-name>
              <string-name>Banerjee, O.</string-name>
              <string-name>Topol, E.J.</string-name>
            </person-group>
            <year>2022</year>
            <article-title>AI in Health and Medicine</article-title>
            <source>Nature Medicine</source>
            <volume>28</volume>
            <pub-id pub-id-type="doi">10.1038/s41591-021-01614-0</pub-id>
            <pub-id pub-id-type="pmid">35058619</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B43">
        <label>43.</label>
        <citation-alternatives>
          <mixed-citation publication-type="book">World Health Organization (2021) Ethics and Governance of Artificial Intelligence for Health. WHO Press.</mixed-citation>
          <element-citation publication-type="book">
            <year>2021</year>
            <article-title>Ethics and Governance of Artificial Intelligence for Health</article-title>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B44">
        <label>44.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Breiman, L. (2001) Random Forests. <italic>Machine</italic><italic>Learning</italic>, 45, 5-32. https://doi.org/10.1023/a:1010933404324 <pub-id pub-id-type="doi">10.1023/a:1010933404324</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1023/a:1010933404324">https://doi.org/10.1023/a:1010933404324</ext-link></mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Breiman, L.</string-name>
            </person-group>
            <year>2001</year>
            <article-title>Random Forests</article-title>
            <source>Machine Learning</source>
            <volume>45</volume>
            <fpage>101093</fpage>
            <pub-id pub-id-type="doi">10.1023/a:1010933404324</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B45">
        <label>45.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Nalluri, M., Mounika, P. and Nageswara, R.E. (2020) A Scalable Tree Boosting System: XG Boost. <italic>International Journal of Research Studies in Science</italic>, <italic>Engineering and Technology</italic>, 7, 36-51.</mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Nalluri, M.</string-name>
              <string-name>Mounika, P.</string-name>
              <string-name>Nageswara, R.E.</string-name>
              <string-name>Science, E</string-name>
            </person-group>
            <year>2020</year>
            <article-title>A Scalable Tree Boosting System: XG Boost</article-title>
            <source>International Journal of Research Studies in Science</source>
            <volume>7</volume>
          </element-citation>
        </citation-alternatives>
      </ref>
    </ref-list>
  </back>
</article>