<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.4 20241031//EN" "JATS-journalpublishing1-4.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article" dtd-version="1.4" xml:lang="en">
  <front>
    <journal-meta>
      <journal-id journal-id-type="publisher-id">jcc</journal-id>
      <journal-title-group>
        <journal-title>Journal of Computer and Communications</journal-title>
      </journal-title-group>
      <issn pub-type="epub">2327-5227</issn>
      <issn pub-type="ppub">2327-5219</issn>
      <publisher>
        <publisher-name>Scientific Research Publishing</publisher-name>
      </publisher>
    </journal-meta>
    <article-meta>
      <article-id pub-id-type="doi">10.4236/jcc.2026.141002</article-id>
      <article-id pub-id-type="publisher-id">jcc-148925</article-id>
      <article-categories>
        <subj-group>
          <subject>Article</subject>
        </subj-group>
        <subj-group>
          <subject>Computer Science</subject>
          <subject>Communications</subject>
        </subj-group>
      </article-categories>
      <title-group>
        <article-title>Comparative Analysis of ML Models for Survival Prediction of Glioblastoma</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <name name-style="western">
            <surname>Awel</surname>
            <given-names>Muna</given-names>
          </name>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <name name-style="western">
            <surname>Rushit</surname>
            <given-names>Dave</given-names>
          </name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <name name-style="western">
            <surname>Katner</surname>
            <given-names>Samantha J.</given-names>
          </name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <name name-style="western">
            <surname>Bhavsar</surname>
            <given-names>Mansi</given-names>
          </name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
      </contrib-group>
      <aff id="aff1"><label>1</label> Department of Biochemistry, Chemistry, and Geology, Minnesota State University, Mankato, Mankato, MN, USA </aff>
      <aff id="aff2"><label>2</label> Department of Computer Information Science, Minnesota State University, Mankato, Mankato, MN, USA </aff>
      <author-notes>
        <fn fn-type="conflict" id="fn-conflict">
          <p>The authors declare no conflicts of interest regarding the publication of this paper.</p>
        </fn>
      </author-notes>
      <pub-date pub-type="epub">
        <day>09</day>
        <month>01</month>
        <year>2026</year>
      </pub-date>
      <pub-date pub-type="collection">
        <month>01</month>
        <year>2026</year>
      </pub-date>
      <volume>14</volume>
      <issue>01</issue>
      <fpage>20</fpage>
      <lpage>32</lpage>
      <history>
        <date date-type="received">
          <day>16</day>
          <month>12</month>
          <year>2025</year>
        </date>
        <date date-type="accepted">
          <day>16</day>
          <month>01</month>
          <year>2026</year>
        </date>
        <date date-type="published">
          <day>19</day>
          <month>01</month>
          <year>2026</year>
        </date>
      </history>
      <permissions>
        <copyright-statement>© 2026 by the authors and Scientific Research Publishing Inc.</copyright-statement>
        <copyright-year>2026</copyright-year>
        <license license-type="open-access">
          <license-p> This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license ( <ext-link ext-link-type="uri" xlink:href="https://creativecommons.org/licenses/by/4.0/">https://creativecommons.org/licenses/by/4.0/</ext-link> ). </license-p>
        </license>
      </permissions>
      <self-uri content-type="doi" xlink:href="https://doi.org/10.4236/jcc.2026.141002">https://doi.org/10.4236/jcc.2026.141002</self-uri>
      <abstract>
        <p>Glioblastoma multiforme (GBM) remains one of the most aggressive brain malignancies, with a median survival of less than 15 months. This study advances glioblastoma multiforme (GBM) survival prediction by developing a comprehensive machine learning (ML) pipeline that integrates four classifiers: Logistic Regression, Random Forest, XGBoost, and Support Vector Machine (SVM) on TCGA-derived multi-omics datasets. Rigorous preprocessing, including missing data assessment (MCAR test), multicollinearity checks, and feature selection, was followed by hyperparameter optimization using GridSearchCV and 10-fold cross-validation to enhance model performance and generalizability. Predictive performance was evaluated with AUC-ROC, precision-recall curves, and classification reports, while interpretability was assessed through SHAP (SHapley Additive exPlanations) analysis to identify the most influential features driving survival predictions. Random Forest achieved the highest predictive accuracy while maintaining strong interpretability, highlighting key drivers of GBM prognosis such as age, MGMT promoter methylation status, and specific gene expression signatures. Despite promising results that demonstrate ML’s ability to handle GBM heterogeneity, limitations include the relatively modest sample size and lack of external validation. Future work will incorporate independent cohorts for external validation, explore advanced ensemble and hybrid modeling strategies, and further optimize models to meet clinical requirements for both accuracy and transparent decision-making.</p>
      </abstract>
      <kwd-group kwd-group-type="author-generated" xml:lang="en">
        <kwd>Glioblastoma Multiforme</kwd>
        <kwd>Machine Learning</kwd>
        <kwd>Survival Prediction</kwd>
        <kwd>Multi-Omics</kwd>
        <kwd>Methylation Biomarkers</kwd>
        <kwd>GridSearchCV</kwd>
        <kwd>SHAP</kwd>
        <kwd>Interpretability</kwd>
        <kwd>ROC Curve</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec1">
      <title>1. Introduction/Background Study</title>
      <p>Glioblastoma multiforme (GBM) is well known as the most aggressive and prevalent primary malignant brain tumor. GBM accounts for approximately 48.6% of all malignant central nervous system tumors [<xref ref-type="bibr" rid="B1">1</xref>][<xref ref-type="bibr" rid="B2">2</xref>]. Recently, glioblastoma was classified as a high-grade glioma (grade IV) based on the molecular mutation profiles such as IDH mutant/wildtype, and CDKN2A/B homozygous deletion, as well as whether necrosis or microvascular proliferation is observed [<xref ref-type="bibr" rid="B3">3</xref>]. Lower-grade diffuse gliomas do not demonstrate the same degree of biologic aggressiveness, rapid progression, or resistance to current therapies as GBM [<xref ref-type="bibr" rid="B4">4</xref>]. Thus, among all malignant primary brain tumors, GBM remains the most common and lethal subtype and contributes to more than 15,000 deaths annually in the United States [<xref ref-type="bibr" rid="B4">4</xref>]. The median overall survival remains dismal at 14 - 15 months after surgical diagnosis, even with standard temozolomide-radiotherapy regimens [<xref ref-type="bibr" rid="B5">5</xref>].</p>
      <p>Despite aggressive multimodal therapy, long-term outcomes for patients with GBM remain extremely poor, highlighting the limitations of current clinical decision-making guided primarily by histopathological classification and treatment response [<xref ref-type="bibr" rid="B3">3</xref>][<xref ref-type="bibr" rid="B4">4</xref>]. Moreover, GBM exhibits marked inter- and intra-tumor heterogeneity driven by complex genetic and epigenetic alterations that contribute to variable disease evolution and therapeutic resistance [<xref ref-type="bibr" rid="B6">6</xref>]. Recent large-scale network-based analyses of tumor genetics have shown that high-dimensional graph representations can more accurately capture survival signatures than traditional diagnostic categories alone [<xref ref-type="bibr" rid="B6">6</xref>]. Complementary work integrating single-cell RNA sequencing, spatial transcriptomics, and deep learning on whole-slide histology images has further demonstrated that spatial cellular architecture and transcriptional subtype composition are strongly associated with prognosis in GBM, and that these features can be inferred directly from routine histology [<xref ref-type="bibr" rid="B7">7</xref>]. Together, these epidemiologic and molecular insights underscore the urgent need for robust prognostic biomarkers and advanced predictive modeling approaches capable of supporting personalized management in GBM.</p>
      <p>One of the most clinically significant sources of heterogeneity is DNA methylation, which strongly influences tumor progression, therapeutic response, and prognosis [<xref ref-type="bibr" rid="B8">8</xref>]. For instance, mutations in the isocitrate dehydrogenase (IDH)1/2 genes promote accumulation of the oncometabolite 2-hydroxyglutarate, inducing genome-wide hypermethylation and improved clinical outcomes in IDH-mutant GBM [<xref ref-type="bibr" rid="B9">9</xref>]. Another subtype, the methylation of the O6-methylguanine-DNA methyltransferase (MGMT) gene promoter, silences this DNA repair enzyme [<xref ref-type="bibr" rid="B9">9</xref>]. Therefore, methylated MGMT subtype GBM tumors are more to alkylating agents and serve as another favorable prognostic biomarker associated with extended survival. Additionally, the methylated MGMT status has a predictive value with higher cutoffs in IDH-mutant tumors [<xref ref-type="bibr" rid="B10">10</xref>]. TMZ, a standard-of-care oral alkylating agent administered concurrently with radiotherapy and followed by maintenance cycles, significantly improves median survival compared to radiotherapy alone, particularly in MGMT-methylated patients where it enhances treatment efficacy by impairing DNA repair mechanisms [<xref ref-type="bibr" rid="B10">10</xref>]. Thus, understanding the effects of different therapies is important to find the best mechanism that is beneficial to the patient.</p>
      <p>With advancement of technologies in machine learning (ML) and deep learning newer mechanisms have developed integrating multi-omics data, radiomics, and clinical variables. Ensemble methods, such as gradient-boosted trees and deep neural networks, have improved predictive performance by addressing data heterogeneity and enhancing interpretability via techniques like SHAP, tackling limitations in traditional models [<xref ref-type="bibr" rid="B11">11</xref>]. However, many existing GBM survival prediction studies primarily emphasize predictive accuracy, often relying on complex model architectures, while providing limited attention to interpretability, reproducibility, and consistent benchmarking across models. In addition, prior studies frequently evaluate models using heterogeneous datasets or preprocessing strategies, making direct comparison and clinical translation challenging. This research builds on these advancements by developing a comprehensive ML pipeline using logistic regression, random forest, SVM, and XGBoost on TCGA-derived multi-omics datasets, incorporating SMOTE for class imbalance and GridSearchCV for optimization, to predict binary survival outcomes.</p>
    </sec>
    <sec id="sec2">
      <title>2. Methodology</title>
      <p>This section explores the comprehensive methodology employed within the study, covering topics on data acquisition, preprocessing, feature engineering, and model development. All analyses were conducted using Python (version 3.12) with libraries including scikit-learn (for logistic regression and imputation), XGBoost (for gradient boosting), SHAP (model interpretation). The methodology adheres to best practices in biomedical data science, prioritizing reproducibility, statistical rigor, and ethical considerations in handling sensitive health data.</p>
      <fig id="fig1">
        <label>Figure 1</label>
        <graphic xlink:href="https://html.scirp.org/file/1733421-rId15.jpeg?20260119032447" />
      </fig>
      <p><bold>Figure 1</bold><bold>.</bold> System architecture of proposed methodology highlights the sequential flow from data input to final model assessment, ensuring clarity in how each stage contributes to the overall system.</p>
      <p>To visualize the workflow, a system architecture diagram (<xref ref-type="fig" rid="fig1">Figure 1</xref>) illustrates the sequential steps of the data pipeline, from raw data ingestion to model interpretation. This architecture ensures a modular, scalable design that can be adapted for similar omics-based studies.</p>
      <sec id="sec2dot1">
        <title>2.1. Data Collection</title>
        <p>The datasets utilized within this study are sourced from The Cancer Genome Atlas (TCGA) via cBioPortal, specifically the Glioblastoma Multiforme (TCGA, Cell 2013 and Firehose Legacy) datasets, comprising 577 samples with multi-omics data. EDA was performed to characterize feature distributions, identify data types, and detect anomalies. </p>
      </sec>
      <sec id="sec2dot2">
        <title>2.2. Data Preprocessing</title>
        <p>The data preprocessing step is the critical step in data analysis to ensure the data quality, integrity and prepare the data inputs for modeling. In fact, this involved several sub-steps executed in a pipeline using scikit-learn’s Pipeline class for reproducibility.</p>
      </sec>
      <sec id="sec2dot3">
        <title>2.3. Missing Values</title>
        <p>Missing values were first quantified to identify the correct mechanism to handle the null values. Features that had missing values 50% or more were dropped to reduce noise and bias for the models. To characterize the missing data mechanism, a global Little’s MCAR test was conducted rejecting the null hypothesis of complete randomness that missingness occurred completely at random [<xref ref-type="bibr" rid="B12">12</xref>]. To further explore dependencies, pairwise independence tests were performed between missingness indicators for all unique feature pairs. For categorical variables <italic>χ</italic><sup>2</sup> tests of independence between the categorical value and the missingness indicators of other features. On the other hand, for continuous variables Welch’s two-sample t-tests were performed comparing means across the missing/not-missing groups of other features as shown in <bold>Table 1</bold> [<xref ref-type="bibr" rid="B13">13</xref>]. </p>
        <p><bold>Table 1</bold><bold>.</bold> t-test results for continuous variables vs. missingness indicators show significant dependencies (P &lt; 0.05). The t-test results assess whether continuous clinical variables differ based on missingness in key molecular markers. The significant P-values indicate that missingness is not completely random and may reflect underlying patient or biological characteristics.</p>
        <table-wrap id="tbl1">
          <label>Table 1</label>
          <table>
            <tbody>
              <tr>
                <td>
                  <bold>Continuous Variable</bold>
                </td>
                <td>
                  <bold>Missingness Variable</bold>
                </td>
                <td>
                  <bold>t-statistics</bold>
                </td>
                <td>
                  <bold>P-value</bold>
                </td>
              </tr>
              <tr>
                <td>Diagnosis Age</td>
                <td>IDH1 Mutation</td>
                <td>−2.0759</td>
                <td>0.0394</td>
              </tr>
              <tr>
                <td>Overall Survival (Months)</td>
                <td>Methylation Status</td>
                <td>2.1364</td>
                <td>0.0333</td>
              </tr>
              <tr>
                <td>Overall Survival (Months)</td>
                <td>MGMT Status</td>
                <td>2.1831</td>
                <td>0.0296</td>
              </tr>
            </tbody>
          </table>
        </table-wrap>
        <p><bold>Table 1</bold> shows the results of two sample t-tests comparing the means of continuous variables against the missingness of other variables, where significant differences (P &lt; 0.05) were observed. The t-statistic indicates the magnitude and direction of the mean difference, with negative values suggesting a higher mean in the “missing” group, and positive values indicating a higher mean in the “not missing” group. </p>
        <p>Overall based on the analysis of handling missing values the non-MCAR, likely MAR mechanism and clinical relevance, imputation was preferred. K-Nearest Neighbors (KNN) imputation was chosen to leverage multivariate relationships, suitable for mixed omics data [<xref ref-type="bibr" rid="B14">14</xref>]. This mechanism was selected over model-based approaches such as Multiple Imputation by Chained Equations (MICE) due to its non-parametric nature and its ability to preserve local multivariate structure without imposing distributional assumptions. This property is particularly advantageous for high-dimensional multi-omics datasets with mixed feature types, where nonlinear relationships and complex dependencies are common. Prior studies have demonstrated that KNN imputation performs competitively for genomic and epigenomic data under Missing at Random (MAR) mechanisms while maintaining computational efficiency and stability [<xref ref-type="bibr" rid="B15">15</xref>].</p>
      </sec>
      <sec id="sec2dot4">
        <title>2.4. Handling Categorical Features</title>
        <p>To enable numerical analysis and avoid introducing ordinal assumptions, one-hot encoding was applied to multi-class categorical variables using scikit-learn’s OneHotEncoder. This process transformed each category into a binary column, creating a sparse matrix of dummy variables [<xref ref-type="bibr" rid="B16">16</xref>]. </p>
      </sec>
      <sec id="sec2dot5">
        <title>2.5. Multicollinearity Assessment</title>
        <p>Variation Inflation Factor (VIF) was calculated for all features. VIF values of greater than 10 were dropped to mitigate high multicollinearity, which can destabilize model coefficients [<xref ref-type="bibr" rid="B17">17</xref>]. L2 regularization method (Ridge regression) was applied, penalizing large coefficients to stabilize model estimates [<xref ref-type="bibr" rid="B18">18</xref>]. This approach was implemented on the numeric dataset, including one-hot encoded variables, reducing the impact of correlated features.</p>
      </sec>
      <sec id="sec2dot6">
        <title>2.6. Model Development</title>
        <p>Prior to model training, Overall Survival (Months) originally recorded as a continuous variable was binarized using the cohort median survival time as the threshold, with patients surviving longer than or equal to the median labeled as long-term survivors (Class 1) and those below the median labeled as short-term survivors (Class 0). This binarization supports reproducible classification modeling and clinically meaningful risk stratification.</p>
        <p>Logistic regression was the benchmark model for its high interpretability. Additional models like SVM, RF and XG BOOST are implemented for comparison to find a most suitable model. Regularization (L2) is applied to reduce multicollinearity and enhance model performance, while features are normalized using StandardScaler to ensure compatibility with scale-sensitive models [<xref ref-type="bibr" rid="B18">18</xref>]. Hyperparameter tuning is conducted using GridSearchCV to optimize model parameters, and 10-fold cross-validation with stratified sampling is employed to handle censored data and improve robustness [<xref ref-type="bibr" rid="B19">19</xref>]. This multi-model strategy, combined with preprocessing and tuning, aims to address gaps in interpretability and generalizability from prior studies.</p>
      </sec>
      <sec id="sec2dot7">
        <title>2.7. Model Evaluation</title>
        <p>This section evaluates model performance for binary GBM survival prediction. To address class imbalance, SMOTE was applied within training folds only to balance minority and majority outcomes. Performance is summarized by the ROC-AUC, reflecting discrimination across thresholds, and a classification report (accuracy, precision, recall, and F1) reported both per class and as macro-averages to capture minority-class behavior. Where relevant, PR-AUC complements ROC-AUC under imbalance. Estimates are obtained via stratified 10-fold cross-validation, preserving class proportions and yielding optimized, generalizable metrics.</p>
      </sec>
    </sec>
    <sec id="sec3">
      <title>3. Results</title>
      <p>This section explores the results of each model and model interpretability of the features that were highly significant to the model’s predictions.</p>
      <fig id="fig2">
        <label>Figure 2</label>
        <graphic xlink:href="https://html.scirp.org/file/1733421-rId16.jpeg?20260119032449" />
      </fig>
      <p><bold>Figure 2</bold><bold>.</bold> A comparison of model performance across multiple evaluation metrics. This model presents CV ROC-AUC, Test ROC-AUC, Accuracy, Precision, Recall, and F1 scores for four classification models: Random Forest, XGBoost, Logistic Regression, and SVM.</p>
      <p><xref ref-type="fig" rid="fig2">Figure 2</xref> compares the performance of four machine learning models Random Forest, XGBoost, Logistic Regression, and SVM across key evaluation metrics, including cross-validated ROC-AUC, test ROC-AUC, accuracy, and the average precision, recall, and F1-scores. Overall, Random Forest achieves the strongest performance, with the highest accuracy, precision, and F1-score, as well as competitive ROC-AUC values. XGBoost shows slightly lower accuracy than Random Forest but maintains comparable ROC-AUC and balanced average precision, recall, and F1. Overall, the chart highlights the trade-offs between models, showing how each one emphasizes different aspects of classification performance depending on whether accuracy, balance, or minority class detection is prioritized.</p>
      <fig id="fig3">
        <label>Figure 3</label>
        <graphic xlink:href="https://html.scirp.org/file/1733421-rId17.jpeg?20260119032449" />
      </fig>
      <p><bold>Figure 3</bold><bold>.</bold> ROC curves for GBM survival classifiers. (A-B) ROC curves for Logistic Regression and SVM models trained with SMOTE to address class imbalance. (C-D) ROC curves for Random Forest and XGBoost models, which manage imbalance internally through their ensemble structures.</p>
      <p>As shown in <xref ref-type="fig" rid="fig3">Figure 3</xref>, XGBoost achieved the highest AUC (0.81), closely followed by Random Forest (0.80), while Logistic Regression + SMOTE and SVM (RBF) + SMOTE both reached 0.72. Curves lie well above the diagonal, indicating discrimination better than chance. </p>
    </sec>
    <sec id="sec4">
      <title>4. Discussion</title>
      <p>This study evaluated machine-learning approaches for binary GBM survival prediction using TCGA data. After rigorous preprocessing XGBoost achieved the highest ROC-AUC (~0.81), closely followed by Random Forest (~0.80), while SVM (RBF) and Logistic Regression reached ~0.72. These results suggest that models capturing nonlinear interactions and higher-order effects better reflect the heterogeneity of GBM.</p>
      <p>Sensitivity (Recall)-Specificity Trade-off; in the aspect balancing between recall and specificity, minimizing false negative is important. Recall is true positive rate TP/(TP + FN) and increases as the threshold lowers, while specifically true negative rate for low-risk patients TN/(TN + FP) typically decrease. Thus, Random Forest supports a sensitivity-first strategy (very high Class-1 recall with acceptable precision) at the cost of more false positives (lower Class-0 recall/specificity), whereas XGBoost offers a more conservative balance (higher specificity at the same target sensitivity, with modestly lower recall).</p>
      <fig id="fig4">
        <label>Figure 4</label>
        <graphic xlink:href="https://html.scirp.org/file/1733421-rId18.jpeg?20260119032450" />
      </fig>
      <p><bold>Figure 4</bold><bold>.</bold> Global feature importance based on SHAP values. This bar chart summarizes the global contribution of each feature to the model’s predictions using the mean absolute SHAP value (mean |SHAP|) across all samples. Features are ranked from top to bottom according to their overall impact on the model, with larger values indicating greater influence on prediction outcomes. It provides an aggregated view of feature importance across the cohort. Additionally, this representation facilitates straightforward comparison of feature importance and highlights the dominant clinical and molecular drivers underlying the model’s prognostic decisions.</p>
      <p>Model interpretation/Genomic Analysis: SHAP indicated that clinical/treatment variables and genome-level summaries contribute meaningfully to discrimination. Survival-proximal covariates overall survival months and disease-free duration, and therapy indicators exert the largest effects, with age and fraction genome altered (FGA) also contributing meaningfully. As shown in <xref ref-type="fig" rid="fig4">Figure 4</xref>, the therapy indicators show the treatment course for GBM patients. The feature “therapy_TMZ Chemoradiation, TMZ Chemo” flags patients who received the modern standard Stupp regimen: concurrent external-beam radiation with <bold>temozolomide (TMZ)</bold> followed by adjuvant TMZ [<xref ref-type="bibr" rid="B20">20</xref>]. The “therapy_Standard Radiation, TMZ” denotes conventional-fractionation radiation co-administered with TMZ but recorded under a separate label; functionally this still reflects radiotherapy + TMZ [<xref ref-type="bibr" rid="B21">21</xref>]. The “therapy Nonstandard Radiation, TMZ Chemo” captures atypical radiotherapy schedules given with TMZ often used in older or frailer patients [<xref ref-type="bibr" rid="B22">22</xref>]. In the SHAP bar plot, therapy-related indicators exhibit high mean absolute SHAP values, indicating a strong global contribution to the model’s predictive performance. Their prominence aligns with the known prognostic relevance of TMZ-based chemoradiation in many GBM cases [<xref ref-type="bibr" rid="B20">20</xref>]. In terms of DNA methylation, although MGMT promoter methylation and G-CIMP are well-established prognostic biomarkers, they show lower standalone SHAP importance in our model. This does not contradict their clinical relevance. Instead, it reflects on their redundancy with global methylation clusters (e.g., CL_3/CL_4/CL_6) that encode genome-wide CpG patterns and therefore absorb much of the MGMT/G-CIMP signal.</p>
      <p>Clinical relevance: Although purely predictive, the models highlight variables aligned with current practice (e.g., age, therapy patterns) and may inform risk stratification for follow-up intensity or trial eligibility. Tree-based SHAP explanations can support clinician review by revealing case-level drivers of risk. For example, at the post-surgical evaluation stage, a patient receiving standard-of-care therapy but predicted by the model to be high risk based on combined molecular and clinical features could be considered for closer surveillance, earlier follow-up imaging, or prioritization for clinical trial enrollment. Tree-based SHAP explanations further support clinician review by revealing patient-specific drivers of risk, enabling predictions to be interpreted in the context of known biological and treatment-related factors rather than used as opaque risk scores.</p>
    </sec>
    <sec id="sec5">
      <title>5. Limitations</title>
      <p>This study on glioblastoma multiforme (GBM) survival prediction using machine learning models presents several limitations that influence both the generalizability and clinical applicability of the findings. First, the dataset is relatively small, consisting of 577 multi-omics samples from TCGA. Although this cohort is well-curated, the limited sample size reduces statistical power and may not capture the full biological and clinical heterogeneity of GBM. Moreover, the study lacks external validation such as evaluation on CGGA or other independent datasets which restrict confidence in the model’s generalizability across populations and sequencing platforms.</p>
      <p>Second, certain preprocessing steps may have unintentionally weakened biologically relevant signals. For example, methylation preprocessing likely absorbed much of the G-CIMP and IDH information, leading these features to appear less significant in the SHAP interpretability analysis. This suggests that feature scaling and transformation should be handled cautiously to preserve biologically meaningful variation. Additionally, the study relied on binary survival outcomes, which oversimplified time-to-event data and failed to account for censoring.</p>
      <p>Third, while SMOTE was used to address class imbalance, synthetic oversampling may not accurately represent the intricate molecular and biological relationships of key biomarkers like IDH and MGMT. The resulting synthetic samples may distort data distribution and introduce artifacts, explaining the model’s relatively weaker performance on the minority class. Future work should explore more advanced imbalance techniques, such as class-weighted learning or focal loss, to better handle uneven survival classes.</p>
      <p>Interpretability also poses an important limitation. Although SHAP values were employed to improve transparency, models such as XGBoost and SVM remain largely “black-box”, and SHAP explanations may not fully bridge the gap between algorithmic predictions and clinical trust. Clinical decision-making requires transparent and stable explanations, and further work integrating causally grounded or rule-based interpretability frameworks could strengthen reliability and physician confidence. </p>
      <p>Another limitation lies in the scope of data modalities. Relying solely on multi-omics features overlooks other crucial determinants of GBM survival, such as radiographic imaging, surgical resection extent, treatment regimens, and patient performance status. Incorporating multimodal data including radiomics and clinical variables could yield a more comprehensive understanding of GBM heterogeneity and enhance model robustness. Overall, these limitations underscore the need for larger, externally validated, and multimodal datasets, more biologically informed preprocessing, survival-specific modeling strategies, and enhanced interpretability methods. Addressing these aspects in future research will be vital to improve transparency, and clinical adoption of machine-learning-based survival prediction models for GBM.</p>
    </sec>
    <sec id="sec6">
      <title>6. Conclusions</title>
      <p>The central takeaway of this work is that predictive accuracy alone is insufficient for clinical relevance. By integrating SHAP-based explanations, the proposed framework links model predictions to biologically and clinically established factors such as IDH mutation and MGMT methylation, enabling model outputs to be reviewed and contextualized by clinicians rather than treated as opaque risk scores. This alignment between predictive modeling and domain knowledge is critical for building trust and supporting clinician-in-the-loop decision-making.</p>
      <p>From a broader perspective, this work highlights how carefully benchmarked, interpretable ML pipelines can bridge the gap between computational performance and practical utility in neuro-oncology. Rather than advocating for fully automated decision-making, the framework illustrates a pathway for ML models to function as auxiliary decision-support tools, informing follow-up intensity, patient stratification, and research trial eligibility alongside standard clinical assessment.</p>
      <p>Overall, these findings underscore the importance of balancing predictive power with interpretability in GBM survival modeling. By prioritizing transparency, reproducibility, and clinical alignment, this study contributes to the development of machine learning approaches that are not only technically robust but also positioned for meaningful integration into future precision oncology workflows.</p>
    </sec>
    <sec id="sec7">
      <title>7. Future Work</title>
      <p>Future work will focus on strengthening the generalizability, interpretability, and clinical readiness of the proposed machine learning framework for GBM survival prediction. One of the most immediate steps is to incorporate external validation datasets and other multi-center cohorts, to evaluate model performance across diverse populations and sequencing platforms. This will help assess the reproducibility of predictive performance and mitigate potential dataset biases introduced by single-source training data. External validation will also enable testing of the pipeline under different demographic, genetic, and clinical distributions, ensuring that the model maintains its predictive reliability beyond the TCGA cohort.</p>
      <p>Future research will also explore deep learning architectures to better model the complex, nonlinear relationships within high-dimensional omics data. Incorporating Convolutional Neural Networks (CNNs) and autoencoders can enable automatic feature extraction from multi-omics inputs, potentially improving representation learning and model scalability. Additionally, hybrid frameworks combining traditional ML algorithms with deep survival models will allow the incorporation of censored time-to-event data, providing more clinically relevant survival probabilities instead of binary outcomes.</p>
      <p>Expanding the dataset beyond methylation and expression profiles is another important step. Integrating multi-modal data sources including radiomics, histopathological imaging, treatment histories, and demographic or clinical factors can improve the model’s ability to capture the full heterogeneity of GBM. Such multimodal fusion can be achieved using late or intermediate fusion strategies, where features from different data types are merged to produce comprehensive survival predictions. This expansion will not only enhance model accuracy but also strengthen the biological interpretability of predictions.</p>
      <p>Finally, emphasis will be placed on improving model deployment and real-world usability. Future iterations will include model calibration and uncertainty estimation through bootstrap confidence intervals and conformal prediction intervals to quantify prediction reliability. Building a lightweight deployment pipeline such as a web-based clinical interface or containerized application can enable real-time prediction and visualization within clinical workflows, particularly in resource-constrained environments where computational and data access limitations are prevalent. Ensuring reproducibility through open-source code, environmental documentation, and model cards will further promote transparency and facilitate collaboration among researchers and clinicians.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <title>References</title>
      <ref id="B1">
        <label>1.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Qi, D., Li, J., Quarles, C.C., Fonkem, E. and Wu, E. (2022) Assessment and Prediction of Glioblastoma Therapy Response: Challenges and Opportunities. <italic>Brain</italic>, 146, 1281-1298. https://doi.org/10.1093/brain/awac450 <pub-id pub-id-type="doi">10.1093/brain/awac450</pub-id><pub-id pub-id-type="pmid">36445396</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1093/brain/awac450">https://doi.org/10.1093/brain/awac450</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Qi, D.</string-name>
              <string-name>Li, J.</string-name>
              <string-name>Quarles, C.C.</string-name>
              <string-name>Fonkem, E.</string-name>
              <string-name>Wu, E.</string-name>
            </person-group>
            <year>2022</year>
            <article-title>Assessment and Prediction of Glioblastoma Therapy Response: Challenges and Opportunities</article-title>
            <source>Brain</source>
            <volume>146</volume>
            <pub-id pub-id-type="doi">10.1093/brain/awac450</pub-id>
            <pub-id pub-id-type="pmid">36445396</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B2">
        <label>2.</label>
        <citation-alternatives>
          <mixed-citation publication-type="report">Ostrom, Q.T., Price, M., Neff, C., Cioffi, G., Waite, K.A., Kruchko, C., <italic>et al.</italic> (2022) CBTRUS Statistical Report: Primary Brain and Other Central Nervous System Tumors Diagnosed in the United States in 2015-2019. <italic>Neuro</italic>- <italic>Oncology</italic>, 24, v1-v95. https://doi.org/10.1093/neuonc/noac202 <pub-id pub-id-type="doi">10.1093/neuonc/noac202</pub-id><pub-id pub-id-type="pmid">36196752</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1093/neuonc/noac202">https://doi.org/10.1093/neuonc/noac202</ext-link></mixed-citation>
          <element-citation publication-type="report">
            <person-group person-group-type="author">
              <string-name>Ostrom, Q.T.</string-name>
              <string-name>Price, M.</string-name>
              <string-name>Neff, C.</string-name>
              <string-name>Cioffi, G.</string-name>
              <string-name>Waite, K.A.</string-name>
              <string-name>Kruchko, C.</string-name>
            </person-group>
            <year>2022</year>
            <article-title>CBTRUS Statistical Report: Primary Brain and Other Central Nervous System Tumors Diagnosed in the United States in 2015-2019</article-title>
            <source>Neuro-Oncology</source>
            <volume>24</volume>
            <pub-id pub-id-type="doi">10.1093/neuonc/noac202</pub-id>
            <pub-id pub-id-type="pmid">36196752</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B3">
        <label>3.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Louis, D.N., Perry, A., Wesseling, P., Brat, D.J., Cree, I.A., Figarella-Branger, D., <italic>et al.</italic> (2021) The 2021 WHO Classification of Tumors of the Central Nervous System: A Summary. <italic>Neuro-Oncology</italic>, 23, 1231-1251. https://doi.org/10.1093/neuonc/noab106 <pub-id pub-id-type="doi">10.1093/neuonc/noab106</pub-id><pub-id pub-id-type="pmid">34185076</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1093/neuonc/noab106">https://doi.org/10.1093/neuonc/noab106</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Louis, D.N.</string-name>
              <string-name>Perry, A.</string-name>
              <string-name>Wesseling, P.</string-name>
              <string-name>Brat, D.J.</string-name>
              <string-name>Cree, I.A.</string-name>
              <string-name>Figarella-Branger, D.</string-name>
            </person-group>
            <year>2021</year>
            <article-title>The 2021 WHO Classification of Tumors of the Central Nervous System: A Summary</article-title>
            <source>Neuro-Oncology</source>
            <volume>23</volume>
            <pub-id pub-id-type="doi">10.1093/neuonc/noab106</pub-id>
            <pub-id pub-id-type="pmid">34185076</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B4">
        <label>4.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Schaff, L.R. and Mellinghoff, I.K. (2023) Glioblastoma and Other Primary Brain Malignancies in Adults. <italic>JAMA</italic>, 329, 574-587. https://doi.org/10.1001/jama.2023.0023 <pub-id pub-id-type="doi">10.1001/jama.2023.0023</pub-id><pub-id pub-id-type="pmid">36809318</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1001/jama.2023.0023">https://doi.org/10.1001/jama.2023.0023</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Schaff, L.R.</string-name>
              <string-name>Mellinghoff, I.K.</string-name>
            </person-group>
            <year>2023</year>
            <article-title>Glioblastoma and Other Primary Brain Malignancies in Adults</article-title>
            <source>JAMA</source>
            <volume>329</volume>
            <pub-id pub-id-type="doi">10.1001/jama.2023.0023</pub-id>
            <pub-id pub-id-type="pmid">36809318</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B5">
        <label>5.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Grochans, S., Cybulska, A.M., Simińska, D., Korbecki, J., Kojder, K., Chlubek, D., <italic>et al.</italic> (2022) Epidemiology of Glioblastoma Multiforme-Literature Review. <italic>Cancers</italic>, 14, Article 2412. https://doi.org/10.3390/cancers14102412 <pub-id pub-id-type="doi">10.3390/cancers14102412</pub-id><pub-id pub-id-type="pmid">35626018</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.3390/cancers14102412">https://doi.org/10.3390/cancers14102412</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Grochans, S.</string-name>
              <string-name>Cybulska, A.M.</string-name>
              <string-name>Korbecki, J.</string-name>
              <string-name>Kojder, K.</string-name>
              <string-name>Chlubek, D.</string-name>
            </person-group>
            <year>2022</year>
            <article-title>Epidemiology of Glioblastoma Multiforme-Literature Review</article-title>
            <source>Cancers</source>
            <volume>14</volume>
            <elocation-id>2412</elocation-id>
            <pub-id pub-id-type="doi">10.3390/cancers14102412</pub-id>
            <pub-id pub-id-type="pmid">35626018</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B6">
        <label>6.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Ruffle, J.K., Mohinta, S., Pombo, G., Gray, R., Kopanitsa, V., Lee, F., <italic>et al.</italic> (2023) Brain Tumour Genetic Network Signatures of Survival. <italic>Brain</italic>, 146, 4736-4754. https://doi.org/10.1093/brain/awad199 <pub-id pub-id-type="doi">10.1093/brain/awad199</pub-id><pub-id pub-id-type="pmid">37665980</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1093/brain/awad199">https://doi.org/10.1093/brain/awad199</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Ruffle, J.K.</string-name>
              <string-name>Mohinta, S.</string-name>
              <string-name>Pombo, G.</string-name>
              <string-name>Gray, R.</string-name>
              <string-name>Kopanitsa, V.</string-name>
              <string-name>Lee, F.</string-name>
            </person-group>
            <year>2023</year>
            <article-title>Brain Tumour Genetic Network Signatures of Survival</article-title>
            <source>Brain</source>
            <volume>146</volume>
            <pub-id pub-id-type="doi">10.1093/brain/awad199</pub-id>
            <pub-id pub-id-type="pmid">37665980</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B7">
        <label>7.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Zheng, Y., Carrillo-Perez, F., Pizurica, M., Heiland, D.H. and Gevaert, O. (2023) Spatial Cellular Architecture Predicts Prognosis in Glioblastoma. <italic>Nature</italic><italic>Communications</italic>, 14, Article No. 4122. https://doi.org/10.1038/s41467-023-39933-0 <pub-id pub-id-type="doi">10.1038/s41467-023-39933-0</pub-id><pub-id pub-id-type="pmid">37433817</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1038/s41467-023-39933-0">https://doi.org/10.1038/s41467-023-39933-0</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Zheng, Y.</string-name>
              <string-name>Carrillo-Perez, F.</string-name>
              <string-name>Pizurica, M.</string-name>
              <string-name>Heiland, D.H.</string-name>
              <string-name>Gevaert, O.</string-name>
            </person-group>
            <year>2023</year>
            <article-title>Spatial Cellular Architecture Predicts Prognosis in Glioblastoma</article-title>
            <source>Nature Communications</source>
            <volume>14</volume>
            <elocation-id>No</elocation-id>
            <pub-id pub-id-type="doi">10.1038/s41467-023-39933-0</pub-id>
            <pub-id pub-id-type="pmid">37433817</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B8">
        <label>8.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Pouyan, A., et al. (2025) Glioblastoma Multiforme: Insights into Pathogenesis, Key Signaling Pathways, and Therapeutic Strategies. <italic>Molecular Cancer</italic>, 24, Article No. 58.</mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Pouyan, A.</string-name>
              <string-name>Pathogenesis, K</string-name>
            </person-group>
            <year>2025</year>
            <article-title>Glioblastoma Multiforme: Insights into Pathogenesis, Key Signaling Pathways, and Therapeutic Strategies</article-title>
            <source>Molecular Cancer</source>
            <volume>24</volume>
            <elocation-id>No</elocation-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B9">
        <label>9.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Noushmehr, H., Weisenberger, D.J., Diefes, K., Phillips, H.S., Pujara, K., Berman, B.P., <italic>et al.</italic> (2010) Identification of a CPG Island Methylator Phenotype That Defines a Distinct Subgroup of Glioma. <italic>Cancer</italic><italic>Cell</italic>, 17, 510-522. https://doi.org/10.1016/j.ccr.2010.03.017 <pub-id pub-id-type="doi">10.1016/j.ccr.2010.03.017</pub-id><pub-id pub-id-type="pmid">20399149</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1016/j.ccr.2010.03.017">https://doi.org/10.1016/j.ccr.2010.03.017</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Noushmehr, H.</string-name>
              <string-name>Weisenberger, D.J.</string-name>
              <string-name>Diefes, K.</string-name>
              <string-name>Phillips, H.S.</string-name>
              <string-name>Pujara, K.</string-name>
              <string-name>Berman, B.P.</string-name>
            </person-group>
            <year>2010</year>
            <article-title>Identification of a CPG Island Methylator Phenotype That Defines a Distinct Subgroup of Glioma</article-title>
            <source>Cancer Cell</source>
            <volume>17</volume>
            <pub-id pub-id-type="doi">10.1016/j.ccr.2010.03.017</pub-id>
            <pub-id pub-id-type="pmid">20399149</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B10">
        <label>10.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Drexler, R., Khatri, R., Schüller, U., Eckhardt, A., Ryba, A., Sauvigny, T., <italic>et al.</italic> (2024) Temporal Change of DNA Methylation Subclasses between Matched Newly Diagnosed and Recurrent Glioblastoma. <italic>Acta</italic><italic>Neuropathologica</italic>, 147, Article No. 21. https://doi.org/10.1007/s00401-023-02677-8 <pub-id pub-id-type="doi">10.1007/s00401-023-02677-8</pub-id><pub-id pub-id-type="pmid">38244080</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1007/s00401-023-02677-8">https://doi.org/10.1007/s00401-023-02677-8</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Drexler, R.</string-name>
              <string-name>Khatri, R.</string-name>
              <string-name>Eckhardt, A.</string-name>
              <string-name>Ryba, A.</string-name>
              <string-name>Sauvigny, T.</string-name>
            </person-group>
            <year>2024</year>
            <article-title>Temporal Change of DNA Methylation Subclasses between Matched Newly Diagnosed and Recurrent Glioblastoma</article-title>
            <source>Acta Neuropathologica</source>
            <volume>147</volume>
            <elocation-id>No</elocation-id>
            <pub-id pub-id-type="doi">10.1007/s00401-023-02677-8</pub-id>
            <pub-id pub-id-type="pmid">38244080</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B11">
        <label>11.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Babaei Rikan, S., Sorayaie Azar, A., Naemi, A., Bagherzadeh Mohasefi, J., Pirnejad, H. and Wiil, U.K. (2024) Survival Prediction of Glioblastoma Patients Using Modern Deep Learning and Machine Learning Techniques. <italic>Scientific</italic><italic>Reports</italic>, 14, Article No. 2371. https://doi.org/10.1038/s41598-024-53006-2 <pub-id pub-id-type="doi">10.1038/s41598-024-53006-2</pub-id><pub-id pub-id-type="pmid">38287149</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1038/s41598-024-53006-2">https://doi.org/10.1038/s41598-024-53006-2</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Rikan, S.</string-name>
              <string-name>Azar, A.</string-name>
              <string-name>Naemi, A.</string-name>
              <string-name>Mohasefi, J.</string-name>
              <string-name>Pirnejad, H.</string-name>
              <string-name>Wiil, U.K.</string-name>
            </person-group>
            <year>2024</year>
            <article-title>Survival Prediction of Glioblastoma Patients Using Modern Deep Learning and Machine Learning Techniques</article-title>
            <source>Scientific Reports</source>
            <volume>14</volume>
            <elocation-id>No</elocation-id>
            <pub-id pub-id-type="doi">10.1038/s41598-024-53006-2</pub-id>
            <pub-id pub-id-type="pmid">38287149</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B12">
        <label>12.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Little, R.J.A. (1988) A Test of Missing Completely at Random for Multivariate Data with Missing Values. <italic>Journal</italic><italic>of</italic><italic>the</italic><italic>American</italic><italic>Statistical</italic><italic>Association</italic>, 83, 1198-1202. https://doi.org/10.1080/01621459.1988.10478722 <pub-id pub-id-type="doi">10.1080/01621459.1988.10478722</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1080/01621459.1988.10478722">https://doi.org/10.1080/01621459.1988.10478722</ext-link></mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Little, R.J.A.</string-name>
            </person-group>
            <year>1988</year>
            <article-title>A Test of Missing Completely at Random for Multivariate Data with Missing Values</article-title>
            <source>Journal of the American Statistical Association</source>
            <volume>83</volume>
            <pub-id pub-id-type="doi">10.1080/01621459.1988.10478722</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B13">
        <label>13.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Keselman, H.J., Othman, A.R., Wilcox, R.R. and Fradette, K. (2004) The New and Improved Two-Sample <italic>t</italic> Test. <italic>Psychological</italic><italic>Science</italic>, 15, 47-51. https://doi.org/10.1111/j.0963-7214.2004.01501008.x <pub-id pub-id-type="doi">10.1111/j.0963-7214.2004.01501008.x</pub-id><pub-id pub-id-type="pmid">14717831</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1111/j.0963-7214.2004.01501008.x">https://doi.org/10.1111/j.0963-7214.2004.01501008.x</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Keselman, H.J.</string-name>
              <string-name>Othman, A.R.</string-name>
              <string-name>Wilcox, R.R.</string-name>
              <string-name>Fradette, K.</string-name>
            </person-group>
            <year>2004</year>
            <article-title>The New and Improved Two-Sample t Test</article-title>
            <source>Psychological Science</source>
            <volume>15</volume>
            <pub-id pub-id-type="doi">10.1111/j.0963-7214.2004.01501008.x</pub-id>
            <pub-id pub-id-type="pmid">14717831</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B14">
        <label>14.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Wang, H., Tang, J., Wu, M., Wang, X. and Zhang, T. (2022) Application of Machine Learning Missing Data Imputation Techniques in Clinical Decision Making: Taking the Discharge Assessment of Patients with Spontaneous Supratentorial Intracerebral Hemorrhage as an Example. <italic>BMC</italic><italic>Medical</italic><italic>Informatics</italic><italic>and</italic><italic>Decision</italic><italic>Making</italic>, 22, Article No. 13. https://doi.org/10.1186/s12911-022-01752-6 <pub-id pub-id-type="doi">10.1186/s12911-022-01752-6</pub-id><pub-id pub-id-type="pmid">35027065</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1186/s12911-022-01752-6">https://doi.org/10.1186/s12911-022-01752-6</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Wang, H.</string-name>
              <string-name>Tang, J.</string-name>
              <string-name>Wu, M.</string-name>
              <string-name>Wang, X.</string-name>
              <string-name>Zhang, T.</string-name>
            </person-group>
            <year>2022</year>
            <article-title>Application of Machine Learning Missing Data Imputation Techniques in Clinical Decision Making: Taking the Discharge Assessment of Patients with Spontaneous Supratentorial Intracerebral Hemorrhage as an Example</article-title>
            <source>BMC Medical Informatics and Decision Making</source>
            <volume>22</volume>
            <elocation-id>No</elocation-id>
            <pub-id pub-id-type="doi">10.1186/s12911-022-01752-6</pub-id>
            <pub-id pub-id-type="pmid">35027065</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B15">
        <label>15.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Dong, X., Lin, L., Zhang, R., Zhao, Y., Christiani, D.C., Wei, Y., <italic>et al.</italic> (2018) TOBMI: Trans-Omics Block Missing Data Imputation Using a K-Nearest Neighbor Weighted Approach. <italic>Bioinformatics</italic>, 35, 1278-1283. https://doi.org/10.1093/bioinformatics/bty796 <pub-id pub-id-type="doi">10.1093/bioinformatics/bty796</pub-id><pub-id pub-id-type="pmid">30202885</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1093/bioinformatics/bty796">https://doi.org/10.1093/bioinformatics/bty796</ext-link></mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Dong, X.</string-name>
              <string-name>Lin, L.</string-name>
              <string-name>Zhang, R.</string-name>
              <string-name>Zhao, Y.</string-name>
              <string-name>Christiani, D.C.</string-name>
              <string-name>Wei, Y.</string-name>
            </person-group>
            <year>2018</year>
            <article-title>TOBMI: Trans-Omics Block Missing Data Imputation Using a K-Nearest Neighbor Weighted Approach</article-title>
            <source>Bioinformatics</source>
            <volume>35</volume>
            <pub-id pub-id-type="doi">10.1093/bioinformatics/bty796</pub-id>
            <pub-id pub-id-type="pmid">30202885</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B16">
        <label>16.</label>
        <citation-alternatives>
          <mixed-citation publication-type="confproc">Al Mamlook, R.E., Nasayreh, A., Gharaibeh, H. and Shrestha, S. (2023) Classification of Cancer Genome Atlas Glioblastoma Multiform (TCGA-GBM) Using Machine Learning Method. 2023 <italic>IEEE International Conference on Electro Information Technology</italic> ( <italic>eIT</italic>), Romeoville, 18-20 May 2023, 265-270. https://doi.org/10.1109/eit57321.2023.10187283 <pub-id pub-id-type="doi">10.1109/eit57321.2023.10187283</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1109/eit57321.2023.10187283">https://doi.org/10.1109/eit57321.2023.10187283</ext-link></mixed-citation>
          <element-citation publication-type="confproc">
            <person-group person-group-type="author">
              <string-name>Mamlook, R.E.</string-name>
              <string-name>Nasayreh, A.</string-name>
              <string-name>Gharaibeh, H.</string-name>
              <string-name>Shrestha, S.</string-name>
            </person-group>
            <year>2023</year>
            <article-title>Classification of Cancer Genome Atlas Glioblastoma Multiform (TCGA-GBM) Using Machine Learning Method</article-title>
            <source>2023 IEEE International Conference on Electro Information Technology (eIT)</source>
            <volume>18</volume>
            <pub-id pub-id-type="doi">10.1109/eit57321.2023.10187283</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B17">
        <label>17.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Alin, A. (2010) Multicollinearity. <italic>WIREs</italic><italic>Computational</italic><italic>Statistics</italic>, 2, 370-374. https://doi.org/10.1002/wics.84 <pub-id pub-id-type="doi">10.1002/wics.84</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1002/wics.84">https://doi.org/10.1002/wics.84</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Alin, A.</string-name>
            </person-group>
            <year>2010</year>
            <article-title>Multicollinearity</article-title>
            <source>WIREs Computational Statistics</source>
            <volume>2</volume>
            <pub-id pub-id-type="doi">10.1002/wics.84</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B18">
        <label>18.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Mohammad-Djafari, A. (2021) Regularization, Bayesian Inference, and Machine Learning Methods for Inverse Problems. <italic>Entropy</italic>, 23, Article 1673. https://doi.org/10.3390/e23121673 <pub-id pub-id-type="doi">10.3390/e23121673</pub-id><pub-id pub-id-type="pmid">34945979</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.3390/e23121673">https://doi.org/10.3390/e23121673</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Mohammad-Djafari, A.</string-name>
              <string-name>Regularization, B</string-name>
            </person-group>
            <year>2021</year>
            <article-title>Regularization, Bayesian Inference, and Machine Learning Methods for Inverse Problems</article-title>
            <source>Entropy</source>
            <volume>23</volume>
            <elocation-id>1673</elocation-id>
            <pub-id pub-id-type="doi">10.3390/e23121673</pub-id>
            <pub-id pub-id-type="pmid">34945979</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B19">
        <label>19.</label>
        <citation-alternatives>
          <mixed-citation publication-type="confproc">Shekar, B.H. and Dagnew, G. (2019) Grid Search-Based Hyperparameter Tuning and Classification of Microarray Cancer Data. 2019 <italic>Second International Conference on Advanced Computational and Communication Paradigms</italic> ( <italic>ICACCP</italic>), Gangtok, 25-28 February 2019, 1-8. https://doi.org/10.1109/icaccp.2019.8882943 <pub-id pub-id-type="doi">10.1109/icaccp.2019.8882943</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1109/icaccp.2019.8882943">https://doi.org/10.1109/icaccp.2019.8882943</ext-link></mixed-citation>
          <element-citation publication-type="confproc">
            <person-group person-group-type="author">
              <string-name>Shekar, B.H.</string-name>
              <string-name>Dagnew, G.</string-name>
            </person-group>
            <year>2019</year>
            <article-title>Grid Search-Based Hyperparameter Tuning and Classification of Microarray Cancer Data</article-title>
            <source>2019 Second International Conference on Advanced Computational and Communication Paradigms (ICACCP)</source>
            <volume>25</volume>
            <pub-id pub-id-type="doi">10.1109/icaccp.2019.8882943</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B20">
        <label>20.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Stupp, R., Mason, W.P., van den Bent, M.J., Weller, M., Fisher, B., Taphoorn, M.J.B., <italic>et al.</italic> (2005) Radiotherapy Plus Concomitant and Adjuvant Temozolomide for Glioblastoma. <italic>New</italic><italic>England</italic><italic>Journal</italic><italic>of</italic><italic>Medicine</italic>, 352, 987-996. https://doi.org/10.1056/nejmoa043330 <pub-id pub-id-type="doi">10.1056/nejmoa043330</pub-id><pub-id pub-id-type="pmid">15758009</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1056/nejmoa043330">https://doi.org/10.1056/nejmoa043330</ext-link></mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Stupp, R.</string-name>
              <string-name>Mason, W.P.</string-name>
              <string-name>Bent, M.J.</string-name>
              <string-name>Weller, M.</string-name>
              <string-name>Fisher, B.</string-name>
              <string-name>Taphoorn, M.J.B.</string-name>
            </person-group>
            <year>2005</year>
            <article-title>Radiotherapy Plus Concomitant and Adjuvant Temozolomide for Glioblastoma</article-title>
            <source>New England Journal of Medicine</source>
            <volume>352</volume>
            <pub-id pub-id-type="doi">10.1056/nejmoa043330</pub-id>
            <pub-id pub-id-type="pmid">15758009</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B21">
        <label>21.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Azoulay, M., Santos, F., Souhami, L., Panet-Raymond, V., Petrecca, K., Owen, S., <italic>et al.</italic> (2015) Comparison of Radiation Regimens in the Treatment of Glioblastoma Multiforme: Results from a Single Institution. <italic>Radiation</italic><italic>Oncology</italic>, 10, Article No. 106. https://doi.org/10.1186/s13014-015-0396-6 <pub-id pub-id-type="doi">10.1186/s13014-015-0396-6</pub-id><pub-id pub-id-type="pmid">25927334</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1186/s13014-015-0396-6">https://doi.org/10.1186/s13014-015-0396-6</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Azoulay, M.</string-name>
              <string-name>Santos, F.</string-name>
              <string-name>Souhami, L.</string-name>
              <string-name>Panet-Raymond, V.</string-name>
              <string-name>Petrecca, K.</string-name>
              <string-name>Owen, S.</string-name>
            </person-group>
            <year>2015</year>
            <article-title>Comparison of Radiation Regimens in the Treatment of Glioblastoma Multiforme: Results from a Single Institution</article-title>
            <source>Radiation Oncology</source>
            <volume>10</volume>
            <elocation-id>No</elocation-id>
            <pub-id pub-id-type="doi">10.1186/s13014-015-0396-6</pub-id>
            <pub-id pub-id-type="pmid">25927334</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B22">
        <label>22.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Jiang, C., Mogilevsky, C., Belal, Z., Kurtz, G. and Alonso-Basanta, M. (2023) Hypofractionation in Glioblastoma: An Overview of Palliative, Definitive, and Exploratory Uses. <italic>Cancers</italic>, 15, Article 5650. https://doi.org/10.3390/cancers15235650 <pub-id pub-id-type="doi">10.3390/cancers15235650</pub-id><pub-id pub-id-type="pmid">38067354</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.3390/cancers15235650">https://doi.org/10.3390/cancers15235650</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Jiang, C.</string-name>
              <string-name>Mogilevsky, C.</string-name>
              <string-name>Belal, Z.</string-name>
              <string-name>Kurtz, G.</string-name>
              <string-name>Alonso-Basanta, M.</string-name>
              <string-name>Palliative, D</string-name>
            </person-group>
            <year>2023</year>
            <article-title>Hypofractionation in Glioblastoma: An Overview of Palliative, Definitive, and Exploratory Uses</article-title>
            <source>Cancers</source>
            <volume>15</volume>
            <elocation-id>5650</elocation-id>
            <pub-id pub-id-type="doi">10.3390/cancers15235650</pub-id>
            <pub-id pub-id-type="pmid">38067354</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
    </ref-list>
  </back>
</article>