<?xml version="1.0" encoding="UTF-8"?><!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v3.0 20080202//EN" "http://dtd.nlm.nih.gov/publishing/3.0/journalpublishing3.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" dtd-version="3.0" xml:lang="en" article-type="research article">
 <front>
  <journal-meta>
   <journal-id journal-id-type="publisher-id">
    ojn
   </journal-id>
   <journal-title-group>
    <journal-title>
     Open Journal of Nursing
    </journal-title>
   </journal-title-group>
   <issn pub-type="epub">
    2162-5336
   </issn>
   <issn publication-format="print">
    2162-5344
   </issn>
   <publisher>
    <publisher-name>
     Scientific Research Publishing
    </publisher-name>
   </publisher>
  </journal-meta>
  <article-meta>
   <article-id pub-id-type="doi">
    10.4236/ojn.2025.158047
   </article-id>
   <article-id pub-id-type="publisher-id">
    ojn-145031
   </article-id>
   <article-categories>
    <subj-group subj-group-type="heading">
     <subject>
      Articles
     </subject>
    </subj-group>
    <subj-group subj-group-type="Discipline-v2">
     <subject>
      Medicine 
     </subject>
     <subject>
       Healthcare
     </subject>
    </subj-group>
   </article-categories>
   <title-group>
    Statistical Analysis of a Diabetes Dataset and the Impact of Principal Component Analysis on Prediction Accuracy
   </title-group>
   <contrib-group>
    <contrib contrib-type="author" xlink:type="simple">
     <name name-style="western">
      <surname>
       Elizabeth
      </surname>
      <given-names>
       Diamond
      </given-names>
     </name>
    </contrib>
    <contrib contrib-type="author" xlink:type="simple">
     <name name-style="western">
      <surname>
       Faith
      </surname>
      <given-names>
       Idoko
      </given-names>
     </name>
    </contrib>
    <contrib contrib-type="author" xlink:type="simple">
     <name name-style="western">
      <surname>
       Michael
      </surname>
      <given-names>
       Olowe
      </given-names>
     </name>
    </contrib>
   </contrib-group> 
   <aff id="affnull">
    <addr-line>
     aDepartment of Industrial and Systems Engineering, North Carolina A&amp;T State University, Greensboro, NC, USA
    </addr-line> 
   </aff> 
   <pub-date pub-type="epub">
    <day>
     06
    </day> 
    <month>
     08
    </month>
    <year>
     2025
    </year>
   </pub-date> 
   <volume>
    15
   </volume> 
   <issue>
    08
   </issue>
   <fpage>
    638
   </fpage>
   <lpage>
    665
   </lpage>
   <history>
    <date date-type="received">
     <day>
      25,
     </day>
     <month>
      May
     </month>
     <year>
      2025
     </year>
    </date>
    <date date-type="published">
     <day>
      19,
     </day>
     <month>
      May
     </month>
     <year>
      2025
     </year> 
    </date> 
    <date date-type="accepted">
     <day>
      19,
     </day>
     <month>
      August
     </month>
     <year>
      2025
     </year> 
    </date>
   </history>
   <permissions>
    <copyright-statement>
     © Copyright 2014 by authors and Scientific Research Publishing Inc. 
    </copyright-statement>
    <copyright-year>
     2014
    </copyright-year>
    <license>
     <license-p>
      This work is licensed under the Creative Commons Attribution International License (CC BY). http://creativecommons.org/licenses/by/4.0/
     </license-p>
    </license>
   </permissions>
   <abstract>
    This paper aims to investigate the effectiveness of logistic regression and discriminant analysis in predicting diabetes in patients using a diabetes dataset. Additionally, the paper explores the impact of principal component analysis (PCA) on the prediction accuracy of these methods. The dataset used for this study contains clinical and demographic information of patients with and without diabetes. Logistic regression (LR) and discriminant analysis (DA) were employed to build predictive models using the dataset. The models were then evaluated using various performance metrics such as sensitivity, specificity, and accuracy. The hypothesis (D
    <sub>0</sub> = patient does NOT have diabetes, whereas D
    <sub>1</sub> = patient HAS diabetes) is determined with statistical analysis. Results show both logistic regression and discriminant analysis can accurately predict diabetes in patients. Performing PCA did not improve the prediction accuracy of these statistical techniques on the diabetes dataset. The analysis dataset contained 390 patient records with 14 clinical variables. While the dataset provides valuable insights, the relatively small sample size may limit the generalization of the results to broader populations. Our findings suggest that logistic regression or discriminant analysis can be a powerful tool for predicting diabetes in patients, aiding in early detection and effective prevention or management of the disease.
   </abstract>
   <kwd-group> 
    <kwd>
     Logistic Regression
    </kwd> 
    <kwd>
      Discriminant Analysis
    </kwd> 
    <kwd>
      Principal Component Analysis
    </kwd>
   </kwd-group>
  </article-meta>
 </front>
 <body>
  <sec id="s1">
   <title>1. Introduction</title>
   <sec id="s1_1">
    <title>1.1. Background</title>
    <p>According to the Centers for Disease Control and Prevention (CDC), diabetes is a long-term health condition that impacts how the body converts food into energy. When food is digested, it is transformed into glucose and absorbed into the bloodstream. Insulin is produced by the body to enable the uptake of blood sugar into the cells, where it is used as energy. However, for individuals with diabetes, the body either fails to produce sufficient insulin or is unable to utilize the insulin effectively to meet the energy needs of the body <xref ref-type="bibr" rid="scirp.145031-1">
      [1]
     </xref>.</p>
    <p>Recent statistics indicate that 37.3 million people in the United States, which represents 11.3% of the population, have diabetes <xref ref-type="bibr" rid="scirp.145031-2">
      [2]
     </xref>. Shockingly, one out of every five people with diabetes are unaware of their condition <xref ref-type="bibr" rid="scirp.145031-1">
      [1]
     </xref>. Diabetes can be classified into three major types, including Type 1, Type 2, and gestational diabetes, which occurs during pregnancy. Any of these variants can affect an individual.</p>
   </sec>
   <sec id="s1_2">
    <title>1.2. Problem Description</title>
    <p>Diabetes is a global health concern, characterized by the body’s inability to regulate blood glucose levels effectively. According to the Centers of Disease Control and Prevention National Diabetes Statistics Report , an estimated 130 million adults in the United States have diabetes or prediabetes. Diabetes is the leading cause of kidney failure, limb amputation, and blindness in adults . Over the past few decades, the prevalence of diagnosed type 2 diabetes in adults has risen dramatically, from 4.5% in 1995 to 10.2% in 2020 . The estimated percentage in 2022 has risen to 14.7% , and 1.4 million new cases of diabetes were diagnosed among people aged 18 and older in 2019. Adults with a family income below the federal poverty level had the highest prevalence for both men (13.7%) and women (14.4%), while individuals with less education were more likely to have diagnosed diabetes .</p>
    <p>Prediabetes is a reversible condition where blood sugar levels are high but not as high as those seen in diabetes. Understanding prediabetes is essential as it can prevent or delay the onset of type 2 diabetes. Diet and exercise have a significant influence on diabetes prevention. The National Diabetes Statistics Report contains full details on prediabetes.</p>
    <p>The increasing prevalence of diabetes is a significant public health concern, given its associated complications and status as the eighth leading cause of death in the United States. As a result, a considerable amount of research is underway to understand the causes of diabetes and how to prevent it. Our project aims to predict the probability of an individual having diabetes based on various health and physical factors, including cholesterol and glucose levels, systolic and diastolic blood pressure, age, height, weight, body mass index (BMI), waist, hip, and waist-hip ratio <xref ref-type="bibr" rid="scirp.145031-5">
      [5]
     </xref>.</p>
   </sec>
   <sec id="s1_3">
    <title>1.3. Project Scope</title>
    <p>The objective of this project is to analyze a diabetes dataset and develop a prediction model that can be used to screen for diabetes in patients based on predominant attributes and patterns. This analysis follows the hypothesis testing structure defined below.</p>
    <p>D<sub>0</sub>: Diabetes is absent.</p>
    <p>D<sub>1</sub>: Diabetes is present.</p>
    <p>Based on the attributes, the null hypothesis (D<sub>0</sub>) states that the patient has no diabetes while the alternate hypothesis (D<sub>1</sub>) states that the patient actually has diabetes.</p>
    <p>Three types of multivariate statistical techniques—principal component analysis, discriminant analysis, and logistic regression analysis will be employed. Publicly available data on diabetes will be used for this analysis. The following research questions serve as a guide for this study.</p>
    <p>The remaining part of this article is organized as follows: Section 2 includes previous work that addressed the same problem statements already discussed. Section 3 introduces the details of the used dataset, the pre-processing and data preparation phase and the machine learning algorithms used. Further, the results of statistical techniques used, and the associated accuracy are presented and discussed in Section 4. We finish with a conclusion in Section 5.</p>
    <sec id="s1">
     <title>2. Literature Review</title>
     <p>The prediction of diabetes is an important topic in medical research and has been the subject of numerous studies. In <xref ref-type="table" rid="table1">
       Table 1
      </xref>, we have highlighted major findings by researchers who have adopted several statistical procedures in the bid to predict diabetes. Some of these techniques include principal component analysis, logistic regression, decision trees, support vector machines (SVM), discriminant analysis and artificial neural networks (ANN) <xref ref-type="bibr" rid="scirp.145031-6">
       [6]
      </xref>.</p>
     <table-wrap id="table1">
      <label>
       <xref ref-type="table" rid="table1">
        Table 1
       </xref></label>
      <caption>
       <title>
        <xref ref-type="bibr" rid="scirp.145031-"></xref>Table 1. Literature review details.</title>
      </caption>
      <table class="MsoTableGrid custom-table" border="0" cellspacing="0" cellpadding="0"> 
       <tr> 
        <td class="custom-top-td acenter" width="4.45%"><p style="text-align:center">S/N</p></td> 
        <td class="custom-bottom-td custom-top-td acenter" width="31.06%"><p style="text-align:center">Technical Paper</p></td> 
        <td class="custom-bottom-td custom-top-td acenter" width="10.36%"><p style="text-align:center">Year of Publication</p></td> 
        <td class="custom-bottom-td custom-top-td acenter" width="10.35%"><p style="text-align:center">Statistical Techniques</p></td> 
        <td class="custom-bottom-td custom-top-td acenter" width="43.78%"><p style="text-align:center">Summary of Findings</p></td> 
       </tr> 
       <tr> 
        <td class="acenter" width="4.45%"><p style="text-align:center">1</p></td> 
        <td class="custom-top-td aleft" width="31.06%"><p style="text-align:left">Zhu, Changsheng, Christian Uwa Idemudia, and Wenfang Feng. “Improved logistic regression model for diabetes prediction by integrating PCA and K-means techniques.” Informatics in Medicine Unlocked 17 (2019): 100179. <xref ref-type="bibr" rid="scirp.145031-7">
           [7]
          </xref></p></td> 
        <td class="custom-top-td acenter" width="10.36%"><p style="text-align:center">2019</p></td> 
        <td class="custom-top-td acenter" width="10.35%"><p style="text-align:center">Logistic Regression</p></td> 
        <td class="custom-top-td aleft" width="43.78%"><p style="text-align:left">They determined ways of improving the k-means clustering and logistic regression accuracy result of Pima Indian Diabetes dataset using PCA (principal component analysis), k-means and logistic regression algorithm. The experimental investigations show that PCA improved the k-means clustering algorithm and logistic regression classifier accuracy versus the result of other published studies, with a k-means output of 25 more correctly classified data, and a logistic regression accuracy of 1.98% higher.</p></td> 
       </tr> 
       <tr> 
        <td class="acenter" width="4.45%"><p style="text-align:center">2</p></td> 
        <td class="aleft" width="31.06%"><p style="text-align:left">Rajendra, Priyanka, and Shahram Latifi. “Prediction of diabetes using logistic regression and ensemble techniques.” Computer Methods and Programs in Biomedicine Update 1 (2021): 100032. <xref ref-type="bibr" rid="scirp.145031-8">
           [8]
          </xref></p></td> 
        <td class="acenter" width="10.36%"><p style="text-align:center">2021</p></td> 
        <td class="acenter" width="10.35%"><p style="text-align:center">Logistic Regression</p></td> 
        <td class="aleft" width="43.78%"><p style="text-align:left">Logistic Regression technique was used to develop a prediction model using two datasets (Pima &amp; Vanderbilt) and two ensemble methods were further employed to improve the model performance by producing better predictions compared to a single model. With the adoption of ensemble methods, performance of the model was found to increase to 78% for dataset 1 and 93% for dataset 2.</p></td> 
       </tr> 
       <tr> 
        <td class="acenter" width="4.45%"><p style="text-align:center">3</p></td> 
        <td class="aleft" width="31.06%"><p style="text-align:left">Joshi, Ram D., and Chandra K. Dhakal. “Predicting type 2 diabetes using logistic regression and machine learning approaches.” International journal of environmental research and public health 18.14 (2021): 7346. <xref ref-type="bibr" rid="scirp.145031-9">
           [9]
          </xref></p></td> 
        <td class="acenter" width="10.36%"><p style="text-align:center">2021</p></td> 
        <td class="acenter" width="10.35%"><p style="text-align:center">Logistic Regression/Decision Tree</p></td> 
        <td class="aleft" width="43.78%"><p style="text-align:left">By employing the use of Logistic Regression and Decision Tree, the researchers sought to predict type 2 diabetes for Pima Indian women. Their analysis found five predominant predictors of type 2 diabetes which are glucose, pregnancy, body mass index (BMI), diabetes pedigree function, and age.</p></td> 
       </tr> 
       <tr> 
        <td class="acenter" width="4.45%"><p style="text-align:center">4</p></td> 
        <td class="aleft" width="31.06%"><p style="text-align:left">Polat, Kemal, Salih Güneş, and Ahmet Arslan. “A cascade learning system for classification of diabetes disease: Generalized discriminant analysis and least square support vector machine.” Expert systems with applications 34.1 (2008): 482-487. <xref ref-type="bibr" rid="scirp.145031-10">
           [10]
          </xref></p></td> 
        <td class="acenter" width="10.36%"><p style="text-align:center">2008</p></td> 
        <td class="acenter" width="10.35%"><p style="text-align:center">Discriminant Analysis/SVM</p></td> 
        <td class="aleft" width="43.78%"><p style="text-align:left">They developed a cascade learning system for accurately predicting diabetes in patients using a combined model called GDA-LS-SVM (Generalized Discriminant Analysis and Least Square Support Vector Machine). This model gave a classification accuracy relatively higher than the conventional LS-SVM model.</p></td> 
       </tr> 
       <tr> 
        <td class="acenter" width="4.45%"><p style="text-align:center">5</p></td> 
        <td class="aleft" width="31.06%"><p style="text-align:left">Alharan, Abbas FH, et al. “Improving classification performance for diabetes with linear discriminant analysis and genetic algorithm.” 2021 Palestinian International Conference on Information and Communication Technology (PICICT). IEEE, 2021. <xref ref-type="bibr" rid="scirp.145031-11">
           [11]
          </xref></p></td> 
        <td class="acenter" width="10.36%"><p style="text-align:center">2021</p></td> 
        <td class="acenter" width="10.35%"><p style="text-align:center">Discriminant Analysis/Genetic Algorithm</p></td> 
        <td class="aleft" width="43.78%"><p style="text-align:left">The researcher employed Discriminant Analysis and Genetic Algorithm on two datasets (Pima Indian Diabetes and Data of Dr John Schorling). They used the two techniques for feature selection and four techniques were used for classification. They found that Random Forest classifiers using DA and GA gave accuracy up to 91%.</p></td> 
       </tr> 
       <tr> 
        <td class="acenter" width="4.45%"><p style="text-align:center">6</p></td> 
        <td class="aleft" width="31.06%"><p style="text-align:left">Abdulhadi, Nour, and Amjed Al-Mousa. “Diabetes detection using machine learning classification methods.” 2021 International Conference on Information Technology (ICIT). IEEE, 2021. <xref ref-type="bibr" rid="scirp.145031-6">
           [6]
          </xref></p></td> 
        <td class="acenter" width="10.36%"><p style="text-align:center">2021</p></td> 
        <td class="acenter" width="10.35%"><p style="text-align:center">Linear Discriminant Analysis/Random Forest</p></td> 
        <td class="aleft" width="43.78%"><p style="text-align:left">The researchers were able to predict type 2 diabetes on female patients using the Pima Indian Diabetes Dataset obtained from National Institute of Diabetes and Digestive and Kidney Diseases were used. The Random Forest Classifier gave an accuracy of 82%.</p></td> 
       </tr> 
       <tr> 
        <td class="acenter" width="4.45%"><p style="text-align:center">7</p></td> 
        <td class="aleft" width="31.06%"><p style="text-align:left">El_Jerjawi, Nesreen Samer, and Samy S. Abu-Naser. “Diabetes prediction using artificial neural network.” (2018). <xref ref-type="bibr" rid="scirp.145031-12">
           [12]
          </xref></p></td> 
        <td class="acenter" width="10.36%"><p style="text-align:center">2018</p></td> 
        <td class="acenter" width="10.35%"><p style="text-align:center">Artificial Neural Network</p></td> 
        <td class="aleft" width="43.78%"><p style="text-align:left">The researchers used artificial neural networks to predict whether a person is diabetic or not. The criterion was to minimize the error function in neural network training using a neural network model. After training the ANN model, the average error function of the neural network was equal to 0.01 and the accuracy of the prediction of whether a person is diabetic or not was 87.3%.</p></td> 
       </tr> 
       <tr> 
        <td class="acenter" width="4.45%"><p style="text-align:center">8</p></td> 
        <td class="aleft" width="31.06%"><p style="text-align:left">Srivastava, Suyash, et al. “Prediction of diabetes using artificial neural network approach.” Engineering Vibration, Communication and Information Processing: ICoEVCI 2018, India. Springer Singapore, 2019. <xref ref-type="bibr" rid="scirp.145031-13">
           [13]
          </xref></p></td> 
        <td class="acenter" width="10.36%"><p style="text-align:center">2019</p></td> 
        <td class="acenter" width="10.35%"><p style="text-align:center">Artificial Neural Network</p></td> 
        <td class="aleft" width="43.78%"><p style="text-align:left">In this research work, a data sample of Pima Indians was taken to predict the possibility of diabetes. Among several algorithms of Machine learning, Artificial Neural Network (ANN) was chosen for building the model to predict diabetes. This model gave a prediction accuracy of 92% with the possibility of achieving higher accuracy if trained with a larger dataset.</p></td> 
       </tr> 
      </table>
     </table-wrap>
     <p>
      <xref ref-type="table" rid="table2">
       Table 2
      </xref> highlights some of the pros and cons of using these identified statistical techniques based on findings.</p>
     <table-wrap id="table2">
      <label>
       <xref ref-type="table" rid="table2">
        Table 2
       </xref></label>
      <caption>
       <title>
        <xref ref-type="bibr" rid="scirp.145031-"></xref>Table 2. Pros and cons of statistical techniques.</title>
      </caption>
      <table class="MsoTableGrid custom-table" border="0" cellspacing="0" cellpadding="0"> 
       <tr> 
        <td class="custom-bottom-td acenter" width="19.12%"><p style="text-align:center">Statistical Technique</p></td> 
        <td class="custom-bottom-td acenter" width="32.67%"><p style="text-align:center">Advantages</p></td> 
        <td class="custom-bottom-td acenter" width="36.98%"><p style="text-align:center">Disadvantages</p></td> 
       </tr> 
       <tr> 
        <td class="custom-bottom-td custom-top-td aleft" width="19.12%"><p style="text-align:left">Logistic Regression</p></td> 
        <td class="custom-bottom-td custom-top-td aleft" width="32.67%"><p style="text-align:left">- Simple and easy to interpret</p><p style="text-align:left">- Can handle categorical and continuous predictors</p><p style="text-align:left">- Can model the probability of an outcome</p></td> 
        <td class="custom-bottom-td custom-top-td aleft" width="36.98%"><p style="text-align:left">- Assumes linear relationship between predictors and outcome</p><p style="text-align:left">- Assumes independence of observations</p><p style="text-align:left">- May not handle complex interactions well</p></td> 
       </tr> 
       <tr> 
        <td class="custom-bottom-td custom-top-td aleft" width="19.12%"><p style="text-align:left">Decision Trees</p></td> 
        <td class="custom-bottom-td custom-top-td aleft" width="32.67%"><p style="text-align:left">- Nonlinear relationships between predictors and outcome can be captured.</p><p style="text-align:left">- Can handle both categorical and continuous predictors</p><p style="text-align:left">- Easy to interpret and visualize</p></td> 
        <td class="custom-bottom-td custom-top-td aleft" width="36.98%"><p style="text-align:left">- Prone to overfitting, especially with noisy data</p><p style="text-align:left">- Unstable: small changes in data can result in large changes in tree structure</p><p style="text-align:left">- Tendency to favor predictors with many categories</p></td> 
       </tr> 
       <tr> 
        <td class="custom-bottom-td custom-top-td aleft" width="19.12%"><p style="text-align:left">Random Forests</p></td> 
        <td class="custom-bottom-td custom-top-td aleft" width="32.67%"><p style="text-align:left">- Nonlinear relationships between predictors and outcome can be captured</p><p style="text-align:left">- Can handle both categorical and continuous predictors</p><p style="text-align:left">- Reduces overfitting by combining multiple decision trees</p><p style="text-align:left">- Can handle missing data</p></td> 
        <td class="custom-bottom-td custom-top-td aleft" width="36.98%"><p style="text-align:left">- Less interpretable than decision trees</p><p style="text-align:left">- Requires larger sample size and longer training time than decision trees</p><p style="text-align:left">- Can be computationally intensive</p></td> 
       </tr> 
       <tr> 
        <td class="custom-bottom-td custom-top-td aleft" width="19.12%"><p style="text-align:left">Support Vector Machines</p></td> 
        <td class="custom-bottom-td custom-top-td aleft" width="32.67%"><p style="text-align:left">- Can model complex, nonlinear relationships between predictors and outcome</p><p style="text-align:left">- Effective for high-dimensional data</p><p style="text-align:left">- Can handle both categorical and continuous predictors</p></td> 
        <td class="custom-bottom-td custom-top-td aleft" width="36.98%"><p style="text-align:left">- Requires careful selection of hyperparameters</p><p style="text-align:left">- Can be sensitive to outliers</p><p style="text-align:left">- Limited interpretability</p></td> 
       </tr> 
       <tr> 
        <td class="custom-bottom-td custom-top-td aleft" width="19.12%"><p style="text-align:left">Neural Networks</p></td> 
        <td class="custom-bottom-td custom-top-td aleft" width="32.67%"><p style="text-align:left">- Can model complex, nonlinear relationships between predictors and outcome</p><p style="text-align:left">- Can handle both categorical and continuous predictors</p><p style="text-align:left">- Can learn from noisy or incomplete data</p></td> 
        <td class="custom-bottom-td custom-top-td aleft" width="36.98%"><p style="text-align:left">- Requires larger sample size and longer training time than other methods</p><p style="text-align:left">- Can be computationally intensive</p><p style="text-align:left">- Less interpretable than other methods</p></td> 
       </tr> 
       <tr> 
        <td class="custom-top-td aleft" width="19.12%"><p style="text-align:left">Discriminant Analysis</p></td> 
        <td class="custom-top-td aleft" width="32.67%"><p style="text-align:left">- Can handle multiple predictors and multiple outcome classes</p><p style="text-align:left">- Assumes normality and equal covariance matrices among predictors</p><p style="text-align:left">- Can provide insight into which predictors are most important</p><p style="text-align:left">- Can handle missing data</p></td> 
        <td class="custom-top-td aleft" width="36.98%"><p style="text-align:left">- Assumes linear relationships between predictors and outcome</p><p style="text-align:left">- Sensitive to outliers and non-normality of data</p><p style="text-align:left">- Limited to two outcome classes (linear discriminant analysis) or assumes equal prior probabilities (quadratic discriminant analysis)</p></td> 
       </tr> 
      </table>
     </table-wrap>
    </sec>
    <sec id="s2_4">
     <title>2.1. Diabetes Logistic Regression Use Case</title>
     <p>A multivariate logistic regression equation can be used to screen for diabetes. In November of 2002, Diabetes Care published an article describing the development and validation of this empirical equation. The research design and methods were unique as subjects were from international sources. A predictive equation was developed with data collected from 1032 Egyptian subjects with no history of diabetes. The equation incorporated age, gender, BMI, postprandial time, and random capillary plasma glucose as independent covariates for prediction of undiagnosed diabetes. These covariates were based on a fasting plasma glucose level of ≥126 mg/dl and/or a plasma glucose level 2 hr after a 75-g oral glucose load ≥200 mg/dl. The equation was validated using data collected from an independent sample of 1065 American subjects <xref ref-type="bibr" rid="scirp.145031-14">
       [14]
      </xref>. The predictive equation was calculated with the following logistic regression parameters:</p>
     <p>
      <xref ref-type="bibr" rid="scirp.145031-"></xref>P = 1/(1 − e<sup>−</sup><sup>x</sup>), where x = −10.0382 + [0.0331*(age in years) + 0.0308*(random plasma glucose in mg/dl) + 0.2500*(postprandial time assessed as 0 to ≥8 hr) + 0.5620*(if female) + 0.0346*(BMI)].</p>
     <p>The cut-off point for the prediction of previously undiagnosed diabetes was defined as a probability value ≥ 0.20. The equation’s sensitivity was 65%, specificity 96%, and positive predictive value (PPV) 67%. When applied to a new sample, the equation’s sensitivity was 62%, specificity 96%, and PPV 63% . The equation improved on recommended methods of screening for undiagnosed diabetes and could be easily implemented in a handheld programmable calculator to predict previously undiagnosed diabetes . Machine learning, using several different models, has been attempted without further model improvement <xref ref-type="bibr" rid="scirp.145031-15">
       [15]
      </xref>.</p>
    </sec>
    <sec id="s2_5">
     <title>2.2. Performance Metrics</title>
     <p>The sensitivity of a test is defined as the proportion of people with disease who will have a positive result. A test with a high sensitivity is useful for “ruling out” a disease if a person tests negative <xref ref-type="bibr" rid="scirp.145031-17">
       [17]
      </xref>.</p>
     <p>The specificity of a test is the proportion of people without the disease who will have a negative result. A test with a high specificity is useful for “ruling in” a disease if a person tests positive <xref ref-type="bibr" rid="scirp.145031-17">
       [17]
      </xref>.</p>
     <fig id="fig1" position="float">
      <label>Figure 1</label>
      <caption>
       <title>Figure 1. Sensitivity, specificity and more (image source <xref ref-type="bibr" rid="scirp.145031-16">
         [16]
        </xref>).</title>
      </caption>
      <graphic mimetype="image" position="float" xlink:type="simple" xlink:href="https://html.scirp.org/file/1442492-rId13.jpeg?20250822030605" />
     </fig>
     <p>The predictive value of a test is determined by the test’s sensitivity and specificity and by the prevalence of the condition for which the test is used. Both PPV and NPV vary with changing prevalence of disease. It will therefore be wrong for clinicians to directly apply published predictive values of a test to their own populations when the prevalence of disease in their population is different from the prevalence of disease in the population in which the published study was carried out <xref ref-type="bibr" rid="scirp.145031-17">
       [17]
      </xref>. <xref ref-type="fig" rid="fig1">
       Figure 1
      </xref> shows a graphical representation of how Sensitivity, Specificity, Negative predictive value (NPV), Positive predictive value (PPV), and Accuracy are calculated.</p>
    </sec>
   </sec>
   <sec id="s3">
    <title>3. Methodology</title>
    <sec id="s3_1">
     <title>Data Collection and Description</title>
     <p>For this work, publicly available data from Kaggle was used, which was initially sourced from the National Institute of Diabetes and Digestive and Kidney Diseases. The dataset contains 390 instances with 14 predictor variables on patients, including their overall cholesterol level, HDL cholesterol level, cholesterol to HDL ratio, glucose levels, systolic and diastolic blood pressure, age, height, weight, BMI, waist and hip measurements, and diabetes status. All variables and their meanings are presented in <xref ref-type="table" rid="table3">
       Table 3
      </xref>. These values represent the ground truth labels in the dataset used for model training and evaluation. Performance metrics were chosen to provide a comprehensive evaluation of model performance, including sensitivity (recall), specificity (true negative rate), precision (positive predictive value), negative predictive value, and overall accuracy. These metrics were selected to reflect the clinical importance of both correctly identifying diabetic patients and minimizing false positives. In addition, a scatter plot matrix, and a correlation matrix were plotted to check for multicollinearity and to identify whether any strongly correlated features were present in the dataset that could impact model performance.</p>
     <p>The dependent variable in this analysis was the binary variable Outcome, representing the presence or absence of diabetes, while the 13 independent variables were the diagnostic measures mentioned above. Before analysis, the data were cleaned to ensure their suitability for the study.</p>
     <p>The National Institute of Diabetes and Digestive and Kidney Diseases collects data from different individuals through various programs and sources, such as surveys, the Diabetes Prevention Program, the National Diabetes Information Clearinghouse, and clinical trials. For example, the National Health and Nutrition Examination Survey (NHANES), conducted by the Centers for Disease Control and Prevention (CDC) in collaboration with the National Institute of Diabetes and Digestive and Kidney Diseases (NIDDK), collects information on diabetes prevalence, risk factors, and treatment. In contrast, the National Diabetes Prevention Program (NDPP) collects data on participants’ demographics, medical history, physical activity, and dietary habits to evaluate the program’s effectiveness. The National Diabetes Information Clearinghouse (NDIC) collects and compiles data on diabetes incidence, prevalence, risk factors, and complications from scientific literature and government reports. Finally, NIDDK conducts and supports clinical trials to test new treatments and interventions for diabetes and related conditions, collecting data on participants’ demographics, medical history, symptoms, and treatment outcomes to evaluate the safety and effectiveness of the intervention.</p>
     <table-wrap id="table3">
      <label>
       <xref ref-type="table" rid="table3">
        Table 3
       </xref></label>
      <caption>
       <title>
        <xref ref-type="bibr" rid="scirp.145031-"></xref>Table 3. Data description.</title>
      </caption>
      <table class="MsoTableGrid custom-table" border="0" cellspacing="0" cellpadding="0"> 
       <tr> 
        <td class="custom-bottom-td acenter" width="7.16%"><p style="text-align:center">S/N</p></td> 
        <td class="custom-bottom-td acenter" width="18.70%"><p style="text-align:center">Attribute Name</p></td> 
        <td class="custom-bottom-td acenter" width="58.20%"><p style="text-align:center">Attribute Description</p></td> 
        <td class="custom-bottom-td acenter" width="15.93%"><p style="text-align:center">Unit</p></td> 
       </tr> 
       <tr> 
        <td class="custom-top-td acenter" width="7.16%"><p style="text-align:center">1</p></td> 
        <td class="custom-top-td acenter" width="18.70%"><p style="text-align:center">Age</p></td> 
        <td class="custom-top-td aleft" width="58.20%"><p style="text-align:left">The patient’s age</p></td> 
        <td class="custom-top-td acenter" width="15.93%"><p style="text-align:center">years</p></td> 
       </tr> 
       <tr> 
        <td class="acenter" width="7.16%"><p style="text-align:center">2</p></td> 
        <td class="acenter" width="18.70%"><p style="text-align:center">Weight</p></td> 
        <td class="aleft" width="58.20%"><p style="text-align:left">The patient’s weight in kilograms</p></td> 
        <td class="acenter" width="15.93%"><p style="text-align:center">pounds (lbs)</p></td> 
       </tr> 
       <tr> 
        <td class="acenter" width="7.16%"><p style="text-align:center">3</p></td> 
        <td class="acenter" width="18.70%"><p style="text-align:center">BMI</p></td> 
        <td class="aleft" width="58.20%"><p style="text-align:left">The patient’s body mass index, calculated as weight divided by height squared</p></td> 
        <td class="acenter" width="15.93%"><p style="text-align:center">lbs/m<sup>2</sup></p></td> 
       </tr> 
       <tr> 
        <td class="acenter" width="7.16%"><p style="text-align:center">4</p></td> 
        <td class="acenter" width="18.70%"><p style="text-align:center">Systolicbp</p></td> 
        <td class="aleft" width="58.20%"><p style="text-align:left">The patient’s systolic blood pressure in millimeters of mercury</p></td> 
        <td class="acenter" width="15.93%"><p style="text-align:center">mmHg</p></td> 
       </tr> 
       <tr> 
        <td class="acenter" width="7.16%"><p style="text-align:center">5</p></td> 
        <td class="acenter" width="18.70%"><p style="text-align:center">Diastolicbp</p></td> 
        <td class="aleft" width="58.20%"><p style="text-align:left">The patient’s diastolic blood pressure level</p></td> 
        <td class="acenter" width="15.93%"><p style="text-align:center">mmHg</p></td> 
       </tr> 
       <tr> 
        <td class="acenter" width="7.16%"><p style="text-align:center">6</p></td> 
        <td class="acenter" width="18.70%"><p style="text-align:center">Waist</p></td> 
        <td class="aleft" width="58.20%"><p style="text-align:left">The patient’s waist circumference</p></td> 
        <td class="acenter" width="15.93%"><p style="text-align:center">Inches</p></td> 
       </tr> 
       <tr> 
        <td class="acenter" width="7.16%"><p style="text-align:center">7</p></td> 
        <td class="acenter" width="18.70%"><p style="text-align:center">Hip</p></td> 
        <td class="aleft" width="58.20%"><p style="text-align:left">The patient’s hip circumference</p></td> 
        <td class="acenter" width="15.93%"><p style="text-align:center">Inches</p></td> 
       </tr> 
       <tr> 
        <td class="acenter" width="7.16%"><p style="text-align:center">8</p></td> 
        <td class="acenter" width="18.70%"><p style="text-align:center">Height</p></td> 
        <td class="aleft" width="58.20%"><p style="text-align:left">The patient’s height</p></td> 
        <td class="acenter" width="15.93%"><p style="text-align:center">Inches</p></td> 
       </tr> 
       <tr> 
        <td class="acenter" width="7.16%"><p style="text-align:center">9</p></td> 
        <td class="acenter" width="18.70%"><p style="text-align:center">Waisthip ratio</p></td> 
        <td class="aleft" width="58.20%"><p style="text-align:left">The patient’s waist to hip ratio</p></td> 
        <td class="acenter" width="15.93%"><p style="text-align:center"></p></td> 
       </tr> 
       <tr> 
        <td class="acenter" width="7.16%"><p style="text-align:center">10</p></td> 
        <td class="acenter" width="18.70%"><p style="text-align:center">Cholesterol</p></td> 
        <td class="aleft" width="58.20%"><p style="text-align:left">The patient’s cholesterol level in milligrams per deciliter</p></td> 
        <td class="acenter" width="15.93%"><p style="text-align:center">mg/dL</p></td> 
       </tr> 
       <tr> 
        <td class="acenter" width="7.16%"><p style="text-align:center">11</p></td> 
        <td class="acenter" width="18.70%"><p style="text-align:center">Glucose</p></td> 
        <td class="aleft" width="58.20%"><p style="text-align:left">The patient’s glucose level</p></td> 
        <td class="acenter" width="15.93%"><p style="text-align:center">mg/dL</p></td> 
       </tr> 
       <tr> 
        <td class="acenter" width="7.16%"><p style="text-align:center">12</p></td> 
        <td class="acenter" width="18.70%"><p style="text-align:center">hdlchol</p></td> 
        <td class="aleft" width="58.20%"><p style="text-align:left">The patient’s HDL (High-density lipoprotein) cholesterol level</p></td> 
        <td class="acenter" width="15.93%"><p style="text-align:center">mg/dL</p></td> 
       </tr> 
       <tr> 
        <td class="acenter" width="7.16%"><p style="text-align:center">13</p></td> 
        <td class="acenter" width="18.70%"><p style="text-align:center">cholhdlratio</p></td> 
        <td class="aleft" width="58.20%"><p style="text-align:left">The patient’s total cholesterol to HDL cholesterol ratio</p></td> 
        <td class="acenter" width="15.93%"><p style="text-align:center"></p></td> 
       </tr> 
       <tr> 
        <td class="acenter" width="7.16%"><p style="text-align:center">14</p></td> 
        <td class="acenter" width="18.70%"><p style="text-align:center">Diabetes</p></td> 
        <td class="aleft" width="58.20%"><p style="text-align:left">A binary variable indicating whether the patient has been diagnosed with diabetes (0 = no, 1 = yes).</p></td> 
        <td class="acenter" width="15.93%"><p style="text-align:center"></p></td> 
       </tr> 
      </table>
     </table-wrap>
     <p>In this paper, we aim to predict the susceptibility of patients to diabetes using multivariate statistical techniques. Logistic Regression (LR) and Discriminant Analysis (DA) were chosen due to their established use in the field of medicine and engineering with categorical or discrete response variables. They are also robust to outliers and missing data, making them suitable for real-world datasets that may be incomplete or contain errors.</p>
     <p>The first phase of our research involves using DA and LR models to determine the dimensions of the cleaned Kaggle diabetes dataset that can reliably and accurately classify subjects into groups (i.e., diabetic or non-diabetic). LR will be used to identify a linear combination of independent variables (IVs) that best predicts membership in the diabetic or non-diabetic group, as measured by a categorical dependent variable (DV). This will involve deriving an equation for predicting the likelihood of diabetes in identified subjects. Additionally, we will fit the DA model to our dataset to derive the two discriminant functions.</p>
     <p>In the second phase of our research, we will perform Principal Component Analysis (PCA) on our dataset and fit the LR and DA models again on the PCA-reduced dataset. We will then check the model fit and accuracy of the LR and DA models on both the original and PCA-reduced datasets to draw insights into the prediction accuracies <xref ref-type="bibr" rid="scirp.145031-8">
       [8]
      </xref>. Finally, we will use the K-fold cross-validation technique to compare the performance of the LR and DA models <xref ref-type="bibr" rid="scirp.145031-18">
       [18]
      </xref>. We used SAS software and Python IDE as the statistical tools for coding. These statistical methods will help us determine the most accurate procedures for predicting diabetes susceptibility and provide insights into the usefulness of PCA as a tool for data reduction in medical studies <xref ref-type="bibr" rid="scirp.145031-7">
       [7]
      </xref>. Principal Component Analysis (PCA) was applied as an exploratory technique for dimensionality reduction. The variance explained by each component was examined through the proportion of the corresponding eigenvalues. The scree plot was used to determine the optimal number of principal components to retain for the model. Each principal component is a linear combination of the original features and the components are orthogonal (uncorrelated), allowing them to represent distinct directions of variance in the data. These components were used for model training to evaluate whether dimensionality reduction would improve classification performance. The component loadings were examined using the Maximum Likelihood Estimates from logistic regression and the Linear Discriminant Function coefficients.</p>
     <p>Model development for both Linear Discriminant Analysis (LDA) and Logistic Regression was performed using SAS software. The LDA model was developed based on the statistical analysis of linear discriminant functions, which aim to find a linear combination of features that best separates the two classes (diabetic and non-diabetic). The Logistic Regression model was developed as a linear model that estimates the probability of diabetes occurrence based on the input features through a logistic function. Both models were trained on the original set of 14 features and, separately, on a reduced set of principal components. Training was performed on both the original feature set and the reduced PCA-based feature set to enable direct comparison of model performance between the full and PCA-based representations.</p>
    </sec>
   </sec>
   <sec id="s4">
    <title>4. Results and Discussion</title>
    <sec id="s4_1">
     <title>4.1. Model Adequacy Check</title>
     <p>To ensure the validity of our analysis, several data pre-processing steps were conducted before assessing the adequacy of our model. Missing data, outliers, and unreadable characters in our dataset were checked for and removed. Also, special characters, such as commas and semicolons were fixed, to ensure the data was properly formatted.</p>
     <p>Next, several tests (in SAS) were conducted to ensure that our dataset met the necessary conditions for conducting multivariate statistical analysis, specifically Logistics Regression (LR) and Discriminant Analysis (DA). As multivariate analysis relies on assumptions of normality and homogeneity of variance, we performed univariate normality checks on randomly selected variables in the dataset to verify that the data followed a normal distribution.</p>
     <fig id="fig2" position="float">
      <label>Figure 2</label>
      <caption>
       <title>Figure 2. (a-f) Q-Q plots and distribution curves of randomly selected variables.</title>
      </caption>
      <graphic mimetype="image" position="float" xlink:type="simple" xlink:href="https://html.scirp.org/file/1442492-rId14.jpeg?20250822030611" />
     </fig>
     <p>Due to the complexity involved in testing multivariate normality, a univariate normality procedure was opted for, instead. Investigations showed that the dataset met the necessary normality assumptions, as confirmed by the Q-Q plots and distribution curves presented in <xref ref-type="fig" rid="fig2">
       Figure 2
      </xref>. These steps ensured that our dataset was suitable for further statistical analysis.</p>
     <p>Next, the test of homoscedasticity (homogeneity of variance) was conducted in SAS. This is given in the scatter matrix provided in <xref ref-type="fig" rid="fig3">
       Figure 3
      </xref>. No known pattern or trend can be seen, implying that the test of homoscedasticity is not violated.</p>
     <fig id="fig3" position="float">
      <label>Figure 3</label>
      <caption>
       <title>Figure 3. Scatter plot matrix.</title>
      </caption>
      <graphic mimetype="image" position="float" xlink:type="simple" xlink:href="https://html.scirp.org/file/1442492-rId15.jpeg?20250822030613" />
     </fig>
     <p>The first step involved conducting a correlation analysis amongst each of the independent variables. From the results obtained and presented in <xref ref-type="table" rid="table4">
       Table 4
      </xref>, there are no such variables with very high correlation among the independent variables. Although, high correlation was recorded between Waist and Weight (0.8478), Hip and Weight (0.8270) and Waist and Hip (0.8352). We will still proceed with our analysis.</p>
     <table-wrap id="table4">
      <label>
       <xref ref-type="table" rid="table4">
        Table 4
       </xref></label>
      <caption>
       <title>
        <xref ref-type="bibr" rid="scirp.145031-"></xref>Table 4. Correlation matrix.</title>
      </caption>
      <table class="MsoTableGrid custom-table" border="0" cellspacing="0" cellpadding="0"> 
       <tr> 
        <td class="custom-bottom-td acenter" width="106.31%" colspan="14"><p style="text-align:center">Correlation Matrix</p></td> 
       </tr> 
       <tr> 
        <td class="custom-bottom-td custom-top-td acenter" width="10.12%"><p style="text-align:center"></p></td> 
        <td class="custom-bottom-td custom-top-td acenter" width="8.87%"><p style="text-align:center">Cholesterol</p></td> 
        <td class="custom-bottom-td custom-top-td acenter" width="6.92%"><p style="text-align:center">Glucose</p></td> 
        <td class="custom-bottom-td custom-top-td acenter" width="6.64%"><p style="text-align:center">hdlchol</p></td> 
        <td class="custom-bottom-td custom-top-td acenter" width="9.29%"><p style="text-align:center">Cholhdl-ratio</p></td> 
        <td class="custom-bottom-td custom-top-td acenter" width="6.12%"><p style="text-align:center">Age</p></td> 
        <td class="custom-bottom-td custom-top-td acenter" width="6.22%"><p style="text-align:center">Height</p></td> 
        <td class="custom-bottom-td custom-top-td acenter" width="6.43%"><p style="text-align:center">Weight</p></td> 
        <td class="custom-bottom-td custom-top-td acenter" width="6.12%"><p style="text-align:center">BMI</p></td> 
        <td class="custom-bottom-td custom-top-td acenter" width="8.31%"><p style="text-align:center">Systolicbp</p></td> 
        <td class="custom-bottom-td custom-top-td acenter" width="8.93%"><p style="text-align:center">Diastolicbp</p></td> 
        <td class="custom-bottom-td custom-top-td acenter" width="6.12%"><p style="text-align:center">Waist</p></td> 
        <td class="custom-bottom-td custom-top-td acenter" width="6.12%"><p style="text-align:center">Hip</p></td> 
        <td class="custom-bottom-td custom-top-td acenter" width="10.12%"><p style="text-align:center">Waisthip-ratio</p></td> 
       </tr> 
       <tr> 
        <td class="custom-top-td acenter" width="10.12%"><p style="text-align:center">Cholesterol</p></td> 
        <td class="custom-top-td acenter" width="8.87%"><p style="text-align:center">1.0000</p></td> 
        <td class="custom-top-td acenter" width="6.92%"><p style="text-align:center">0.1581</p></td> 
        <td class="custom-top-td acenter" width="6.64%"><p style="text-align:center">0.1932</p></td> 
        <td class="custom-top-td acenter" width="9.29%"><p style="text-align:center">0.4759</p></td> 
        <td class="custom-top-td acenter" width="6.12%"><p style="text-align:center">0.2473</p></td> 
        <td class="custom-top-td acenter" width="6.22%"><p style="text-align:center">−0.0636</p></td> 
        <td class="custom-top-td acenter" width="6.43%"><p style="text-align:center">0.0624</p></td> 
        <td class="custom-top-td acenter" width="6.12%"><p style="text-align:center">0.0917</p></td> 
        <td class="custom-top-td acenter" width="8.31%"><p style="text-align:center">0.2077</p></td> 
        <td class="custom-top-td acenter" width="8.93%"><p style="text-align:center">0.1662</p></td> 
        <td class="custom-top-td acenter" width="6.12%"><p style="text-align:center">0.1340</p></td> 
        <td class="custom-top-td acenter" width="6.12%"><p style="text-align:center">0.0934</p></td> 
        <td class="custom-top-td acenter" width="10.12%"><p style="text-align:center">0.0918</p></td> 
       </tr> 
       <tr> 
        <td class="acenter" width="10.12%"><p style="text-align:center">Glucose</p></td> 
        <td class="acenter" width="8.87%"><p style="text-align:center">0.1581</p></td> 
        <td class="acenter" width="6.92%"><p style="text-align:center">1.0000</p></td> 
        <td class="acenter" width="6.64%"><p style="text-align:center">−0.1583</p></td> 
        <td class="acenter" width="9.29%"><p style="text-align:center">0.2822</p></td> 
        <td class="acenter" width="6.12%"><p style="text-align:center">0.2944</p></td> 
        <td class="acenter" width="6.22%"><p style="text-align:center">0.0981</p></td> 
        <td class="acenter" width="6.43%"><p style="text-align:center">0.1904</p></td> 
        <td class="acenter" width="6.12%"><p style="text-align:center">0.1293</p></td> 
        <td class="acenter" width="8.31%"><p style="text-align:center">0.1628</p></td> 
        <td class="acenter" width="8.93%"><p style="text-align:center">0.0203</p></td> 
        <td class="acenter" width="6.12%"><p style="text-align:center">0.2223</p></td> 
        <td class="acenter" width="6.12%"><p style="text-align:center">0.1382</p></td> 
        <td class="acenter" width="10.12%"><p style="text-align:center">0.1851</p></td> 
       </tr> 
       <tr> 
        <td class="acenter" width="10.12%"><p style="text-align:center">hdlchol</p></td> 
        <td class="acenter" width="8.87%"><p style="text-align:center">0.1932</p></td> 
        <td class="acenter" width="6.92%"><p style="text-align:center">−0.1583</p></td> 
        <td class="acenter" width="6.64%"><p style="text-align:center">1.0000</p></td> 
        <td class="acenter" width="9.29%"><p style="text-align:center">−0.6819</p></td> 
        <td class="acenter" width="6.12%"><p style="text-align:center">0.0282</p></td> 
        <td class="acenter" width="6.22%"><p style="text-align:center">−0.0872</p></td> 
        <td class="acenter" width="6.43%"><p style="text-align:center">−0.2919</p></td> 
        <td class="acenter" width="6.12%"><p style="text-align:center">−0.2419</p></td> 
        <td class="acenter" width="8.31%"><p style="text-align:center">0.0318</p></td> 
        <td class="acenter" width="8.93%"><p style="text-align:center">0.0783</p></td> 
        <td class="acenter" width="6.12%"><p style="text-align:center">−0.2767</p></td> 
        <td class="acenter" width="6.12%"><p style="text-align:center">−0.2238</p></td> 
        <td class="acenter" width="10.12%"><p style="text-align:center">−0.1588</p></td> 
       </tr> 
       <tr> 
        <td class="acenter" width="10.12%"><p style="text-align:center">Cholhdl-ratio</p></td> 
        <td class="acenter" width="8.87%"><p style="text-align:center">0.4759</p></td> 
        <td class="acenter" width="6.92%"><p style="text-align:center">0.2822</p></td> 
        <td class="acenter" width="6.64%"><p style="text-align:center">−0.6819</p></td> 
        <td class="acenter" width="9.29%"><p style="text-align:center">1.0000</p></td> 
        <td class="acenter" width="6.12%"><p style="text-align:center">0.1632</p></td> 
        <td class="acenter" width="6.22%"><p style="text-align:center">0.0812</p></td> 
        <td class="acenter" width="6.43%"><p style="text-align:center">0.2788</p></td> 
        <td class="acenter" width="6.12%"><p style="text-align:center">0.2284</p></td> 
        <td class="acenter" width="8.31%"><p style="text-align:center">0.1155</p></td> 
        <td class="acenter" width="8.93%"><p style="text-align:center">0.0382</p></td> 
        <td class="acenter" width="6.12%"><p style="text-align:center">0.3133</p></td> 
        <td class="acenter" width="6.12%"><p style="text-align:center">0.2089</p></td> 
        <td class="acenter" width="10.12%"><p style="text-align:center">0.2433</p></td> 
       </tr> 
       <tr> 
        <td class="acenter" width="10.12%"><p style="text-align:center">Age</p></td> 
        <td class="acenter" width="8.87%"><p style="text-align:center">0.2473</p></td> 
        <td class="acenter" width="6.92%"><p style="text-align:center">0.2944</p></td> 
        <td class="acenter" width="6.64%"><p style="text-align:center">0.0282</p></td> 
        <td class="acenter" width="9.29%"><p style="text-align:center">0.1632</p></td> 
        <td class="acenter" width="6.12%"><p style="text-align:center">1.0000</p></td> 
        <td class="acenter" width="6.22%"><p style="text-align:center">−0.0822</p></td> 
        <td class="acenter" width="6.43%"><p style="text-align:center">−0.0568</p></td> 
        <td class="acenter" width="6.12%"><p style="text-align:center">−0.0092</p></td> 
        <td class="acenter" width="8.31%"><p style="text-align:center">0.4534</p></td> 
        <td class="acenter" width="8.93%"><p style="text-align:center">0.0686</p></td> 
        <td class="acenter" width="6.12%"><p style="text-align:center">0.1506</p></td> 
        <td class="acenter" width="6.12%"><p style="text-align:center">0.0047</p></td> 
        <td class="acenter" width="10.12%"><p style="text-align:center">0.2752</p></td> 
       </tr> 
       <tr> 
        <td class="acenter" width="10.12%"><p style="text-align:center">Height</p></td> 
        <td class="acenter" width="8.87%"><p style="text-align:center">−0.0636</p></td> 
        <td class="acenter" width="6.92%"><p style="text-align:center">0.0981</p></td> 
        <td class="acenter" width="6.64%"><p style="text-align:center">−0.0872</p></td> 
        <td class="acenter" width="9.29%"><p style="text-align:center">0.0812</p></td> 
        <td class="acenter" width="6.12%"><p style="text-align:center">−0.0822</p></td> 
        <td class="acenter" width="6.22%"><p style="text-align:center">1.0000</p></td> 
        <td class="acenter" width="6.43%"><p style="text-align:center">0.2554</p></td> 
        <td class="acenter" width="6.12%"><p style="text-align:center">−0.2596</p></td> 
        <td class="acenter" width="8.31%"><p style="text-align:center">−0.0407</p></td> 
        <td class="acenter" width="8.93%"><p style="text-align:center">0.0436</p></td> 
        <td class="acenter" width="6.12%"><p style="text-align:center">0.0574</p></td> 
        <td class="acenter" width="6.12%"><p style="text-align:center">−0.0959</p></td> 
        <td class="acenter" width="10.12%"><p style="text-align:center">0.2525</p></td> 
       </tr> 
       <tr> 
        <td class="acenter" width="10.12%"><p style="text-align:center">Weight</p></td> 
        <td class="acenter" width="8.87%"><p style="text-align:center">0.0624</p></td> 
        <td class="acenter" width="6.92%"><p style="text-align:center">0.1904</p></td> 
        <td class="acenter" width="6.64%"><p style="text-align:center">−0.2919</p></td> 
        <td class="acenter" width="9.29%"><p style="text-align:center">0.2788</p></td> 
        <td class="acenter" width="6.12%"><p style="text-align:center">−0.0568</p></td> 
        <td class="acenter" width="6.22%"><p style="text-align:center">0.2554</p></td> 
        <td class="acenter" width="6.43%"><p style="text-align:center">1.0000</p></td> 
        <td class="acenter" width="6.12%"><p style="text-align:center">0.8601</p></td> 
        <td class="acenter" width="8.31%"><p style="text-align:center">0.0975</p></td> 
        <td class="acenter" width="8.93%"><p style="text-align:center">0.1665</p></td> 
        <td class="acenter" width="6.12%"><p style="text-align:center">0.8478</p></td> 
        <td class="acenter" width="6.12%"><p style="text-align:center">0.8270</p></td> 
        <td class="acenter" width="10.12%"><p style="text-align:center">0.2505</p></td> 
       </tr> 
       <tr> 
        <td class="acenter" width="10.12%"><p style="text-align:center">BMI</p></td> 
        <td class="acenter" width="8.87%"><p style="text-align:center">0.0917</p></td> 
        <td class="acenter" width="6.92%"><p style="text-align:center">0.1293</p></td> 
        <td class="acenter" width="6.64%"><p style="text-align:center">−0.2419</p></td> 
        <td class="acenter" width="9.29%"><p style="text-align:center">0.2284</p></td> 
        <td class="acenter" width="6.12%"><p style="text-align:center">−0.0092</p></td> 
        <td class="acenter" width="6.22%"><p style="text-align:center">−0.2596</p></td> 
        <td class="acenter" width="6.43%"><p style="text-align:center">0.8601</p></td> 
        <td class="acenter" width="6.12%"><p style="text-align:center">1.0000</p></td> 
        <td class="acenter" width="8.31%"><p style="text-align:center">0.1214</p></td> 
        <td class="acenter" width="8.93%"><p style="text-align:center">0.1453</p></td> 
        <td class="acenter" width="6.12%"><p style="text-align:center">0.8107</p></td> 
        <td class="acenter" width="6.12%"><p style="text-align:center">0.8817</p></td> 
        <td class="acenter" width="10.12%"><p style="text-align:center">0.1009</p></td> 
       </tr> 
       <tr> 
        <td class="acenter" width="10.12%"><p style="text-align:center">Systolicbp</p></td> 
        <td class="acenter" width="8.87%"><p style="text-align:center">0.2077</p></td> 
        <td class="acenter" width="6.92%"><p style="text-align:center">0.1628</p></td> 
        <td class="acenter" width="6.64%"><p style="text-align:center">0.0318</p></td> 
        <td class="acenter" width="9.29%"><p style="text-align:center">0.1155</p></td> 
        <td class="acenter" width="6.12%"><p style="text-align:center">0.4534</p></td> 
        <td class="acenter" width="6.22%"><p style="text-align:center">−0.0407</p></td> 
        <td class="acenter" width="6.43%"><p style="text-align:center">0.0975</p></td> 
        <td class="acenter" width="6.12%"><p style="text-align:center">0.1214</p></td> 
        <td class="acenter" width="8.31%"><p style="text-align:center">1.0000</p></td> 
        <td class="acenter" width="8.93%"><p style="text-align:center">0.6037</p></td> 
        <td class="acenter" width="6.12%"><p style="text-align:center">0.2109</p></td> 
        <td class="acenter" width="6.12%"><p style="text-align:center">0.1553</p></td> 
        <td class="acenter" width="10.12%"><p style="text-align:center">0.1379</p></td> 
       </tr> 
       <tr> 
        <td class="acenter" width="10.12%"><p style="text-align:center">Diastolicbp</p></td> 
        <td class="acenter" width="8.87%"><p style="text-align:center">0.1662</p></td> 
        <td class="acenter" width="6.92%"><p style="text-align:center">0.0203</p></td> 
        <td class="acenter" width="6.64%"><p style="text-align:center">0.0783</p></td> 
        <td class="acenter" width="9.29%"><p style="text-align:center">0.0382</p></td> 
        <td class="acenter" width="6.12%"><p style="text-align:center">0.0686</p></td> 
        <td class="acenter" width="6.22%"><p style="text-align:center">0.0436</p></td> 
        <td class="acenter" width="6.43%"><p style="text-align:center">0.1665</p></td> 
        <td class="acenter" width="6.12%"><p style="text-align:center">0.1453</p></td> 
        <td class="acenter" width="8.31%"><p style="text-align:center">0.6037</p></td> 
        <td class="acenter" width="8.93%"><p style="text-align:center">1.0000</p></td> 
        <td class="acenter" width="6.12%"><p style="text-align:center">0.1658</p></td> 
        <td class="acenter" width="6.12%"><p style="text-align:center">0.1439</p></td> 
        <td class="acenter" width="10.12%"><p style="text-align:center">0.0779</p></td> 
       </tr> 
       <tr> 
        <td class="acenter" width="10.12%"><p style="text-align:center">Waist</p></td> 
        <td class="acenter" width="8.87%"><p style="text-align:center">0.1340</p></td> 
        <td class="acenter" width="6.92%"><p style="text-align:center">0.2223</p></td> 
        <td class="acenter" width="6.64%"><p style="text-align:center">−0.2767</p></td> 
        <td class="acenter" width="9.29%"><p style="text-align:center">0.3133</p></td> 
        <td class="acenter" width="6.12%"><p style="text-align:center">0.1506</p></td> 
        <td class="acenter" width="6.22%"><p style="text-align:center">0.0574</p></td> 
        <td class="acenter" width="6.43%"><p style="text-align:center">0.8478</p></td> 
        <td class="acenter" width="6.12%"><p style="text-align:center">0.8107</p></td> 
        <td class="acenter" width="8.31%"><p style="text-align:center">0.2109</p></td> 
        <td class="acenter" width="8.93%"><p style="text-align:center">0.1658</p></td> 
        <td class="acenter" width="6.12%"><p style="text-align:center">1.0000</p></td> 
        <td class="acenter" width="6.12%"><p style="text-align:center">0.8352</p></td> 
        <td class="acenter" width="10.12%"><p style="text-align:center">0.5142</p></td> 
       </tr> 
       <tr> 
        <td class="acenter" width="10.12%"><p style="text-align:center">Hip</p></td> 
        <td class="acenter" width="8.87%"><p style="text-align:center">0.0934</p></td> 
        <td class="acenter" width="6.92%"><p style="text-align:center">0.1382</p></td> 
        <td class="acenter" width="6.64%"><p style="text-align:center">−0.2238</p></td> 
        <td class="acenter" width="9.29%"><p style="text-align:center">0.2089</p></td> 
        <td class="acenter" width="6.12%"><p style="text-align:center">0.0047</p></td> 
        <td class="acenter" width="6.22%"><p style="text-align:center">−0.0959</p></td> 
        <td class="acenter" width="6.43%"><p style="text-align:center">0.8270</p></td> 
        <td class="acenter" width="6.12%"><p style="text-align:center">0.8817</p></td> 
        <td class="acenter" width="8.31%"><p style="text-align:center">0.1553</p></td> 
        <td class="acenter" width="8.93%"><p style="text-align:center">0.1439</p></td> 
        <td class="acenter" width="6.12%"><p style="text-align:center">0.8352</p></td> 
        <td class="acenter" width="6.12%"><p style="text-align:center">1.0000</p></td> 
        <td class="acenter" width="10.12%"><p style="text-align:center">−0.0377</p></td> 
       </tr> 
       <tr> 
        <td class="acenter" width="10.12%"><p style="text-align:center">Waisthip-ratio</p></td> 
        <td class="acenter" width="8.87%"><p style="text-align:center">0.0918</p></td> 
        <td class="acenter" width="6.92%"><p style="text-align:center">0.1851</p></td> 
        <td class="acenter" width="6.64%"><p style="text-align:center">−0.1588</p></td> 
        <td class="acenter" width="9.29%"><p style="text-align:center">0.2433</p></td> 
        <td class="acenter" width="6.12%"><p style="text-align:center">0.2752</p></td> 
        <td class="acenter" width="6.22%"><p style="text-align:center">0.2525</p></td> 
        <td class="acenter" width="6.43%"><p style="text-align:center">0.2505</p></td> 
        <td class="acenter" width="6.12%"><p style="text-align:center">0.1009</p></td> 
        <td class="acenter" width="8.31%"><p style="text-align:center">0.1379</p></td> 
        <td class="acenter" width="8.93%"><p style="text-align:center">0.0779</p></td> 
        <td class="acenter" width="6.12%"><p style="text-align:center">0.5142</p></td> 
        <td class="acenter" width="6.12%"><p style="text-align:center">−0.0377</p></td> 
        <td class="acenter" width="10.12%"><p style="text-align:center">1.0000</p></td> 
       </tr> 
      </table>
     </table-wrap>
    </sec>
    <sec id="s4_2">
     <title>4.2. Descriptive Statistics</title>
     <p>As can be seen from the result of the simple statistics conducted on the Diabetes dataset (please see <xref ref-type="table" rid="table5">
       Table 5
      </xref>), the mean age of the patients is 47 years, which implies that the population sampled is an aging population with a standard deviation of approximately 16.4 years. The average height and weight of the population are 66 inches and 177.4 lbs respectively (corresponding standard deviations were measured as 3.9 inches and 40.4 lbs respectively). The systolic/diastolic blood pressure averaged 137/83. According to the American Heart Association, the systolic to diastolic blood pressure is ideal around 120/80, although this could vary depending on age, sex and health condition. The glucose level average for all the patients was 107.3 mg/dL but the dispersion or spread was relatively high (standard deviation = 53.80 mg/dL). The descriptive statistics are also shown in the visualization plot (error plot) shown in <xref ref-type="fig" rid="fig4">
       Figure 4
      </xref>.</p>
     <table-wrap id="table5">
      <label>
       <xref ref-type="table" rid="table5">
        Table 5
       </xref></label>
      <caption>
       <title>
        <xref ref-type="bibr" rid="scirp.145031-"></xref>Table 5. Simple statistics of the dataset.</title>
      </caption>
      <table class="MsoTableGrid custom-table" border="0" cellspacing="0" cellpadding="0"> 
       <tr> 
        <td class="custom-bottom-td acenter" width="100.00%" colspan="14"><p style="text-align:center">Simple Statistics</p></td> 
       </tr> 
       <tr> 
        <td class="custom-bottom-td custom-top-td acenter" width="6.28%"><p style="text-align:center"></p></td> 
        <td class="custom-bottom-td custom-top-td acenter" width="8.52%"><p style="text-align:center">Cholesterol</p></td> 
        <td class="custom-bottom-td custom-top-td acenter" width="5.92%"><p style="text-align:center">Glucose</p></td> 
        <td class="custom-bottom-td custom-top-td acenter" width="5.92%"><p style="text-align:center">hdlchol</p></td> 
        <td class="custom-bottom-td custom-top-td acenter" width="10.35%"><p style="text-align:center">Cholhdl-ratio</p></td> 
        <td class="custom-bottom-td custom-top-td acenter" width="5.54%"><p style="text-align:center">Age</p></td> 
        <td class="custom-bottom-td custom-top-td acenter" width="5.55%"><p style="text-align:center">Height</p></td> 
        <td class="custom-bottom-td custom-top-td acenter" width="5.54%"><p style="text-align:center">Weight</p></td> 
        <td class="custom-bottom-td custom-top-td acenter" width="5.55%"><p style="text-align:center">BMI</p></td> 
        <td class="custom-bottom-td custom-top-td acenter" width="9.61%"><p style="text-align:center">Systolic-bp</p></td> 
        <td class="custom-bottom-td custom-top-td acenter" width="9.62%"><p style="text-align:center">Diastolic-bp</p></td> 
        <td class="custom-bottom-td custom-top-td acenter" width="5.18%"><p style="text-align:center">Waist</p></td> 
        <td class="custom-bottom-td custom-top-td acenter" width="5.18%"><p style="text-align:center">Hip</p></td> 
        <td class="custom-bottom-td custom-top-td acenter" width="11.24%"><p style="text-align:center">Waisthip-ratio</p></td> 
       </tr> 
       <tr> 
        <td class="custom-top-td acenter" width="6.28%"><p style="text-align:center">Mean</p></td> 
        <td class="custom-top-td acenter" width="8.52%"><p style="text-align:center">207.23</p></td> 
        <td class="custom-top-td acenter" width="5.92%"><p style="text-align:center">107.34</p></td> 
        <td class="custom-top-td acenter" width="5.92%"><p style="text-align:center">50.27</p></td> 
        <td class="custom-top-td acenter" width="10.35%"><p style="text-align:center">4.52</p></td> 
        <td class="custom-top-td acenter" width="5.54%"><p style="text-align:center">46.77</p></td> 
        <td class="custom-top-td acenter" width="5.55%"><p style="text-align:center">65.95</p></td> 
        <td class="custom-top-td acenter" width="5.54%"><p style="text-align:center">177.41</p></td> 
        <td class="custom-top-td acenter" width="5.55%"><p style="text-align:center">28.78</p></td> 
        <td class="custom-top-td acenter" width="9.61%"><p style="text-align:center">137.13</p></td> 
        <td class="custom-top-td acenter" width="9.62%"><p style="text-align:center">83.29</p></td> 
        <td class="custom-top-td acenter" width="5.18%"><p style="text-align:center">37.87</p></td> 
        <td class="custom-top-td acenter" width="5.18%"><p style="text-align:center">42.99</p></td> 
        <td class="custom-top-td acenter" width="11.24%"><p style="text-align:center">0.88</p></td> 
       </tr> 
       <tr> 
        <td class="acenter" width="6.28%"><p style="text-align:center">StD</p></td> 
        <td class="acenter" width="8.52%"><p style="text-align:center">44.67</p></td> 
        <td class="acenter" width="5.92%"><p style="text-align:center">53.80</p></td> 
        <td class="acenter" width="5.92%"><p style="text-align:center">17.28</p></td> 
        <td class="acenter" width="10.35%"><p style="text-align:center">1.74</p></td> 
        <td class="acenter" width="5.54%"><p style="text-align:center">16.44</p></td> 
        <td class="acenter" width="5.55%"><p style="text-align:center">3.92</p></td> 
        <td class="acenter" width="5.54%"><p style="text-align:center">40.41</p></td> 
        <td class="acenter" width="5.55%"><p style="text-align:center">6.60</p></td> 
        <td class="acenter" width="9.61%"><p style="text-align:center">22.86</p></td> 
        <td class="acenter" width="9.62%"><p style="text-align:center">13.50</p></td> 
        <td class="acenter" width="5.18%"><p style="text-align:center">5.76</p></td> 
        <td class="acenter" width="5.18%"><p style="text-align:center">5.66</p></td> 
        <td class="acenter" width="11.24%"><p style="text-align:center">0.07</p></td> 
       </tr> 
      </table>
     </table-wrap>
     <fig id="fig4" position="float">
      <label>Figure 4</label>
      <caption>
       <title>Figure 4. Error plots of the predictor variables.</title>
      </caption>
      <graphic mimetype="image" position="float" xlink:type="simple" xlink:href="https://html.scirp.org/file/1442492-rId16.jpeg?20250822030613" />
     </fig>
    </sec>
    <sec id="s4_3">
     <title>4.3. Logistic Regression Results</title>
     <p>The choice of Logistic Regression (LR) was motivated by the fact that the dependent variables in our dataset were categorical (binary). We fitted the pre-processed dataset to a Logistic Regression model and checked the output file for the final results in SAS.</p>
     <p>To assess the validity and goodness of fit of our model, the −2LogL (where L is the likelihood function), AIC (Akaike Information Criterion), and SC (Schwarz’s Bayesian Criterion) values were evaluated and shown in <xref ref-type="table" rid="table6">
       Table 6
      </xref>. Lower values of −2LogL, AIC, and SC typically indicate better model fitting, and we will discuss this in more detail in another section.</p>
     <p>The Chi-square test statistics yielded p-values of ≤0.001, which is less than the level of significance of 0.05, indicating that the model is statistically significant.</p>
     <table-wrap id="table6">
      <label>
       <xref ref-type="table" rid="table6">
        Table 6
       </xref></label>
      <caption>
       <title>
        <xref ref-type="bibr" rid="scirp.145031-"></xref>Table 6. Model fit statistics.</title>
      </caption>
      <table class="MsoTableGrid custom-table" border="0" cellspacing="0" cellpadding="0"> 
       <tr> 
        <td class="custom-bottom-td acenter" width="100.00%" colspan="4"><p style="text-align:center">Model Fit Statistics</p></td> 
       </tr> 
       <tr> 
        <td class="custom-bottom-td custom-top-td acenter" width="25.86%"><p style="text-align:center">Criterion</p></td> 
        <td class="custom-bottom-td custom-top-td acenter" width="30.18%"><p style="text-align:center">Intercept Only</p></td> 
        <td class="custom-bottom-td custom-top-td acenter" width="43.96%" colspan="2"><p style="text-align:center">Intercept and Covariates</p></td> 
       </tr> 
       <tr> 
        <td class="custom-top-td acenter" width="25.86%"><p style="text-align:center">AIC</p></td> 
        <td class="custom-top-td acenter" width="30.18%"><p style="text-align:center">336.87</p></td> 
        <td class="custom-top-td acenter" width="43.96%" colspan="2"><p style="text-align:center">191.24</p></td> 
       </tr> 
       <tr> 
        <td class="acenter" width="25.86%"><p style="text-align:center">SC</p></td> 
        <td class="acenter" width="30.18%"><p style="text-align:center">340.84</p></td> 
        <td class="acenter" width="43.96%" colspan="2"><p style="text-align:center">246.76</p></td> 
       </tr> 
       <tr> 
        <td class="custom-bottom-td acenter" width="25.86%"><p style="text-align:center">−2 Log L</p></td> 
        <td class="custom-bottom-td acenter" width="30.18%"><p style="text-align:center">334.87</p></td> 
        <td class="custom-bottom-td acenter" width="43.96%" colspan="2"><p style="text-align:center">163.24</p></td> 
       </tr> 
       <tr> 
        <td class="custom-bottom-td custom-top-td acenter" width="100.00%" colspan="4"><p style="text-align:center">Testing Global Null Hypothesis: BETA = 0</p></td> 
       </tr> 
       <tr> 
        <td class="custom-bottom-td custom-top-td acenter" width="25.86%"><p style="text-align:center">Test</p></td> 
        <td class="custom-bottom-td custom-top-td acenter" width="30.18%"><p style="text-align:center">Chi-Square</p></td> 
        <td class="custom-bottom-td custom-top-td acenter" width="18.94%"><p style="text-align:center">DF</p></td> 
        <td class="custom-bottom-td custom-top-td acenter" width="25.01%"><p style="text-align:center">Pr &gt; ChiSq</p></td> 
       </tr> 
       <tr> 
        <td class="custom-top-td acenter" width="25.86%"><p style="text-align:center">Likelihood Ratio</p></td> 
        <td class="custom-top-td acenter" width="30.18%"><p style="text-align:center">171.64</p></td> 
        <td class="custom-top-td acenter" width="18.94%"><p style="text-align:center">13</p></td> 
        <td class="custom-top-td acenter" width="25.01%"><p style="text-align:center">&lt;0.0001</p></td> 
       </tr> 
       <tr> 
        <td class="acenter" width="25.86%"><p style="text-align:center">Score</p></td> 
        <td class="acenter" width="30.18%"><p style="text-align:center">194.80</p></td> 
        <td class="acenter" width="18.94%"><p style="text-align:center">13</p></td> 
        <td class="acenter" width="25.01%"><p style="text-align:center">&lt;0.0001</p></td> 
       </tr> 
       <tr> 
        <td class="acenter" width="25.86%"><p style="text-align:center">Wald</p></td> 
        <td class="acenter" width="30.18%"><p style="text-align:center">66.21</p></td> 
        <td class="acenter" width="18.94%"><p style="text-align:center">13</p></td> 
        <td class="acenter" width="25.01%"><p style="text-align:center">&lt;0.0001</p></td> 
       </tr> 
      </table>
     </table-wrap>
     <p>Overall, these results suggest that our fitted Logistic Regression model fits our dataset well and provides a statistically significant explanation for the relationship between our independent and dependent variables. A subsequent section will provide further details on the model performance and interpretation. The Logistic Regression model below is formulated according to the SAS output file shown:</p>
     <p>
      <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"> <mtable> 
        <mtr> 
         <mtd> 
          <mi>
            ln 
          </mi> 
          <mrow> 
           <mo>
             ( 
           </mo> 
           <mrow> 
            <mfrac> 
             <mi>
               p 
             </mi> 
             <mrow> 
              <mn>
                1 
              </mn> 
              <mo>
                − 
              </mo> 
              <mi>
                p 
              </mi> 
             </mrow> 
            </mfrac> 
           </mrow> 
           <mo>
             ) 
           </mo> 
          </mrow> 
          <mo>
            = 
          </mo> 
          <mo>
            − 
          </mo> 
          <mn>
            11.4431 
          </mn> 
          <mo>
            + 
          </mo> 
          <mn>
            0.0121 
          </mn> 
          <mo>
            ∗ 
          </mo> 
          <mtext>
            cholesterol 
          </mtext> 
          <mo>
            + 
          </mo> 
          <mn>
            0.0372 
          </mn> 
          <mo>
            ∗ 
          </mo> 
          <mtext>
            glucose 
          </mtext> 
         </mtd> 
        </mtr> 
        <mtr> 
         <mtd> 
          <mtext>
              
          </mtext> 
          <mtext>
              
          </mtext> 
          <mo>
            − 
          </mo> 
          <mn>
            0.0278 
          </mn> 
          <mo>
            ∗ 
          </mo> 
          <mtext>
            hdlchol 
          </mtext> 
          <mo>
            − 
          </mo> 
          <mn>
            0.1552 
          </mn> 
          <mo>
            ∗ 
          </mo> 
          <mtext>
            cholhdlratio 
          </mtext> 
          <mo>
            + 
          </mo> 
          <mn>
            0.0315 
          </mn> 
          <mo>
            ∗ 
          </mo> 
          <mtext>
            Age 
          </mtext> 
         </mtd> 
        </mtr> 
        <mtr> 
         <mtd> 
          <mtext>
              
          </mtext> 
          <mtext>
              
          </mtext> 
          <mo>
            + 
          </mo> 
          <mn>
            0.0392 
          </mn> 
          <mo>
            ∗ 
          </mo> 
          <mtext>
            Height 
          </mtext> 
          <mo>
            − 
          </mo> 
          <mn>
            0.0174 
          </mn> 
          <mo>
            ∗ 
          </mo> 
          <mtext>
            Weight 
          </mtext> 
          <mo>
            + 
          </mo> 
          <mn>
            0.1258 
          </mn> 
          <mo>
            ∗ 
          </mo> 
          <mtext>
            BMI 
          </mtext> 
         </mtd> 
        </mtr> 
        <mtr> 
         <mtd> 
          <mtext>
              
          </mtext> 
          <mtext>
              
          </mtext> 
          <mo>
            + 
          </mo> 
          <mn>
            0.00639 
          </mn> 
          <mo>
            ∗ 
          </mo> 
          <mtext>
            Systolicbp 
          </mtext> 
          <mo>
            + 
          </mo> 
          <mn>
            0.00868 
          </mn> 
          <mo>
            ∗ 
          </mo> 
          <mtext>
            Diastolicbp 
          </mtext> 
         </mtd> 
        </mtr> 
        <mtr> 
         <mtd> 
          <mtext>
              
          </mtext> 
          <mtext>
              
          </mtext> 
          <mo>
            + 
          </mo> 
          <mn>
            0.1098 
          </mn> 
          <mo>
            ∗ 
          </mo> 
          <mtext>
            Waist 
          </mtext> 
          <mo>
            − 
          </mo> 
          <mn>
            0.08333 
          </mn> 
          <mo>
            ∗ 
          </mo> 
          <mtext>
            Hip 
          </mtext> 
          <mo>
            − 
          </mo> 
          <mn>
            2.6863 
          </mn> 
          <mo>
            ∗ 
          </mo> 
          <mtext>
            Waisthipratio 
          </mtext> 
         </mtd> 
        </mtr> 
       </mtable> 
      </math></p>
     <p>where p is the probability of a patient having diabetes based on the maximum likelihood estimate approach (MLE). From <xref ref-type="table" rid="table7">
       Table 7
      </xref>, the p-value of glucose being less than 0.05 indicates that the glucose level is a predominant predictor variable from the LR model. The medical implication of this result is that higher glucose levels of patients could be a pointer to a higher risk of diabetes.</p>
     <table-wrap id="table7">
      <label>
       <xref ref-type="table" rid="table7">
        Table 7
       </xref></label>
      <caption>
       <title>
        <xref ref-type="bibr" rid="scirp.145031-"></xref>Table 7. Maximum likelihood.</title>
      </caption>
      <table class="MsoTableGrid custom-table" border="0" cellspacing="0" cellpadding="0"> 
       <tr> 
        <td class="custom-bottom-td acenter" width="100.00%" colspan="6"><p style="text-align:center">Analysis of Maximum Likelihood Estimates</p></td> 
       </tr> 
       <tr> 
        <td class="custom-bottom-td custom-top-td acenter" width="17.24%"><p style="text-align:center">Parameter</p></td> 
        <td class="custom-bottom-td custom-top-td acenter" width="10.78%"><p style="text-align:center">DF</p></td> 
        <td class="custom-bottom-td custom-top-td acenter" width="15.08%"><p style="text-align:center">Estimate</p></td> 
        <td class="custom-bottom-td custom-top-td acenter" width="19.40%"><p style="text-align:center">Standard Error</p></td> 
        <td class="custom-bottom-td custom-top-td acenter" width="20.89%"><p style="text-align:center">Wald Chi-Square</p></td> 
        <td class="custom-bottom-td custom-top-td acenter" width="16.60%"><p style="text-align:center">Pr &gt; ChiSq</p></td> 
       </tr> 
       <tr> 
        <td class="custom-top-td acenter" width="17.24%"><p style="text-align:center">Intercept</p></td> 
        <td class="custom-top-td acenter" width="10.78%"><p style="text-align:center">1</p></td> 
        <td class="custom-top-td acenter" width="15.08%"><p style="text-align:center">−11.4431</p></td> 
        <td class="custom-top-td acenter" width="19.40%"><p style="text-align:center">26.1632</p></td> 
        <td class="custom-top-td acenter" width="20.89%"><p style="text-align:center">0.1913</p></td> 
        <td class="custom-top-td acenter" width="16.60%"><p style="text-align:center">0.6618</p></td> 
       </tr> 
       <tr> 
        <td class="acenter" width="17.24%"><p style="text-align:center">Cholesterol</p></td> 
        <td class="acenter" width="10.78%"><p style="text-align:center">1</p></td> 
        <td class="acenter" width="15.08%"><p style="text-align:center">0.0121</p></td> 
        <td class="acenter" width="19.40%"><p style="text-align:center">0.00856</p></td> 
        <td class="acenter" width="20.89%"><p style="text-align:center">1.9913</p></td> 
        <td class="acenter" width="16.60%"><p style="text-align:center">0.1582</p></td> 
       </tr> 
       <tr> 
        <td class="acenter" width="17.24%"><p style="text-align:center">Glucose</p></td> 
        <td class="acenter" width="10.78%"><p style="text-align:center">1</p></td> 
        <td class="acenter" width="15.08%"><p style="text-align:center">0.0372</p></td> 
        <td class="acenter" width="19.40%"><p style="text-align:center">0.00544</p></td> 
        <td class="acenter" width="20.89%"><p style="text-align:center">46.7283</p></td> 
        <td class="acenter" width="16.60%"><p style="text-align:center">&lt;0.0001</p></td> 
       </tr> 
       <tr> 
        <td class="acenter" width="17.24%"><p style="text-align:center">hdlchol</p></td> 
        <td class="acenter" width="10.78%"><p style="text-align:center">1</p></td> 
        <td class="acenter" width="15.08%"><p style="text-align:center">−0.0278</p></td> 
        <td class="acenter" width="19.40%"><p style="text-align:center">0.0274</p></td> 
        <td class="acenter" width="20.89%"><p style="text-align:center">1.0318</p></td> 
        <td class="acenter" width="16.60%"><p style="text-align:center">0.3097</p></td> 
       </tr> 
       <tr> 
        <td class="acenter" width="17.24%"><p style="text-align:center">Cholhdl-ratio</p></td> 
        <td class="acenter" width="10.78%"><p style="text-align:center">1</p></td> 
        <td class="acenter" width="15.08%"><p style="text-align:center">−0.1552</p></td> 
        <td class="acenter" width="19.40%"><p style="text-align:center">0.2601</p></td> 
        <td class="acenter" width="20.89%"><p style="text-align:center">0.3559</p></td> 
        <td class="acenter" width="16.60%"><p style="text-align:center">0.5508</p></td> 
       </tr> 
       <tr> 
        <td class="acenter" width="17.24%"><p style="text-align:center">Age</p></td> 
        <td class="acenter" width="10.78%"><p style="text-align:center">1</p></td> 
        <td class="acenter" width="15.08%"><p style="text-align:center">0.0315</p></td> 
        <td class="acenter" width="19.40%"><p style="text-align:center">0.0171</p></td> 
        <td class="acenter" width="20.89%"><p style="text-align:center">3.3992</p></td> 
        <td class="acenter" width="16.60%"><p style="text-align:center">0.0652</p></td> 
       </tr> 
       <tr> 
        <td class="acenter" width="17.24%"><p style="text-align:center">Height</p></td> 
        <td class="acenter" width="10.78%"><p style="text-align:center">1</p></td> 
        <td class="acenter" width="15.08%"><p style="text-align:center">0.0392</p></td> 
        <td class="acenter" width="19.40%"><p style="text-align:center">0.2513</p></td> 
        <td class="acenter" width="20.89%"><p style="text-align:center">0.0243</p></td> 
        <td class="acenter" width="16.60%"><p style="text-align:center">0.8761</p></td> 
       </tr> 
       <tr> 
        <td class="acenter" width="17.24%"><p style="text-align:center">Weight</p></td> 
        <td class="acenter" width="10.78%"><p style="text-align:center">1</p></td> 
        <td class="acenter" width="15.08%"><p style="text-align:center">−0.0174</p></td> 
        <td class="acenter" width="19.40%"><p style="text-align:center">0.0433</p></td> 
        <td class="acenter" width="20.89%"><p style="text-align:center">0.1616</p></td> 
        <td class="acenter" width="16.60%"><p style="text-align:center">0.6876</p></td> 
       </tr> 
       <tr> 
        <td class="acenter" width="17.24%"><p style="text-align:center">BMI</p></td> 
        <td class="acenter" width="10.78%"><p style="text-align:center">1</p></td> 
        <td class="acenter" width="15.08%"><p style="text-align:center">0.1258</p></td> 
        <td class="acenter" width="19.40%"><p style="text-align:center">0.2538</p></td> 
        <td class="acenter" width="20.89%"><p style="text-align:center">0.2456</p></td> 
        <td class="acenter" width="16.60%"><p style="text-align:center">0.6202</p></td> 
       </tr> 
       <tr> 
        <td class="acenter" width="17.24%"><p style="text-align:center">Systolicbp</p></td> 
        <td class="acenter" width="10.78%"><p style="text-align:center">1</p></td> 
        <td class="acenter" width="15.08%"><p style="text-align:center">0.00639</p></td> 
        <td class="acenter" width="19.40%"><p style="text-align:center">0.0121</p></td> 
        <td class="acenter" width="20.89%"><p style="text-align:center">0.2794</p></td> 
        <td class="acenter" width="16.60%"><p style="text-align:center">0.5971</p></td> 
       </tr> 
       <tr> 
        <td class="acenter" width="17.24%"><p style="text-align:center">Diastolicbp</p></td> 
        <td class="acenter" width="10.78%"><p style="text-align:center">1</p></td> 
        <td class="acenter" width="15.08%"><p style="text-align:center">0.00868</p></td> 
        <td class="acenter" width="19.40%"><p style="text-align:center">0.0208</p></td> 
        <td class="acenter" width="20.89%"><p style="text-align:center">0.1744</p></td> 
        <td class="acenter" width="16.60%"><p style="text-align:center">0.6762</p></td> 
       </tr> 
       <tr> 
        <td class="acenter" width="17.24%"><p style="text-align:center">Waist</p></td> 
        <td class="acenter" width="10.78%"><p style="text-align:center">1</p></td> 
        <td class="acenter" width="15.08%"><p style="text-align:center">0.1098</p></td> 
        <td class="acenter" width="19.40%"><p style="text-align:center">0.5561</p></td> 
        <td class="acenter" width="20.89%"><p style="text-align:center">0.0390</p></td> 
        <td class="acenter" width="16.60%"><p style="text-align:center">0.8435</p></td> 
       </tr> 
       <tr> 
        <td class="acenter" width="17.24%"><p style="text-align:center">Hip</p></td> 
        <td class="acenter" width="10.78%"><p style="text-align:center">1</p></td> 
        <td class="acenter" width="15.08%"><p style="text-align:center">−0.0833</p></td> 
        <td class="acenter" width="19.40%"><p style="text-align:center">0.4790</p></td> 
        <td class="acenter" width="20.89%"><p style="text-align:center">0.0302</p></td> 
        <td class="acenter" width="16.60%"><p style="text-align:center">0.8620</p></td> 
       </tr> 
       <tr> 
        <td class="acenter" width="17.24%"><p style="text-align:center">Waisthip-ratio</p></td> 
        <td class="acenter" width="10.78%"><p style="text-align:center">1</p></td> 
        <td class="acenter" width="15.08%"><p style="text-align:center">−2.6864</p></td> 
        <td class="acenter" width="19.40%"><p style="text-align:center">24.7747</p></td> 
        <td class="acenter" width="20.89%"><p style="text-align:center">0.0118</p></td> 
        <td class="acenter" width="16.60%"><p style="text-align:center">0.9137</p></td> 
       </tr> 
      </table>
     </table-wrap>
     <p>Additionally, the Odds ratio estimates other than 1 indicates that the independent variables have some form of relationship with the outcome variables. The Odds ratio estimate is shown in <xref ref-type="table" rid="table8">
       Table 8
      </xref>.</p>
     <table-wrap id="table8">
      <label>
       <xref ref-type="table" rid="table8">
        Table 8
       </xref></label>
      <caption>
       <title>
        <xref ref-type="bibr" rid="scirp.145031-"></xref>Table 8. Odds ratio estimates.</title>
      </caption>
      <table class="MsoTableGrid custom-table" border="0" cellspacing="0" cellpadding="0"> 
       <tr> 
        <td class="custom-bottom-td acenter" width="100.00%" colspan="4"><p style="text-align:center">Odds Ratio Estimates</p></td> 
       </tr> 
       <tr> 
        <td class="custom-bottom-td custom-top-td acenter" width="24.98%"><p style="text-align:center">Effect</p></td> 
        <td class="custom-bottom-td custom-top-td acenter" width="26.74%"><p style="text-align:center">Point Estimate</p></td> 
        <td class="custom-bottom-td custom-top-td acenter" width="48.27%" colspan="2"><p style="text-align:center">95% Wald Confidence Limits</p></td> 
       </tr> 
       <tr> 
        <td class="custom-top-td acenter" width="24.98%"><p style="text-align:center">Cholesterol</p></td> 
        <td class="custom-top-td acenter" width="26.74%"><p style="text-align:center">1.012</p></td> 
        <td class="custom-top-td acenter" width="24.13%"><p style="text-align:center">0.995</p></td> 
        <td class="custom-top-td acenter" width="24.14%"><p style="text-align:center">1.029</p></td> 
       </tr> 
       <tr> 
        <td class="acenter" width="24.98%"><p style="text-align:center">Glucose</p></td> 
        <td class="acenter" width="26.74%"><p style="text-align:center">1.038</p></td> 
        <td class="acenter" width="24.13%"><p style="text-align:center">1.027</p></td> 
        <td class="acenter" width="24.14%"><p style="text-align:center">1.049</p></td> 
       </tr> 
       <tr> 
        <td class="acenter" width="24.98%"><p style="text-align:center">hdlchol</p></td> 
        <td class="acenter" width="26.74%"><p style="text-align:center">0.973</p></td> 
        <td class="acenter" width="24.13%"><p style="text-align:center">0.922</p></td> 
        <td class="acenter" width="24.14%"><p style="text-align:center">1.026</p></td> 
       </tr> 
       <tr> 
        <td class="acenter" width="24.98%"><p style="text-align:center">Cholhdl-ratio</p></td> 
        <td class="acenter" width="26.74%"><p style="text-align:center">0.856</p></td> 
        <td class="acenter" width="24.13%"><p style="text-align:center">0.514</p></td> 
        <td class="acenter" width="24.14%"><p style="text-align:center">1.426</p></td> 
       </tr> 
       <tr> 
        <td class="acenter" width="24.98%"><p style="text-align:center">Age</p></td> 
        <td class="acenter" width="26.74%"><p style="text-align:center">1.032</p></td> 
        <td class="acenter" width="24.13%"><p style="text-align:center">0.998</p></td> 
        <td class="acenter" width="24.14%"><p style="text-align:center">1.067</p></td> 
       </tr> 
       <tr> 
        <td class="acenter" width="24.98%"><p style="text-align:center">Height</p></td> 
        <td class="acenter" width="26.74%"><p style="text-align:center">1.040</p></td> 
        <td class="acenter" width="24.13%"><p style="text-align:center">0.636</p></td> 
        <td class="acenter" width="24.14%"><p style="text-align:center">1.702</p></td> 
       </tr> 
       <tr> 
        <td class="acenter" width="24.98%"><p style="text-align:center">Weight</p></td> 
        <td class="acenter" width="26.74%"><p style="text-align:center">0.983</p></td> 
        <td class="acenter" width="24.13%"><p style="text-align:center">0.903</p></td> 
        <td class="acenter" width="24.14%"><p style="text-align:center">1.070</p></td> 
       </tr> 
       <tr> 
        <td class="acenter" width="24.98%"><p style="text-align:center">BMI</p></td> 
        <td class="acenter" width="26.74%"><p style="text-align:center">1.134</p></td> 
        <td class="acenter" width="24.13%"><p style="text-align:center">0.690</p></td> 
        <td class="acenter" width="24.14%"><p style="text-align:center">1.865</p></td> 
       </tr> 
       <tr> 
        <td class="acenter" width="24.98%"><p style="text-align:center">Systolicbp</p></td> 
        <td class="acenter" width="26.74%"><p style="text-align:center">1.006</p></td> 
        <td class="acenter" width="24.13%"><p style="text-align:center">0.983</p></td> 
        <td class="acenter" width="24.14%"><p style="text-align:center">1.031</p></td> 
       </tr> 
       <tr> 
        <td class="acenter" width="24.98%"><p style="text-align:center">Diastolicbp</p></td> 
        <td class="acenter" width="26.74%"><p style="text-align:center">1.009</p></td> 
        <td class="acenter" width="24.13%"><p style="text-align:center">0.968</p></td> 
        <td class="acenter" width="24.14%"><p style="text-align:center">1.051</p></td> 
       </tr> 
       <tr> 
        <td class="acenter" width="24.98%"><p style="text-align:center">Waist</p></td> 
        <td class="acenter" width="26.74%"><p style="text-align:center">1.116</p></td> 
        <td class="acenter" width="24.13%"><p style="text-align:center">0.375</p></td> 
        <td class="acenter" width="24.14%"><p style="text-align:center">3.319</p></td> 
       </tr> 
       <tr> 
        <td class="acenter" width="24.98%"><p style="text-align:center">Hip</p></td> 
        <td class="acenter" width="26.74%"><p style="text-align:center">0.920</p></td> 
        <td class="acenter" width="24.13%"><p style="text-align:center">0.360</p></td> 
        <td class="acenter" width="24.14%"><p style="text-align:center">2.353</p></td> 
       </tr> 
       <tr> 
        <td class="acenter" width="24.98%"><p style="text-align:center">Waisthip-ratio</p></td> 
        <td class="acenter" width="26.74%"><p style="text-align:center">0.068</p></td> 
        <td class="acenter" width="24.13%"><p style="text-align:center">&lt;0.001</p></td> 
        <td class="acenter" width="24.14%"><p style="text-align:center">&gt;999.999</p></td> 
       </tr> 
      </table>
     </table-wrap>
     <p>Given the results above, the following inferences can be deduced:</p>
     <p>1) The odds ratio of 1.134 for BMI suggests that the odds of having diabetes are about 13.4% higher in obese individuals compared to non-obese individuals.</p>
     <p>2) The odds ratio of 1.032 for Age suggests that the odds of having diabetes are about 3.2% higher in older individuals compared to younger individuals.</p>
     <p>3) The odds ratios of 1.006 and 1.009, respectively, for Systolic and Diastolic blood pressure suggest that the odds of having diabetes are almost the same (about 0.6% and 0.9% higher) in groups with high and low blood pressure.</p>
     <p>4) The odds ratio of 1.038 for glucose levels suggests that the odds of having diabetes are about 3.8% higher in individuals with higher glucose levels compared to those with lower glucose levels.</p>
     <p>Overall, the odds ratios provide insight into the relationship between each independent variable and the dependent variable (diabetes) in the logistic regression model. These interpretations can be used to better understand the factors that are associated with the presence of diabetes in the population under study.</p>
    </sec>
    <sec id="s4_4">
     <title>4.4. Discriminant Analysis Results</title>
     <p>Next, we performed a Discriminant Analysis (DA) on the same pre-processed Kaggle diabetes dataset that was used for the Logistic Regression (LR) model. DA is a similar statistical procedure to LR, and it was carried out using the SAS software.</p>
     <p>The results obtained from the DA model were analyzed, and the confusion matrix for the statistical technique is presented in <xref ref-type="table" rid="table9">
       Table 9
      </xref>, also obtained from SAS.</p>
     <p>This additional analysis using DA provides further insights into the relationship between the independent variables and the dependent variable (diabetes) in our dataset <xref ref-type="bibr" rid="scirp.145031-11">
       [11]
      </xref>. By comparing the results of the DA model to those of the LR model, we can better understand the robustness and generalizability of our findings.</p>
     <p>The confusion table (refer to <xref ref-type="table" rid="table9">
       Table 9
      </xref>) shows that there is a total of 390 patients (330 of which are patients without diabetes while 60 are patients with diabetes). Of the 330 patients without diabetes, 310 patients were correctly predicted as non-diabetic while 20 were wrongly classified as diabetic), we can state that the model gave correct prediction for approximately 94% of the patients without diabetes. Additionally, the model wrongly predicted 11 diabetes patients as “non-diabetic” and appropriately predicted 49 patients as actually having diabetes. We can say that the model prediction here is around 82%.</p>
     <table-wrap id="table9">
      <label>
       <xref ref-type="table" rid="table9">
        Table 9
       </xref></label>
      <caption>
       <title>
        <xref ref-type="bibr" rid="scirp.145031-"></xref>Table 9. Classification summary for discriminant analysis.</title>
      </caption>
      <table class="MsoTableGrid custom-table" border="0" cellspacing="0" cellpadding="0"> 
       <tr> 
        <td class="acenter" width="119.48%" colspan="4"><p style="text-align:center">Number of Observations and Percent Classified into Diabetes</p></td> 
       </tr> 
       <tr> 
        <td class="custom-bottom-td custom-top-td acenter" width="29.86%"><p style="text-align:center">From Diabetes</p></td> 
        <td class="custom-bottom-td custom-top-td acenter" width="29.86%"><p style="text-align:center">0</p></td> 
        <td class="custom-bottom-td custom-top-td acenter" width="29.88%"><p style="text-align:center">1</p></td> 
        <td class="custom-bottom-td custom-top-td acenter" width="29.88%"><p style="text-align:center">Total</p></td> 
       </tr> 
       <tr> 
        <td class="custom-top-td acenter" width="29.86%"><p style="text-align:center">0</p></td> 
        <td class="custom-top-td acenter" width="29.86%"><p style="text-align:center">310</p><p style="text-align:center">93.94</p></td> 
        <td class="custom-top-td acenter" width="29.88%"><p style="text-align:center">20</p><p style="text-align:center">6.06</p></td> 
        <td class="custom-top-td acenter" width="29.88%"><p style="text-align:center">330</p><p style="text-align:center">100.00</p></td> 
       </tr> 
       <tr> 
        <td class="acenter" width="29.86%"><p style="text-align:center">1</p></td> 
        <td class="acenter" width="29.86%"><p style="text-align:center">11</p><p style="text-align:center">18.33</p></td> 
        <td class="acenter" width="29.88%"><p style="text-align:center">49</p><p style="text-align:center">81.67</p></td> 
        <td class="acenter" width="29.88%"><p style="text-align:center">60</p><p style="text-align:center">100.00</p></td> 
       </tr> 
       <tr> 
        <td class="acenter" width="29.86%"><p style="text-align:center">Total</p></td> 
        <td class="acenter" width="29.86%"><p style="text-align:center">321</p><p style="text-align:center">82.31</p></td> 
        <td class="acenter" width="29.88%"><p style="text-align:center">69</p><p style="text-align:center">17.69</p></td> 
        <td class="acenter" width="29.88%"><p style="text-align:center">390</p><p style="text-align:center">100.00</p></td> 
       </tr> 
       <tr> 
        <td class="acenter" width="29.86%"><p style="text-align:center">Priors</p></td> 
        <td class="acenter" width="29.86%"><p style="text-align:center">0.5</p></td> 
        <td class="acenter" width="29.88%"><p style="text-align:center">0.5</p></td> 
        <td class="acenter" width="29.88%"><p style="text-align:center"></p></td> 
       </tr> 
      </table>
     </table-wrap>
     <p>The overall accuracy of a binary classification is a measure of how often the model correctly predicts the class of new observations. It is calculated as follows:</p>
     <p>
      <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
        <mtext>
          Overall accuracy 
        </mtext> 
        <mo>
          = 
        </mo> 
        <mfrac> 
         <mrow> 
          <mtext>
            TP 
          </mtext> 
          <mo>
            + 
          </mo> 
          <mtext>
            TN 
          </mtext> 
         </mrow> 
         <mrow> 
          <mtext>
            TP 
          </mtext> 
          <mo>
            + 
          </mo> 
          <mtext>
            FP 
          </mtext> 
          <mo>
            + 
          </mo> 
          <mtext>
            TN 
          </mtext> 
          <mo>
            + 
          </mo> 
          <mtext>
            FN 
          </mtext> 
         </mrow> 
        </mfrac> 
       </mrow> 
      </math>,</p>
     <p>where TP = True Positive, TN = True Negative, FP = False Positive and FN = False Negative.</p>
     <p>The overall prediction of the DA model can therefore be computed as follows:</p>
     <p>TN = 310, TP = 49, FN = 20 and FP = 11. The prior probabilities were 50:50 (an equally likely chance scenario)</p>
     <p>
      <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
        <mfrac> 
         <mrow> 
          <mn>
            310 
          </mn> 
          <mo>
            + 
          </mo> 
          <mn>
            49 
          </mn> 
         </mrow> 
         <mrow> 
          <mn>
            310 
          </mn> 
          <mo>
            + 
          </mo> 
          <mn>
            49 
          </mn> 
          <mo>
            + 
          </mo> 
          <mn>
            20 
          </mn> 
          <mo>
            + 
          </mo> 
          <mn>
            11 
          </mn> 
         </mrow> 
        </mfrac> 
        <mtext> 
        </mtext> 
        <mo>
          × 
        </mo> 
        <mn>
          100 
        </mn> 
        <mo>
          = 
        </mo> 
        <mn>
          92 
        </mn> 
        <mi>
          % 
        </mi> 
       </mrow> 
      </math></p>
     <p>That means that the DA model correctly predicted outcomes (class labels) for 92% of the sample population provided.</p>
     <table-wrap id="table10">
      <label>
       <xref ref-type="table" rid="table10">
        Table 10
       </xref></label>
      <caption>
       <title>
        <xref ref-type="bibr" rid="scirp.145031-"></xref>Table 10. Performance metrics.</title>
      </caption>
      <table class="MsoTableGrid custom-table" border="0" cellspacing="0" cellpadding="0"> 
       <tr> 
        <td class="custom-bottom-td acenter" width="42.88%"><p style="text-align:center">Performance Metric</p></td> 
        <td class="custom-bottom-td acenter" width="27.37%"><p style="text-align:center">Score</p></td> 
       </tr> 
       <tr> 
        <td class="custom-top-td acenter" width="42.88%"><p style="text-align:center">Overall accuracy</p></td> 
        <td class="custom-top-td acenter" width="27.37%"><p style="text-align:center">92%</p></td> 
       </tr> 
       <tr> 
        <td class="acenter" width="42.88%"><p style="text-align:center">Specificity</p></td> 
        <td class="acenter" width="27.37%"><p style="text-align:center">97%</p></td> 
       </tr> 
       <tr> 
        <td class="acenter" width="42.88%"><p style="text-align:center">Sensitivity</p></td> 
        <td class="acenter" width="27.37%"><p style="text-align:center">71%</p></td> 
       </tr> 
       <tr> 
        <td class="acenter" width="42.88%"><p style="text-align:center">Negative Predictive Value</p></td> 
        <td class="acenter" width="27.37%"><p style="text-align:center">94%</p></td> 
       </tr> 
       <tr> 
        <td class="acenter" width="42.88%"><p style="text-align:center">Positive Predictive Value</p></td> 
        <td class="acenter" width="27.37%"><p style="text-align:center">82%</p></td> 
       </tr> 
      </table>
     </table-wrap>
     <p>Performance metrics were computed from the confusion matrix presented in <xref ref-type="table" rid="table10">
       Table 10
      </xref>, using the following standard formulas: Metrics Formula Value Overall Accuracy (TP + TN)/(TP + TN + FP + FN) 92.05%, Sensitivity (Recall) TP/(TP + FN) 71.01%, Specificity TN/(TN + FP) 96.57%, Precision (PPV) TP/(TP + FP) 81.67%, Negative Predictive Value (NPV) TN/(TN + FN) 93.94%.</p>
     <p>The high overall accuracy and specificity suggests that the DA model effectively identifies non-diabetic individuals <xref ref-type="bibr" rid="scirp.145031-10">
       [10]
      </xref>. At the same time, the lower sensitivity implies a higher chance of false negatives in identifying diabetic individuals. The high negative predictive value means that the model effectively identifies individuals who do not have diabetes. In contrast, the positive predictive value indicates that the model can correctly identify 8 out of every 10 individuals who have diabetes.</p>
    </sec>
    <sec id="s4_5">
     <title>4.5. Principal Component Analysis</title>
     <p>The dataset was subjected to Principal Component Analysis (PCA) to reduce the number of attributes to the optimal number of variables that capture the variability in the data <xref ref-type="bibr" rid="scirp.145031-19">
       [19]
      </xref>. From the analysis done in SAS, it was observed that five principal components are sufficient to run the model, as their eigenvalues were greater than 1. The results indicate that the Logistics Regression and Discriminant Analysis procedures can be run using these five principal components, which explain approximately 78.1% of the variance in the dataset. <xref ref-type="table" rid="table11">
       Table 11
      </xref> displays the eigenvalues and the proportion of variance explained by each principal component.</p>
     <table-wrap id="table11">
      <label>
       <xref ref-type="table" rid="table11">
        Table 11
       </xref></label>
      <caption>
       <title>
        <xref ref-type="bibr" rid="scirp.145031-"></xref>Table 11. Eigenvalues of the correlation matrix.</title>
      </caption>
      <table class="MsoTableGrid custom-table" border="0" cellspacing="0" cellpadding="0"> 
       <tr> 
        <td class="custom-bottom-td acenter" width="100.00%" colspan="5"><p style="text-align:center">Eigenvalues of the Correlation Matrix</p></td> 
       </tr> 
       <tr> 
        <td class="custom-bottom-td custom-top-td acenter" width="10.80%"><p style="text-align:center"></p></td> 
        <td class="custom-bottom-td custom-top-td acenter" width="22.29%"><p style="text-align:center">Eigenvalue</p></td> 
        <td class="custom-bottom-td custom-top-td acenter" width="22.31%"><p style="text-align:center">Difference</p></td> 
        <td class="custom-bottom-td custom-top-td acenter" width="22.29%"><p style="text-align:center">Proportion</p></td> 
        <td class="custom-bottom-td custom-top-td acenter" width="22.31%"><p style="text-align:center">Cumulative</p></td> 
       </tr> 
       <tr> 
        <td class="custom-top-td acenter" width="10.80%"><p style="text-align:center">1</p></td> 
        <td class="custom-top-td acenter" width="22.29%"><p style="text-align:center">4.08527421</p></td> 
        <td class="custom-top-td acenter" width="22.31%"><p style="text-align:center">2.06260964</p></td> 
        <td class="custom-top-td acenter" width="22.29%"><p style="text-align:center">0.3143</p></td> 
        <td class="custom-top-td acenter" width="22.31%"><p style="text-align:center">0.3143</p></td> 
       </tr> 
       <tr> 
        <td class="acenter" width="10.80%"><p style="text-align:center">2</p></td> 
        <td class="acenter" width="22.29%"><p style="text-align:center">2.02266457</p></td> 
        <td class="acenter" width="22.31%"><p style="text-align:center">0.33165910</p></td> 
        <td class="acenter" width="22.29%"><p style="text-align:center">0.1556</p></td> 
        <td class="acenter" width="22.31%"><p style="text-align:center">0.4698</p></td> 
       </tr> 
       <tr> 
        <td class="acenter" width="10.80%"><p style="text-align:center">3</p></td> 
        <td class="acenter" width="22.29%"><p style="text-align:center">1.69100547</p></td> 
        <td class="acenter" width="22.31%"><p style="text-align:center">0.39560790</p></td> 
        <td class="acenter" width="22.29%"><p style="text-align:center">0.1301</p></td> 
        <td class="acenter" width="22.31%"><p style="text-align:center">0.5999</p></td> 
       </tr> 
       <tr> 
        <td class="acenter" width="10.80%"><p style="text-align:center">4</p></td> 
        <td class="acenter" width="22.29%"><p style="text-align:center">1.29539756</p></td> 
        <td class="acenter" width="22.31%"><p style="text-align:center">0.23773573</p></td> 
        <td class="acenter" width="22.29%"><p style="text-align:center">0.0996</p></td> 
        <td class="acenter" width="22.31%"><p style="text-align:center">0.6996</p></td> 
       </tr> 
       <tr> 
        <td class="acenter" width="10.80%"><p style="text-align:center">5</p></td> 
        <td class="acenter" width="22.29%"><p style="text-align:center">1.05766184</p></td> 
        <td class="acenter" width="22.31%"><p style="text-align:center">0.09126156</p></td> 
        <td class="acenter" width="22.29%"><p style="text-align:center">0.0814</p></td> 
        <td class="acenter" width="22.31%"><p style="text-align:center">0.7809</p></td> 
       </tr> 
       <tr> 
        <td class="acenter" width="10.80%"><p style="text-align:center">6</p></td> 
        <td class="acenter" width="22.29%"><p style="text-align:center">0.96640027</p></td> 
        <td class="acenter" width="22.31%"><p style="text-align:center">0.14253078</p></td> 
        <td class="acenter" width="22.29%"><p style="text-align:center">0.0743</p></td> 
        <td class="acenter" width="22.31%"><p style="text-align:center">0.8553</p></td> 
       </tr> 
       <tr> 
        <td class="acenter" width="10.80%"><p style="text-align:center">7</p></td> 
        <td class="acenter" width="22.29%"><p style="text-align:center">0.82386949</p></td> 
        <td class="acenter" width="22.31%"><p style="text-align:center">0.25632787</p></td> 
        <td class="acenter" width="22.29%"><p style="text-align:center">0.0634</p></td> 
        <td class="acenter" width="22.31%"><p style="text-align:center">0.9186</p></td> 
       </tr> 
       <tr> 
        <td class="acenter" width="10.80%"><p style="text-align:center">8</p></td> 
        <td class="acenter" width="22.29%"><p style="text-align:center">0.56754162</p></td> 
        <td class="acenter" width="22.31%"><p style="text-align:center">0.29866870</p></td> 
        <td class="acenter" width="22.29%"><p style="text-align:center">0.0437</p></td> 
        <td class="acenter" width="22.31%"><p style="text-align:center">0.9623</p></td> 
       </tr> 
       <tr> 
        <td class="acenter" width="10.80%"><p style="text-align:center">9</p></td> 
        <td class="acenter" width="22.29%"><p style="text-align:center">0.26887292</p></td> 
        <td class="acenter" width="22.31%"><p style="text-align:center">0.12323370</p></td> 
        <td class="acenter" width="22.29%"><p style="text-align:center">0.0207</p></td> 
        <td class="acenter" width="22.31%"><p style="text-align:center">0.9830</p></td> 
       </tr> 
       <tr> 
        <td class="acenter" width="10.80%"><p style="text-align:center">10</p></td> 
        <td class="acenter" width="22.29%"><p style="text-align:center">0.14563921</p></td> 
        <td class="acenter" width="22.31%"><p style="text-align:center">0.07814038</p></td> 
        <td class="acenter" width="22.29%"><p style="text-align:center">0.0112</p></td> 
        <td class="acenter" width="22.31%"><p style="text-align:center">0.9942</p></td> 
       </tr> 
       <tr> 
        <td class="acenter" width="10.80%"><p style="text-align:center">11</p></td> 
        <td class="acenter" width="22.29%"><p style="text-align:center">0.06749884</p></td> 
        <td class="acenter" width="22.31%"><p style="text-align:center">0.06132666</p></td> 
        <td class="acenter" width="22.29%"><p style="text-align:center">0.0052</p></td> 
        <td class="acenter" width="22.31%"><p style="text-align:center">0.9994</p></td> 
       </tr> 
       <tr> 
        <td class="acenter" width="10.80%"><p style="text-align:center">12</p></td> 
        <td class="acenter" width="22.29%"><p style="text-align:center">0.00617218</p></td> 
        <td class="acenter" width="22.31%"><p style="text-align:center">0.00417034</p></td> 
        <td class="acenter" width="22.29%"><p style="text-align:center">0.0005</p></td> 
        <td class="acenter" width="22.31%"><p style="text-align:center">0.9998</p></td> 
       </tr> 
       <tr> 
        <td class="acenter" width="10.80%"><p style="text-align:center">13</p></td> 
        <td class="acenter" width="22.29%"><p style="text-align:center">0.00200184</p></td> 
        <td class="acenter" width="22.31%"><p style="text-align:center"></p></td> 
        <td class="acenter" width="22.29%"><p style="text-align:center">0.0002</p></td> 
        <td class="acenter" width="22.31%"><p style="text-align:center">1.0000</p></td> 
       </tr> 
      </table>
     </table-wrap>
     <p>The scree plot elbow also establishes that five principal components are statistically sufficient to explain the variance (please see <xref ref-type="fig" rid="fig5">
       Figure 5
      </xref>).</p>
     <fig id="fig5" position="float">
      <label>Figure 5</label>
      <caption>
       <title>Figure 5. Scree plot and variance explanation.</title>
      </caption>
      <graphic mimetype="image" position="float" xlink:type="simple" xlink:href="https://html.scirp.org/file/1442492-rId23.jpeg?20250822030619" />
     </fig>
    </sec>
    <sec id="s4_6">
     <title>4.6. Logistic Regression Results (after PCA)</title>
     <p>We fitted Logistic Regression to the PCA-reduced dataset. We checked the output file for the final results and the goodness of fit by inspecting the -2LogL values, the AIC (Akaike Information Criterion), SC (Schwarz’s Bayesian Criterion). We checked the difference in the values as the intercepts, prin1, prin2, prin3, prin4 and prin5 using the stepwise procedure (forward selection). <xref ref-type="fig" rid="fig6">
       Figure 6
      </xref> presents screenshots results of the step by step process utilized. The p-values (≤0.001) are less than 0.05 (the significance level), indicating that the model is statistically significant at every level of variable addition. The result section is given below:</p>
     <fig-group id="fig6" position="float">
      <fig id="fig6" position="float">
       <label>Figure 6</label>
       <caption>
        <title>Figure 6. (a-f) Screen shot of model convergence status.</title>
       </caption>
       <graphic mimetype="image" position="float" xlink:type="simple" xlink:href="https://html.scirp.org/file/1442492-rId24.jpeg?20250822030620" />
      </fig>
      <fig id="fig6" position="float">
       <label>Figure 6</label>
       <caption>
        <title>Figure 6. (a-f) Screen shot of model convergence status.</title>
       </caption>
       <graphic mimetype="image" position="float" xlink:type="simple" xlink:href="https://html.scirp.org/file/1442492-rId25.jpeg?20250822030621" />
      </fig>
      <fig id="fig6" position="float">
       <label>Figure 6</label>
       <caption>
        <title>Figure 6. (a-f) Screen shot of model convergence status.</title>
       </caption>
       <graphic mimetype="image" position="float" xlink:type="simple" xlink:href="https://html.scirp.org/file/1442492-rId26.jpeg?20250822030621" />
      </fig>
      <fig id="fig6" position="float">
       <label>Figure 6</label>
       <caption>
        <title>Figure 6. (a-f) Screen shot of model convergence status.</title>
       </caption>
       <graphic mimetype="image" position="float" xlink:type="simple" xlink:href="https://html.scirp.org/file/1442492-rId28.jpeg?20250822030620" />
      </fig>
     </fig-group>
     <table-wrap id="table12">
      <label>
       <xref ref-type="table" rid="table12">
        Table 12
       </xref></label>
      <caption>
       <title>
        <xref ref-type="bibr" rid="scirp.145031-"></xref>Table 12. Validating model fit using principal component analysis (PCA).</title>
      </caption>
      <table class="MsoTableGrid custom-table" border="0" cellspacing="0" cellpadding="0"> 
       <tr> 
        <td class="custom-bottom-td acenter" width="20.94%"><p style="text-align:center">Procedure</p></td> 
        <td class="custom-bottom-td acenter" width="16.21%"><p style="text-align:center">0</p></td> 
        <td class="custom-bottom-td acenter" width="16.21%"><p style="text-align:center">1</p></td> 
        <td class="custom-bottom-td acenter" width="16.21%"><p style="text-align:center">2</p></td> 
        <td class="custom-bottom-td acenter" width="16.21%"><p style="text-align:center">3</p></td> 
        <td class="custom-bottom-td acenter" width="16.21%"><p style="text-align:center">4</p></td> 
        <td class="custom-bottom-td acenter" width="16.21%"><p style="text-align:center">5</p></td> 
       </tr> 
       <tr> 
        <td class="custom-top-td acenter" width="20.94%"><p style="text-align:center">AIC</p></td> 
        <td class="custom-top-td acenter" width="16.21%"><p style="text-align:center"></p></td> 
        <td class="custom-top-td acenter" width="16.21%"><p style="text-align:center">336.872</p></td> 
        <td class="custom-top-td acenter" width="16.21%"><p style="text-align:center">336.872</p></td> 
        <td class="custom-top-td acenter" width="16.21%"><p style="text-align:center">336.872</p></td> 
        <td class="custom-top-td acenter" width="16.21%"><p style="text-align:center">336.872</p></td> 
        <td class="custom-top-td acenter" width="16.21%"><p style="text-align:center">336.872</p></td> 
       </tr> 
       <tr> 
        <td class="acenter" width="20.94%"><p style="text-align:center">SC</p></td> 
        <td class="acenter" width="16.21%"><p style="text-align:center"></p></td> 
        <td class="acenter" width="16.21%"><p style="text-align:center">340.838</p></td> 
        <td class="acenter" width="16.21%"><p style="text-align:center">340.838</p></td> 
        <td class="acenter" width="16.21%"><p style="text-align:center">340.838</p></td> 
        <td class="acenter" width="16.21%"><p style="text-align:center">340.838</p></td> 
        <td class="acenter" width="16.21%"><p style="text-align:center">340.838</p></td> 
       </tr> 
       <tr> 
        <td class="acenter" width="20.94%"><p style="text-align:center">−2logL</p></td> 
        <td class="acenter" width="16.21%"><p style="text-align:center">334.872</p></td> 
        <td class="acenter" width="16.21%"><p style="text-align:center">334.872</p></td> 
        <td class="acenter" width="16.21%"><p style="text-align:center">334.872</p></td> 
        <td class="acenter" width="16.21%"><p style="text-align:center">334.872</p></td> 
        <td class="acenter" width="16.21%"><p style="text-align:center">334.872</p></td> 
        <td class="acenter" width="16.21%"><p style="text-align:center">334.872</p></td> 
       </tr> 
      </table>
     </table-wrap>
     <p>As seen in <xref ref-type="fig" rid="figFigures 6(a)-(f)">
       Figures 6(a)-(f)
      </xref> and <xref ref-type="table" rid="table12">
       Table 12
      </xref>, the values of the parameters (AIC, SC, and −2logL) remained unchanged. Theoretically, this implies that the addition of new variables (principal components) did not directly improve the model, but this is not to say that the principal components are not outrightly irrelevant.</p>
     <p>The Logistic Regression model below is formulated according to the maximum likelihood estimates obtained from SAS and presented in <xref ref-type="table" rid="table13">
       Table 13
      </xref>.</p>
     <table-wrap id="table13">
      <label>
       <xref ref-type="table" rid="table13">
        Table 13
       </xref></label>
      <caption>
       <title>
        <xref ref-type="bibr" rid="scirp.145031-"></xref>Table 13. Maximum likelihood.</title>
      </caption>
      <table class="MsoTableGrid custom-table" border="0" cellspacing="0" cellpadding="0"> 
       <tr> 
        <td class="custom-bottom-td acenter" width="100.00%" colspan="6"><p style="text-align:center">Analysis of Maximum Likelihood Estimates</p></td> 
       </tr> 
       <tr> 
        <td class="custom-bottom-td custom-top-td acenter" width="16.66%"><p style="text-align:center">Parameter</p></td> 
        <td class="custom-bottom-td custom-top-td acenter" width="11.36%"><p style="text-align:center">DF</p></td> 
        <td class="custom-bottom-td custom-top-td acenter" width="15.08%"><p style="text-align:center">Estimate</p></td> 
        <td class="custom-bottom-td custom-top-td acenter" width="19.40%"><p style="text-align:center">Standard Error</p></td> 
        <td class="custom-bottom-td custom-top-td acenter" width="20.83%"><p style="text-align:center">Wald Chi-Square</p></td> 
        <td class="custom-bottom-td custom-top-td acenter" width="16.66%"><p style="text-align:center">Pr &gt; ChiSq</p></td> 
       </tr> 
       <tr> 
        <td class="custom-top-td acenter" width="16.66%"><p style="text-align:center">Intercept</p></td> 
        <td class="custom-top-td acenter" width="11.36%"><p style="text-align:center">1</p></td> 
        <td class="custom-top-td acenter" width="15.08%"><p style="text-align:center">−2.5661</p></td> 
        <td class="custom-top-td acenter" width="19.40%"><p style="text-align:center">0.2432</p></td> 
        <td class="custom-top-td acenter" width="20.83%"><p style="text-align:center">111.2905</p></td> 
        <td class="custom-top-td acenter" width="16.66%"><p style="text-align:center">&lt;0.0001</p></td> 
       </tr> 
       <tr> 
        <td class="acenter" width="16.66%"><p style="text-align:center">Prin1</p></td> 
        <td class="acenter" width="11.36%"><p style="text-align:center">1</p></td> 
        <td class="acenter" width="15.08%"><p style="text-align:center">0.5689</p></td> 
        <td class="acenter" width="19.40%"><p style="text-align:center">0.099</p></td> 
        <td class="acenter" width="20.83%"><p style="text-align:center">32.4574</p></td> 
        <td class="acenter" width="16.66%"><p style="text-align:center">&lt;0.0001</p></td> 
       </tr> 
       <tr> 
        <td class="acenter" width="16.66%"><p style="text-align:center">Prin2</p></td> 
        <td class="acenter" width="11.36%"><p style="text-align:center">1</p></td> 
        <td class="acenter" width="15.08%"><p style="text-align:center">0.7415</p></td> 
        <td class="acenter" width="19.40%"><p style="text-align:center">0.1314</p></td> 
        <td class="acenter" width="20.83%"><p style="text-align:center">31.8303</p></td> 
        <td class="acenter" width="16.66%"><p style="text-align:center">&lt;0.0001</p></td> 
       </tr> 
       <tr> 
        <td class="acenter" width="16.66%"><p style="text-align:center">Prin3</p></td> 
        <td class="acenter" width="11.36%"><p style="text-align:center">1</p></td> 
        <td class="acenter" width="15.08%"><p style="text-align:center">−0.3527</p></td> 
        <td class="acenter" width="19.40%"><p style="text-align:center">0.1235</p></td> 
        <td class="acenter" width="20.83%"><p style="text-align:center">8.1559</p></td> 
        <td class="acenter" width="16.66%"><p style="text-align:center">0.0043</p></td> 
       </tr> 
       <tr> 
        <td class="acenter" width="16.66%"><p style="text-align:center">Prin4</p></td> 
        <td class="acenter" width="11.36%"><p style="text-align:center">1</p></td> 
        <td class="acenter" width="15.08%"><p style="text-align:center">−0.3473</p></td> 
        <td class="acenter" width="19.40%"><p style="text-align:center">0.1578</p></td> 
        <td class="acenter" width="20.83%"><p style="text-align:center">4.8446</p></td> 
        <td class="acenter" width="16.66%"><p style="text-align:center">0.0277</p></td> 
       </tr> 
       <tr> 
        <td class="acenter" width="16.66%"><p style="text-align:center">Prin5</p></td> 
        <td class="acenter" width="11.36%"><p style="text-align:center">1</p></td> 
        <td class="acenter" width="15.08%"><p style="text-align:center">0.6877</p></td> 
        <td class="acenter" width="19.40%"><p style="text-align:center">0.1612</p></td> 
        <td class="acenter" width="20.83%"><p style="text-align:center">18.1996</p></td> 
        <td class="acenter" width="16.66%"><p style="text-align:center">&lt;0.0001</p></td> 
       </tr> 
      </table>
     </table-wrap>
     <p>
      <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"> <mtable> 
        <mtr> 
         <mtd> 
          <mi>
            ln 
          </mi> 
          <mrow> 
           <mo>
             ( 
           </mo> 
           <mrow> 
            <mfrac> 
             <mi>
               p 
             </mi> 
             <mrow> 
              <mn>
                1 
              </mn> 
              <mo>
                − 
              </mo> 
              <mi>
                p 
              </mi> 
             </mrow> 
            </mfrac> 
           </mrow> 
           <mo>
             ) 
           </mo> 
          </mrow> 
          <mo>
            = 
          </mo> 
          <mo>
            − 
          </mo> 
          <mn>
            2.5661 
          </mn> 
          <mo>
            + 
          </mo> 
          <mn>
            0.5689 
          </mn> 
          <mo>
            × 
          </mo> 
          <mtext>
            Prin 
          </mtext> 
          <mn>
            1 
          </mn> 
          <mo>
            + 
          </mo> 
          <mn>
            0.7415 
          </mn> 
          <mo>
            × 
          </mo> 
          <mtext>
            Prin 
          </mtext> 
          <mn>
            2 
          </mn> 
          <mo>
            − 
          </mo> 
          <mn>
            0.3527 
          </mn> 
          <mo>
            × 
          </mo> 
          <mtext>
            Prin 
          </mtext> 
          <mn>
            3 
          </mn> 
         </mtd> 
        </mtr> 
        <mtr> 
         <mtd> 
          <mtext>
              
          </mtext> 
          <mtext>
              
          </mtext> 
          <mo>
            − 
          </mo> 
          <mn>
            0.3473 
          </mn> 
          <mo>
            × 
          </mo> 
          <mtext>
            Prin 
          </mtext> 
          <mn>
            4 
          </mn> 
          <mo>
            + 
          </mo> 
          <mn>
            0.6877 
          </mn> 
          <mo>
            × 
          </mo> 
          <mtext>
            Prin 
          </mtext> 
          <mn>
            5 
          </mn> 
         </mtd> 
        </mtr> 
       </mtable> 
      </math></p>
    </sec>
    <sec id="s4_7">
     <title>4.7. Discriminant Analysis (with PCA)</title>
     <p>Using the PCA-reduced dataset, the DA model produced classification results (<xref ref-type="table" rid="table14">
       Table 14
      </xref>), and two discriminant functions (<xref ref-type="table" rid="table15">
       Table 15
      </xref>) which were derived in SAS.</p>
     <table-wrap id="table14">
      <label>
       <xref ref-type="table" rid="table14">
        Table 14
       </xref></label>
      <caption>
       <title>
        <xref ref-type="bibr" rid="scirp.145031-"></xref>Table 14. Classification summary for discriminant analysis with PCA.</title>
      </caption>
      <table class="MsoTableGrid custom-table" border="0" cellspacing="0" cellpadding="0"> 
       <tr> 
        <td class="custom-bottom-td acenter" width="115.55%" colspan="4"><p style="text-align:center">Number of Observations and Percent Classified into Diabetes</p></td> 
       </tr> 
       <tr> 
        <td class="custom-bottom-td custom-top-td acenter" width="22.14%"><p style="text-align:center">From Diabetes</p></td> 
        <td class="custom-bottom-td custom-top-td acenter" width="31.14%"><p style="text-align:center">0</p></td> 
        <td class="custom-bottom-td custom-top-td acenter" width="31.14%"><p style="text-align:center">1</p></td> 
        <td class="custom-bottom-td custom-top-td acenter" width="31.14%"><p style="text-align:center">Total</p></td> 
       </tr> 
       <tr> 
        <td class="custom-top-td acenter" width="22.14%"><p style="text-align:center">0</p></td> 
        <td class="custom-top-td acenter" width="31.14%"><p style="text-align:center">27081.82</p></td> 
        <td class="custom-top-td acenter" width="31.14%"><p style="text-align:center">6018.18</p></td> 
        <td class="custom-top-td acenter" width="31.14%"><p style="text-align:center">330100.00</p></td> 
       </tr> 
       <tr> 
        <td class="acenter" width="22.14%"><p style="text-align:center">1</p></td> 
        <td class="acenter" width="31.14%"><p style="text-align:center">1525.00</p></td> 
        <td class="acenter" width="31.14%"><p style="text-align:center">4575.00</p></td> 
        <td class="acenter" width="31.14%"><p style="text-align:center">60100.00</p></td> 
       </tr> 
       <tr> 
        <td class="acenter" width="22.14%"><p style="text-align:center">Total</p></td> 
        <td class="acenter" width="31.14%"><p style="text-align:center">28573.08</p></td> 
        <td class="acenter" width="31.14%"><p style="text-align:center">10526.92</p></td> 
        <td class="acenter" width="31.14%"><p style="text-align:center">390100.00</p></td> 
       </tr> 
       <tr> 
        <td class="acenter" width="22.14%"><p style="text-align:center">Priors</p></td> 
        <td class="acenter" width="31.14%"><p style="text-align:center">0.5</p></td> 
        <td class="acenter" width="31.14%"><p style="text-align:center">0.5</p></td> 
        <td class="acenter" width="31.14%"><p style="text-align:center"></p></td> 
       </tr> 
      </table>
     </table-wrap>
     <table-wrap id="table15">
      <label>
       <xref ref-type="table" rid="table15">
        Table 15
       </xref></label>
      <caption>
       <title>
        <xref ref-type="bibr" rid="scirp.145031-"></xref>Table 15. Linear discriminant function.</title>
      </caption>
      <table class="MsoTableGrid custom-table" border="0" cellspacing="0" cellpadding="0"> 
       <tr> 
        <td class="custom-bottom-td acenter" width="119.48%" colspan="3"><p style="text-align:center">Linear Discriminant Function for Diabetes</p></td> 
       </tr> 
       <tr> 
        <td class="custom-bottom-td custom-top-td acenter" width="39.84%"><p style="text-align:center">Variable</p></td> 
        <td class="custom-bottom-td custom-top-td acenter" width="39.82%"><p style="text-align:center">0</p></td> 
        <td class="custom-bottom-td custom-top-td acenter" width="39.82%"><p style="text-align:center">1</p></td> 
       </tr> 
       <tr> 
        <td class="custom-top-td acenter" width="39.84%"><p style="text-align:center">Constant</p></td> 
        <td class="custom-top-td acenter" width="39.82%"><p style="text-align:center">−0.03684</p></td> 
        <td class="custom-top-td acenter" width="39.82%"><p style="text-align:center">−1.11450</p></td> 
       </tr> 
       <tr> 
        <td class="acenter" width="39.84%"><p style="text-align:center">Prin1</p></td> 
        <td class="acenter" width="39.82%"><p style="text-align:center">−0.09219</p></td> 
        <td class="acenter" width="39.82%"><p style="text-align:center">0.50703</p></td> 
       </tr> 
       <tr> 
        <td class="acenter" width="39.84%"><p style="text-align:center">Prin2</p></td> 
        <td class="acenter" width="39.82%"><p style="text-align:center">−0.14026</p></td> 
        <td class="acenter" width="39.82%"><p style="text-align:center">0.77140</p></td> 
       </tr> 
       <tr> 
        <td class="acenter" width="39.84%"><p style="text-align:center">Prin3</p></td> 
        <td class="acenter" width="39.82%"><p style="text-align:center">0.07197</p></td> 
        <td class="acenter" width="39.82%"><p style="text-align:center">−0.39586</p></td> 
       </tr> 
       <tr> 
        <td class="acenter" width="39.84%"><p style="text-align:center">Prin4</p></td> 
        <td class="acenter" width="39.82%"><p style="text-align:center">0.05707</p></td> 
        <td class="acenter" width="39.82%"><p style="text-align:center">−0.31388</p></td> 
       </tr> 
       <tr> 
        <td class="acenter" width="39.84%"><p style="text-align:center">Prin5</p></td> 
        <td class="acenter" width="39.82%"><p style="text-align:center">−0.12280</p></td> 
        <td class="acenter" width="39.82%"><p style="text-align:center">0.67540</p></td> 
       </tr> 
      </table>
     </table-wrap>
     <p>
      <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"> <mtable> 
        <mtr> 
         <mtd> 
          <msub> 
           <mi>
             D 
           </mi> 
           <mn>
             0 
           </mn> 
          </msub> 
          <mrow> 
           <mo>
             ( 
           </mo> 
           <mrow> 
            <mtext>
              No 
            </mtext> 
            <mtext>
                
            </mtext> 
            <mtext>
              Diabetes 
            </mtext> 
           </mrow> 
           <mo>
             ) 
           </mo> 
          </mrow> 
          <mo>
            = 
          </mo> 
          <mo>
            − 
          </mo> 
          <mn>
            0.03684 
          </mn> 
          <mo>
            − 
          </mo> 
          <mn>
            0.09219 
          </mn> 
          <mo>
            × 
          </mo> 
          <mtext>
            Prin 
          </mtext> 
          <mn>
            1 
          </mn> 
          <mo>
            − 
          </mo> 
          <mn>
            0.14026 
          </mn> 
          <mo>
            × 
          </mo> 
          <mtext>
            Prin 
          </mtext> 
          <mn>
            2 
          </mn> 
         </mtd> 
        </mtr> 
        <mtr> 
         <mtd> 
          <mtext>
              
          </mtext> 
          <mtext>
              
          </mtext> 
          <mo>
            + 
          </mo> 
          <mn>
            0.07197 
          </mn> 
          <mo>
            × 
          </mo> 
          <mtext>
            Prin 
          </mtext> 
          <mn>
            3 
          </mn> 
          <mo>
            + 
          </mo> 
          <mn>
            0.05707 
          </mn> 
          <mo>
            × 
          </mo> 
          <mtext>
            Prin 
          </mtext> 
          <mn>
            4 
          </mn> 
          <mo>
            − 
          </mo> 
          <mn>
            0.12280 
          </mn> 
          <mo>
            × 
          </mo> 
          <mtext>
            Prin5 
          </mtext> 
         </mtd> 
        </mtr> 
       </mtable> 
      </math></p>
     <p>
      <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"> <mtable> 
        <mtr> 
         <mtd> 
          <msub> 
           <mi>
             D 
           </mi> 
           <mn>
             1 
           </mn> 
          </msub> 
          <mrow> 
           <mo>
             ( 
           </mo> 
           <mrow> 
            <mtext>
              Diabetes 
            </mtext> 
            <mtext>
                
            </mtext> 
            <mtext>
              present 
            </mtext> 
           </mrow> 
           <mo>
             ) 
           </mo> 
          </mrow> 
          <mo>
            = 
          </mo> 
          <mo>
            − 
          </mo> 
          <mn>
            1.11450 
          </mn> 
          <mo>
            + 
          </mo> 
          <mn>
            0.50703 
          </mn> 
          <mo>
            × 
          </mo> 
          <mtext>
            Prin 
          </mtext> 
          <mn>
            1 
          </mn> 
          <mo>
            + 
          </mo> 
          <mn>
            0.77140 
          </mn> 
          <mo>
            × 
          </mo> 
          <mtext>
            Prin 
          </mtext> 
          <mn>
            2 
          </mn> 
         </mtd> 
        </mtr> 
        <mtr> 
         <mtd> 
          <mtext>
              
          </mtext> 
          <mtext>
              
          </mtext> 
          <mo>
            − 
          </mo> 
          <mn>
            0.39586 
          </mn> 
          <mo>
            × 
          </mo> 
          <mtext>
            Prin 
          </mtext> 
          <mn>
            3 
          </mn> 
          <mo>
            − 
          </mo> 
          <mn>
            0.31388 
          </mn> 
          <mo>
            × 
          </mo> 
          <mtext>
            Prin 
          </mtext> 
          <mn>
            4 
          </mn> 
          <mo>
            + 
          </mo> 
          <mn>
            0.67540 
          </mn> 
          <mo>
            × 
          </mo> 
          <mtext>
            Prin 
          </mtext> 
          <mn>
            5 
          </mn> 
         </mtd> 
        </mtr> 
       </mtable> 
      </math></p>
    </sec>
    <sec id="s4_8">
     <title>4.8. Comparing both Statistical Techniques</title>
     <p>The values of the goodness of fit (−2logL, AIC, and SC) were examined for the LR model outcomes on the original diabetes dataset and after the PCA transformation. The results, presented in <xref ref-type="table" rid="table16">
       Table 16
      </xref>, showed that these values remained unchanged for both procedures, indicating that PCA did not enhance the model’s performance. PCA is typically used to improve model accuracy in datasets with highly correlated independent variables. However, in this study, PCA did not significantly affect the model performance, likely because only a few variables had high correlation values. Additionally, comparing the AIC and SC values before and after PCA transformation, which were 336.872 and 340.838, respectively, further confirmed that the procedure did not improve the model fit. In general, to achieve a better model fit, the values of −2logL, AIC, and SC should decrease after the PCA transformation, with smaller values indicating better performance.</p>
     <table-wrap id="table16">
      <label>
       <xref ref-type="table" rid="table16">
        Table 16
       </xref></label>
      <caption>
       <title>
        <xref ref-type="bibr" rid="scirp.145031-"></xref>Table 16. Goodness of fit variables.</title>
      </caption>
      <table class="MsoTableGrid custom-table" border="0" cellspacing="0" cellpadding="0"> 
       <tr> 
        <td class="custom-bottom-td acenter" width="25.86%"><p style="text-align:center"></p></td> 
        <td class="custom-bottom-td acenter" width="37.07%"><p style="text-align:center">Before PCA</p></td> 
        <td class="custom-bottom-td acenter" width="37.07%"><p style="text-align:center">After PCA</p></td> 
       </tr> 
       <tr> 
        <td class="custom-top-td acenter" width="25.86%"><p style="text-align:center">−2logL</p></td> 
        <td class="custom-top-td acenter" width="37.07%"><p style="text-align:center">334.872</p></td> 
        <td class="custom-top-td acenter" width="37.07%"><p style="text-align:center">334.872</p></td> 
       </tr> 
       <tr> 
        <td class="acenter" width="25.86%"><p style="text-align:center">AIC</p></td> 
        <td class="acenter" width="37.07%"><p style="text-align:center">336.872</p></td> 
        <td class="acenter" width="37.07%"><p style="text-align:center">336.872</p></td> 
       </tr> 
       <tr> 
        <td class="acenter" width="25.86%"><p style="text-align:center">SC</p></td> 
        <td class="acenter" width="37.07%"><p style="text-align:center">340.878</p></td> 
        <td class="acenter" width="37.07%"><p style="text-align:center">340.878</p></td> 
       </tr> 
      </table>
     </table-wrap>
     <p>
      <xref ref-type="table" rid="table17">
       Table 17
      </xref> shows the comparison of the results obtained for the DA procedure before the PCA was done and after. It was observed that the overall model accuracy reduced from 92% to 81% including specificity and sensitivity.</p>
     <table-wrap id="table17">
      <label>
       <xref ref-type="table" rid="table17">
        Table 17
       </xref></label>
      <caption>
       <title>
        <xref ref-type="bibr" rid="scirp.145031-"></xref>Table 17. Before and after DA.</title>
      </caption>
      <table class="MsoTableGrid custom-table" border="0" cellspacing="0" cellpadding="0"> 
       <tr> 
        <td class="custom-bottom-td acenter" width="34.48%"><p style="text-align:center">Performance Metric</p></td> 
        <td class="custom-bottom-td acenter" width="23.70%"><p style="text-align:center">Before PCA</p></td> 
        <td class="custom-bottom-td acenter" width="23.72%"><p style="text-align:center">After PCA</p></td> 
        <td class="custom-bottom-td acenter" width="18.09%"><p style="text-align:center"></p></td> 
       </tr> 
       <tr> 
        <td class="custom-top-td acenter" width="34.48%"><p style="text-align:center">Overall Accuracy</p></td> 
        <td class="custom-top-td acenter" width="23.70%"><p style="text-align:center">92%</p></td> 
        <td class="custom-top-td acenter" width="23.72%"><p style="text-align:center">81%</p></td> 
        <td class="custom-top-td acenter" width="18.09%"><p style="text-align:center">−11%</p></td> 
       </tr> 
       <tr> 
        <td class="acenter" width="34.48%"><p style="text-align:center">Specificity</p></td> 
        <td class="acenter" width="23.70%"><p style="text-align:center">97%</p></td> 
        <td class="acenter" width="23.72%"><p style="text-align:center">95%</p></td> 
        <td class="acenter" width="18.09%"><p style="text-align:center">−2%</p></td> 
       </tr> 
       <tr> 
        <td class="acenter" width="34.48%"><p style="text-align:center">Sensitivity</p></td> 
        <td class="acenter" width="23.70%"><p style="text-align:center">71%</p></td> 
        <td class="acenter" width="23.72%"><p style="text-align:center">43%</p></td> 
        <td class="acenter" width="18.09%"><p style="text-align:center">−28%</p></td> 
       </tr> 
       <tr> 
        <td class="acenter" width="34.48%"><p style="text-align:center">Negative Predictive Value</p></td> 
        <td class="acenter" width="23.70%"><p style="text-align:center">94%</p></td> 
        <td class="acenter" width="23.72%"><p style="text-align:center">82%</p></td> 
        <td class="acenter" width="18.09%"><p style="text-align:center">−12%</p></td> 
       </tr> 
       <tr> 
        <td class="acenter" width="34.48%"><p style="text-align:center">Positive Predictive Value</p></td> 
        <td class="acenter" width="23.70%"><p style="text-align:center">82%</p></td> 
        <td class="acenter" width="23.72%"><p style="text-align:center">75%</p></td> 
        <td class="acenter" width="18.09%"><p style="text-align:center">−7%</p></td> 
       </tr> 
      </table>
     </table-wrap>
     <p>In order to compare the accuracy of the Logistic Regression and Discriminant Analysis and check which procedure is better, cross-validation procedure was</p>
     <fig id="fig7" position="float">
      <label>Figure 7</label>
      <caption>
       <title>Figure 7. Snapshot of cross validation result for Logistic Regression.</title>
      </caption>
      <graphic mimetype="image" position="float" xlink:type="simple" xlink:href="https://html.scirp.org/file/1442492-rId36.jpeg?20250822030628" />
     </fig>
     <fig id="fig8" position="float">
      <label>Figure 8</label>
      <caption>
       <title>Figure 8. Snapshot of cross validation result for discriminant analysis.</title>
      </caption>
      <graphic mimetype="image" position="float" xlink:type="simple" xlink:href="https://html.scirp.org/file/1442492-rId37.jpeg?20250822030628" />
     </fig>
     <p>done on the original diabetes dataset using Python Integrated Development Environment (code snippets for this procedure are presented in <xref ref-type="fig" rid="fig7">
       Figure 7
      </xref> and <xref ref-type="fig" rid="fig8">
       Figure 8
      </xref> respectively). From the sklearn.model selection library, Kfold and cross_val_score are imported.</p>
     <p>The result obtained for the k-fold cross validation shows a mean score of 92.67% for the LR model and 92.39% for the DA model. Both statistical techniques gave very good results.</p>
    </sec>
    <sec id="s4_9">
     <title>4.9. Implications for Clinical Practice</title>
     <p>The hazard risk for diabetes mellitus is rising <xref ref-type="bibr" rid="scirp.145031-20">
       [20]
      </xref>. Accurate prediction analysis can significantly benefit clinical practice. Other studies have indicated large waist circumference, high HgbA1C, and high Fatty Liver Index (FLI) all indicate the greater possibility of contracting diabetes <xref ref-type="bibr" rid="scirp.145031-21">
       [21]
      </xref>. Our dataset has one similar variable, the waist-to-hip ratio and others like glucose and BMI. Our list of variables may be more reflective of collection site limited data. Regardless, we are able to provide a valid prediction model.</p>
     <p>Uncontrolled glucose levels in the bloodstream can increase damage to the body. Serious damage to the eyes, kidneys, and vasculature (heart) require extensive medical intervention. Predicting the likelihood of becoming a diabetic can help direct educational resources to those most in need of reversing this trend. Education focuses on a healthy diet and frequent exercise to improve long-term health outcomes <xref ref-type="bibr" rid="scirp.145031-22">
       [22]
      </xref>. Computational intensity of neural networks and the need for large datasets, limit complex model usefulness in a clinical practice environment. More likely, these techniques can be found in data analyst environments. Given the small size of our dataset, LR and DA adequately supported prediction analysis without extensive time and costs. Having a data analyst on the healthcare team could be beneficial.</p>
    </sec>
    <sec id="s4_10">
     <title>4.10. Limitations of LR and DA</title>
     <p>There are at least four critical assumptions to consider when choosing LR and DA. We rely on assumptions about the data. Violating these assumptions can lead to inaccurate results. The assumptions are:</p>
     <p>1) Linearity of data</p>
     <p>2) Normality of residuals</p>
     <p>3) Homogeneity of residuals variance</p>
     <p>4) Independence of residuals error terms</p>
     <p>Before applying a technique, it is crucial to ensure that the assumptions are met. Regression beta coefficients and R2 can be used to tell how well the LR model fits to the data.</p>
     <p>There are potential problems to address:</p>
     <p>1) Non-linearity</p>
     <p>2) Heteroscedasticity</p>
     <p>3) Presence of influential values</p>
     <p>These problems can be checked with diagnostic plots in R software. Moderate multicollinearity may not be problematic. However, severe multicollinearity is a problem because it can increase the variance of the coefficient estimates and make the estimates very sensitive to minor changes in the model. The result is that the coefficient estimates are unstable and difficult to interpret. Multicollinearity saps the statistical power of the analysis, can cause the coefficients to switch signs, and makes it more difficult to specify the correct model (Regression Analysis blog). High predictor correlation or multicollinearity can impact the interpretability of the regression model. However, a reasonable prediction can still be made, if attention is directed toward eliminating inflated standard errors and preventing overfitting.</p>
    </sec>
   </sec>
   <sec id="s5">
    <title>5. Conclusions</title>
    <p>In this study, Principal Component Analysis (PCA) was applied as an exploratory step for training the Logistic Regression (LR) and Discriminant Analysis (DA) models. The goal was to observe whether reducing the number of input features in the form of principal components could improve model performance. However, no improvement was observed after PCA was applied. This outcome can be explained by the nature of both the dataset and the models used. The correlation matrix of the original Diabetes dataset did not reveal strongly correlated features, which reduced the need for dimensionality reduction. Although some high correlation was present between Waist and Weight (0.8478), Hip and Weight (0.8270), and Waist and Hip (0.8352), the majority of variables did not exhibit strong collinearity. Furthermore, both LR and DA are linear models and are capable of handling moderately correlated features. Applying a linear transformation such as PCA did not introduce any advantage and may have discarded useful information. As a result, PCA was not used further for developing and tuning the final models. Post-PCA correlation structure was not computed, as PCA by design generates uncorrelated principal components. Since PCA did not improve model performance, and the original dataset showed limited strong correlations, further correlation analysis post-PCA was not pursued. For the exploratory analysis, an optimal number of principal components were selected based on the scree plot presented in <xref ref-type="fig" rid="fig4">
      Figure 4
     </xref>. The values of the principal components are the loading coefficients obtained from the Maximum Likelihood Estimates and Linear Discriminant Analysis procedures performed using SAS and are presented in <xref ref-type="table" rid="table7">
      Table 7
     </xref> and <xref ref-type="table" rid="table9">
      Table 9
     </xref>.</p>
    <p>The effectiveness of Logistic Regression and Discriminant Analysis techniques in accurately predicting diabetes in individuals was demonstrated. The results obtained from our Diabetes dataset show that the use of PCA did not enhance the performance of these models. This reflects a weak relationship between variables. According to Praxis Business School, PCA is ineffectual when the correlation matrix is below 0.3 as is the case for most of the correlation coefficients that were observed. However, it is important to note that our study is limited by the relatively small size of our dataset, which may restrict the generalization of our findings.</p>
    <p>To build on these results, future research can focus on training these models to adapt to larger datasets, as this may provide more robust and reliable results. Additionally, the use of more advanced intelligent systems, such as Artificial Neural Networks (ANN) <xref ref-type="bibr" rid="scirp.145031-13">
      [13]
     </xref> and ensemble models, can be explored to further improve the accuracy of predicting diabetes in individuals <xref ref-type="bibr" rid="scirp.145031-12">
      [12]
     </xref>.</p>
    <p>Furthermore, it would be interesting to investigate the performance of these models based on gender demographic. This could provide useful insights into how diabetes manifests differently in males and females and could lead to more tailored and effective interventions for diabetes prevention and management.</p>
    <p>Overall, while our study has some limitations, our findings underscore the usefulness of Logistic Regression and Discriminant Analysis techniques in predicting diabetes and highlight the need for further research to continue to improve the accuracy and effectiveness of these models.</p>
   </sec>
   <sec id="s6">
    <title>Author Contributions</title>
    <p>E.D., F.I. and M.O. contributed equally to the conceptualization and methodology of this original research. All authors have read and agreed to the published version of the manuscript.</p>
   </sec>
   <sec id="s7">
    <title>Funding</title>
    <p>The authors would like to express their gratitude for the funding support from the Department of Industrial and Systems Engineering, North Carolina Agricultural and Technical State University.</p>
   </sec>
  </sec>
 </body><back>
  <ref-list>
   <title>References</title>
   <ref id="scirp.145031-ref1">
    <label>1</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Centers for Disease Control (2024) CDC Diabetes Basics. &gt;https://www.cdc.gov/diabetes/about/index.html 
    </mixed-citation>
   </ref>
   <ref id="scirp.145031-ref2">
    <label>2</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Gwira, J.A., Fryar, C.D. and Gu, Q.P. (2024) Prevalence of Total, Diagnosed, and Undiagnosed Diabetes in Adults: United States, August 2021-August 2023. NCHS Data Briefs, National Center for Health Statistics (US). 
    </mixed-citation>
   </ref>
   <ref id="scirp.145031-ref3">
    <label>3</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Centers for Disease Control (2024) National Diabetes Statistics Report. &gt;https://www.cdc.gov/diabetes/php/data-research/methods.html?CDC_AAref_Val=&gt;https://www.cdc.gov/diabetes/data/statistics-report/index.html?ACSTrackingID=DM72996&amp;ACSTrackingLabel=New%2520Report%2520Shares%2520Latest%2520Diabetes%2520Stats%2520&amp;deliveryName=DM72996
    </mixed-citation>
   </ref>
   <ref id="scirp.145031-ref4">
    <label>4</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Walker, R.J., Garacci, E., Ozieh, M. and Egede, L.E. (2021) Food Insecurity and Glycemic Control in Individuals with Diagnosed and Undiagnosed Diabetes in the United States. Primary Care Diabetes, 15, 813-818. &gt;https://doi.org/10.1016/j.pcd.2021.05.003
    </mixed-citation>
   </ref>
   <ref id="scirp.145031-ref5">
    <label>5</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Berkowitz, S.A., Karter, A.J., Corbie-Smith, G., Seligman, H.K., Ackroyd, S.A., Barnard, L.S., et al. (2018) Food Insecurity, Food “Deserts,” and Glycemic Control in Patients with Diabetes: A Longitudinal Analysis. Diabetes Care, 41, 1188-1195. &gt;https://doi.org/10.2337/dc17-1981 
    </mixed-citation>
   </ref>
   <ref id="scirp.145031-ref6">
    <label>6</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Abdulhadi, N. and Al-Mousa, A. (2021) Diabetes Detection Using Machine Learning Classification Methods. 2021 International Conference on Information Technology (ICIT), Amman, 14-15 July 2021, 350-354. &gt;https://doi.org/10.1109/icit52682.2021.9491788
    </mixed-citation>
   </ref>
   <ref id="scirp.145031-ref7">
    <label>7</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Zhu, C., Idemudia, C.U. and Feng, W. (2019) Improved Logistic Regression Model for Diabetes Prediction by Integrating PCA and K-Means Techniques. Informatics in Medicine Unlocked, 17, Article 100179. &gt;https://doi.org/10.1016/j.imu.2019.100179
    </mixed-citation>
   </ref>
   <ref id="scirp.145031-ref8">
    <label>8</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Rajendra, P. and Latifi, S. (2021) Prediction of Diabetes Using Logistic Regression and Ensemble Techniques. Computer Methods and Programs in Biomedicine Update, 1, Article 100032. &gt;https://doi.org/10.1016/j.cmpbup.2021.100032
    </mixed-citation>
   </ref>
   <ref id="scirp.145031-ref9">
    <label>9</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Joshi, R.D. and Dhakal, C.K. (2021) Predicting Type 2 Diabetes Using Logistic Regression and Machine Learning Approaches. International Journal of Environmental Research and Public Health, 18, Article 7346. &gt;https://doi.org/10.3390/ijerph18147346
    </mixed-citation>
   </ref>
   <ref id="scirp.145031-ref10">
    <label>10</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Polat, K., Güneş, S. and Arslan, A. (2008) A Cascade Learning System for Classification of Diabetes Disease: Generalized Discriminant Analysis and Least Square Support Vector Machine. Expert Systems with Applications, 34, 482-487. &gt;https://doi.org/10.1016/j.eswa.2006.09.012
    </mixed-citation>
   </ref>
   <ref id="scirp.145031-ref11">
    <label>11</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Alharan, A.F.H., Algelal, Z.M., Ali, N.S. and Al-Garaawi, N. (2021) Improving Classification Performance for Diabetes with Linear Discriminant Analysis and Genetic Algorithm. 2021 Palestinian International Conference on Information and Communication Technology (PICICT), Gaza, 28-29 September 2021, 38-44. &gt;https://doi.org/10.1109/picict53635.2021.00019
    </mixed-citation>
   </ref>
   <ref id="scirp.145031-ref12">
    <label>12</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     El Jerjawi, N.S. and Abu-Naser, S.S. (2018) Diabetes Prediction Using Artificial Neural Network. International Journal of Advanced Science and Technology, 121, 55-64.
    </mixed-citation>
   </ref>
   <ref id="scirp.145031-ref13">
    <label>13</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Srivastava, S., Sharma, L., Sharma, V., Kumar, A. and Darbari, H. (2018) Prediction of Diabetes Using Artificial Neural Network Approach. In: Ray, K., Sharan, S., Rawat, S., Jain, S., Srivastava, S. and Bandyopadhyay, A., Eds., Lecture Notes in Electrical Engineering, Springer, 679-687. &gt;https://doi.org/10.1007/978-981-13-1642-5_59
    </mixed-citation>
   </ref>
   <ref id="scirp.145031-ref14">
    <label>14</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Tabaei, B.P. and Herman, W.H. (2002) A Multivariate Logistic Regression Equation to Screen for Diabetes. Diabetes Care, 25, 1999-2003. &gt;https://doi.org/10.2337/diacare.25.11.1999&gt;http://diabetesjournals.org/care/article-pdf/25/11/1999/588640/dc1102001999.pdf 
    </mixed-citation>
   </ref>
   <ref id="scirp.145031-ref15">
    <label>15</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Ye, Y., Xiong, Y., Zhou, Q., Wu, J., Li, X. and Xiao, X. (2020) Comparison of Machine Learning Methods and Conventional Logistic Regressions for Predicting Gestational Diabetes Using Routine Clinical Data: A Retrospective Cohort Study. Journal of Diabetes Research, 2020, Article ID: 4168340. &gt;https://doi.org/10.1155/2020/4168340
    </mixed-citation>
   </ref>
   <ref id="scirp.145031-ref16">
    <label>16</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Vihinen, M. (2012) How to Evaluate Performance of Prediction Methods? Measures and Their Interpretation in Variation Effect Analysis. BMC Genomics, 13, S2. &gt;https://doi.org/10.1186/1471-2164-13-s4-s2
    </mixed-citation>
   </ref>
   <ref id="scirp.145031-ref17">
    <label>17</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Akobeng, A.K. (2007) Understanding Diagnostic Tests 1: Sensitivity, Specificity and Predictive Values. Acta Paediatrica, 96, 338-341. &gt;https://doi.org/10.1111/j.1651-2227.2006.00180.x
    </mixed-citation>
   </ref>
   <ref id="scirp.145031-ref18">
    <label>18</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Nti, I.K., Nyarko-Boateng, O. and Aning, J. (2021) Performance of Machine Learning Algorithms with Different K Values in K-Fold Cross-Validation. International Journal of Information Technology and Computer Science, 13, 61-71.
    </mixed-citation>
   </ref>
   <ref id="scirp.145031-ref19">
    <label>19</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Rahimloo, P. and Jafarian, A. (2016) Prediction of Diabetes by Using Artificial Neural Network, Logistic Regression Statistical Model and Combination of Them. Bulletin de la Société Royale des Sciences de Liège, 85, 1148-1164. 
    </mixed-citation>
   </ref>
   <ref id="scirp.145031-ref20">
    <label>20</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Hounguè, P. and Bigirimana, A.G. (2022) Leveraging Pima Dataset to Diabetes Prediction: Case Study of Deep Neural Network. Journal of Computer and Communications, 10, 15-28. &gt;https://doi.org/10.4236/jcc.2022.1011002
    </mixed-citation>
   </ref>
   <ref id="scirp.145031-ref21">
    <label>21</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Tanaka, M., Akiyama, Y., Mori, K., Hosaka, I., Kato, K., Endo, K., et al. (2024) Predictive Modeling for the Development of Diabetes Mellitus Using Key Factors in Various Machine Learning Approaches. Diabetes Epidemiology and Management, 13, Article 100191. &gt;https://doi.org/10.1016/j.deman.2023.100191
    </mixed-citation>
   </ref>
   <ref id="scirp.145031-ref22">
    <label>22</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Feinberg, A., et al. (2018) Prescribing Food as a Specialty Drug. NEJM Catalyst.
    </mixed-citation>
   </ref>
  </ref-list>
 </back>
</article>