<?xml version="1.0" encoding="UTF-8"?><!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v3.0 20080202//EN" "http://dtd.nlm.nih.gov/publishing/3.0/journalpublishing3.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" dtd-version="3.0" xml:lang="en" article-type="research article">
 <front>
  <journal-meta>
   <journal-id journal-id-type="publisher-id">
    jdaip
   </journal-id>
   <journal-title-group>
    <journal-title>
     Journal of Data Analysis and Information Processing
    </journal-title>
   </journal-title-group>
   <issn pub-type="epub">
    2327-7211
   </issn>
   <issn publication-format="print">
    2327-7203
   </issn>
   <publisher>
    <publisher-name>
     Scientific Research Publishing
    </publisher-name>
   </publisher>
  </journal-meta>
  <article-meta>
   <article-id pub-id-type="doi">
    10.4236/jdaip.2025.132011
   </article-id>
   <article-id pub-id-type="publisher-id">
    jdaip-142785
   </article-id>
   <article-categories>
    <subj-group subj-group-type="heading">
     <subject>
      Articles
     </subject>
    </subj-group>
    <subj-group subj-group-type="Discipline-v2">
     <subject>
      Computer Science 
     </subject>
     <subject>
       Communications, Physics 
     </subject>
     <subject>
       Mathematics
     </subject>
    </subj-group>
   </article-categories>
   <title-group>
    Cumulative Link Ordinal Outcome Neural Networks: An Evaluation of Current Methodology 
   </title-group>
   <contrib-group>
    <contrib contrib-type="author" xlink:type="simple">
     <name name-style="western">
      <surname>
       Andre
      </surname>
      <given-names>
       Williams
      </given-names>
     </name>
    </contrib>
   </contrib-group> 
   <aff id="affnull">
    <addr-line>
     aChristine E. Lynn College of Nursing, Florida Atlantic University, Boca Raton, FL, USA
    </addr-line> 
   </aff> 
   <pub-date pub-type="epub">
    <day>
     11
    </day> 
    <month>
     04
    </month>
    <year>
     2025
    </year>
   </pub-date> 
   <volume>
    13
   </volume> 
   <issue>
    02
   </issue>
   <fpage>
    182
   </fpage>
   <lpage>
    198
   </lpage>
   <history>
    <date date-type="received">
     <day>
      10,
     </day>
     <month>
      April
     </month>
     <year>
      2025
     </year>
    </date>
    <date date-type="published">
     <day>
      23,
     </day>
     <month>
      April
     </month>
     <year>
      2025
     </year> 
    </date> 
    <date date-type="accepted">
     <day>
      23,
     </day>
     <month>
      May
     </month>
     <year>
      2025
     </year> 
    </date>
   </history>
   <permissions>
    <copyright-statement>
     © Copyright 2014 by authors and Scientific Research Publishing Inc. 
    </copyright-statement>
    <copyright-year>
     2014
    </copyright-year>
    <license>
     <license-p>
      This work is licensed under the Creative Commons Attribution International License (CC BY). http://creativecommons.org/licenses/by/4.0/
     </license-p>
    </license>
   </permissions>
   <abstract>
    Ordinal outcome neural networks represent an innovative and robust methodology for analyzing high-dimensional health data characterized by ordinal outcomes. This study offers a comparative analysis of several ordinal activation functions that utilize four key cumulative link functions used in ordinal regression: logit (logistic distribution), probit (normal distribution), complementary log log (cloglog), based on the Gompertz distribution, and log log (loglog), derived from the Gumbel distribution. The objective of this study was to systematically assess the performance of various activation functions in relation to the four link functions, focusing on how these functions affect the predictive accuracy of ordinal outcome neural networks. The findings indicate that the logit link consistently outperformed the other cumulative links across different datasets, achieving the highest percentage of correctly classified observations. The results of this study contribute to model selection considerations when analyzing complex health-related data, leading to more accurate predictions and an improved understanding of biomedical phenomena.
   </abstract>
   <kwd-group> 
    <kwd>
     Neural Network
    </kwd> 
    <kwd>
      Cumulative Link Model
    </kwd> 
    <kwd>
      Ordinal Outcome
    </kwd> 
    <kwd>
      Activation Function
    </kwd> 
    <kwd>
      Machine Learning
    </kwd>
   </kwd-group>
  </article-meta>
 </front>
 <body>
  <sec id="s1">
   <title>1. Introduction</title>
   <p>A substantial amount of data is currently being collected from diverse facets of our daily existence <xref ref-type="bibr" rid="scirp.142785-1">
     [1]
    </xref>. This encompasses information about dietary consumption, transactions undertaken in retail locations, imagery captured at traffic intersections, locational data derived from mobile devices, and even biometric information. In addition, we have entered an era in which data gathering is particularly abundant in domains related to health <xref ref-type="bibr" rid="scirp.142785-2">
     [2]
    </xref>.</p>
   <p>Several technological advancements have led to the development of a comprehensive framework for gathering health-related information. These include incorporating wearable devices, studies involving genomic data, and the evolution of Electronic Health Records (EHRs), among other data sources <xref ref-type="bibr" rid="scirp.142785-2">
     [2]
    </xref> <xref ref-type="bibr" rid="scirp.142785-3">
     [3]
    </xref>. This framework includes a wide range of patient information, including sociodemographic characteristics and treatment outcomes <xref ref-type="bibr" rid="scirp.142785-4">
     [4]
    </xref> <xref ref-type="bibr" rid="scirp.142785-5">
     [5]
    </xref>.</p>
   <p>This data-rich environment holds immense promise for the future of health care. This paves the way for personalized medicine, enabling the development of advanced predictive models to forecast individual patient outcomes and tailor treatment plans accordingly <xref ref-type="bibr" rid="scirp.142785-6">
     [6]
    </xref>-<xref ref-type="bibr" rid="scirp.142785-8">
     [8]
    </xref>. These models are crucial for anticipating adverse events, optimizing resource allocation, informing clinical trial designs, and guiding personalized patient counseling and research efforts <xref ref-type="bibr" rid="scirp.142785-9">
     [9]
    </xref>.</p>
   <p>However, the current data collection landscape is not without its challenges. The dynamic nature of these data presents a hurdle for many traditional statistical and machine learning methods, making it challenging to analyze and extract value from them effectively. The gathering of genetic, demographic, and clinical data leads to an analytical dataset with a large number of variables compared to the number of observations, further complicating the analysis process <xref ref-type="bibr" rid="scirp.142785-10">
     [10]
    </xref>.</p>
   <p>Much of the collected information, such as treatment outcomes, is ordinal in health-related data. Ordinal variables are categorical variables with a natural order or ranking; however, the differences between categories are not measurable <xref ref-type="bibr" rid="scirp.142785-11">
     [11]
    </xref> <xref ref-type="bibr" rid="scirp.142785-12">
     [12]
    </xref>. Examples of ordinal outcomes include:</p>
   <p>1) Treatment response (minimal, mild, moderate, moderate-severe, and severe)</p>
   <p>2) Tumor stage (I, II, III, IV)</p>
   <p>3) Severity of side effects (None, Mild, Moderate, Severe, Life-threatening)</p>
   <p>4) Performance Status, ECOG (0: fully active to 5: dead)</p>
   <p>Employing ordinal outcome models is advantageous and essential for comprehensively harnessing the information related to ordinal outcomes. In prior analyses, numerous researchers transformed ordinal outcomes into a continuous or binary variable <xref ref-type="bibr" rid="scirp.142785-13">
     [13]
    </xref>. However, these methods are often not ideal because they do not consider the natural structure of the outcomes. Ordinal data must be modeled as ordinal data rather than converted into binary or continuous formats <xref ref-type="bibr" rid="scirp.142785-14">
     [14]
    </xref>. This preserves the natural representation of the variable. A more appropriate analysis can be performed by maintaining the ordinal nature of the data <xref ref-type="bibr" rid="scirp.142785-15">
     [15]
    </xref> <xref ref-type="bibr" rid="scirp.142785-16">
     [16]
    </xref>.</p>
   <p>Ordinal neural networks are viable candidates for analyzing high-dimensional health data with ordinal outcomes. These models facilitate the prediction of outcomes using a specified covariate set. Furthermore, neural networks provide a robust framework for addressing ordinal classification problems in this domain, owing to their capacity to learn complex patterns from high-dimensional data <xref ref-type="bibr" rid="scirp.142785-17">
     [17]
    </xref>. These models integrate the strengths of neural networks and traditional ordinal regression techniques. Vargas et al. <xref ref-type="bibr" rid="scirp.142785-18">
     [18]
    </xref> proposed a deep neural network for ordinal regression based on the proportional odds model, which projects patterns into a one-dimensional space using non-linear deep learning. They combined this approach with a loss function that considers class distances, demonstrating improved performance on ordinal classification problems. Kook et al. <xref ref-type="bibr" rid="scirp.142785-19">
     [19]
    </xref> <xref ref-type="bibr" rid="scirp.142785-20">
     [20]
    </xref> introduced ordinal neural network transformation models (ONTRAMs), which can handle tabular and complex data such as images while maintaining interpretability. Moayed &amp; Shell <xref ref-type="bibr" rid="scirp.142785-21">
     [21]
    </xref> compared neural networks with logistic regression to predict occupational health and safety outcomes using ordinal variables. Their study, utilizing data from construction workers, demonstrates that neural networks significantly outperform logistic regression when working with datasets comprised entirely of ordinal variables. However, the selection of the activation function, which maps the network’s output to ordered categories, significantly impacts the model’s performance and interpretability <xref ref-type="bibr" rid="scirp.142785-22">
     [22]
    </xref>.</p>
   <p>As such, this manuscript presents a comparative analysis of various ordinal activation functions based on four prominent cumulative link functions employed in ordinal regression: logit, which is based on the logistic distribution; probit, which is based on the normal distribution; complementary log log (cloglog) which is based on the Gompertz distribution; and log log (loglog) function which is based on the Gumbel distribution <xref ref-type="bibr" rid="scirp.142785-23">
     [23]
    </xref>. The logit model assumes proportional odds ratios across categories and provides a widely adopted and interpretable foundation for biomedical research <xref ref-type="bibr" rid="scirp.142785-23">
     [23]
    </xref>. The normal distribution (probit) is undoubtedly the most popular and easily understood distribution <xref ref-type="bibr" rid="scirp.142785-24">
     [24]
    </xref>. The Gompertz distribution (cloglog), which is often utilized in survival analysis and scenarios involving rapidly changing odds (e.g., rapid disease progression or treatment response), offers a viable alternative <xref ref-type="bibr" rid="scirp.142785-25">
     [25]
    </xref>. Furthermore, the Gumbel distribution (loglog), which is commonly used in extreme value theory, introduces another perspective by modeling the distribution of maximum or minimum values, which can be relevant for capturing extreme events in biomedical phenomena <xref ref-type="bibr" rid="scirp.142785-26">
     [26]
    </xref>.</p>
   <p>This study formally describes neural networks that employ the four activation functions. The four resulting models were tested on both simulated and real-world datasets <xref ref-type="bibr" rid="scirp.142785-27">
     [27]
    </xref>. The objective of this study was to systematically assess the performance of various activation functions in relation to the four link functions, focusing on how these functions affect the predictive accuracy of ordinal neural networks. Additionally, this study examines whether certain activation functions consistently outperform others across different datasets or whether their effectiveness depends on the unique features of the data.</p>
   <p>The insights gleaned from this investigation will deepen our understanding of ordinal neural networks and provide practical guidance for biomedical researchers in selecting optimal activation functions according to their specific needs. This work further illuminates the intricate relationship between activation functions, distributional assumptions, and the nature of health data, ultimately contributing to the better application of ordinal neural networks to draw valuable information from health data.</p>
  </sec><sec id="s2">
   <title>2. Materials and Methods</title>
   <p>For a given observation i, denote 
    <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
      <msub> 
       <mstyle mathvariant="bold" mathsize="normal"> 
        <mi>
          x 
        </mi> 
       </mstyle> 
       <mi>
         i 
       </mi> 
      </msub> 
      <mo>
        = 
      </mo> 
      <mrow> 
       <mo>
         ( 
       </mo> 
       <mrow> 
        <msub> 
         <mi>
           x 
         </mi> 
         <mrow> 
          <mi>
            i 
          </mi> 
          <mn>
            1 
          </mn> 
         </mrow> 
        </msub> 
        <mo>
          , 
        </mo> 
        <msub> 
         <mi>
           x 
         </mi> 
         <mrow> 
          <mi>
            i 
          </mi> 
          <mn>
            2 
          </mn> 
         </mrow> 
        </msub> 
        <mo>
          , 
        </mo> 
        <mo>
          ⋯ 
        </mo> 
        <mo>
          , 
        </mo> 
        <msub> 
         <mi>
           x 
         </mi> 
         <mrow> 
          <mi>
            i 
          </mi> 
          <mi>
            p 
          </mi> 
         </mrow> 
        </msub> 
       </mrow> 
       <mo>
         ) 
       </mo> 
      </mrow> 
     </mrow> 
    </math> as a vector of p covariates, with 
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
      <mi>
        i 
      </mi> 
      <mo>
        = 
      </mo> 
      <mn>
        1 
      </mn> 
      <mo>
        , 
      </mo> 
      <mn>
        2 
      </mn> 
      <mo>
        , 
      </mo> 
      <mo>
        ⋯ 
      </mo> 
      <mo>
        , 
      </mo> 
      <mi>
        n 
      </mi> 
     </mrow> 
    </math>. This can be interpreted in biomedical research as the vector of independent variables collected on a research participant throughout the study’s course. In addition, we assumed that an ordinal outcome would be recorded for this study. As such, we denote the outcome vector 
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
      <msub> 
       <mi>
         y 
       </mi> 
       <mrow> 
        <mi>
          i 
        </mi> 
        <mo> 
        </mo> 
       </mrow> 
      </msub> 
      <mo>
        = 
      </mo> 
      <mrow> 
       <mo>
         ( 
       </mo> 
       <mrow> 
        <msub> 
         <mi>
           y 
         </mi> 
         <mrow> 
          <mi>
            i 
          </mi> 
          <mn>
            1 
          </mn> 
         </mrow> 
        </msub> 
        <mo>
          , 
        </mo> 
        <msub> 
         <mi>
           y 
         </mi> 
         <mrow> 
          <mi>
            i 
          </mi> 
          <mn>
            2 
          </mn> 
         </mrow> 
        </msub> 
        <mo>
          , 
        </mo> 
        <mo>
          ⋯ 
        </mo> 
        <mo>
          , 
        </mo> 
        <msub> 
         <mi>
           y 
         </mi> 
         <mrow> 
          <mi>
            i 
          </mi> 
          <mi>
            J 
          </mi> 
         </mrow> 
        </msub> 
       </mrow> 
       <mo>
         ) 
       </mo> 
      </mrow> 
     </mrow> 
    </math> where 
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
      <msub> 
       <mi>
         y 
       </mi> 
       <mrow> 
        <mi>
          i 
        </mi> 
        <mi>
          j 
        </mi> 
       </mrow> 
      </msub> 
      <mo>
        = 
      </mo> 
      <mn>
        1 
      </mn> 
     </mrow> 
    </math> if, for observation i, the outcome is in the j<sup>th</sup> category, all other entries are set to 0. There are J possible levels for the outcome. An example of a relevant outcome could be the response to treatment at the following levels:</p>
   <p>1) Complete Response,</p>
   <p>2) Partial Response,</p>
   <p>3) Stable Disease,</p>
   <p>4) Progressive Disease.</p>
   <p>The ordering of categories is evident. For any given study, the goal is to use the covariate vector 
    <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
      <msub> 
       <mstyle mathvariant="bold" mathsize="normal"> 
        <mi>
          x 
        </mi> 
       </mstyle> 
       <mi>
         i 
       </mi> 
      </msub> 
     </mrow> 
    </math> to predict the outcome vector of interest, 
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
      <msub> 
       <mstyle mathvariant="bold" mathsize="normal"> 
        <mi>
          y 
        </mi> 
       </mstyle> 
       <mi>
         i 
       </mi> 
      </msub> 
     </mrow> 
    </math>, for all study participants. The aggregation of vectors for all subjects yields the covariate matrix 
    <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"> <mstyle mathvariant="bold" mathsize="normal"> 
      <mi>
        X 
      </mi> 
     </mstyle> 
    </math>, a p by n matrix where the i<sup>th</sup> column is set to 
    <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
      <msub> 
       <mi>
         x 
       </mi> 
       <mi>
         i 
       </mi> 
      </msub> 
     </mrow> 
    </math>. We also have a J by n matrix denoted 
    <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"> <mstyle mathvariant="bold" mathsize="normal"> 
      <mi>
        Y 
      </mi> 
     </mstyle> 
    </math>, where the i<sup>th</sup> column is set to 
    <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
      <msub> 
       <mi>
         y 
       </mi> 
       <mi>
         i 
       </mi> 
      </msub> 
     </mrow> 
    </math>. The goal is to develop a function 
    <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
      <mi>
        g 
      </mi> 
      <mo>
        : 
      </mo> 
      <mstyle mathvariant="bold" mathsize="normal"> 
       <mi>
         X 
       </mi> 
      </mstyle> 
      <mo>
        → 
      </mo> 
      <mstyle mathvariant="bold" mathsize="normal"> 
       <mi>
         Y 
       </mi> 
      </mstyle> 
     </mrow> 
    </math>, that aims to predict the 
    <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"> <mstyle mathvariant="bold" mathsize="normal"> 
      <mi>
        Y 
      </mi> 
     </mstyle> 
    </math> matrix using the 
    <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"> <mstyle mathvariant="bold" mathsize="normal"> 
      <mi>
        X 
      </mi> 
     </mstyle> 
    </math> matrix. To accomplish this task, an ordinal outcome neural network framework was employed as the machine learning model. Initially, cost functions based on four activation functions are presented. Subsequently, the overall architecture of the neural networks is delineated. The application of the methodology to the two datasets is described, along with the metrics collected. A description of the statistics used to compare the differences in the classification rates for the four methods is presented.</p>
   <sec id="s2_1">
    <title>2.1. Cumulative Link Based Outcome Functions</title>
    <p>
     <xref ref-type="bibr" rid="scirp.142785-"></xref>The cost function for the neural network is derived from the natural logarithm of the likelihood of a multinomial distribution ( 
     <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mi>
         log 
       </mi> 
       <mi>
         L 
       </mi> 
      </mrow> 
     </math>). This function is represented as</p>
    <p>
     <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mi>
         log 
       </mi> 
       <mi>
         L 
       </mi> 
       <mo>
         = 
       </mo> 
       <munderover> 
        <mstyle mathsize="140%" displaystyle="true"> 
         <mo>
           ∑ 
         </mo> 
        </mstyle> 
        <mrow> 
         <mi>
           i 
         </mi> 
         <mo>
           = 
         </mo> 
         <mn>
           1 
         </mn> 
        </mrow> 
        <mi>
          n 
        </mi> 
       </munderover> 
       <munderover> 
        <mstyle mathsize="140%" displaystyle="true"> 
         <mo>
           ∑ 
         </mo> 
        </mstyle> 
        <mrow> 
         <mi>
           j 
         </mi> 
         <mo>
           = 
         </mo> 
         <mn>
           1 
         </mn> 
        </mrow> 
        <mi>
          J 
        </mi> 
       </munderover> 
       <mtext>
           
       </mtext> 
       <msub> 
        <mi>
          y 
        </mi> 
        <mrow> 
         <mi>
           i 
         </mi> 
         <mi>
           j 
         </mi> 
        </mrow> 
       </msub> 
       <mo>
         × 
       </mo> 
       <mi>
         log 
       </mi> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <msub> 
          <mi>
            π 
          </mi> 
          <mrow> 
           <mi>
             i 
           </mi> 
           <mi>
             j 
           </mi> 
          </mrow> 
         </msub> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
      </mrow> 
     </math>,(1)</p>
    <p>where 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          π 
        </mi> 
        <mrow> 
         <mi>
           i 
         </mi> 
         <mi>
           j 
         </mi> 
        </mrow> 
       </msub> 
       <mo>
         = 
       </mo> 
       <mi>
         P 
       </mi> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <msub> 
          <mi>
            y 
          </mi> 
          <mrow> 
           <mi>
             i 
           </mi> 
           <mi>
             j 
           </mi> 
          </mrow> 
         </msub> 
         <mo>
           = 
         </mo> 
         <mn>
           1 
         </mn> 
         <mo>
           | 
         </mo> 
         <msub> 
          <mi>
            x 
          </mi> 
          <mi>
            i 
          </mi> 
         </msub> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
      </mrow> 
     </math>. Define the cumulative probabilities 
     <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          p 
        </mi> 
        <mrow> 
         <mi>
           i 
         </mi> 
         <mi>
           j 
         </mi> 
        </mrow> 
       </msub> 
      </mrow> 
     </math> as 
     <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          p 
        </mi> 
        <mrow> 
         <mi>
           i 
         </mi> 
         <mi>
           j 
         </mi> 
        </mrow> 
       </msub> 
       <mo>
         = 
       </mo> 
       <mi>
         P 
       </mi> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <msub> 
          <mi>
            y 
          </mi> 
          <mrow> 
           <mi>
             i 
           </mi> 
           <mi>
             j 
           </mi> 
          </mrow> 
         </msub> 
         <mo>
           = 
         </mo> 
         <mn>
           1 
         </mn> 
         <mo>
           | 
         </mo> 
         <msub> 
          <mi>
            x 
          </mi> 
          <mi>
            i 
          </mi> 
         </msub> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
       <mo>
         + 
       </mo> 
       <mi>
         P 
       </mi> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <msub> 
          <mi>
            y 
          </mi> 
          <mrow> 
           <mi>
             i 
           </mi> 
           <mi>
             j 
           </mi> 
           <mo>
             − 
           </mo> 
           <mn>
             1 
           </mn> 
          </mrow> 
         </msub> 
         <mo>
           = 
         </mo> 
         <mn>
           1 
         </mn> 
         <mo>
           | 
         </mo> 
         <msub> 
          <mi>
            x 
          </mi> 
          <mi>
            i 
          </mi> 
         </msub> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
       <mo>
         + 
       </mo> 
       <mo>
         ⋯ 
       </mo> 
       <mo>
         + 
       </mo> 
       <mi>
         P 
       </mi> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <msub> 
          <mi>
            y 
          </mi> 
          <mrow> 
           <mi>
             i 
           </mi> 
           <mn>
             1 
           </mn> 
          </mrow> 
         </msub> 
         <mo>
           = 
         </mo> 
         <mn>
           1 
         </mn> 
         <mo>
           | 
         </mo> 
         <msub> 
          <mi>
            x 
          </mi> 
          <mi>
            i 
          </mi> 
         </msub> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
      </mrow> 
     </math>. In addition, 
     <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          p 
        </mi> 
        <mrow> 
         <mi>
           i 
         </mi> 
         <mi>
           J 
         </mi> 
        </mrow> 
       </msub> 
       <mo>
         = 
       </mo> 
       <mn>
         1 
       </mn> 
      </mrow> 
     </math> and 
     <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          p 
        </mi> 
        <mrow> 
         <mi>
           i 
         </mi> 
         <mn>
           0 
         </mn> 
        </mrow> 
       </msub> 
       <mo>
         = 
       </mo> 
       <mn>
         0 
       </mn> 
      </mrow> 
     </math>. Replacing 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          π 
        </mi> 
        <mrow> 
         <mi>
           i 
         </mi> 
         <mi>
           j 
         </mi> 
        </mrow> 
       </msub> 
      </mrow> 
     </math> with the cumulative probabilities 
     <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          p 
        </mi> 
        <mrow> 
         <mi>
           i 
         </mi> 
         <mi>
           j 
         </mi> 
        </mrow> 
       </msub> 
      </mrow> 
     </math> in 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mi>
         log 
       </mi> 
       <mi>
         L 
       </mi> 
      </mrow> 
     </math> allows us to rewrite Equation (1) as</p>
    <p>
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mi>
         log 
       </mi> 
       <mi>
         L 
       </mi> 
       <mo>
         = 
       </mo> 
       <munderover> 
        <mstyle mathsize="140%" displaystyle="true"> 
         <mo>
           ∑ 
         </mo> 
        </mstyle> 
        <mrow> 
         <mi>
           i 
         </mi> 
         <mo>
           = 
         </mo> 
         <mn>
           1 
         </mn> 
        </mrow> 
        <mi>
          n 
        </mi> 
       </munderover> 
       <munderover> 
        <mstyle mathsize="140%" displaystyle="true"> 
         <mo>
           ∑ 
         </mo> 
        </mstyle> 
        <mrow> 
         <mi>
           j 
         </mi> 
         <mo>
           = 
         </mo> 
         <mn>
           1 
         </mn> 
        </mrow> 
        <mi>
          J 
        </mi> 
       </munderover> 
       <mtext>
           
       </mtext> 
       <msub> 
        <mi>
          y 
        </mi> 
        <mrow> 
         <mi>
           i 
         </mi> 
         <mi>
           j 
         </mi> 
        </mrow> 
       </msub> 
       <mo>
         × 
       </mo> 
       <mi>
         log 
       </mi> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <msub> 
          <mi>
            p 
          </mi> 
          <mrow> 
           <mi>
             i 
           </mi> 
           <mi>
             j 
           </mi> 
          </mrow> 
         </msub> 
         <mo>
           − 
         </mo> 
         <msub> 
          <mi>
            p 
          </mi> 
          <mrow> 
           <mi>
             i 
           </mi> 
           <mi>
             j 
           </mi> 
           <mo>
             − 
           </mo> 
           <mn>
             1 
           </mn> 
          </mrow> 
         </msub> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
      </mrow> 
     </math>.(2)</p>
    <p>We are primarily concerned with directly modeling cumulative probabilities, and to accomplish this, we employed cumulative link models (CLMs).</p>
    <p>CLMs are statistical models designed to analyze data with ordinal outcomes. As such, we are concerned with modeling the cumulative probabilities 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          p 
        </mi> 
        <mrow> 
         <mi>
           i 
         </mi> 
         <mi>
           j 
         </mi> 
        </mrow> 
       </msub> 
      </mrow> 
     </math>. In CLMs, 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          p 
        </mi> 
        <mrow> 
         <mi>
           i 
         </mi> 
         <mi>
           j 
         </mi> 
        </mrow> 
       </msub> 
      </mrow> 
     </math> is linked to a function of predictor variables through a link function. This approach allows for the incorporation of the covariate vector 
     <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          x 
        </mi> 
        <mi>
          i 
        </mi> 
       </msub> 
      </mrow> 
     </math> to explain the variation in the ordinal outcome while maintaining the inherent order of the response categories. The four-link functions considered in this study are as follows:</p>
    <p>1) Logit (logistic Distribution)</p>
    <p>2) Probit (normal distribution)</p>
    <p>3) Complementary log-log (Gompertz distribution)</p>
    <p>4) Log Log (Gumbel Distribution)</p>
    <p>By employing different link functions within a neural network framework, we can leverage the flexibility and power of machine-learning techniques while maintaining the structured approach of CLMs.</p>
    <p>By employing the logit link, the cumulative probabilities are modeled as:</p>
    <p>
     <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msubsup> 
        <mi>
          p 
        </mi> 
        <mrow> 
         <mi>
           i 
         </mi> 
         <mi>
           j 
         </mi> 
        </mrow> 
        <mrow> 
         <mrow> 
          <mo>
            [ 
          </mo> 
          <mn>
            1 
          </mn> 
          <mo>
            ] 
          </mo> 
         </mrow> 
        </mrow> 
       </msubsup> 
       <mo>
         = 
       </mo> 
       <mfrac> 
        <mn>
          1 
        </mn> 
        <mrow> 
         <mn>
           1 
         </mn> 
         <mo>
           + 
         </mo> 
         <msup> 
          <mtext>
            e 
          </mtext> 
          <mrow> 
           <mo>
             − 
           </mo> 
           <mrow> 
            <mo>
              ( 
            </mo> 
            <mrow> 
             <mi>
               f 
             </mi> 
             <mrow> 
              <mo>
                ( 
              </mo> 
              <mrow> 
               <msub> 
                <mi>
                  x 
                </mi> 
                <mi>
                  i 
                </mi> 
               </msub> 
              </mrow> 
              <mo>
                ) 
              </mo> 
             </mrow> 
             <mo>
               + 
             </mo> 
             <msub> 
              <mi>
                b 
              </mi> 
              <mi>
                j 
              </mi> 
             </msub> 
            </mrow> 
            <mo>
              ) 
            </mo> 
           </mrow> 
          </mrow> 
         </msup> 
        </mrow> 
       </mfrac> 
      </mrow> 
     </math>(3)</p>
    <p>where 
     <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          b 
        </mi> 
        <mi>
          j 
        </mi> 
       </msub> 
      </mrow> 
     </math> can be thought of as the intercept parameter for the j<sup>th</sup> level and 
     <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mi>
         f 
       </mi> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <msub> 
          <mi>
            x 
          </mi> 
          <mi>
            i 
          </mi> 
         </msub> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
      </mrow> 
     </math> is a non-linear function of 
     <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          x 
        </mi> 
        <mi>
          i 
        </mi> 
       </msub> 
      </mrow> 
     </math>. Considering the probit link to model cumulative probabilities leads to:</p>
    <p>
     <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msubsup> 
        <mi>
          p 
        </mi> 
        <mrow> 
         <mi>
           i 
         </mi> 
         <mi>
           j 
         </mi> 
        </mrow> 
        <mrow> 
         <mrow> 
          <mo>
            [ 
          </mo> 
          <mn>
            2 
          </mn> 
          <mo>
            ] 
          </mo> 
         </mrow> 
        </mrow> 
       </msubsup> 
       <mo>
         = 
       </mo> 
       <mstyle displaystyle="true"> 
        <mrow> 
         <msubsup> 
          <mo>
            ∫ 
          </mo> 
          <mrow> 
           <mo>
             − 
           </mo> 
           <mi>
             ∞ 
           </mi> 
          </mrow> 
          <mrow> 
           <mi>
             f 
           </mi> 
           <mrow> 
            <mo>
              ( 
            </mo> 
            <mrow> 
             <msub> 
              <mi>
                x 
              </mi> 
              <mi>
                i 
              </mi> 
             </msub> 
            </mrow> 
            <mo>
              ) 
            </mo> 
           </mrow> 
           <mo>
             + 
           </mo> 
           <msub> 
            <mi>
              b 
            </mi> 
            <mi>
              j 
            </mi> 
           </msub> 
          </mrow> 
         </msubsup> 
         <mrow> 
          <mfrac> 
           <mn>
             1 
           </mn> 
           <mrow> 
            <msqrt> 
             <mrow> 
              <mn>
                2 
              </mn> 
              <mi>
                π 
              </mi> 
             </mrow> 
            </msqrt> 
           </mrow> 
          </mfrac> 
          <msup> 
           <mtext>
             e 
           </mtext> 
           <mrow> 
            <mo>
              − 
            </mo> 
            <mfrac> 
             <mn>
               1 
             </mn> 
             <mn>
               2 
             </mn> 
            </mfrac> 
            <msup> 
             <mi>
               x 
             </mi> 
             <mn>
               2 
             </mn> 
            </msup> 
           </mrow> 
          </msup> 
          <mtext>
            d 
          </mtext> 
          <mi>
            x 
          </mi> 
         </mrow> 
        </mrow> 
       </mstyle> 
      </mrow> 
     </math>.(4)</p>
    <p>When the cloglog link is considered, we have</p>
    <p>
     <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msubsup> 
        <mi>
          p 
        </mi> 
        <mrow> 
         <mi>
           i 
         </mi> 
         <mi>
           j 
         </mi> 
        </mrow> 
        <mrow> 
         <mrow> 
          <mo>
            [ 
          </mo> 
          <mn>
            3 
          </mn> 
          <mo>
            ] 
          </mo> 
         </mrow> 
        </mrow> 
       </msubsup> 
       <mo>
         = 
       </mo> 
       <mn>
         1 
       </mn> 
       <mo>
         − 
       </mo> 
       <msup> 
        <mtext>
          e 
        </mtext> 
        <mrow> 
         <mo>
           − 
         </mo> 
         <msup> 
          <mtext>
            e 
          </mtext> 
          <mrow> 
           <mi>
             f 
           </mi> 
           <mrow> 
            <mo>
              ( 
            </mo> 
            <mrow> 
             <msub> 
              <mi>
                x 
              </mi> 
              <mi>
                i 
              </mi> 
             </msub> 
            </mrow> 
            <mo>
              ) 
            </mo> 
           </mrow> 
           <mo>
             + 
           </mo> 
           <msub> 
            <mi>
              b 
            </mi> 
            <mi>
              j 
            </mi> 
           </msub> 
          </mrow> 
         </msup> 
        </mrow> 
       </msup> 
      </mrow> 
     </math>(5)</p>
    <p>Finally, utilizing the loglog link leads to:</p>
    <p>
     <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msubsup> 
        <mi>
          p 
        </mi> 
        <mrow> 
         <mi>
           i 
         </mi> 
         <mi>
           j 
         </mi> 
        </mrow> 
        <mrow> 
         <mrow> 
          <mo>
            [ 
          </mo> 
          <mn>
            4 
          </mn> 
          <mo>
            ] 
          </mo> 
         </mrow> 
        </mrow> 
       </msubsup> 
       <mo>
         = 
       </mo> 
       <msup> 
        <mtext>
          e 
        </mtext> 
        <mrow> 
         <mo>
           − 
         </mo> 
         <msup> 
          <mtext>
            e 
          </mtext> 
          <mrow> 
           <mi>
             f 
           </mi> 
           <mrow> 
            <mo>
              ( 
            </mo> 
            <mrow> 
             <msub> 
              <mi>
                x 
              </mi> 
              <mi>
                i 
              </mi> 
             </msub> 
            </mrow> 
            <mo>
              ) 
            </mo> 
           </mrow> 
           <mo>
             + 
           </mo> 
           <msub> 
            <mi>
              b 
            </mi> 
            <mi>
              j 
            </mi> 
           </msub> 
          </mrow> 
         </msup> 
        </mrow> 
       </msup> 
      </mrow> 
     </math>(6)</p>
    <p>These cumulative probabilities can be substituted into Equation (2) and solved accordingly.</p>
   </sec>
   <sec id="s2_2">
    <title>2.2. Neural Network Architecture</title>
    <p>
     <xref ref-type="bibr" rid="scirp.142785-"></xref>A neural network was used in this investigation. The neural network comprised one hidden layer with six nodes. A neural network with a single hidden layer balances simplicity and capability, leveraging universal approximation to approximate functions while minimizing complexity and reducing overfitting. The architecture is subsequently delineated. For the first layer, we applied the function</p>
    <p>
     <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msup> 
        <mi>
          Z 
        </mi> 
        <mrow> 
         <mrow> 
          <mo>
            [ 
          </mo> 
          <mn>
            1 
          </mn> 
          <mo>
            ] 
          </mo> 
         </mrow> 
        </mrow> 
       </msup> 
       <mo>
         = 
       </mo> 
       <msup> 
        <mi>
          W 
        </mi> 
        <mrow> 
         <mrow> 
          <mo>
            [ 
          </mo> 
          <mn>
            1 
          </mn> 
          <mo>
            ] 
          </mo> 
         </mrow> 
        </mrow> 
       </msup> 
       <mi>
         X 
       </mi> 
       <mo>
         + 
       </mo> 
       <msup> 
        <mi>
          b 
        </mi> 
        <mrow> 
         <mrow> 
          <mo>
            [ 
          </mo> 
          <mn>
            1 
          </mn> 
          <mo>
            ] 
          </mo> 
         </mrow> 
        </mrow> 
       </msup> 
      </mrow> 
     </math>(7)</p>
    <p>where 
     <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msup> 
        <mi>
          W 
        </mi> 
        <mrow> 
         <mrow> 
          <mo>
            [ 
          </mo> 
          <mn>
            1 
          </mn> 
          <mo>
            ] 
          </mo> 
         </mrow> 
        </mrow> 
       </msup> 
      </mrow> 
     </math> is a 6 by p matrix and 
     <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msup> 
        <mi>
          b 
        </mi> 
        <mrow> 
         <mrow> 
          <mo>
            [ 
          </mo> 
          <mn>
            1 
          </mn> 
          <mo>
            ] 
          </mo> 
         </mrow> 
        </mrow> 
       </msup> 
      </mrow> 
     </math> is a column vector of length 6. As, such 
     <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msup> 
        <mi>
          Z 
        </mi> 
        <mrow> 
         <mrow> 
          <mo>
            [ 
          </mo> 
          <mn>
            1 
          </mn> 
          <mo>
            ] 
          </mo> 
         </mrow> 
        </mrow> 
       </msup> 
      </mrow> 
     </math> is of dimension 6 by n. After this, the first activation function was applied and is defined as</p>
    <p>
     <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msup> 
        <mi>
          A 
        </mi> 
        <mrow> 
         <mrow> 
          <mo>
            [ 
          </mo> 
          <mn>
            1 
          </mn> 
          <mo>
            ] 
          </mo> 
         </mrow> 
        </mrow> 
       </msup> 
       <mo>
         = 
       </mo> 
       <msup> 
        <mi>
          g 
        </mi> 
        <mrow> 
         <mrow> 
          <mo>
            [ 
          </mo> 
          <mn>
            1 
          </mn> 
          <mo>
            ] 
          </mo> 
         </mrow> 
        </mrow> 
       </msup> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <msup> 
          <mi>
            Z 
          </mi> 
          <mrow> 
           <mrow> 
            <mo>
              [ 
            </mo> 
            <mn>
              1 
            </mn> 
            <mo>
              ] 
            </mo> 
           </mrow> 
          </mrow> 
         </msup> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
      </mrow> 
     </math>(8)</p>
    <p>where 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msup> 
        <mi>
          g 
        </mi> 
        <mrow> 
         <mrow> 
          <mo>
            [ 
          </mo> 
          <mn>
            1 
          </mn> 
          <mo>
            ] 
          </mo> 
         </mrow> 
        </mrow> 
       </msup> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mi>
          z 
        </mi> 
        <mo>
          ) 
        </mo> 
       </mrow> 
      </mrow> 
     </math> is called the Leaky Rectified Linear Unit (RELU) <xref ref-type="bibr" rid="scirp.142785-28">
      [28]
     </xref> and is defined as 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mi>
         max 
       </mi> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <mn>
           0.01 
         </mn> 
         <mi>
           z 
         </mi> 
         <mo>
           , 
         </mo> 
         <mi>
           z 
         </mi> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
      </mrow> 
     </math>. For the second layer, we applied the function</p>
    <p>
     <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msup> 
        <mi>
          Z 
        </mi> 
        <mrow> 
         <mrow> 
          <mo>
            [ 
          </mo> 
          <mn>
            2 
          </mn> 
          <mo>
            ] 
          </mo> 
         </mrow> 
        </mrow> 
       </msup> 
       <mo>
         = 
       </mo> 
       <msup> 
        <mi>
          W 
        </mi> 
        <mrow> 
         <mrow> 
          <mo>
            [ 
          </mo> 
          <mn>
            2 
          </mn> 
          <mo>
            ] 
          </mo> 
         </mrow> 
        </mrow> 
       </msup> 
       <msup> 
        <mi>
          A 
        </mi> 
        <mrow> 
         <mrow> 
          <mo>
            [ 
          </mo> 
          <mn>
            1 
          </mn> 
          <mo>
            ] 
          </mo> 
         </mrow> 
        </mrow> 
       </msup> 
       <mo>
         + 
       </mo> 
       <msup> 
        <mi>
          b 
        </mi> 
        <mrow> 
         <mrow> 
          <mo>
            [ 
          </mo> 
          <mn>
            2 
          </mn> 
          <mo>
            ] 
          </mo> 
         </mrow> 
        </mrow> 
       </msup> 
      </mrow> 
     </math>(9)</p>
    <p>where 
     <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msup> 
        <mi>
          W 
        </mi> 
        <mrow> 
         <mrow> 
          <mo>
            [ 
          </mo> 
          <mn>
            2 
          </mn> 
          <mo>
            ] 
          </mo> 
         </mrow> 
        </mrow> 
       </msup> 
      </mrow> 
     </math> is a J-1 by 6 matrix and 
     <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msup> 
        <mi>
          b 
        </mi> 
        <mrow> 
         <mrow> 
          <mo>
            [ 
          </mo> 
          <mn>
            2 
          </mn> 
          <mo>
            ] 
          </mo> 
         </mrow> 
        </mrow> 
       </msup> 
      </mrow> 
     </math> is a column vector of length J-1 with the j<sup>th</sup> element defined as 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          b 
        </mi> 
        <mi>
          j 
        </mi> 
       </msub> 
      </mrow> 
     </math> from the previous section. As such 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mi>
         f 
       </mi> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <msub> 
          <mi>
            x 
          </mi> 
          <mi>
            i 
          </mi> 
         </msub> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
      </mrow> 
     </math>, from the previous section, is defined as the value from the i<sup>th</sup> column of 
     <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msup> 
        <mi>
          W 
        </mi> 
        <mrow> 
         <mrow> 
          <mo>
            [ 
          </mo> 
          <mn>
            2 
          </mn> 
          <mo>
            ] 
          </mo> 
         </mrow> 
        </mrow> 
       </msup> 
       <msup> 
        <mi>
          A 
        </mi> 
        <mrow> 
         <mrow> 
          <mo>
            [ 
          </mo> 
          <mn>
            1 
          </mn> 
          <mo>
            ] 
          </mo> 
         </mrow> 
        </mrow> 
       </msup> 
      </mrow> 
     </math>. After this, the four link functions from Equations (3) through (6) were applied, elementwise, to 
     <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msup> 
        <mi>
          Z 
        </mi> 
        <mrow> 
         <mrow> 
          <mo>
            [ 
          </mo> 
          <mn>
            2 
          </mn> 
          <mo>
            ] 
          </mo> 
         </mrow> 
        </mrow> 
       </msup> 
      </mrow> 
     </math>. This is presented as</p>
    <p>
     <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msup> 
        <mi>
          A 
        </mi> 
        <mrow> 
         <mrow> 
          <mo>
            [ 
          </mo> 
          <mn>
            2 
          </mn> 
          <mo>
            ] 
          </mo> 
         </mrow> 
         <mrow> 
          <mo>
            [ 
          </mo> 
          <mi>
            k 
          </mi> 
          <mo>
            ] 
          </mo> 
         </mrow> 
        </mrow> 
       </msup> 
       <mo>
         = 
       </mo> 
       <msup> 
        <mi>
          p 
        </mi> 
        <mrow> 
         <mrow> 
          <mo>
            [ 
          </mo> 
          <mi>
            k 
          </mi> 
          <mo>
            ] 
          </mo> 
         </mrow> 
        </mrow> 
       </msup> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <msup> 
          <mi>
            Z 
          </mi> 
          <mrow> 
           <mrow> 
            <mo>
              [ 
            </mo> 
            <mn>
              2 
            </mn> 
            <mo>
              ] 
            </mo> 
           </mrow> 
          </mrow> 
         </msup> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
      </mrow> 
     </math>,(10)</p>
    <p>where the function 
     <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msup> 
        <mi>
          p 
        </mi> 
        <mrow> 
         <mrow> 
          <mo>
            [ 
          </mo> 
          <mi>
            k 
          </mi> 
          <mo>
            ] 
          </mo> 
         </mrow> 
        </mrow> 
       </msup> 
      </mrow> 
     </math> is the element-wise application of Equations (3) through (6) to the matrix 
     <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msup> 
        <mi>
          Z 
        </mi> 
        <mrow> 
         <mrow> 
          <mo>
            [ 
          </mo> 
          <mn>
            2 
          </mn> 
          <mo>
            ] 
          </mo> 
         </mrow> 
        </mrow> 
       </msup> 
      </mrow> 
     </math>. After this, the cost function from Equation (2) was constructed, where 
     <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          p 
        </mi> 
        <mrow> 
         <mi>
           i 
         </mi> 
         <mi>
           j 
         </mi> 
        </mrow> 
       </msub> 
       <mo>
         = 
       </mo> 
       <msubsup> 
        <mi>
          A 
        </mi> 
        <mrow> 
         <mi>
           j 
         </mi> 
         <mi>
           i 
         </mi> 
        </mrow> 
        <mrow> 
         <mrow> 
          <mo>
            [ 
          </mo> 
          <mn>
            2 
          </mn> 
          <mo>
            ] 
          </mo> 
         </mrow> 
         <mrow> 
          <mo>
            [ 
          </mo> 
          <mi>
            k 
          </mi> 
          <mo>
            ] 
          </mo> 
         </mrow> 
        </mrow> 
       </msubsup> 
      </mrow> 
     </math>.</p>
    <p>
     <xref ref-type="bibr" rid="scirp.142785-"></xref>To obtain the optimal values for 
     <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msup> 
        <mi>
          W 
        </mi> 
        <mrow> 
         <mrow> 
          <mo>
            [ 
          </mo> 
          <mn>
            1 
          </mn> 
          <mo>
            ] 
          </mo> 
         </mrow> 
        </mrow> 
       </msup> 
      </mrow> 
     </math>, 
     <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msup> 
        <mi>
          b 
        </mi> 
        <mrow> 
         <mrow> 
          <mo>
            [ 
          </mo> 
          <mn>
            1 
          </mn> 
          <mo>
            ] 
          </mo> 
         </mrow> 
        </mrow> 
       </msup> 
      </mrow> 
     </math>, 
     <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msup> 
        <mi>
          W 
        </mi> 
        <mrow> 
         <mrow> 
          <mo>
            [ 
          </mo> 
          <mn>
            2 
          </mn> 
          <mo>
            ] 
          </mo> 
         </mrow> 
        </mrow> 
       </msup> 
      </mrow> 
     </math>, and 
     <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msup> 
        <mi>
          b 
        </mi> 
        <mrow> 
         <mrow> 
          <mo>
            [ 
          </mo> 
          <mn>
            2 
          </mn> 
          <mo>
            ] 
          </mo> 
         </mrow> 
        </mrow> 
       </msup> 
      </mrow> 
     </math>, the Adam optimization algorithm was applied <xref ref-type="bibr" rid="scirp.142785-11">
      [11]
     </xref> <xref ref-type="bibr" rid="scirp.142785-29">
      [29]
     </xref>. Once these optimal parameters were obtained, they were used to evaluate the neural network’s predictive capabilities on validation data. Hyperparameter tuning was not performed.</p>
   </sec>
   <sec id="s2_3">
    <title>2.3. Application to Simulated Data</title>
    <p>The simulated data procedure has been previously introduced <xref ref-type="bibr" rid="scirp.142785-11">
      [11]
     </xref>. The data were simulated according to four covariate scenarios. The covariate scenarios are as follows:</p>
    <p>1) Autoregressive (1)</p>
    <p>2) Compound Symmetric</p>
    <p>3) Toeplitz</p>
    <p>4) Unstructured</p>
    <p>
     <xref ref-type="bibr" rid="scirp.142785-"></xref>For each scenario, a total of 100 datasets were simulated. Each dataset was divided into subsets comprising 80% for training and 20% for validation. Each dataset was simulated to yield an ordinal outcome comprising four categories (levels). Due to the nature of the simulation, the sample size for each category varies, resulting in an imbalance among the classes in most cases. The four neural networks, predicated on the four link functions delineated in Equations (3) through (6), were implemented on each training dataset. After optimizing these neural networks with respect to their parameters of interest, the parameters were subsequently applied to the validation data, and the percentage of correctly classified observations was documented. The medians and interquartile ranges are reported for each neural network employed within each covariate scenario. Assuming and verifying non-normality via a Kolmogorov-Smirnov test, the Kruskal-Wallis test was used to determine the equality of medians among the four neural networks applied within each scenario level. The Wilcoxon rank-sum test was implemented to facilitate pairwise comparisons of the neural networks within each data scenario. Furthermore, the percentage of observations correctly classified per forward propagation iteration of the Adam optimizer was reported for a subset of the training data. All analyses were conducted using R statistical software <xref ref-type="bibr" rid="scirp.142785-30">
      [30]
     </xref>.</p>
   </sec>
   <sec id="s2_4">
    <title>2.4. Application to Genomic Data</title>
    <p>Hepatocellular carcinoma (HCC) is the most common type of primary liver cancer, accounting for approximately 90% of all liver cancers. It is the fifth most common cancer in men and the seventh most common cancer in women globally, with over half a million new cases diagnosed annually <xref ref-type="bibr" rid="scirp.142785-31">
      [31]
     </xref> <xref ref-type="bibr" rid="scirp.142785-32">
      [32]
     </xref>. The incidence and mortality rates of hepatocellular carcinoma vary significantly across different geographical regions, with the highest rates observed in sub-Saharan Africa and Southeast Asia, where hepatitis B virus infection is endemic <xref ref-type="bibr" rid="scirp.142785-31">
      [31]
     </xref> <xref ref-type="bibr" rid="scirp.142785-33">
      [33]
     </xref>. In contrast, while increasing, the incidence and mortality rates in Europe and the United States remain relatively low compared with other regions <xref ref-type="bibr" rid="scirp.142785-31">
      [31]
     </xref>.</p>
    <p>Several factors contribute to the development of hepatocellular carcinoma, including viral hepatitis (hepatitis B and C), cirrhosis, nonalcoholic fatty liver disease, and various genetic and environmental factors <xref ref-type="bibr" rid="scirp.142785-34">
      [34]
     </xref> <xref ref-type="bibr" rid="scirp.142785-35">
      [35]
     </xref>. Understanding the epidemiology, risk factors, and treatment strategies of hepatocellular carcinoma is crucial for improving patient outcomes and reducing the burden of this disease worldwide.</p>
    <p>
     <xref ref-type="bibr" rid="scirp.142785-"></xref>The real-world dataset employed in this study to predict an HCC-related ordinal outcome is a subset of the genomic dataset found in the Gene Expression Omnibus data repository labeled GSE18081 <xref ref-type="bibr" rid="scirp.142785-27">
      [27]
     </xref>. The data has a sample size of 56. The ordinal outcome of interest is the classification of tissue samples as normal (n = 20), pre-malignant (cirrhosis; n = 16), or malignant (hepatocellular carcinoma; n = 20) using a small set of 46 molecular features. The study utilized data from the BeadArray Cancer Panel I. All technical replicate samples and matched cirrhotic samples from subjects with HCC were removed from the dataset. As such, all the samples in the reduced dataset were independent. The dataset is called the hccframe and is located in the ordinalgmifs R package <xref ref-type="bibr" rid="scirp.142785-27">
      [27]
     </xref>. Data were obtained using the following R code:</p>
    <p>&gt;install.packages("ordinalgmifs")</p>
    <p>&gt;library(ordinalgmifs)</p>
    <p>&gt;data(hccframe).</p>
    <p>
     <xref ref-type="bibr" rid="scirp.142785-"></xref>The input features were standardized to be between zero and one. To evaluate the neural networks against the dataset, a 10-fold cross-validation approach was employed. The dataset was randomly partitioned into ten segments. Nine segments (approximately 90%) were utilized for training the neural network, and the remaining 10% served as validation data. Subsequently, the percentage of correctly classified instances across the entire dataset was computed. This procedure was executed 100 times for each neural network. The median and interquartile range of 100 percent of the correctly classified observations were subsequently calculated for each neural network. Upon assuming and confirming non-normality via the Kolmogorov-Smirnov test, the Kruskal-Wallis test was employed to assess the equality of medians among the four neural networks. The Wilcoxon rank-sum test was also conducted to facilitate pairwise comparisons among the neural networks.</p>
   </sec>
  </sec><sec id="s3">
   <title>3. Results</title>
   <sec id="s3_1">
    <title>3.1. Simulated Data</title>
    <p>The proposed methodology, comprising four cumulative link models, was applied to the simulated data. The objective was to compare the performance of the four distinct cumulative link models. The models were evaluated based on four criteria: optimization, median percent correctly classified, IQR, and nonparametric tests to determine statistically significant differences across groups.</p>
    <p>
     <xref ref-type="fig" rid="fig1">
      Figure 1
     </xref> visually presents the accuracy (percentage of observations correctly classified) achieved by different models when classifying observations across iterations of the applied Adam optimization algorithm, specifically for four distinct covariate scenarios: Autoregressive (1), Compound Symmetric, Toeplitz, and Unstructured. The results are presented for the training datasets. The performance oscillated for all models across all covariate scenarios and then consistently increased for the given datasets. Once the method maximizes the proportion it can correctly estimate, it oscillates slightly around that value. The results indicated that the four models were suitable for optimizing the test datasets. The logit link model outperformed the other models regarding prediction accuracy across all covariate scenarios. However, it requires the most iterations to reach its maximum accuracy level, highlighting a trade-off between accuracy and computational complexity.</p>
    <fig id="fig1" position="float">
     <label>Figure 1</label>
     <caption>
      <title>Figure 1. Classification accuracy across iterations for four covariate scenarios: Autoregressive (1), Compound Symmetric, Toeplitz, and Unstructured-using cumulative link functions (Cloglog, Logit, Loglog, Probit). The x-axis represents the iteration number, while the y-axis showsthe percentage correctly classified.</title>
     </caption>
     <graphic mimetype="image" position="float" xlink:type="simple" xlink:href="https://html.scirp.org/file/2870805-rId128.jpeg?20250526020834" />
    </fig>
    <p>Focusing on the Autoregressive (1) covariate scenario, it is evident that the logit model, represented by the green line, exhibits the highest classification accuracy, attaining a maximum of 90% of observations correctly classified. The other models exhibited a gradual upward trend in their performance metrics. Notably, the probit and loglog models, illustrated in gray and orange, respectively, closely followed the logit model in terms of accuracy at a maximum of 86%. The cloglog (blue line) model had a maximum accuracy of 84%. It is also worth mentioning that the probit model required the most significant number of iterations to achieve optimization at 8861.</p>
    <p>Similarly, the logit model exhibits superior performance for the Compound Symmetric covariate scenario, maintaining a high percentage of correct classifications, with a maximum of 89% correctly classified. The other models demonstrated comparable performance, with a maximum correctly classified rate of 86% for the probit, 84% for the cloglog, and 83% for the loglog. The cloglog model requires the most iterations to be optimized at 8357 iterations.</p>
    <p>The logit model outperforms the others for the Toeplitz covariate scenario, with a maximum correctly classified percentage of 92%. Loglog, cloglog, and probit maximize accuracy at 85%, 83%, and 82%, respectively. The Logit requires the most iterations (15,158) to reach an optimal value of the percentage of observations correctly classified.</p>
    <p>The logit model demonstrated the highest accuracy (90%) for the unstructured covariate scenario, while the cloglog, probit, and loglog derived models exhibited slower but steady improvements with maximum accuracies of 82%, 83%, and 81%, respectively. The logit model required the most iterations to be optimized at 8725.</p>
    <p>The logit model consistently achieved the highest percentage of correctly classified data across all the scenarios. In addition, on average, the logit model requires the most iterations (9750) to attain its maximal value. In contrast, the probit model required the lowest number of iterations (6725) to achieve its maximum value.</p>
    <p>
     <xref ref-type="table" rid="table1">
      Table 1
     </xref> presents the performance of each method as applied to the validation datasets. The medians and interquartile ranges for all methods are reported. The validation data constituted 20% of the original dataset and were not used to develop the final models. The Kruskal-Wallis rank sum test, with a significance level of 0.05, was used to compare the methods’ median accuracy. The Wilcoxon rank-sum test was used for all pairwise comparisons within each covariate scenario.</p>
    <p>
     <xref ref-type="fig" rid="fig2">
      Figure 2
     </xref> graphically shows the same information as <xref ref-type="table" rid="table1">
      Table 1
     </xref>. It displays the percentage of correctly classified data across iterations for four covariate scenarios.</p>
    <p>For the Autoregressive (1) covariate scenario, the logit link model (green boxplot) demonstrated the highest accuracy, rapidly attaining a median accuracy of 87%. Other link models also achieved over 80% accuracy (81% - 82%) but exhibited significantly lower accuracy than the logit model. Furthermore, the logit model displayed fewer outlying values than the other three cumulative link models. For</p>
    <table-wrap id="table1">
     <label>
      <xref ref-type="table" rid="table1">
       Table 1
      </xref></label>
     <caption>
      <title>
       <xref ref-type="bibr" rid="scirp.142785-"></xref>Table 1. Comparison of the four neural networks in ordinal outcome classification.</title>
     </caption>
     <table class="MsoTableGrid custom-table" border="0" cellspacing="0" cellpadding="0"> 
      <tr> 
       <td rowspan="2" class="custom-top-td aleft"><p style="text-align:left"></p></td> 
       <td class="custom-bottom-td custom-top-td acenter" colspan="4"><p style="text-align:center">Cumulative Link</p></td> 
       <td rowspan="2" class="custom-top-td acenter"><p style="text-align:center">p-value<sup>1</sup></p></td> 
      </tr> 
      <tr> 
       <td class="custom-bottom-td custom-top-td acenter"><p style="text-align:center">Cloglog</p></td> 
       <td class="custom-bottom-td custom-top-td acenter"><p style="text-align:center">Logit</p></td> 
       <td class="custom-bottom-td custom-top-td acenter"><p style="text-align:center">Loglog</p></td> 
       <td class="custom-bottom-td custom-top-td acenter"><p style="text-align:center">Probit</p></td> 
      </tr> 
      <tr> 
       <td class="custom-top-td aleft" colspan="6"><p style="text-align:left">Autoregressive (1)</p></td> 
      </tr> 
      <tr> 
       <td class="aleft"><p style="text-align:left">Correctly Classified</p></td> 
       <td class="acenter"><p style="text-align:center"></p></td> 
       <td class="acenter"><p style="text-align:center"></p></td> 
       <td class="acenter"><p style="text-align:center"></p></td> 
       <td class="acenter"><p style="text-align:center"></p></td> 
       <td class="acenter"><p style="text-align:center">&lt;0.001</p></td> 
      </tr> 
      <tr> 
       <td class="aleft"><p style="text-align:left">Median (Q1, Q3)</p></td> 
       <td class="acenter"><p style="text-align:center">0.81 (0.77, 0.83)</p></td> 
       <td class="acenter"><p style="text-align:center">0.87 (0.85, 0.88)</p></td> 
       <td class="acenter"><p style="text-align:center">0.81 (0.78, 0.84)</p></td> 
       <td class="acenter"><p style="text-align:center">0.82 (0.80, 0.84)</p></td> 
       <td class="acenter"><p style="text-align:center"></p></td> 
      </tr> 
      <tr> 
       <td class="aleft" colspan="6"><p style="text-align:left">Compound Symmetric</p></td> 
      </tr> 
      <tr> 
       <td class="aleft"><p style="text-align:left">Correctly Classified</p></td> 
       <td class="acenter"><p style="text-align:center"></p></td> 
       <td class="acenter"><p style="text-align:center"></p></td> 
       <td class="acenter"><p style="text-align:center"></p></td> 
       <td class="acenter"><p style="text-align:center"></p></td> 
       <td class="acenter"><p style="text-align:center">&lt;0.001</p></td> 
      </tr> 
      <tr> 
       <td class="aleft"><p style="text-align:left">Median (Q1, Q3)</p></td> 
       <td class="acenter"><p style="text-align:center">0.82 (0.80, 0.83)</p></td> 
       <td class="acenter"><p style="text-align:center">0.86 (0.84, 0.87)</p></td> 
       <td class="acenter"><p style="text-align:center">0.82 (0.79, 0.84)</p></td> 
       <td class="acenter"><p style="text-align:center">0.82 (0.79, 0.84)</p></td> 
       <td class="acenter"><p style="text-align:center"></p></td> 
      </tr> 
      <tr> 
       <td class="aleft" colspan="6"><p style="text-align:left">Toeplitz</p></td> 
      </tr> 
      <tr> 
       <td class="aleft"><p style="text-align:left">Correctly Classified</p></td> 
       <td class="acenter"><p style="text-align:center"></p></td> 
       <td class="acenter"><p style="text-align:center"></p></td> 
       <td class="acenter"><p style="text-align:center"></p></td> 
       <td class="acenter"><p style="text-align:center"></p></td> 
       <td class="acenter"><p style="text-align:center">&lt;0.001</p></td> 
      </tr> 
      <tr> 
       <td class="aleft"><p style="text-align:left">Median (Q1, Q3)</p></td> 
       <td class="acenter"><p style="text-align:center">0.81 (0.78, 0.83)</p></td> 
       <td class="acenter"><p style="text-align:center">0.86 (0.83, 0.87)</p></td> 
       <td class="acenter"><p style="text-align:center">0.81 (0.79, 0.83)</p></td> 
       <td class="acenter"><p style="text-align:center">0.81 (0.79, 0.84)</p></td> 
       <td class="acenter"><p style="text-align:center"></p></td> 
      </tr> 
      <tr> 
       <td class="aleft" colspan="6"><p style="text-align:left">Unstructured</p></td> 
      </tr> 
      <tr> 
       <td class="aleft"><p style="text-align:left">Correctly Classified</p></td> 
       <td class="acenter"><p style="text-align:center"></p></td> 
       <td class="acenter"><p style="text-align:center"></p></td> 
       <td class="acenter"><p style="text-align:center"></p></td> 
       <td class="acenter"><p style="text-align:center"></p></td> 
       <td class="acenter"><p style="text-align:center">&lt;0.001</p></td> 
      </tr> 
      <tr> 
       <td class="custom-bottom-td aleft"><p style="text-align:left">Median (Q1, Q3)</p></td> 
       <td class="custom-bottom-td acenter"><p style="text-align:center">0.82 (0.79, 0.84)</p></td> 
       <td class="custom-bottom-td acenter"><p style="text-align:center">0.86 (0.85, 0.88)</p></td> 
       <td class="custom-bottom-td acenter"><p style="text-align:center">0.82 (0.80, 0.84)</p></td> 
       <td class="custom-bottom-td acenter"><p style="text-align:center">0.83 (0.80, 0.84)</p></td> 
       <td class="custom-bottom-td acenter"><p style="text-align:center"></p></td> 
      </tr> 
     </table>
    </table-wrap>
    <p><sup>1</sup>Kruskal-Wallis rank sum test. For each covariate scenario, 100 data sets were created. Each dataset was divided into 80% for training and 20% for validation. After training the four Neural Networks, the models were applied to the validation datasets. The median proportion and interquartile range of correctly classified observations for the 100 datasets are reported. The results of the Kruskal-Wallis rank-sum test are reported for the three models.</p>
    <fig id="fig2" position="float">
     <label>Figure 2</label>
     <caption>
      <title>Figure 2. Comparative Analysis of Activation Functions in Ordinal Outcome Classification Box plots illustrate the percent correctly classified across four activation functions (Cloglog, Logit, Loglog, Probit) under different correlation structures: Autoregressive (1). Compound Symmetric, Toeplitz, and Unstructured. Each box plot represents results from 100 datasets, split into 80% training and 20% validation.</title>
     </caption>
     <graphic mimetype="image" position="float" xlink:type="simple" xlink:href="https://html.scirp.org/file/2870805-rId129.jpeg?20250526020834" />
    </fig>
    <p>the Compound Symmetric covariate scenario, among the models evaluated, the logit approach exhibited the highest accuracy regarding correct classifications. The cloglog, probit, and loglog models yielded comparable results, with a median of 82% correctly classified. For the Toeplitz covariate scenario, the logit model performs superiorly to other models, with the cloglog, probit, and loglog cumulative link models exhibiting moderate improvement. For the Unstructured covariate scenario, the logit model leads in accuracy, seconded by the probit model with a median accuracy of 83%. Cloglog and loglog have a median correctly classified accuracy of 82%. The logit model consistently achieved the highest percentage of correctly classified data across all scenarios, with the fewest outliers.</p>
    <p>
     <xref ref-type="table" rid="table1">
      Table 1
     </xref> compares the performances of four cumulative link models—cloglog, logit, loglog, and probit—to correctly classify the ordinal outcome from the validation data across different covariate scenarios. The logit model consistently demonstrated the highest median accuracy across all covariate scenarios and exhibited the smallest interquartile ranges. The median of correctly classified observations for the logit model was 86% - 87%, whereas the median correct for the other cumulative link models was 81% - 83%. For all covariate scenarios, the p-value for all Kruskal-Wallis rank sum tests was less than 0.001, indicating statistically significant differences in performance among the models. Pairwise comparisons were performed to elucidate these differences. Significant pairwise differences for the Autoregressive 1 covariate scenario include: logit-cloglog at &lt;0.001, logit-loglog at &lt;0.001, logit-probit at &lt;0.001, and probit-cloglog at 0.01. Significant pairwise differences for the Compound Symmetric covariate scenario included logit-cloglog at &lt;0.001, logit-loglog at &lt;0.001, and logit-probit at &lt;0.001. Significant pairwise differences for the Unstructured covariate scenario included logit-cloglog at &lt;0.001, logit-loglog at &lt;0.001, and logit-probit at &lt;0.001. Significant pairwise differences for the Toeplitz Covariate Scenario included logit-cloglog at &lt;0.001, logit-loglog at &lt;0.001, and logit-probit at &lt;0.001. The logit model’s performance is significantly different from that of the other three models across all pairwise comparisons.</p>
   </sec>
   <sec id="s3_2">
    <title>3.2. Genomic Data</title>
    <p>The proposed methodology and the four cumulative link models were applied to the hccframe dataset. The goal was to compare the performance of the four different cumulative link models. <xref ref-type="table" rid="table2">
      Table 2
     </xref> presents the performance of each method as measured by the median and interquartile range for all methods; 10-fold cross-validation was applied 100 times, yielding 100 estimates of percent correctly classified for the dataset. Kruskal-Wallis Rank sum test, with a significance level of 0.05, was used to compare median accuracy for the methods. In addition, the Wilcoxon rank-sum test was used to compare all the pairwise comparisons. <xref ref-type="fig" rid="fig3">
      Figure 3
     </xref> visually presents the same information as <xref ref-type="table" rid="table2">
      Table 2
     </xref>. <xref ref-type="table" rid="table2">
      Table 2
     </xref> compares the performance of four cumulative link models, cloglog, logit, loglog, and probit, in correctly classifying ordinal outcomes. The logit model showed the highest median accuracy across all scenarios and the smallest interquartile range. The median number of correctly classified observations for the logit model was 86%. The loglog model comes in second with a median percentage correctly classified at</p>
    <table-wrap id="table2">
     <label>
      <xref ref-type="table" rid="table2">
       Table 2
      </xref></label>
     <caption>
      <title>
       <xref ref-type="bibr" rid="scirp.142785-"></xref>Table 2. Ordinal neural network accuracy.</title>
     </caption>
     <table class="MsoTableGrid custom-table" border="0" cellspacing="0" cellpadding="0"> 
      <tr> 
       <td rowspan="2" class="custom-top-td acenter" width="24.56%"><p style="text-align:center"></p></td> 
       <td class="custom-bottom-td custom-top-td acenter" width="64.35%" colspan="4"><p style="text-align:center">Model Type</p></td> 
       <td rowspan="2" class="custom-top-td acenter" width="11.09%"><p style="text-align:center">p-value<sup>1</sup></p></td> 
      </tr> 
      <tr> 
       <td class="custom-bottom-td custom-top-td acenter" width="16.36%"><p style="text-align:center">Cloglog</p></td> 
       <td class="custom-bottom-td custom-top-td acenter" width="16.36%"><p style="text-align:center">Logit</p></td> 
       <td class="custom-bottom-td custom-top-td acenter" width="16.36%"><p style="text-align:center">Loglog</p></td> 
       <td class="custom-bottom-td custom-top-td acenter" width="15.28%"><p style="text-align:center">Probit</p></td> 
      </tr> 
      <tr> 
       <td class="custom-top-td acenter" width="24.56%"><p style="text-align:center">Correctly Classified</p></td> 
       <td class="custom-top-td acenter" width="16.36%"><p style="text-align:center"></p></td> 
       <td class="custom-top-td acenter" width="16.36%"><p style="text-align:center"></p></td> 
       <td class="custom-top-td acenter" width="16.36%"><p style="text-align:center"></p></td> 
       <td class="custom-top-td acenter" width="15.28%"><p style="text-align:center"></p></td> 
       <td class="custom-top-td acenter" width="11.09%"><p style="text-align:center">&lt;0.001</p></td> 
      </tr> 
      <tr> 
       <td class="custom-bottom-td acenter" width="24.56%"><p style="text-align:center">Median (Q1, Q3)</p></td> 
       <td class="custom-bottom-td acenter" width="16.36%"><p style="text-align:center">0.66 (0.63, 0.70)</p></td> 
       <td class="custom-bottom-td acenter" width="16.36%"><p style="text-align:center">0.86 (0.82, 0.88)</p></td> 
       <td class="custom-bottom-td acenter" width="16.36%"><p style="text-align:center">0.71 (0.66, 0.73)</p></td> 
       <td class="custom-bottom-td acenter" width="15.28%"><p style="text-align:center">0.66 (0.64, 0.71)</p></td> 
       <td class="custom-bottom-td acenter" width="11.09%"><p style="text-align:center"></p></td> 
      </tr> 
     </table>
    </table-wrap>
    <p><sup>1</sup>Kruskal-Wallis rank sum test. The percent correctly classified for the four models (Cloglog, Logit, Loglog, and Probit) applied to the hccframe dataset. The results are based on 10-fold cross-validation, repeated 100 times.</p>
    <fig id="fig3" position="float">
     <label>Figure 3</label>
     <caption>
      <title>Figure 3. Comparative accuracy of neural network models on the hccframe dataset. Boxplots illustrate the distribution of classification accuracy for Cloglog, Logit, Loglog, and Probit models, evaluated using 10-fold cross-validation over 100 iterations.</title>
     </caption>
     <graphic mimetype="image" position="float" xlink:type="simple" xlink:href="https://html.scirp.org/file/2870805-rId130.jpeg?20250526020835" />
    </fig>
    <p>71%, with the cloglog and probit having values of 66% and 66%, respectively. The interquartile ranges were similar across all four models. The p-value for all Kruskal-Wallis rank-sum tests was less than 0.001, indicating significant differences in performance among the models. Pairwise comparisons were performed to determine these differences. All pairwise comparisons were considered statistically significant at a significance level of 0.05.</p>
   </sec>
  </sec><sec id="s4">
   <title>4. Discussion</title>
   <p>This study aimed to compare the performance of four distinct CLMs in analyzing simulated and real-world datasets, namely the cloglog, cogit, loglog, and probit. The models were assessed based on their accuracy in classifying the ordinal outcomes. The findings consistently demonstrated that the logit model outperformed the other three CLMs across all datasets. The logit model achieved the highest percentage of correctly classified data in both the simulated and real-world datasets, with all pairwise comparisons between it and the other three models being statistically significantly different. These findings suggest that the logit model effectively captures the underlying relationships between the covariates and ordinal outcomes, leading to more accurate predictions. Although the logit model requires more iterations to reach its maximum accuracy level in the simulated datasets, its superior predictive performance often outweighs this computational cost. Though requiring the fewest iterations, the probit model consistently demonstrated lower accuracy than the logit model.</p>
   <p>Several factors may contribute to the logit model’s superior performance and generally preferred biomedical research utilization:</p>
   <p>1) Proportional odds assumption: The logit model assumes proportional odds ratios across categories. This assumption implies that the relationship between covariates and cumulative probabilities remains consistent across different outcome levels.</p>
   <p>2) Broad adoption and interpretability: The logit model is widely adopted and interpretable in biomedical research and provides a familiar framework for researchers.</p>
   <p>3) Applicability to diverse datasets: The consistently high performance of the logit model across various covariate scenarios in the simulated and real-world genomic data indicates its robustness and generalizability.</p>
   <p>The findings of this study have significant implications for biomedical research. For model selection, the results provide strong evidence supporting the logit model as the primary choice when analyzing data with ordinal outcomes. Its robust performance across diverse datasets makes it a reliable and effective tool for predicting ordinal outcomes in biomedical research. This study underscores the importance of carefully considering the CLMs’ distributional assumptions. The selection of an appropriate link function, such as a logit link, can significantly impact the accuracy and interpretability of the results. Ordinal neural networks represent an innovative and robust methodology for analyzing high-dimensional health data characterized by ordinal outcomes. Researchers can significantly improve their analytical capabilities by integrating these networks with various activation functions, each tied to a different statistical link function.</p>
   <p>Combining ordinal neural networks and cumulative link activation functions equips researchers with tools to analyze health data with ordinal outcomes effectively. Future research should explore the following.</p>
   <p>1) Performance of other activation functions: This study focused on four prominent cumulative link functions. Further investigation of other activation functions may reveal additional insights and identify alternative functions that outperform the logit model in specific contexts.</p>
   <p>2) Impact of model architecture: The neural network architecture, including the number of hidden layers and nodes, can influence model performance. Exploring different architectures and delving into deep learning can further enhance the predictive accuracy of ordinal neural networks.</p>
   <p>3) Applications to other biomedical datasets: Validating this study’s findings on other biomedical datasets with ordinal outcomes is crucial for establishing the generalizability of the observed patterns and confirming the CLMs’ performance.</p>
   <p>Even though the logit cumulative link function outperforms the probit, cloglog, and loglog functions, these three functions still have merit. They can be employed in sensitivity analysis to confirm the assumption of the latent variable coming from the logistic distribution. The logit link is the most widely used option for ordinal regression in health sciences <xref ref-type="bibr" rid="scirp.142785-36">
     [36]
    </xref>, providing interpretable regression coefficients regarding odds ratios. However, some researchers may prefer the log link, which yields coefficients that are functions of probabilities rather than odds <xref ref-type="bibr" rid="scirp.142785-36">
     [36]
    </xref>.</p>
   <p>Ultimately, the appropriate link function selection should be determined by the research objectives, distributional assumptions, dataset characteristics, and interpretability of the resultant findings. In the probit link, covariates were not easily interpretable. However, aligning the latent variable to a well-known distribution benefits the probit model. The probit link assumes a normal latent distribution to interpret predictor changes. The cloglog link associates coefficients with changes in the complement of the log of the negative log of cumulative probability, which is less straightforward. By contrast, the logit link associates coefficients with log-odds changes in the outcome for a unit change in the predictor <xref ref-type="bibr" rid="scirp.142785-12">
     [12]
    </xref>, a concept most biomedical researchers understand. The loglog link represents regression coefficients as changes in the log of the negative log of cumulative probability for ordinal outcomes, holding other variables constant <xref ref-type="bibr" rid="scirp.142785-36">
     [36]
    </xref> <xref ref-type="bibr" rid="scirp.142785-37">
     [37]
    </xref>, differing from the log odds interpretation of the logit link.</p>
   <p>This study faces several limitations. The neural network architecture, featuring one hidden layer and six nodes, affects performance; different architectures and deep learning techniques could enhance accuracy. In addition, only one real-world genomic dataset and one simulated dataset were used, highlighting the need for further validation with other biomedical datasets to improve generalizability. Although the logit model performs well, it requires more iterations, indicating a trade-off between accuracy and computational cost. The logit model assumes proportional odds, which may not always hold, thus limiting its appropriateness; other models can be tested when this assumption is violated.</p>
   <p>This study contributes to the growing literature on the application of ordinal neural networks to health-related data analysis. Demonstrating the advantages of employing CLM-based neural networks provides a framework for biomedical researchers to improve the accuracy of predictive models. Future research should explore the evaluation of additional CLMs and compare their performance using the logit link.</p>
  </sec><sec id="s5">
   <title>5. Conclusion</title>
   <p>In conclusion, this study demonstrated that the logit link consistently outperforms other CLMs when analyzing simulated and real-world health data with ordinal outcomes. The findings highlight the importance of careful model selection and consideration of distributional assumptions in ordinal outcome modeling. The study’s insights contribute to model selection considerations when analyzing complex health data, leading to more accurate predictions and an improved understanding of biomedical phenomena.</p>
  </sec><sec id="s6">
   <title>Data Availability Statement</title>
   <p>The dataset is called the hccframe and is located in the ordinalgmifs R package <xref ref-type="bibr" rid="scirp.142785-27">
     [27]
    </xref>. Data were obtained using the following R code:</p>
   <p>&gt;install.packages("ordinalgmifs")</p>
   <p>&gt;library(ordinalgmifs)</p>
   <p>&gt;data(hccframe).</p>
  </sec><sec id="s7">
   <title>Acknowledgements</title>
   <p>The author expresses gratitude to the anonymous referees for their constructive feedback.</p>
  </sec><sec id="s8">
   <title>Funding</title>
   <p>The author did not receive financial support for this article’s research, writing, or publication.</p>
  </sec>
 </body><back>
  <ref-list>
   <title>References</title>
   <ref id="scirp.142785-ref1">
    <label>1</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Douglas, K. (2014) How Data Is Changing Our World. Nurse Leader, 12, 37-67. &gt;https://doi.org/10.1016/j.mnl.2014.07.003
    </mixed-citation>
   </ref>
   <ref id="scirp.142785-ref2">
    <label>2</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Francis, J.G. and Francis, L.P. (2021) Enhancing Surveillance: New Data, New Technologies, and New Actors. In: Francis, J.G. and Francis, L.P., Eds., Sustaining Surveillance: The Importance of Information for Public Health, Springer, 119-158. &gt;https://doi.org/10.1007/978-3-030-63928-0_5
    </mixed-citation>
   </ref>
   <ref id="scirp.142785-ref3">
    <label>3</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Gambhir, S.S., Ge, T.J., Vermesh, O. and Spitler, R. (2018) Toward Achieving Precision Health. Science Translational Medicine, 10, eaao3612. &gt;https://doi.org/10.1126/scitranslmed.aao3612
    </mixed-citation>
   </ref>
   <ref id="scirp.142785-ref4">
    <label>4</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Feldman, K., Johnson, R.A. and Chawla, N.V. (2018) The State of Data in Healthcare: Path Towards Standardization. Journal of Healthcare Informatics Research, 2, 248-271. &gt;https://doi.org/10.1007/s41666-018-0019-8
    </mixed-citation>
   </ref>
   <ref id="scirp.142785-ref5">
    <label>5</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Hull, S. (2015) Patient-Generated Health Data Foundation for Personalized Collaborative Care. CIN: Computers, Informatics, Nursing, 33, 177-180. &gt;https://doi.org/10.1097/cin.0000000000000159
    </mixed-citation>
   </ref>
   <ref id="scirp.142785-ref6">
    <label>6</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Cantor, M.N. (2009) Enabling Personalized Medicine through the Use of Healthcare Information Technology. Personalized Medicine, 6, 589-594. &gt;https://doi.org/10.2217/pme.09.35
    </mixed-citation>
   </ref>
   <ref id="scirp.142785-ref7">
    <label>7</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Reza Soroushmehr, S.M. and Najarian, K. (2016) Transforming Big Data into Computational Models for Personalized Medicine and Health Care. Dialogues in Clinical Neuroscience, 18, 339-343. &gt;https://doi.org/10.31887/dcns.2016.18.3/ssoroushmehr
    </mixed-citation>
   </ref>
   <ref id="scirp.142785-ref8">
    <label>8</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Fröhlich, H., Balling, R., Beerenwinkel, N., Kohlbacher, O., Kumar, S., Lengauer, T., et al. (2018) From Hype to Reality: Data Science Enabling Personalized Medicine. BMC Medicine, 16, Article No. 150. &gt;https://doi.org/10.1186/s12916-018-1122-7
    </mixed-citation>
   </ref>
   <ref id="scirp.142785-ref9">
    <label>9</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Amarasingham, R., Patzer, R.E., Huesch, M., Nguyen, N.Q. and Xie, B. (2014) Implementing Electronic Health Care Predictive Analytics: Considerations and Challenges. Health Affairs, 33, 1148-1154. &gt;https://doi.org/10.1377/hlthaff.2014.0352
    </mixed-citation>
   </ref>
   <ref id="scirp.142785-ref10">
    <label>10</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Feldman, K., Faust, L., Wu, X., Huang, C. and Chawla, N.V. (2017) Beyond Volume: The Impact of Complex Healthcare Data on the Machine Learning Pipeline. In: Holzinger, A., Goebel, R., Ferri, M. and Palade, V., Eds., Towards Integrative Machine Learning and Knowledge Extraction, Springer, 150-169. &gt;https://doi.org/10.1007/978-3-319-69775-8_9
    </mixed-citation>
   </ref>
   <ref id="scirp.142785-ref11">
    <label>11</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Williams, A.A.A. (2019) Ordinal Outcome Modeling: The Application of the Adaptive Moment Estimation Optimizer to the Elastic Net Penalized Stereotype Logit. Journal of Data Analysis and Information Processing, 7, 14-27. &gt;https://doi.org/10.4236/jdaip.2019.71002
    </mixed-citation>
   </ref>
   <ref id="scirp.142785-ref12">
    <label>12</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Agresti, A. (2010). Analysis of Ordinal Categorical Data. Wiley. &gt;https://doi.org/10.1002/9780470594001
    </mixed-citation>
   </ref>
   <ref id="scirp.142785-ref13">
    <label>13</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Jakobsson, U. (2004) Statistical Presentation and Analysis of Ordinal Data in Nursing Research. Scandinavian Journal of Caring Sciences, 18, 437-440. &gt;https://doi.org/10.1111/j.1471-6712.2004.00305.x
    </mixed-citation>
   </ref>
   <ref id="scirp.142785-ref14">
    <label>14</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Larasati, A., DeYong, C. and Slevitch, L. (2011) Comparing Neural Network and Ordinal Logistic Regression to Analyze Attitude Responses. Service Science, 3, 304-312. &gt;https://doi.org/10.1287/serv.3.4.304
    </mixed-citation>
   </ref>
   <ref id="scirp.142785-ref15">
    <label>15</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Verhulst, B. and Neale, M.C. (2021) Best Practices for Binary and Ordinal Data Analyses. Behavior Genetics, 51, 204-214. &gt;https://doi.org/10.1007/s10519-020-10031-x
    </mixed-citation>
   </ref>
   <ref id="scirp.142785-ref16">
    <label>16</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Nwankpa, C., Ijomah, W., Gachagan, A., et al. (2018) Activation Functions: Comparison of Trends in Practice and Research for Deep Learning. arXiv: 1811.03378
    </mixed-citation>
   </ref>
   <ref id="scirp.142785-ref17">
    <label>17</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Wang, L. and Zhu, D. (2021) Tackling Ordinal Regression Problem for Heterogeneous Data: Sparse and Deep Multi-Task Learning Approaches. Data Mining and Knowledge Discovery, 35, 1134-1161. &gt;https://doi.org/10.1007/s10618-021-00746-8
    </mixed-citation>
   </ref>
   <ref id="scirp.142785-ref18">
    <label>18</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Vargas, V.M., Gutiérrez, P.A. and Hervás, C. (2019) Deep Ordinal Classification Based on the Proportional Odds Model. In: Ferrández Vicente, J., Álvarez-Sánchez, J., de la Paz López, F., Toledo Moreo, J. and Adeli, H., Eds., From Bioinspired Systems and Biomedical Applications to Machine Learning, Springer, 441-451. &gt;https://doi.org/10.1007/978-3-030-19651-6_43
    </mixed-citation>
   </ref>
   <ref id="scirp.142785-ref19">
    <label>19</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Kook, L., Herzog, L., Hothorn, T., et al (2020) Ordinal Neural Network Transformation Models: Deep and Interpretable Regression Models for Ordinal Outcomes. arXiv: 2010.08376&gt;https://api.semanticscholar.org/CorpusID:223953349 
    </mixed-citation>
   </ref>
   <ref id="scirp.142785-ref20">
    <label>20</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Kook, L., Herzog, L., Hothorn, T., Dürr, O. and Sick, B. (2022) Deep and Interpretable Regression Models for Ordinal Outcomes. Pattern Recognition, 122, Article ID: 108263. &gt;https://doi.org/10.1016/j.patcog.2021.108263
    </mixed-citation>
   </ref>
   <ref id="scirp.142785-ref21">
    <label>21</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Moayed, F.A. and Shell, R.L. (2010) Application of Artificial Neural Network Models in Occupational Safety and Health Utilizing Ordinal Variables. The Annals of Occupational Hygiene, 55, 132-142.
    </mixed-citation>
   </ref>
   <ref id="scirp.142785-ref22">
    <label>22</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Vargas, V.M., Gutiérrez, P.A. and Hervás-Martínez, C. (2020) Cumulative Link Models for Deep Ordinal Classification. Neurocomputing, 401, 48-58. &gt;https://doi.org/10.1016/j.neucom.2020.03.034
    </mixed-citation>
   </ref>
   <ref id="scirp.142785-ref23">
    <label>23</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Tutz, G. (2011) Regression for Categorical Data. Cambridge University Press. &gt;https://doi.org/10.1017/cbo9780511842061
    </mixed-citation>
   </ref>
   <ref id="scirp.142785-ref24">
    <label>24</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Daykin, A.R. and Moffatt, P.G. (2002) Analyzing Ordered Responses: A Review of the Ordered Probit Model. Understanding Statistics, 1, 157-166. &gt;https://doi.org/10.1207/s15328031us0103_02
    </mixed-citation>
   </ref>
   <ref id="scirp.142785-ref25">
    <label>25</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Wilson, D.L. (1994) The Analysis of Survival (Mortality) Data: Fitting Gompertz, Weibull, and Logistic Functions. Mechanisms of Ageing and Development, 74, 15-33. &gt;https://doi.org/10.1016/0047-6374(94)90095-7
    </mixed-citation>
   </ref>
   <ref id="scirp.142785-ref26">
    <label>26</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     González, R.G., Parra, M.I., Acero, F.J., et al. (2019) An Improved Method for the Estimation of the Gumbel Distribution Parameters. arXiv: 1902.07963
    </mixed-citation>
   </ref>
   <ref id="scirp.142785-ref27">
    <label>27</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Archer, K.J., Hou, J., Zhou, Q., Ferber, K., Layne, J.G. and Gentry, A.E. (2014) Ordinalgmifs: An R Package for Ordinal Regression in High-Dimensional Data Settings. Cancer Informatics, 13, 187-195. &gt;https://doi.org/10.4137/cin.s20806
    </mixed-citation>
   </ref>
   <ref id="scirp.142785-ref28">
    <label>28</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Ramachandran, P., Zoph, B. and Le, Q.V. (2017) Searching for Activation Functions. arXiv: 171005941
    </mixed-citation>
   </ref>
   <ref id="scirp.142785-ref29">
    <label>29</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Kingma, D.P. and Ba, J. (2014) Adam: A Method for Stochastic Optimization. arXiv: 1412.6980
    </mixed-citation>
   </ref>
   <ref id="scirp.142785-ref30">
    <label>30</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     R Core Team (2023) R: A Language and Environment for Statistical Computing. R Foundation for Statistical Computing. &gt;https://www.R-project.org/ 
    </mixed-citation>
   </ref>
   <ref id="scirp.142785-ref31">
    <label>31</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Gomes, M.A., Priolli, D.G., Tralhão, J.G. and Botelho, M.F. (2013) Carcinoma hepatocelular: Epidemiologia, biologia, diagnóstico e terapias. Revista da Associação Médica Brasileira, 59, 514-524. &gt;https://doi.org/10.1016/j.ramb.2013.03.005
    </mixed-citation>
   </ref>
   <ref id="scirp.142785-ref32">
    <label>32</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Afonso Rabelo-Gonalves, E.M., Maria, B. and Robilotta Zeitune, J.M. (2014) Helicobacter pylori and Liver—Detection of Bacteria in Liver Tissue from Patients with Hepatocellular Carcinoma Using Laser Capture Microdissection Technique (LCM). In: Roesler, B.M., Ed., Trends in Helicobacter pylori Infection, InTech Open, 1. &gt;https://doi.org/10.5772/57080
    </mixed-citation>
   </ref>
   <ref id="scirp.142785-ref33">
    <label>33</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Uzzau, A., Laura, M., Cherchi, V., Bertozzi, S., Avellini, C. and Soardo, G. (2013) Surgical Treatment Strategies and Prognosis of Hepatocellular Carcinoma. In: Kaseb, A.O., Ed., Hepatocellular Carcinoma—Future Outlook, InTech Open, 1. &gt;https://doi.org/10.5772/55890
    </mixed-citation>
   </ref>
   <ref id="scirp.142785-ref34">
    <label>34</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Alberts, C.J., Clifford, G.M., Georges, D., Negro, F., Lesi, O.A., Hutin, Y.J., et al. (2022) Worldwide Prevalence of Hepatitis B Virus and Hepatitis C Virus among Patients with Cirrhosis at Country, Region, and Global Levels: A Systematic Review. The Lancet Gastroenterology&amp;Hepatology, 7, 724-735. &gt;https://doi.org/10.1016/s2468-1253(22)00050-4
    </mixed-citation>
   </ref>
   <ref id="scirp.142785-ref35">
    <label>35</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Rawla, P., Sunkara, T., Muralidharan, P. and Raj, J.P. (2018) Update in Global Trends and Aetiology of Hepatocellular Carcinoma. Współczesna Onkologia, 22, 141-150. &gt;https://doi.org/10.5114/wo.2018.78941
    </mixed-citation>
   </ref>
   <ref id="scirp.142785-ref36">
    <label>36</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Singh, G. and Hilton Fick, G. (2020) Ordinal Outcomes: A Cumulative Probability Model with the Log Link and an Assumption of Proportionality. Statistics in Medicine, 39, 1343-1361. &gt;https://doi.org/10.1002/sim.8479
    </mixed-citation>
   </ref>
   <ref id="scirp.142785-ref37">
    <label>37</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Smith, T.J., Walker, D.A. and McKenna, C.M. (2020) An Exploration of Link Functions Used in Ordinal Regression. Journal of Modern Applied Statistical Methods, 18, 2-15. &gt;https://doi.org/10.22237/jmasm/1556669640
    </mixed-citation>
   </ref>
  </ref-list>
 </back>
</article>