<?xml version="1.0" encoding="UTF-8"?><!DOCTYPE article  PUBLIC "-//NLM//DTD Journal Publishing DTD v3.0 20080202//EN" "http://dtd.nlm.nih.gov/publishing/3.0/journalpublishing3.dtd"><article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" dtd-version="3.0" xml:lang="en" article-type="research article"><front><journal-meta><journal-id journal-id-type="publisher-id">JDAIP</journal-id><journal-title-group><journal-title>Journal of Data Analysis and Information Processing</journal-title></journal-title-group><issn pub-type="epub">2327-7211</issn><publisher><publisher-name>Scientific Research Publishing</publisher-name></publisher></journal-meta><article-meta><article-id pub-id-type="doi">10.4236/jdaip.2019.71002</article-id><article-id pub-id-type="publisher-id">JDAIP-90673</article-id><article-categories><subj-group subj-group-type="heading"><subject>Articles</subject></subj-group><subj-group subj-group-type="Discipline-v2"><subject>Computer Science&amp;Communications</subject><subject> Physics&amp;Mathematics</subject></subj-group></article-categories><title-group><article-title>
 
 
  Ordinal Outcome Modeling: The Application of the Adaptive Moment Estimation Optimizer to the Elastic Net Penalized Stereotype Logit
 
</article-title></title-group><contrib-group><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>André</surname><given-names>A. A. Williams</given-names></name><xref ref-type="aff" rid="aff1"><sub>1</sub></xref><xref ref-type="corresp" rid="cor1"><sup>*</sup></xref></contrib></contrib-group><aff id="aff1"><label>1</label><addr-line>Center for Healthcare Delivery Science, Nemours Children’s Specialty Care, Jacksonville, USA</addr-line></aff><pub-date pub-type="epub"><day>14</day><month>12</month><year>2018</year></pub-date><volume>07</volume><issue>01</issue><fpage>14</fpage><lpage>27</lpage><history><date date-type="received"><day>17,</day>	<month>December</month>	<year>2018</year></date><date date-type="rev-recd"><day>19,</day>	<month>February</month>	<year>2019</year>	</date><date date-type="accepted"><day>22,</day>	<month>February</month>	<year>2019</year></date></history><permissions><copyright-statement>&#169; Copyright  2014 by authors and Scientific Research Publishing Inc. </copyright-statement><copyright-year>2014</copyright-year><license><license-p>This work is licensed under the Creative Commons Attribution International License (CC BY). http://creativecommons.org/licenses/by/4.0/</license-p></license></permissions><abstract><p>
 
 
  Penalized ordinal outcome models were developed to model high dimensional data with ordinal outcomes. One option is the penalized stereotype logit, which includes nonlinear combinations of parameter estimates. Optimization algorithms assuming linearity and function convexity were applied to fit this model. In this study the application of the adaptive moment estimation (Adam) optimizer, suited for nonlinear optimization, to the elastic net penalized stereotype logit model is proposed. The proposed model is compared to the L1 penalized ordinalgmifs stereotype model. Both methods were applied to simulated and real data, with non-Hodgkin lymphoma (NHL) cancer subtypes as the outcome, with results presented and discussed.
 
</p></abstract><kwd-group><kwd>Stereotype Logit</kwd><kwd> Elastic Net Penalty</kwd><kwd> Adam Optimizer</kwd><kwd> Non-Hodgkin Lymphoma</kwd></kwd-group></article-meta></front><body><sec id="s1"><title>1. Introduction</title><p>Many research studies seek to predict related outcomes given a set of independent variables or to quantify the relationship between them. In certain instances, the outcome of interest is ordinal. Ordinal variables are defined as having distinct ordered levels; however, the distance between the levels cannot be ascertained. An example of an ordinal variable is cancer stage. Take, for instance, testicular seminoma, a germ cell tumor in the sperm of the testes [<xref ref-type="bibr" rid="scirp.90673-ref1">1</xref>] . This cancer can be classified according to stage. The stages are:</p><p>1) Tumor stage 1, cancer has not spread beyond the testicle.</p><p>2) Tumor stage 2, cancer has spread to the blood or lymphatic vessels.</p><p>3) Tumor stage 3, cancer has spread beyond the lymphatic and blood vessels nodes to the spermatic cord.</p><p>4) Tumor stage 4, cancer has spread beyond previously mentioned areas to other parts of the body.</p><p>The ordering of categories is evident. The aim of statistical and machine learning models is to quantify the relationship between covariates and associated outcome so that one can predict the outcome variable and assess the relationship between the two with statistical significance. The range of ordinal outcome models includes cumulative logit, proportional odds model, adjacent-category logit [<xref ref-type="bibr" rid="scirp.90673-ref2">2</xref>] , and stereotype logit [<xref ref-type="bibr" rid="scirp.90673-ref3">3</xref>] . These procedures assume there are more observations than independent variables, or covariates. Another assumption is the resulting parameter estimates follow a normal distribution.</p><p>In addition, we now live in an era of high dimensional data, and massive amounts of information are being collected [<xref ref-type="bibr" rid="scirp.90673-ref4">4</xref>] . These data are used to better understand and analyze related issues. However, this comes at a cost, and traditional methods are ill-equipped to utilize these datasets. These data may have more variables than observations. The distributions of the parameter estimates may not follow a normal distribution. We may collect genetic data, demographic data, and clinical data, resulting in an analysis data set containing thousands of variables with a few hundred observations when evaluating health conditions with associated ordinal outcomes [<xref ref-type="bibr" rid="scirp.90673-ref5">5</xref>] ; the distribution of the parameter estimates may not follow a normal, or other known, distribution.</p><p>Penalized ordinal outcome models were developed to analyze high dimensional data with ordinal outcomes. Some of these modeling schemes are glmnetcr [<xref ref-type="bibr" rid="scirp.90673-ref6">6</xref>] , ordinalgmifs [<xref ref-type="bibr" rid="scirp.90673-ref7">7</xref>] , and penalized stereotype logit models [<xref ref-type="bibr" rid="scirp.90673-ref8">8</xref>] . Apart from the stereotype logit, these models are linear in that the objective cost functions can be represented as a linear combination of parameters estimates; these models also assume the cost function is convex. For the stereotype logit, which includes nonlinear combinations of parameter estimates, optimization algorithms that assumed linearity, and function convexity, were applied [<xref ref-type="bibr" rid="scirp.90673-ref7">7</xref>] [<xref ref-type="bibr" rid="scirp.90673-ref9">9</xref>] . A nonlinear and nonconvex approach to optimize the cost function of the penalized stereotype logit should be explored.</p><p>This study investigates the extension of a previously developed elastic net penalized stereotype logit [<xref ref-type="bibr" rid="scirp.90673-ref8">8</xref>] . We add an elastic net penalty [<xref ref-type="bibr" rid="scirp.90673-ref10">10</xref>] to the stereotype logit model [<xref ref-type="bibr" rid="scirp.90673-ref3">3</xref>] . To optimize the penalized function, we use the Adam optimizer [<xref ref-type="bibr" rid="scirp.90673-ref11">11</xref>] which is suited to nonlinear functions. This, in turn, allows us to evaluate the prediction accuracy of the model, which we were not able to do previously. The updated modeling procedure is presented with first order derivatives, optimization procedure, and a bootstrap resampling scheme to assess variable importance. Said modeling procedure is applied to simulated and real-world datasets with reported results. The proposed method is compared with the ordinalgmifs implemented L1 penalized stereotype logit [<xref ref-type="bibr" rid="scirp.90673-ref7">7</xref>] , an existing method for analyzing high dimensional data with an ordinal outcome.</p></sec><sec id="s2"><title>2. Method</title><p>For a given observation i (there are a total of n observations), denote the outcome vector y i as ( y i 1 , y i 2 , ⋯ , y i J ) where y i j = 1 if for that observation, the outcome is in the j<sup>th</sup> category, and all other entries are set to 0. There are J possible outcomes. Denote the vector x i = ( x i 1 x i 2 ⋯ x i p ) ′ as the covariate vector consisting of p values. The log of the information entropy, based on a multinomial distribution, is represented as</p><p>L ( θ | y , x ) = 1 n ∑ i = 1 n [ ∑ j = 1 J − 1 y i j θ i j + log π J ( x i ) ] (1)</p><p>where</p><p>θ i j = log π j ( x i ) π J ( x i ) , (2)</p><p>and</p><p>π j ( x i ) = e θ i j ∑ j = 1 J e θ i j . (3)</p><p>The log of the odds ratio, with level J being the reference level, θ i j , is represented as α j + ϕ j { x ′ i β } . Therefore, π j ( x i ) are now modeled as</p><p>exp ( α j + ϕ j x ′ i β ) 1 + ∑ j ' = 1 J − 1 exp ( α j ′ + ϕ j ′ x ′ i β ) . (4)</p><p>This representation is known as the stereotype logit [<xref ref-type="bibr" rid="scirp.90673-ref3">3</xref>] . For each ordered level, the effect of the independent variables is equal to an overall effect, x ′ i β , multiplied by a value ϕ j , which is referred to as the intensity parameter. Primarily, as we are concerned with modeling the log of the odds ratios, θ i j , the log of the information entropy can now be written as</p><p>L ( β , α , ϕ | y , x ) = 1 n ∑ i = 1 n ∑ i = 1 J − 1 y i j ( α j + ϕ j { x ′ i β } ) − log 1 + ∑ j = 1 J − 1 e α j + ϕ j { x ′ i β } . (5)</p><sec id="s2_1"><title>2.1. Elastic Net Penalized Stereotype Logit</title><p>We take the log of the information entropy, with a stereotype logit parameterization, and add an elastic net penalty [<xref ref-type="bibr" rid="scirp.90673-ref10">10</xref>] . A penalty on the sum of the squared and absolute values of the parameters is enforced [<xref ref-type="bibr" rid="scirp.90673-ref12">12</xref>] [<xref ref-type="bibr" rid="scirp.90673-ref13">13</xref>] . For a set of parameters, represented in a p length vector β , the elastic net penalty is defined as</p><p>ι = λ 2 n ∑ k = 1 p ( ς β k 2 + ( 1 − ς ) | β k | ) , (6)</p><p>where 0 &lt; λ &lt; ∞ can vary and n is the sample size of the dataset. For this study, ς was set to 0.5. The goal of the elastic net is to penalize large values of the parameter estimates, forcing their magnitude to decrease in proportion to their size. During optimization, the system is forced shrink the parameter estimates’ size when finding an optimal solution.</p><p>Based on the log of the information entropy, we are concerned with finding estimates for parameters, β ^ , α ^ , and ϕ ^ such that</p><p>( β ^ , α ^ , ϕ ^ ) = arg max β , α , ϕ L ( β , α , ϕ | y , x ) (7)</p><p>where α ^ denotes the vector of length J − 1 containing the intercepts for the J − 1 logits and ϕ ^ denotes the vector on length J − 1 containing the intensity parameters. In addition, minimizing the negative log entropy is equivalent to maximizing the log entropy, and we will work with the negative representation. Therefore, after imposing the elastic net penalty, we are concerned with finding parameter estimates such that:</p><p>( β ^ , α ^ , ϕ ^ ) = arg min β , α , ϕ { − L ( β , α , ϕ | y , x ) + λ 2 n ∑ k = 1 p ( ς β k 2 + ( 1 − ς ) | β k | ) } (8)</p><p>In machine learning, for any given model there is usually a hyperparameter set, a set of parameters that is not optimized over but whose choice of values affects the final solution. In machine learning, if there are multiple hyperparameters for any optimization procedure, there is still no established method to select these to optimize the function with respect to the parameters [<xref ref-type="bibr" rid="scirp.90673-ref14">14</xref>] . For the purposes of this manuscript, the hyperparameters were set to given values (a small range of values was considered with the optimal set being selected); λ was set to 0.001. The partial first derivatives with respect to α<sub>j</sub>, β<sub>k</sub>, and ϕ<sub>j</sub> are presented.</p><p>∂ L ∂ α j = 1 n ∑ i = 1 n ( y i j − π i j ) (9)</p><p>∂ L ∂ β k = 1 n ( ∑ i = 1 n x i k ∑ j = 1 J − 1 ϕ j ( y i j − π i j ) − λ ( ς β k + ( 1 − ς ) s i g n ( β k ) / 2 ) ) (10)</p><p>∂ L ∂ ϕ j = 1 n ∑ i = 1 n ( y i j − π i j ) ∑ k = 1 p x i k β k (11)</p><p>Denote the full parameter set β, α, and ϕ as ψ. The partial derivatives are vectorized (placed into one vector) and are represented by a derivative vector, denoted ∇ L ( ψ ) .</p></sec><sec id="s2_2"><title>2.2. Adam Optimization</title><p>The implemented Adam algorithm [<xref ref-type="bibr" rid="scirp.90673-ref11">11</xref>] attempts to find a parameter set that will minimize model Equation (8). The approach uses non-linear programs to find optimal solutions given a hyperparameter set. The Adam optimizer combines the idea of momentum optimization [<xref ref-type="bibr" rid="scirp.90673-ref15">15</xref>] and RMSProp [<xref ref-type="bibr" rid="scirp.90673-ref16">16</xref>] . Adam keeps track of an exponentially decaying average of past gradients from previous iterations (momentum optimization). Adam also tracks the exponentially decaying average of past squared gradients from previous iterations (RMSProp). The applied algorithm is listed below.</p><p>1) Initialize m and s to have all zero entries; these vectors are of length p + 2 ( J − 1 ) .</p><p>2) Initialize ψ using He Initialization [<xref ref-type="bibr" rid="scirp.90673-ref17">17</xref>] .</p><p>3) Compute ∇ L ( ψ ) .</p><p>4) m ← ν 1 m + ( 1 − ν 1 ) ∇ L ( ψ ) .</p><p>5) s ← τ 2 s + ( 1 − τ 2 ) ∇ L ( ψ ) 2 .</p><p>6) ψ ← ψ − η m &#247; s + ε .</p><p>7) Repeat steps 3 through 6 until L ( ψ | y , x ) i + i − L ( ψ | y , x ) i &lt; ζ , where i references the iteration number, or until a prespecified number of iterations are reached.</p><p>The vectors m and s contain the exponentially decaying averages of ∇ L ( ψ ) and ∇ L ( ψ ) 2 . For the He initialization [<xref ref-type="bibr" rid="scirp.90673-ref17">17</xref>] , the parameter estimate vector ψ is initialized using the random normal function of the form</p><p>N o r m ( 0 , 1 ) * 2 / p , (12)</p><p>where N o r m ( 0 , 1 ) are randomly generated values from a normal distribution with a mean 0 and standard deviation 1 and p is the number of covariates in the dataset. The hyperparameter set consists of ν<sub>1</sub>, τ<sub>2</sub>, η, ε, ζ, and ς. For this study, after considering a small range of candidate values, ν<sub>1</sub> was set to 0.5, τ<sub>2</sub> was set to 0.8, η was set to 0.008, ε was set to 1E-7, ζ was set to 1E-5, λ was set to 0.001 and, ς was set to 0.5. Steps three through six are repeated until a specified number of iterations is reached (800 for this study) or until L ( ψ | y , x ) i + i − L ( ψ | y , x ) i &lt; ζ . In applying this algorithm, we need to include an adjustment for the elastic net penalty. When taking derivatives with respect to β, we adjust these functions by subtracting the derivatives of the elastic net penalty. This is not done for α or ϕ. As a result, when computing the derivatives for the β<sub>k</sub>, where k references the iteration, we subtract from that derivative term ( λ / n ) ( ς β k + ( 1 − ς ) s i g n ( β k ) / 2 ) which leads to the derivatives for each β subject to the elastic net penalty. For each iteration, in addition to modifying β by subtracting a function of its derivative we also shrink the parameters by a factor of λ ( ς β k + ( 1 − ς ) s i g n ( β k ) / 2 ) . This method was implemented in the R programming environment [<xref ref-type="bibr" rid="scirp.90673-ref18">18</xref>] . Functions from the MASS [<xref ref-type="bibr" rid="scirp.90673-ref19">19</xref>] and matrixcalc [<xref ref-type="bibr" rid="scirp.90673-ref20">20</xref>] R packages were used to implement the proposed model.</p></sec><sec id="s2_3"><title>2.3. Applied Bootstrap Resampling Procedure</title><p>For the proposed model, the standard errors of our parameter estimates are computed using a bootstrapping pairs design [<xref ref-type="bibr" rid="scirp.90673-ref21">21</xref>] . Denote B as the number of resamples without replacement. For this study, B is set to 200 [<xref ref-type="bibr" rid="scirp.90673-ref21">21</xref>] . For each bootstrap resample, we resample n tuples with replacement, which gives us the dataset X b and y b , b = 1 , 2 , ⋯ , B . The proposed model is then fit to each resampled data set. Once the B models are fit, the corresponding parameter estimates are obtained. Denote the bth bootstrap parameter estimates as ( α ^ , β ^ , ϕ ^ ) b . Having these B parameter estimates allows us to estimate their standard errors and construct confidence intervals.</p><p>The bootstrap-t confidence interval method is used to construct confidence intervals. The bootstrap-t confidence intervals are of the form</p><p>[ β ^ k − t ^ ( 1 − α ) &#215; s e ( β ^ k ) , β ^ k + t ^ ( 1 − α ) &#215; s e ( β ^ k ) ] (13)</p><p>where</p><p>s e ( β ^ k ) = V ( β ^ k ) / B (14)</p><p>with V ( β ^ k ) being defined as</p><p>V ( β ^ k ) = 1 B − 1 ∑ b = 1 B ( β ^ ( . ) k − β ^ * k b ) 2 (15)</p><p>where k = 1 , 2 , ⋯ , p and</p><p>β ^ ( . ) k = 1 B ∑ b = 1 B β ^ k , b * , (16)</p><p>where β ^ k , b * denotes the estimate of β ^ k from the b<sup>th</sup> bootstrapped resamples dataset. In addition, t ^ ( α ) is chosen from the standard normal distribution such that</p><p>∑ b = 1 B { Z * ( b ) ≤ t ^ ( α ) } / B = α , (17)</p><p>where Z * ( b ) is defined as</p><p>Z * ( b ) = β ^ * k , p − β ^ ( . ) k s e ( β ^ ( . ) k ) (18)</p><p>The R programming environment [<xref ref-type="bibr" rid="scirp.90673-ref18">18</xref>] was used to implement this procedure.</p></sec></sec><sec id="s3"><title>3. Application to Simulated Data</title><p>The simulation procedure used is the same as previously presented [<xref ref-type="bibr" rid="scirp.90673-ref8">8</xref>] with a few noted changes. In that study, one dataset was simulated using a compound symmetric correlation structure for the covariates with ρ = 0.01. In addition, we simulated three additional dataset types. The second dataset type was simulated using a first order autoregressive, AR(1), correlation structure with ρ set to 0.1. The third dataset type has a Toeplitz correlation structure with each ρ generated randomly using a uniform distribution U (0, 0.4). The fourth dataset type has an unstructured correlation structure with each ~U (0, 0.4). For all simulated datatypes, 20 covariates (10 are significant) were generated for 1000 observations. Among the 10 significant parameter estimates, 5 were randomly set at 0.5 and 5 at −0.5 for all datasets. Each dataset was centered and scaled before the proposed model was fit. For each datatype, 100 datasets were simulated. The described bootstrap resampling technique was used to provide 95% confidence intervals; B = 200 resamples were used. In the model fitting process, the data were split into training: test data in the ratio 8:2. The model was developed using the training data. The final model was applied to the test dataset (independent data not used to build the model). The main criteria examined were the number of significant covariates that have non-zero parameter estimates, the number of non-significant coefficients that have estimates close to zero within a threshold, the accuracy of predictions when the model is applied to the test data set, and execution times. Functions from the R packages MASS [<xref ref-type="bibr" rid="scirp.90673-ref19">19</xref>] , mvtnorm [<xref ref-type="bibr" rid="scirp.90673-ref22">22</xref>] [<xref ref-type="bibr" rid="scirp.90673-ref23">23</xref>] , futility [<xref ref-type="bibr" rid="scirp.90673-ref24">24</xref>] , MBESS [<xref ref-type="bibr" rid="scirp.90673-ref25">25</xref>] , Matrix [<xref ref-type="bibr" rid="scirp.90673-ref26">26</xref>] , and corpcor [<xref ref-type="bibr" rid="scirp.90673-ref27">27</xref>] were employed to implement this procedure.</p>Results<p>The proposed methodology, and ordinalgmifs method with the option probability, model = “Stereotype”, was applied to the simulated data. The goal was to compare two implementations of penalized stereotype logit models. <xref ref-type="table" rid="table1">Table 1</xref> presents the mean, and standard deviations, for prediction accuracy (determined by the test data) and execution times for both methods. Two-sided, two sample Welch’s t tests, with significance level of 0.05, were used to compare mean accuracy and execution times for both methods. For the proposed method, the average prediction accuracy for the compound symmetric simulated data is 96.1%; for AR(1) correlated data, 96.52% of the observations were correctly classified. This rate was 96.24% for the Toeplitz correlated data and 96.49% for the unstructured correlated data. Regarding classification, the proposed method outperformed the ordinalgmifs method on all datasets as determined by the t tests (all p values &lt; 0.001). Regarding execution times, the proposed method executed faster for all datasets considered (all p values &lt; 0.001). The average execution times for the proposed method range from 17.21 to 23.63 seconds, for the ordinalgmifs method the range is 166.98 to 200.06 seconds. The proposed method executed ~10-fold faster on average.</p><p>Tables 2-5 present the parameter estimates for the 10 significant parameters based on the proposed method. For all simulated datasets, the 10 non-significant parameters (not shown in the tables) had a maximum absolute value of 0.04; these values were close to 0. For the ordinalgmifs method, all non-significant parameters had estimates of 0. The proposed method selects the significant parameters that are truly related to the outcome while setting estimates of the non-significant parameters close to 0. In comparison, the ordinalgmifs methods set these values to 0. The confidence intervals are somewhat narrow (~0.5) for parameter estimates of significant covariates.</p><p><xref ref-type="fig" rid="fig1">Figure 1</xref> presents the percent of observations correctly classified per iteration for the training data of the simulated datasets. The method performance never decreases for all simulated datasets. Once the method maximizes the proportion that it can correctly estimate, it oscillates around that value. This could be due in part to the use of the Adam optimization algorithm [<xref ref-type="bibr" rid="scirp.90673-ref11">11</xref>] . The results indicate that the proposed model framework is adept at variable selection and classification capabilities when applied to independent datasets. The proposed model outperforms the ordinalgmifs implementation with regards to prediction accuracy and execution times. Computationally, the proposed model executes faster than the ordinalgmifs implementation by ~10-fold. The analysis was performed in the R programming environment [<xref ref-type="bibr" rid="scirp.90673-ref18">18</xref>] .</p><table-wrap id="table1" ><label><xref ref-type="table" rid="table1">Table 1</xref></label><caption><title> Average accuracy (% correctly classified) and execution times for proposed and ordinalgmifs methods, along with standard deviations</title></caption><table><tbody><thead><tr><th align="center" valign="middle" ></th><th align="center" valign="middle"  colspan="2"  >Proposed Method</th><th align="center" valign="middle"  colspan="2"  >Ordinalgmifs</th></tr></thead><tr><td align="center" valign="middle" >Dataset (Correlation Type)</td><td align="center" valign="middle" >Accuracy (%)</td><td align="center" valign="middle" >Execution Time (Seconds)</td><td align="center" valign="middle" >Accuracy (%)</td><td align="center" valign="middle" >Execution Time (Seconds)</td></tr><tr><td align="center" valign="middle" >Compound Symmetric</td><td align="center" valign="middle" >96.1 (1.37)*</td><td align="center" valign="middle" >23.63 (17.50)*</td><td align="center" valign="middle" >93.51 (2.1)*</td><td align="center" valign="middle" >187.38 (50.28)*</td></tr><tr><td align="center" valign="middle" >First Order Autoregressive</td><td align="center" valign="middle" >96.52 (1.46)*</td><td align="center" valign="middle" >17.90 (3.23)*</td><td align="center" valign="middle" >94.13 (2.19)*</td><td align="center" valign="middle" >183.59 (39.30)*</td></tr><tr><td align="center" valign="middle" >Toeplitz</td><td align="center" valign="middle" >96.24 (1.33)*</td><td align="center" valign="middle" >18.21 (3.89)*</td><td align="center" valign="middle" >93.01 (2.46)*</td><td align="center" valign="middle" >200.06 (55.93)*</td></tr><tr><td align="center" valign="middle" >Unstructured</td><td align="center" valign="middle" >96.49 (1.32)*</td><td align="center" valign="middle" >17.21 (0.13)*</td><td align="center" valign="middle" >93.95 (2.1)*</td><td align="center" valign="middle" >166.98 (41.73)*</td></tr></tbody></table></table-wrap><p>For the two methods, accuracy and executions times were compared using a two-sided Welch’s two sample t test, with significance level of 0.05. “*” indicates a statistically significant difference.</p><table-wrap id="table2" ><label><xref ref-type="table" rid="table2">Table 2</xref></label><caption><title> Parameter estimates and 95% confidence intervals for truly important variables included in the final model of the compound symmetric correlated data</title></caption><table><tbody><thead><tr><th align="center" valign="middle" >Truly Important Variable</th><th align="center" valign="middle" >Parameter Estimate</th><th align="center" valign="middle" >95% Confidence Interval</th></tr></thead><tr><td align="center" valign="middle" >V1</td><td align="center" valign="middle" >−1.548</td><td align="center" valign="middle" >(−1.725, −1.372)</td></tr><tr><td align="center" valign="middle" >V2</td><td align="center" valign="middle" >−1.574</td><td align="center" valign="middle" >(−1.754, −1.394)</td></tr><tr><td align="center" valign="middle" >V3</td><td align="center" valign="middle" >−1.62</td><td align="center" valign="middle" >(−1.805, −1.434)</td></tr><tr><td align="center" valign="middle" >V4</td><td align="center" valign="middle" >−1.629</td><td align="center" valign="middle" >(−1.816, −1443)</td></tr><tr><td align="center" valign="middle" >V5</td><td align="center" valign="middle" >−1.55</td><td align="center" valign="middle" >(−1.727, −1.373)</td></tr><tr><td align="center" valign="middle" >V6</td><td align="center" valign="middle" >1.625</td><td align="center" valign="middle" >(1.44, 1.81)</td></tr><tr><td align="center" valign="middle" >V7</td><td align="center" valign="middle" >1.618</td><td align="center" valign="middle" >(1.433, 1.803)</td></tr><tr><td align="center" valign="middle" >V8</td><td align="center" valign="middle" >1.639</td><td align="center" valign="middle" >(1.452, 1.826)</td></tr><tr><td align="center" valign="middle" >V9</td><td align="center" valign="middle" >1.701</td><td align="center" valign="middle" >(1.506, 1.896)</td></tr><tr><td align="center" valign="middle" >V10</td><td align="center" valign="middle" >1.691</td><td align="center" valign="middle" >(1.497, 1.884)</td></tr></tbody></table></table-wrap><table-wrap id="table3" ><label><xref ref-type="table" rid="table3">Table 3</xref></label><caption><title> Parameter estimates and 95% confidence intervals for truly important variables included in the final model of the AR(1) correlated data</title></caption><table><tbody><thead><tr><th align="center" valign="middle" >Truly Important Variable</th><th align="center" valign="middle" >Parameter Estimate</th><th align="center" valign="middle" >95% Confidence Interval</th></tr></thead><tr><td align="center" valign="middle" >V1</td><td align="center" valign="middle" >1.497</td><td align="center" valign="middle" >(1.317, 1.676)</td></tr><tr><td align="center" valign="middle" >V2</td><td align="center" valign="middle" >1.466</td><td align="center" valign="middle" >(1.29, 1.642)</td></tr><tr><td align="center" valign="middle" >V3</td><td align="center" valign="middle" >1.416</td><td align="center" valign="middle" >(1.246, 1.585)</td></tr><tr><td align="center" valign="middle" >V4</td><td align="center" valign="middle" >1.531</td><td align="center" valign="middle" >(1.347, 1.715)</td></tr><tr><td align="center" valign="middle" >V5</td><td align="center" valign="middle" >1.446</td><td align="center" valign="middle" >(1.273, 1.62)</td></tr><tr><td align="center" valign="middle" >V6</td><td align="center" valign="middle" >−1.525</td><td align="center" valign="middle" >(−1.709, −1.341)</td></tr><tr><td align="center" valign="middle" >V7</td><td align="center" valign="middle" >−1.533</td><td align="center" valign="middle" >(−1.716, −1.349)</td></tr><tr><td align="center" valign="middle" >V8</td><td align="center" valign="middle" >−1.537</td><td align="center" valign="middle" >(−1.723, −1.352)</td></tr><tr><td align="center" valign="middle" >V9</td><td align="center" valign="middle" >−1.587</td><td align="center" valign="middle" >(−1.778, −1.397)</td></tr><tr><td align="center" valign="middle" >V10</td><td align="center" valign="middle" >−1.533</td><td align="center" valign="middle" >(−1.717, −1.35)</td></tr></tbody></table></table-wrap><table-wrap id="table4" ><label><xref ref-type="table" rid="table4">Table 4</xref></label><caption><title> Parameter estimates and 95% confidence intervals for truly important variables included in the final model of the Toeplitz correlated data</title></caption><table><tbody><thead><tr><th align="center" valign="middle" >Truly Important Variable</th><th align="center" valign="middle" >Parameter Estimate</th><th align="center" valign="middle" >95% Confidence Interval</th></tr></thead><tr><td align="center" valign="middle" >V1</td><td align="center" valign="middle" >1.466</td><td align="center" valign="middle" >(1.296, 1.636)</td></tr><tr><td align="center" valign="middle" >V2</td><td align="center" valign="middle" >1.522</td><td align="center" valign="middle" >(1.347, 1.698)</td></tr><tr><td align="center" valign="middle" >V3</td><td align="center" valign="middle" >1.485</td><td align="center" valign="middle" >(1.313, 1.657)</td></tr><tr><td align="center" valign="middle" >V4</td><td align="center" valign="middle" >1.521</td><td align="center" valign="middle" >(1.346, 1.697)</td></tr><tr><td align="center" valign="middle" >V5</td><td align="center" valign="middle" >1.507</td><td align="center" valign="middle" >(1.334, 1.681)</td></tr><tr><td align="center" valign="middle" >V6</td><td align="center" valign="middle" >−1.579</td><td align="center" valign="middle" >(−1.761, −1.396)</td></tr><tr><td align="center" valign="middle" >V7</td><td align="center" valign="middle" >−1.547</td><td align="center" valign="middle" >(−1.726, −1.368)</td></tr><tr><td align="center" valign="middle" >V8</td><td align="center" valign="middle" >−1.569</td><td align="center" valign="middle" >(−1.75, −1.388)</td></tr><tr><td align="center" valign="middle" >V9</td><td align="center" valign="middle" >−1.582</td><td align="center" valign="middle" >(−1.765, −1.399)</td></tr><tr><td align="center" valign="middle" >V10</td><td align="center" valign="middle" >−1.588</td><td align="center" valign="middle" >(−1.771, −1.405)</td></tr></tbody></table></table-wrap><table-wrap id="table5" ><label><xref ref-type="table" rid="table5">Table 5</xref></label><caption><title> Parameter estimates and 95% confidence intervals for truly important variables included in the final model of the unstructured correlated data</title></caption><table><tbody><thead><tr><th align="center" valign="middle" >Truly Important Variable</th><th align="center" valign="middle" >Parameter Estimate</th><th align="center" valign="middle" >95% Confidence Interval</th></tr></thead><tr><td align="center" valign="middle" >V1</td><td align="center" valign="middle" >−1.479</td><td align="center" valign="middle" >(−1.653, −1.305)</td></tr><tr><td align="center" valign="middle" >V2</td><td align="center" valign="middle" >−1.602</td><td align="center" valign="middle" >(−1.774 −1.43)</td></tr><tr><td align="center" valign="middle" >V3</td><td align="center" valign="middle" >−1.63</td><td align="center" valign="middle" >(−1.808, −1.453)</td></tr><tr><td align="center" valign="middle" >V4</td><td align="center" valign="middle" >−1.708</td><td align="center" valign="middle" >(−1.895, −1521)</td></tr><tr><td align="center" valign="middle" >V5</td><td align="center" valign="middle" >−1.7</td><td align="center" valign="middle" >(−1.887, −1.514)</td></tr><tr><td align="center" valign="middle" >V6</td><td align="center" valign="middle" >1.729</td><td align="center" valign="middle" >(1.539, 1.919)</td></tr><tr><td align="center" valign="middle" >V7</td><td align="center" valign="middle" >1.703</td><td align="center" valign="middle" >(1.516, 1.889)</td></tr><tr><td align="center" valign="middle" >V8</td><td align="center" valign="middle" >1.654</td><td align="center" valign="middle" >(1.468, 1.84)</td></tr><tr><td align="center" valign="middle" >V9</td><td align="center" valign="middle" >1.806</td><td align="center" valign="middle" >(1.607, 2.005)</td></tr><tr><td align="center" valign="middle" >V10</td><td align="center" valign="middle" >1.623</td><td align="center" valign="middle" >(1.442, 1.804)</td></tr></tbody></table></table-wrap></sec><sec id="s4"><title>4. Application to NHL Data</title><p>The data came from a study titled “Subclass Mapping: Identifying Common Subtypes in Independent Disease Data Sets” [<xref ref-type="bibr" rid="scirp.90673-ref28">28</xref>] . The primary aim of this manuscript was to find gene expression profiles that could predict, with some degree of error, molecular subtypes of diseases. The cancer types evaluated by this manuscript were lymphoma and breast cancer. Lymphoma is defined as a cancer of the lymphatic system. The following review is taken from Cancer Stat Facts [<xref ref-type="bibr" rid="scirp.90673-ref29">29</xref>] . NHL make up approximately 90% of all malignant lymphomas, with the Hodgkin lymphomas accounting for the remaining 10%. NHL “is a heterogeneous disease resulting from the malignant transformation of lymphocytes and includes multiple subtypes each with specific molecular and clinical characteristics” [<xref ref-type="bibr" rid="scirp.90673-ref29">29</xref>] . NHL can either start in the B-lymphocytes or the T-lymphocytes. Among B-cell lymphomas, diffuse large B-cell lymphomas are the most common. T-cell lymphomas account for 15% of NHL in the United States. NHL account for 4.3% of all new cancer cases. There were 72,240 estimated cases for 2017 and 20,140 estimated deaths. The median age of diagnosis was 67, with the highest proportion of new cases occurring in the 65 - 74 age group. The estimated 5-year survival rate was 71.0%. The issue of stage prediction with NHL, using a set of covariates, provides an opportunity to evaluate the performance of the proposed model framework.</p><p>The raw data, DLBCL-A: data set and DLBCL-A: class labels, were downloaded from http://portals.broadinstitute.org/cgi-bin/cancer/datasets.cgi [<xref ref-type="bibr" rid="scirp.90673-ref28">28</xref>] . The data were generated using one-channel oligonucleotide microarrays. The data have three subtypes, designated as oxidative phosphorylation (OxPhos), B-cell response (BCR), and host response (HR) [<xref ref-type="bibr" rid="scirp.90673-ref28">28</xref>] . The independent variables are gene expression values. The R package CePa [<xref ref-type="bibr" rid="scirp.90673-ref30">30</xref>] was used to read in the datasets. All variables were input into the model. The gene expression values were standardized (centered and scaled) prior to model fitting. There were 661 genes in the dataset. Among the 141 samples, 49 were OxPhos, 50 were BCR, and 42 were HR. The proposed model, along with the ordinalgmifs implementation of the stereotype logit, was applied to the gene expression dataset, with associated outcome vectors, to select genes associated with NHL subtypes. The described bootstrap resampling procedure was applied, yielding estimates of standard errors that were used to compute 95% confidence intervals. B = 200 resamples were used. Due to the small sample size, leave-one-out cross validation was used to estimate the predictive capabilities of the model. The analysis was performed in the R programming environment [<xref ref-type="bibr" rid="scirp.90673-ref18">18</xref>] .</p>Results<p><xref ref-type="table" rid="table6">Table 6</xref> shows selected genes along with parameter estimates and confidence intervals. The displayed genes are those with the largest absolute value of the coefficient of variation, with standard deviations provided by the bootstrap resampling scheme. Only the top 20 were displayed. The corresponding gene names are also presented. The names of the genes are provided by the HUGO Gene Nomenclature Committee (https://www.genenames.org/) [<xref ref-type="bibr" rid="scirp.90673-ref31">31</xref>] . As with the results in the simulation section, the 95% confidence intervals are narrow, and prediction accuracy was 73%. When applying the stereotype logit-based ordinalgmifs function to the NHL data, parameter estimates could not be obtained due to optimization issues with that implementation. The following error was reported by the R programming environment “Error in optim(c(alpha, phi), fn.stereo, w = w, x = x, beta = beta, y = y: L-BFGS-B needs finite values of ‘fn’”. The above error relates to the optimization function being passed infinite values during the optimization process. All covariate values were centered and scaled prior to analysis. To correct the error, multiple values of the hyperparameters were passed to the ordinalgmifs function in R with no success. The data was checked for missing values; there were none. In addition, all the data were numeric. As a result, no comparison could be made with the ordinalgmifs implementation of the L1 penalized stereotype logit.</p><table-wrap id="table6" ><label><xref ref-type="table" rid="table6">Table 6</xref></label><caption><title> Variable importance based on the application of the proposed model to the NHL dataset. The topmost 20 genes, in terms of variable importance, are presented. The model achieved a prediction accuracy of 73%</title></caption><table><tbody><thead><tr><th align="center" valign="middle" >Gene Name</th><th align="center" valign="middle" >Definition</th><th align="center" valign="middle" >Parameter Estimate</th><th align="center" valign="middle" >95% Confidence Interval</th></tr></thead><tr><td align="center" valign="middle" >TCF7</td><td align="center" valign="middle" >transcription factor 7</td><td align="center" valign="middle" >−2.666</td><td align="center" valign="middle" >(−2.845, −2.486)</td></tr><tr><td align="center" valign="middle" >GSTM2</td><td align="center" valign="middle" >glutathione S-transferase mu 2</td><td align="center" valign="middle" >−2.629</td><td align="center" valign="middle" >(−2.813, −2.444)</td></tr><tr><td align="center" valign="middle" >ITGB7</td><td align="center" valign="middle" >integrin subunit beta 7</td><td align="center" valign="middle" >−2.598</td><td align="center" valign="middle" >(−2.756, −2.439)</td></tr><tr><td align="center" valign="middle" >EEF1A1</td><td align="center" valign="middle" >eukaryotic translation elongation factor 1 alpha 1</td><td align="center" valign="middle" >−2.59</td><td align="center" valign="middle" >(−2.839, −2.342)</td></tr><tr><td align="center" valign="middle" >LOC220594</td><td align="center" valign="middle" >NA</td><td align="center" valign="middle" >−2.572</td><td align="center" valign="middle" >(−2.736, −2.408)</td></tr><tr><td align="center" valign="middle" >DEK</td><td align="center" valign="middle" >DEK proto-oncogene</td><td align="center" valign="middle" >2.57</td><td align="center" valign="middle" >(2.446, 2.694)</td></tr><tr><td align="center" valign="middle" >ITGAL</td><td align="center" valign="middle" >integrin subunit alpha L</td><td align="center" valign="middle" >−2.561</td><td align="center" valign="middle" >(−2.717, −2.405)</td></tr><tr><td align="center" valign="middle" >BIN1</td><td align="center" valign="middle" >bridging integrator 1</td><td align="center" valign="middle" >−2.556</td><td align="center" valign="middle" >(−2.704, −2.409)</td></tr><tr><td align="center" valign="middle" >RPL21</td><td align="center" valign="middle" >ribosomal protein L21</td><td align="center" valign="middle" >−2.503</td><td align="center" valign="middle" >(−2.671, −2.335)</td></tr><tr><td align="center" valign="middle" >RBPSUH</td><td align="center" valign="middle" >recombination signal binding protein for immunoglobulin kappa J region</td><td align="center" valign="middle" >−2.503</td><td align="center" valign="middle" >(−2.65, −2.356)</td></tr><tr><td align="center" valign="middle" >NCOA1</td><td align="center" valign="middle" >nuclear receptor coactivator 1</td><td align="center" valign="middle" >−2.494</td><td align="center" valign="middle" >(−2.648, −2.34)</td></tr><tr><td align="center" valign="middle" >MYCBP2</td><td align="center" valign="middle" >MYC binding protein 2, E3 ubiquitin protein ligase</td><td align="center" valign="middle" >−2.47</td><td align="center" valign="middle" >(−2.614, −2.326)</td></tr><tr><td align="center" valign="middle" >A2M</td><td align="center" valign="middle" >alpha-2-macroglobulin</td><td align="center" valign="middle" >−2.416</td><td align="center" valign="middle" >(−2.595, −2.238)</td></tr><tr><td align="center" valign="middle" >IL10RA</td><td align="center" valign="middle" >interleukin 10 receptor subunit alpha</td><td align="center" valign="middle" >−2.413</td><td align="center" valign="middle" >(−2.564, −2.262)</td></tr><tr><td align="center" valign="middle" >SLC25A5</td><td align="center" valign="middle" >solute carrier family 25 member 5</td><td align="center" valign="middle" >−2.383</td><td align="center" valign="middle" >(−2.539, −2.227)</td></tr><tr><td align="center" valign="middle" >CCL21</td><td align="center" valign="middle" >C−C motif chemokine ligand 21</td><td align="center" valign="middle" >−2.378</td><td align="center" valign="middle" >(−2.53, −2.227)</td></tr><tr><td align="center" valign="middle" >KPNB1</td><td align="center" valign="middle" >karyopherin subunit beta 1</td><td align="center" valign="middle" >−2.377</td><td align="center" valign="middle" >(−2.536, −2.217)</td></tr><tr><td align="center" valign="middle" >COL9A2</td><td align="center" valign="middle" >collagen type IX alpha 2 chain</td><td align="center" valign="middle" >−2.374</td><td align="center" valign="middle" >(−2.552, −2.196)</td></tr><tr><td align="center" valign="middle" >RPS21</td><td align="center" valign="middle" >ribosomal protein S21</td><td align="center" valign="middle" >−2.363</td><td align="center" valign="middle" >(−2.524, −2.202)</td></tr><tr><td align="center" valign="middle" >ACP1</td><td align="center" valign="middle" >acid phosphatase 1</td><td align="center" valign="middle" >−2.361</td><td align="center" valign="middle" >(−2.501, −2.221)</td></tr></tbody></table></table-wrap></sec><sec id="s5"><title>5. Discussion</title><p>There are multiple hyperparameters in the elastic net constrained stereotype logit, optimal values for these must be explored as this has the potential to increase variable selection and classification capabilities. This is usually accomplished with a grid search and is an open problem in machine learning for multiple hyperparameters [<xref ref-type="bibr" rid="scirp.90673-ref14">14</xref>] . For this study, a small selection of values was considered for each hyperparameter. The hyperparameter of interest is λ but the choice for the remaining hyperparameters is also very important in determining the optimal solution.</p><p>A bootstrap resampling procedure was used to estimate the 95% confidence intervals. The main drawback is the computational time required to produce the confidence intervals with 200 additional models being fit. It may be advisable to perform a closed form estimate of the parameter variance matrix [<xref ref-type="bibr" rid="scirp.90673-ref2">2</xref>] .</p><p>Although the stereotype logit is considered by many a generalized linear model, it is not. As such, an optimal solution may not exist, or there may be inflexion points. As a result, different starting values may yield different solutions. In this study, applying the method to a given dataset does not exhibit a great deal of variation in results, and the results of the applied bootstrap procedure confirm this. To address this, we applied a variable initialization scheme proposed by He [<xref ref-type="bibr" rid="scirp.90673-ref17">17</xref>] . In addition, the Adam optimization function is well suited to dealing with non-convex functions [<xref ref-type="bibr" rid="scirp.90673-ref11">11</xref>] . The combination of these two factors addresses this issue.</p></sec><sec id="s6"><title>6. Conclusion</title><p>A proposed model for the elastic net penalized stereotype logit model, with optimization provided by the Adam optimizer, to analyze ordinal outcome data was presented. The proposed method was applied to simulated and NHL data with reported results. For the simulated data, variable selection was perfect, and only significant variables had parameter estimates not close to 0. The classifications ranged from 96.1% to 96.52% on the test datasets. For the NHL data, 73% were correctly classified. The 20 topmost genes in terms of absolute value of the coefficient of variation were presented. Our evaluation study shows that the proposed method outperforms the ordinalgmifs penalized stereotype logit model; no comparison could be made with the NHL data analysis as the ordinalgmifs implemented stereotype logit was not able to produce parameter estimates. This manuscript is an extension of previous work [<xref ref-type="bibr" rid="scirp.90673-ref8">8</xref>] . In the previous study, variable selection was adequate, but the classification capabilities were lacking. This work improves the prediction accuracy when applied to simulated and NHL data (ranging from 73% to 96.52%). In addition, the variable importance also improved with only the significant parameters having non-zero estimates. This study demonstrates, with success, the application of the Adam optimizer to the elastic net penalized stereotype logit model to analyze ordinal outcome data with promising results, as demonstrated on the simulated and NHL datasets.</p></sec><sec id="s7"><title>Acknowledgements</title><p>First and foremost, I would like to thank God from whom all blessings flow. I would also like to express special thanks to Timothy Wysocki, the Co-director of the Center for Healthcare Delivery Science, at Nemours Children’s Specialty Care for allowing me to work on this project.</p></sec><sec id="s8"><title>Conflicts of Interest</title><p>The author declares no conflicts of interest regarding the publication of this paper.</p></sec><sec id="s9"><title>Cite this paper</title><p>Williams, A.A.A. (2019) Ordinal Outcome Modeling: The Application of the Adaptive Moment Estimation Optimizer to the Elastic Net Penalized Stereotype Logit. Journal of Data Analysis and Information Processing, 7, 14-27. https://doi.org/10.4236/jdaip.2019.71002</p></sec></body><back><ref-list><title>References</title><ref id="scirp.90673-ref1"><label>1</label><mixed-citation publication-type="other" xlink:type="simple">Williams, M. and Schnellhammer, P. (2016) Testicular Seminoma. http://emedicine.medscape.com/article/437966-overview</mixed-citation></ref><ref id="scirp.90673-ref2"><label>2</label><mixed-citation publication-type="other" xlink:type="simple">Agresti, A. (2014) Categorical Data Analysis. John Wiley &amp; Sons, New York.</mixed-citation></ref><ref id="scirp.90673-ref3"><label>3</label><mixed-citation publication-type="other" xlink:type="simple">Anderson, J.A. (1984) Regression and Ordered Categorical Variables. Journal of the Royal Statistical Society: Series B (Methodological), 46, 1-30. http://www.jstor.org/stable/2345457 https://doi.org/10.1111/j.2517-6161.1984.tb01270.x</mixed-citation></ref><ref id="scirp.90673-ref4"><label>4</label><mixed-citation publication-type="other" xlink:type="simple">Bellazzi, R. (2014) Big Data and Biomedical Informatics: A Challenging Opportunity. Yearbook of Medical Informatics, 9, 8-13. https://doi.org/10.15265/IY-2014-0024</mixed-citation></ref><ref id="scirp.90673-ref5"><label>5</label><mixed-citation publication-type="other" xlink:type="simple">Murdoch, T.B. and Detsky, A.S. (2013) The Inevitable Application of Big Data to Health Care. JAMA, 309, 1351-1352. https://doi.org/10.1001/jama.2013.393</mixed-citation></ref><ref id="scirp.90673-ref6"><label>6</label><mixed-citation publication-type="other" xlink:type="simple">Archer, K.J. and Williams, A.A.A. (2012) L1 Penalized Continuation Ratio Models for Ordinal Response Prediction Using High-Dimensional Datasets. Statistics in Medicine, 31, 1464-1474. https://doi.org/10.1002/sim.4484</mixed-citation></ref><ref id="scirp.90673-ref7"><label>7</label><mixed-citation publication-type="other" xlink:type="simple">Archer, K.J., Hou, J., Zhou, Q., Ferber, K., Layne, J.G. and Gentry, A.E. (2014) Ordinalgmifs: An R Package for Ordinal Regression in High-Dimensional Data Settings. Cancer Informatics, 13, CIN.S20806. https://doi.org/10.4137/CIN.S20806</mixed-citation></ref><ref id="scirp.90673-ref8"><label>8</label><mixed-citation publication-type="other" xlink:type="simple">Williams, A.A. and Archer, K.J. (2015) Elastic Net Constrained Stereotype Logit Model for Ordered Categorical Data. Biometrics &amp; Biostatistics International Journal, 2, Article ID: 00049. https://doi.org/10.15406/bbij.2015.02.00049</mixed-citation></ref><ref id="scirp.90673-ref9"><label>9</label><mixed-citation publication-type="other" xlink:type="simple">Hastie, T., Taylor, J., Tibshirani, R., Walther, G., Boyd, S., Friedman, J., et al. (2007) Forward Stagewise Regression and the Monotone Lasso. Electronic Journal of Statistics, 1, 1-29. https://doi.org/10.1214/07-EJS004</mixed-citation></ref><ref id="scirp.90673-ref10"><label>10</label><mixed-citation publication-type="other" xlink:type="simple">Zou, H. and Hastie, T. (2005) Regularization and Variable Selection via the Elastic Net. Journal of the Royal Statistical Society: Series B (Methodological), 67, 301-320. https://web.stanford.edu/~hastie/Papers/B67.2%20%282005%29%20301-320%20Zou%20&amp;%20Hastie.pdf  https://doi.org/10.1111/j.1467-9868.2005.00503.x</mixed-citation></ref><ref id="scirp.90673-ref11"><label>11</label><mixed-citation publication-type="other" xlink:type="simple">Kingma, D.P. and Ba, J. (2014) Adam: A Method for Stochastic Optimization. 1-15.</mixed-citation></ref><ref id="scirp.90673-ref12"><label>12</label><mixed-citation publication-type="other" xlink:type="simple">Hoerl, A.E. and Kennard, R.W. (1970) Ridge Regression: Biased Estimation for Nonorthogonal Problems. Technometrics, 12, 55-67. https://doi.org/10.1080/00401706.1970.10488634</mixed-citation></ref><ref id="scirp.90673-ref13"><label>13</label><mixed-citation publication-type="other" xlink:type="simple">Tibshirani, R. (1996) Regression Shrinkage and Selection via the Lasso. Journal of the Royal Statistical Society: Series B (Methodological), 58, 267-288. http://www.jstor.org/stable/2346178https://doi.org/10.1111/j.2517-6161.1996.tb02080.x</mixed-citation></ref><ref id="scirp.90673-ref14"><label>14</label><mixed-citation publication-type="other" xlink:type="simple">Kaiser, L., Gomez, A.N., Shazeer, N., Vaswani, A., Parmar, N., Jones, L. and Uszkoreit, J. (2017) One Model To Learn Them All. https://arxiv.org/pdf/1706.05137.pdf</mixed-citation></ref><ref id="scirp.90673-ref15"><label>15</label><mixed-citation publication-type="other" xlink:type="simple">Polyak, B.T. (1964). Some Methods of Speeding up the Convergence of Iteration Methods. USSR Computational Mathematics and Mathematical Physics, 4, 1-17. http://www.sciencedirect.com/science/article/pii/0041555364901375 https://doi.org/10.1016/0041-5553(64)90137-5</mixed-citation></ref><ref id="scirp.90673-ref16"><label>16</label><mixed-citation publication-type="other" xlink:type="simple">Hinton, G., Srivastava, N. and Swersky, K. (n.d.) Neural Networks for Machine Learning. http://www.cs.toronto.edu/~tijmen/csc321/slides/lecture_slides_lec6.pdf</mixed-citation></ref><ref id="scirp.90673-ref17"><label>17</label><mixed-citation publication-type="other" xlink:type="simple">He, K., Zhang, X., Ren, S. and Sun, J. (2014) Delving Deep into Rectifiers: Surpassing Human-Level Performance on ImageNet Classification. 1026-1034.</mixed-citation></ref><ref id="scirp.90673-ref18"><label>18</label><mixed-citation publication-type="other" xlink:type="simple">R Core Team (2017) R: A Language and Environment for Statistical Computing. R Foundation for Statistical Computing, Vienna. https://www.r-project.org/</mixed-citation></ref><ref id="scirp.90673-ref19"><label>19</label><mixed-citation publication-type="other" xlink:type="simple">Venables, W.N. and Ripley, B.D. (2002) Modern Applied Statistics with S. Fourth Edition, Springer, New York. https://doi.org/10.1007/978-0-387-21706-2</mixed-citation></ref><ref id="scirp.90673-ref20"><label>20</label><mixed-citation publication-type="other" xlink:type="simple">Novomestky, F. (2012) Matrixcalc: Collection of Functions for Matrix Calculations. https://cran.r-project.org/package=matrixcalc</mixed-citation></ref><ref id="scirp.90673-ref21"><label>21</label><mixed-citation publication-type="other" xlink:type="simple">Efron, B. and Tibshirani, R. (1986) Bootstrap Methods for Standard Errors, Confidence Intervals, and Other Methods of Statistical Accuracy. Statistical Science, 1, 54-75. https://doi.org/10.1214/ss/1177013815</mixed-citation></ref><ref id="scirp.90673-ref22"><label>22</label><mixed-citation publication-type="other" xlink:type="simple">Genz, A., Bretz, F., Miwa, T., Mi, X., Leisch, F., Scheipl, F. and Hothorn, T. (2017) Mvtnorm: Multivariate Normal and t Distributions. http://cran.r-project.org/package=mvtnorm</mixed-citation></ref><ref id="scirp.90673-ref23"><label>23</label><mixed-citation publication-type="other" xlink:type="simple">Genz, A. and Bretz, F. (2009) Computation of Multivariate Normal and t Probabilities. Lecture Notes in Statistics Vol. 195, Springer-Verlag, Heidelberg. https://doi.org/10.1007/978-3-642-01689-9</mixed-citation></ref><ref id="scirp.90673-ref24"><label>24</label><mixed-citation publication-type="other" xlink:type="simple">Zhuang, Y., Juraska, M., Grove, D., Gilbert, P. and Luedtke, A. (2017) Futility: Interim Analysis of Operational Futility in Randomized Trials with Time-to-Event Endpoints and Fixed Follow-Up. https://cran.r-project.org/package=futility</mixed-citation></ref><ref id="scirp.90673-ref25"><label>25</label><mixed-citation publication-type="other" xlink:type="simple">Kelley, K. (2017) MBESS: The MBESS R Package. https://cran.r-project.org/package=MBESS</mixed-citation></ref><ref id="scirp.90673-ref26"><label>26</label><mixed-citation publication-type="other" xlink:type="simple">Bates, D. and Maechler, M. (2017) Matrix: Sparse and Dense Matrix Classes and Methods. https://cran.r-project.org/package=Matrix</mixed-citation></ref><ref id="scirp.90673-ref27"><label>27</label><mixed-citation publication-type="other" xlink:type="simple">Schafer, J., Opgen-Rhein, R., Zuber, V., Ahdesmaki, M., Silva, A.P.D. and Strimmer, K. (2017) Corpcor: Efficient Estimation of Covariance and (Partial) Correlation. https://cran.r-project.org/package=corpcor</mixed-citation></ref><ref id="scirp.90673-ref28"><label>28</label><mixed-citation publication-type="other" xlink:type="simple">Hoshida, Y., Brunet, J.P., Tamayo, P., Golub, T.R. and Mesirov, J.P. (2007) Subclass Mapping: Identifying Common Subtypes in Independent Disease Data Sets. PLoS ONE, 2, e1195. https://doi.org/10.1371/journal.pone.0001195</mixed-citation></ref><ref id="scirp.90673-ref29"><label>29</label><mixed-citation publication-type="other" xlink:type="simple">National Cancer Institute (2017) Cancer Stat Fact: Non-Hodgkin Lymphoma. https://seer.cancer.gov/statfacts/html/nhl.html</mixed-citation></ref><ref id="scirp.90673-ref30"><label>30</label><mixed-citation publication-type="other" xlink:type="simple">Gu, Z. (2012) CePa: Centrality-Based Pathway Enrichment. https://cran.r-project.org/package=CePa</mixed-citation></ref><ref id="scirp.90673-ref31"><label>31</label><mixed-citation publication-type="other" xlink:type="simple">HUGO Gene Nomenclature Committee (n.d.) HGNC Database of Human Genes. https://www.genenames.org/</mixed-citation></ref></ref-list></back></article>