<?xml version="1.0" encoding="UTF-8"?><!DOCTYPE article  PUBLIC "-//NLM//DTD Journal Publishing DTD v3.0 20080202//EN" "http://dtd.nlm.nih.gov/publishing/3.0/journalpublishing3.dtd"><article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" dtd-version="3.0" xml:lang="en" article-type="research article"><front><journal-meta><journal-id journal-id-type="publisher-id">JTTs</journal-id><journal-title-group><journal-title>Journal of Transportation Technologies</journal-title></journal-title-group><issn pub-type="epub">2160-0473</issn><publisher><publisher-name>Scientific Research Publishing</publisher-name></publisher></journal-meta><article-meta><article-id pub-id-type="doi">10.4236/jtts.2017.73019</article-id><article-id pub-id-type="publisher-id">JTTs-77395</article-id><article-categories><subj-group subj-group-type="heading"><subject>Articles</subject></subj-group><subj-group subj-group-type="Discipline-v2"><subject>Engineering</subject></subj-group></article-categories><title-group><article-title>
 
 
  Incorporating the Multinomial Logistic Regression in Vehicle Crash Severity Modeling: A Detailed Overview
 
</article-title></title-group><contrib-group><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Azad</surname><given-names>Abdulhafedh</given-names></name><xref ref-type="aff" rid="aff1"><sub>1</sub></xref></contrib></contrib-group><aff id="aff1"><label>1</label><addr-line>University of Missouri-Columbia, MO, USA</addr-line></aff><author-notes><corresp id="cor1">* E-mail:</corresp></author-notes><pub-date pub-type="epub"><day>01</day><month>06</month><year>2017</year></pub-date><volume>07</volume><issue>03</issue><fpage>279</fpage><lpage>303</lpage><history><date date-type="received"><day>March</day>	<month>31,</month>	<year>2017</year></date><date date-type="rev-recd"><day>Accepted:</day>	<month>July</month>	<year>1,</year>	</date><date date-type="accepted"><day>July</day>	<month>4,</month>	<year>2017</year></date></history><permissions><copyright-statement>&#169; Copyright  2014 by authors and Scientific Research Publishing Inc. </copyright-statement><copyright-year>2014</copyright-year><license><license-p>This work is licensed under the Creative Commons Attribution International License (CC BY). http://creativecommons.org/licenses/by/4.0/</license-p></license></permissions><abstract><p>
 
 
  Multinomial logistic regression (MNL) is an attractive statistical approach in modeling the vehicle crash severity as it does not require the assumption of normality, linearity, or homoscedasticity compared to other approaches, such as the discriminant analysis which requires these assumptions to be met. Moreover, it produces sound estimates by changing the probability range between 0.0 and 1.0 to log odds ranging from negative infinity to positive infinity, as it applies transformation of the dependent variable to a continuous variable. The estimates are asymptotically consistent with the requirements of the nonlinear regression process. The results of MNL can be interpreted by both the regression coefficient estimates and/or the odd ratios (the exponentiated coefficients) as well. In addition, the MNL can be used to improve the fitted model by comparing the full model that includes all predictors to a chosen restricted model by excluding the non-significant predictors. As such, this paper presents a detailed step by step overview of incorporating the MNL in crash severity modeling, using vehicle crash data of the Interstate I70 in the State of Missouri, USA for the years (2013-2015).
 
</p></abstract><kwd-group><kwd>Multinomial Logistic Regression</kwd><kwd> Odd Ratio</kwd><kwd> The Independence of Irrelevant Alternatives</kwd><kwd> The Hausman Specification Test</kwd><kwd> The Hosmer-Lemeshow Test</kwd><kwd> Pseudo R Squares</kwd><kwd> Crash Severity Models</kwd></kwd-group></article-meta></front><body><sec id="s1"><title>1. Introduction</title><p>Since the dependent variable in vehicle crash severity modeling (i.e. crash severity) usually has two or more outcome categories (i.e. fatal, injury, property-damage-only), therefore, logit and probit models are often used to model the severity of crash data. Binary models consider two response outcomes (i.e. fatal vs. non-fatal or injury vs. property-damage-only), and multinomial models consider three or more response outcomes. The multinomial logistic regression (MNL) does not require the assumption of normality, linearity, or homoscedasticity (i.e. the homogeneity of variances) compared to the discriminant analysis which requires these assumptions to be met, and therefore, the MNL is used more frequently than the discriminant analysis. The MNL is used to model the relationships between a polytomous (multinomial) dependent variable (with more than two outcomes) and a set of independent variables (predictors). It is an extension of the binary logistic regression, which analyzes dichotomous (binary) dependent variables with only two outcomes. The multinomial logistic model may be used to handle a dependent variable that is a categorical, unordered variable (i.e. cannot be ordered in any logical way). Ordered logistic regression is used in cases where the dependent variable is ordered in a certain way. The MNL works by choosing one group as the base (reference) category for the other groups. Then MNL contrasts all the outcomes of the dependent variable with this common reference category, which serves as the contrast point for all analyses, and the effects of the analysis are always in reference to the contrast category [<xref ref-type="bibr" rid="scirp.77395-ref1">1</xref>] . The MNL applies the assumption of the independence of irrelevant alternatives (IIA), which means that adding or deleting alternative outcome categories does not affect the prediction among the remaining outcomes. In other words, the odd ratios produced by the logit function for any pair of outcomes are determined without reference to the other categories that might be available [<xref ref-type="bibr" rid="scirp.77395-ref2">2</xref>] [<xref ref-type="bibr" rid="scirp.77395-ref3">3</xref>] , and therefore it must be checked in the modeling process. The MNL has many advantages in modeling vehicle crash severity, such as [<xref ref-type="bibr" rid="scirp.77395-ref1">1</xref>] [<xref ref-type="bibr" rid="scirp.77395-ref4">4</xref>] [<xref ref-type="bibr" rid="scirp.77395-ref5">5</xref>] :</p><p>・ It produces sound estimates as it applies transformation of the multinomial dependent variable to a continuous variable ranging from negative infinity to positive infinity. It is usually difficult to model a variable which has restricted range, such as probability. This transformation attempts to overcome this problem. It changes probability ranging between 0.0 and 1.0 to log odds ranging from negative infinity to positive infinity.</p><p>・ Among all of the many choices of transformation, the log of odds in MNL is one of the easiest to understand and interpret.</p><p>・ The results of MNL can be interpreted by both the regression coefficient estimates and/or the odd ratios (the exponentiated coefficients) as well.</p><p>・ The estimates are asymptotically consistent with the requirements of the nonlinear regression process.</p><p>・ MNL can be used to improve the fitted model by comparing the full model that includes all predictors to a chosen restricted model by excluding the non-significant predictors, and then picks up the best fit.</p></sec><sec id="s2"><title>2. Methodology</title><p>The dependent variable (i.e. crash severity) in this paper consists of four outcome categories (i.e. fatal, disabling injury, minor injury, property-damage-only), and is assumed to be nominal (i.e. unordered), therefore it is modeled by the multinomial logistic regression (MNL). Since the MNL works by choosing one outcome category as the base (reference) category for the other categories, hence, the property damage is considered as the reference group (i.e. base category), because it is the most frequent outcome of crash severity data, and the other outcome levels (i.e. minor injury, disabling injury, and fatal) are estimated relative to the property damage. There are a few applications of the MNL in vehicle crash severity modeling. For example, Abdel-Aty [<xref ref-type="bibr" rid="scirp.77395-ref6">6</xref>] applied the ordered probit model and the ordered MNL to predict crash severity on roadway sections, signalized intersections and toll plazas by using the Florida crash database. Bham et al. [<xref ref-type="bibr" rid="scirp.77395-ref7">7</xref>] applied a multinomial logistic regression to model the severity injury of different vehicle collision patterns in urban highways in Arkansas, and recommended the use of the MNL over other models. Despite these few applications of the MNL, this paper seeks to introduce a variety of new procedures in presenting the results of the MNL applications that have not been reported in other crash severity research. First, the use of odd ratios as regression estimates is explored to interpret the results of prediction instead of regression coefficients. Second, a greater focus is place on the assumption of the independence of irrelevant alternatives (IIA), which is very crucial in the MNL modeling, using the Hausman specification test. Third, the generalized Hosmer-Lemeshow test is used as an important goodness of fit measure to assess whether or not the observed incidents match the predicted incidents. Fourth, the concept of the classification table is evaluated as a measure of goodness of fit to determine the percent of corrected prediction cases. Next, tests for the multicollinearity among the independent variables as precondition assumption are conducted. The pseudo R square measure is used as a potential goodness of fit instead of the classical measures, such as the Deviance, the Akaike Information Criteria (AIC), and the Bayesian Information Criteria (BIC). Lastly, the marginal effects of all independent variables upon the dependent variable are presented. The following sections illustrate the assumptions of the MNL, the concept of logit functions and odd ratios, several methodological procedures that should be used in testing the assumptions of the MNL, and the MNL goodness of fit tests.</p></sec><sec id="s3"><title>3. Data</title><p>Missouri crash data as reported by the Missouri State Highway Patrol (MSHP) and recorded in the Missouri Statewide Traffic Accident Records System (STARS) for the Interstate I70 in the State of Missouri, USA for the years (2013- 2015) were used in the analysis. The I-70 corridor in MO is a multi-lane divided highway that traverses the State of Missouri west to east with a total length of 403 km (250 mile). The STARS and roadway data were carefully examined, labelled, filtered, and outliers and missing data were excluded from the analysis. The total numbers of the observed crashes within the three years (2013-2015) were 5869.0 along the I-70 corridor. In the state of Missouri, the STARS data includes four severity injury categories (i.e. property damage, minor injury, disabled injury, and fatal). As such, crash severity (i.e. the dependent variable) is modeled in this paper using the following four STARS severity categories:</p><p>・ Property-Damage-Only: A property damage crash that includes any crash in which no person was killed or injured but property was damaged in the incident.</p><p>・ Minor Injury: An injury crash in which one or more persons received an evident injury but not disabling in the incident.</p><p>・ Disabled Injury: An injury crash in which one or more persons received a disabling in the incident.</p><p>・ Fatal: A fatal crash includes any crash in which one or more persons were killed and their death occurred within 30 days of the incident.</p><p>If a crash result in more than one injury severity category, then the most severe category would be considered for reporting. For instance, if a crash resulted in fatal, and property damage, then this crash would be reported as fatal [<xref ref-type="bibr" rid="scirp.77395-ref8">8</xref>] . The STARS system provides the latitude and longitude coordinates of each reported crash, rather than reporting the crash characteristics by road segment as is done by reporting agencies in other states. The STARS crash data were partitioned into training and testing datasets. The STARS data for the entire period (2013-2015) was randomly partitioned into two parts, a training dataset that contains 70% of the observations, and a testing dataset that contains 30% of the observations. The training dataset includes 4108 observed crashes for I-70 corridor, and the testing dataset includes 1644 observed crashes. The occurrence of crashes and their degrees of severity can be attributed to different risk factors associated with road geometry, traffic operations, vehicle types, driver factors, and the environment. Given that past research has only made use of limited numbers/types of independent variables, this paper investigated the use of a wide range of independent variables (i.e. risk factors) for estimating the parameters and inferences. The following group factors are included in the analysis:</p><p>・ Road geometry (grade or level; number of lanes);</p><p>・ Road classification (rural or urban; existing of construction zones);</p><p>・ Environment (light conditions);</p><p>・ Traffic operation (annual average daily traffic, AADT);</p><p>・ Driver factors (driver’s age; speeding; aggressive driving; driver intoxicated conditions; the use of cell phone or texting);</p><p>・ Vehicle type (passenger car; motorcycles; truck);</p><p>・ Number of vehicles involved in the crash;</p><p>・ Time factors (hour of crash occurrence; weekday; month);</p><p>・ Accident type (animal; fixed object; overturn; pedestrian; vehicle in transport).</p></sec><sec id="s4"><title>4. The Logit Function and Odd Ratios of the MNL</title><p>The MNL tries to find the best fitted model to describe the relationship between the polytomous dependent variable with more than two categories and a set of independent variables. The logistic regression model is a non-linear transforma-</p><fig id="fig1"  position="float"><label><xref ref-type="fig" rid="fig1">Figure 1</xref></label><caption><title> Comparison of linear and logistic regression</title></caption><graphic mimetype="image"   position="float"  xlink:type="simple"  xlink:href="http://html.scirp.org/file/4-3500367x3.png"/></fig><p>tion of the linear regression model, as it consists of an S-shaped distribution function, and it’s very easy to work with in most applications [<xref ref-type="bibr" rid="scirp.77395-ref9">9</xref>] . The logit distribution constrains the estimated probabilities that lie between 0.0 and 1.0, as shown in <xref ref-type="fig" rid="fig1">Figure 1</xref>. The logistic regression function is bounded by 0.0 and 1.0, whereas the linear regression function may predict values above 1.0 and below 0.0.</p><p>The logistic (logit) function can be expressed as:</p><disp-formula id="scirp.77395-formula60"><label>(1)</label><graphic position="anchor" xlink:href="http://html.scirp.org/file/4-3500367x4.png"  xlink:type="simple"/></disp-formula><p>where,</p><p>p: the probability of presence of an outcome of interest,</p><p>X<sub>k</sub>: the vector of k independent variables,</p><p>b<sub>0</sub>: the regression coefficient on the constant term (intercept),</p><p>b<sub>k</sub>: the vector of regression coefficients on the independent variables X<sub>k</sub>.</p><p>The odd ratio is the probability of the event divided by the probability of the nonevent, and is defined as follows [<xref ref-type="bibr" rid="scirp.77395-ref1">1</xref>] [<xref ref-type="bibr" rid="scirp.77395-ref9">9</xref>] :</p><disp-formula id="scirp.77395-formula61"><label>(2)</label><graphic position="anchor" xlink:href="http://html.scirp.org/file/4-3500367x5.png"  xlink:type="simple"/></disp-formula><p>When p = 0, then odd (p) = 0, when p = 0.5, then odd (p) = 1.0, and when p = 1.0, then odd (p) =<inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-3500367x6.png" xlink:type="simple"/></inline-formula>.</p><p>The logit transformation is defined as the logged odds:</p><disp-formula id="scirp.77395-formula62"><label>(3)</label><graphic position="anchor" xlink:href="http://html.scirp.org/file/4-3500367x7.png"  xlink:type="simple"/></disp-formula><p>The transformation from odds to log of odds is the log transformation, and this is a monotonic transformation. That is, the greater the odds, the greater the log of odds and vice versa. Logit (p) can be back-transformed to p by the following formula:</p><disp-formula id="scirp.77395-formula63"><label>(4)</label><graphic position="anchor" xlink:href="http://html.scirp.org/file/4-3500367x8.png"  xlink:type="simple"/></disp-formula><p>The transformation from probability to odds is a monotonic transformation as well, meaning the odds increase as the probability increases or vice versa. Probability ranges from 0.0 and 1.0. Odds range from 0.0 and positive infinity [<xref ref-type="bibr" rid="scirp.77395-ref5">5</xref>] [<xref ref-type="bibr" rid="scirp.77395-ref9">9</xref>] .</p></sec><sec id="s5"><title>5. The Maximum Likelihood Estimation (MLE)</title><p>The multinomial logistic regression uses the maximum likelihood estimation (MLE) to produce the regression parameters. Assuming that the random variables <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-3500367x9.png" xlink:type="simple"/></inline-formula> form a random sample from a distribution<inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-3500367x10.png" xlink:type="simple"/></inline-formula>; if X is continuous random variable, <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-3500367x11.png" xlink:type="simple"/></inline-formula>is probability density function (pdf), if X is discrete random variable, <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-3500367x12.png" xlink:type="simple"/></inline-formula>is point mass function (pmf). The distribution depends on a parameter θ, where θ could be a real unknown parameter or a vector of parameters. For every observed random sample<inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-3500367x13.png" xlink:type="simple"/></inline-formula>, we define [<xref ref-type="bibr" rid="scirp.77395-ref10">10</xref>] :</p><disp-formula id="scirp.77395-formula64"><label>(5)</label><graphic position="anchor" xlink:href="http://html.scirp.org/file/4-3500367x14.png"  xlink:type="simple"/></disp-formula><p>If <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-3500367x15.png" xlink:type="simple"/></inline-formula> is pdf, <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-3500367x16.png" xlink:type="simple"/></inline-formula>is the joint density function; if <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-3500367x17.png" xlink:type="simple"/></inline-formula> is pmf, <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-3500367x18.png" xlink:type="simple"/></inline-formula>is the joint probability. The function <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-3500367x19.png" xlink:type="simple"/></inline-formula> is the likelihood function, which depends on the unknown parameter θ, and it is denoted as L(θ). In order to get the maximum likelihood function, a value of θ for which the likelihood function L(θ) is a maximum is used as an estimate of θ. Maximizing L(θ) with a product of n terms is equivalent to maximizing logL(θ) because log is a monotonic increasing function. logL(θ) is a log likelihood function, and is denoted as LL(θ), as follows [<xref ref-type="bibr" rid="scirp.77395-ref10">10</xref>] :</p><disp-formula id="scirp.77395-formula65"><label>(6)</label><graphic position="anchor" xlink:href="http://html.scirp.org/file/4-3500367x20.png"  xlink:type="simple"/></disp-formula></sec><sec id="s6"><title>6. The Effect of Independent Variables</title><p>The effect of any independent variable on the outcome can be tested using the likelihood ratio (LR) statistic test. If the dependent variable has M categories, then there are M − 1 non redundant coefficients (β<sub>n</sub>) associated with each independent variable x<sub>n</sub>. The null hypothesis that x<sub>n</sub> does not affect the dependent variable can be written as:</p><disp-formula id="scirp.77395-formula66"><label>(7)</label><graphic position="anchor" xlink:href="http://html.scirp.org/file/4-3500367x21.png"  xlink:type="simple"/></disp-formula><p>where Base is the base category used in the model. The hypothesis can be tested with the LR test. First, the LR estimates the full model that contains all of the independent variables with the resulting LR statistic LR<sub>F</sub>. Second, the LR estimates the restricted model formed by excluding the independent variable x<sub>n</sub> with the resulting LR statistic LR<sub>R</sub>. Finally, the LR estimates the difference between LR<sub>F</sub> and LR<sub>R</sub> which is distributed as chi-square with n degrees of freedom (the number of independent variables). The LR statistic is computed in terms of log likelihood (LL) as follows [<xref ref-type="bibr" rid="scirp.77395-ref5">5</xref>] [<xref ref-type="bibr" rid="scirp.77395-ref10">10</xref>] :</p><disp-formula id="scirp.77395-formula67"><label>(8)</label><graphic position="anchor" xlink:href="http://html.scirp.org/file/4-3500367x22.png"  xlink:type="simple"/></disp-formula><disp-formula id="scirp.77395-formula68"><label>(9)</label><graphic position="anchor" xlink:href="http://html.scirp.org/file/4-3500367x23.png"  xlink:type="simple"/></disp-formula><p>Alternatively, the null model is given by (−2log(L<sub>0</sub>)) where L<sub>0</sub> is the likelihood of obtaining the observations if the independent variables had no effect on the outcome (i.e. model with intercept alone). The full model is given by (−2log(L)) where L is the likelihood of obtaining the observations with all independent variables incorporated in the model. The difference of these two yields a Chi-Squared statistic which is a measure of how well the independent variables affect the outcome or dependent variable [<xref ref-type="bibr" rid="scirp.77395-ref1">1</xref>] . If the LR statistic for the overall model is significant, then there is evidence that the independent variables have contributed to the prediction of the outcome.</p></sec><sec id="s7"><title>7. The Independence of Irrelevant Alternatives (IIA)</title><p>The MNL assumes that the odd ratios for any pair of outcomes (i.e. any pair of the dependent variable categories) are determined without reference to the other categories that might be available [<xref ref-type="bibr" rid="scirp.77395-ref2">2</xref>] [<xref ref-type="bibr" rid="scirp.77395-ref3">3</xref>] . This assumption is called the independence of irrelevant alternatives (IIA), which is very crucial in the MNL modeling. If the IIA holds, then the MNL model can be used, if the IIA does not hold, then the MNL cannot be used and alternative models should be utilized such as, the nested MNL. The IIA can be tested by the Hausman specification test, proposed by Hausman and McFadden [<xref ref-type="bibr" rid="scirp.77395-ref11">11</xref>] , which proceeds by estimating the error coefficients of the full model with all categories of the dependent variable included, then estimating the error coefficients of a restricted model by eliminating one or more outcome categories. The null hypothesis of the test is that the IIA does not exist and estimators of the full and restricted models are consistent, and under the alternative hypothesis the IIA does exist and only the estimators of the restricted model are consistent. The test statistic H<sub>IIA</sub> is asymptotically distributed as chi square, and significant values of H<sub>IIA</sub> indicate that the IIA assumption is violated [<xref ref-type="bibr" rid="scirp.77395-ref11">11</xref>] . The Hausman specification test involves the following steps:</p><p>1) Estimate the error coefficients of the full model with all M categories of the dependent variable included; these coefficients are contained in<inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-3500367x24.png" xlink:type="simple"/></inline-formula>.</p><p>2) Estimate the error coefficients of a restricted model by eliminating one or more outcome categories; theses coefficients are contained in<inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-3500367x25.png" xlink:type="simple"/></inline-formula>.</p><p>3) Let <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-3500367x26.png" xlink:type="simple"/></inline-formula> represents <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-3500367x26.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-3500367x27.png" xlink:type="simple"/></inline-formula> after eliminating all coefficients not estimated in the restricted model. The Hausman specification test of IIA is defined as [<xref ref-type="bibr" rid="scirp.77395-ref11">11</xref>] :</p><disp-formula id="scirp.77395-formula69"><label>(10)</label><graphic position="anchor" xlink:href="http://html.scirp.org/file/4-3500367x28.png"  xlink:type="simple"/></disp-formula><p>H<sub>IIA</sub> is asymptotically distributed as chi square with degrees of freedom equal to the rows in<inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-3500367x29.png" xlink:type="simple"/></inline-formula>. In this dissertation, the Hausman specification test will be applied on each outcome pair of the dependent variable (i.e. crash severity) separately, excluding the other category of the dependent variable. Since the property damage is assumed to be the base category, as it is the most frequent occurred category, therefore the test will be applied on the minor injury vs. disabled injury first, and second; it will be applied on the minor injury vs. fatal injury, and lastly; it will be applied on the disabled injury vs. fatal injury. For each outcome pair, the test statistic H<sub>IIA</sub> will be obtained and compared to the full model with all outcomes. If the value of H<sub>IIA</sub> for any pair is significant, then the IIA assumption is violated and the MNL cannot be used in the modeling process. If the values of H<sub>IIA</sub> for all pairs are insignificant, then the IIA assumption holds and the MNL can be used in the modeling process.</p></sec><sec id="s8"><title>8. Multicollinearity</title><p>Multi-collinearity is the existence of linear relationships among the independent variables that can create inaccurate estimates of the regression coefficients, inflate the standard errors of the regression coefficients, give false, non-significant p-values, and degrade the predictability of the model [<xref ref-type="bibr" rid="scirp.77395-ref1">1</xref>] . The source of the multi-collinearity might come from data collection, sampling techniques, political or legal constraints, and outliers. Testing the multi-collinearity can be achieved by: (1) visual inspection of pairwise scatter plots of independent variables, and looking for near-perfect linear relationships between them; (2) Eigenvalues and Condition Indices; and (3) considering the variance inflation factors (VIF). The VIF is the most widely used test to measure how much the variance of the estimated regression coefficients are inflated as compared to when the predictor variables are not linearly related. The VIF may be calculated for each predictor by doing a linear regression of that predictor on all the other predictors, and then obtaining the R<sup>2</sup> from that regression. The VIFs obtained by the linear regression can still be used in logistic regression models, because the concern is with the relationship among the independent variables included in the model, not with the functional form of the model [<xref ref-type="bibr" rid="scirp.77395-ref12">12</xref>] . Thus, a VIF of 1.6 tells us that the variance (the square of the standard error) of a particular coefficient is 60% larger than it would be if that predictor was completely uncorrelated with all other predictors. The VIF has a lower value of 1.0 but no upper bound. As a rule of thumb, if VIF is more than 10.0, then multicollinearity is considered a serious problem, and must be corrected [<xref ref-type="bibr" rid="scirp.77395-ref1">1</xref>] [<xref ref-type="bibr" rid="scirp.77395-ref12">12</xref>] . Variance inflation factors are scaled measures of the correlation coefficient between variable j and the rest of the independent variables. Specifically:</p><disp-formula id="scirp.77395-formula70"><label>(11)</label><graphic position="anchor" xlink:href="http://html.scirp.org/file/4-3500367x30.png"  xlink:type="simple"/></disp-formula><p>where,</p><p><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-3500367x31.png" xlink:type="simple"/></inline-formula>: is the coefficient of determination of the regression model that includes all predictors except the j<sup>th</sup> predictor.</p><p>Variance inflation factors are often given as the reciprocal of the above formula. In this case, they are referred to as the tolerances. If <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-3500367x32.png" xlink:type="simple"/></inline-formula> equals zero (i.e. no correlation between j and the remaining independent variables), then VIF<sub>j</sub> equals 1.0, and this is the minimum value.</p></sec><sec id="s9"><title>9. The Generalized Hosmer-Lemeshow Statistic</title><p>The generalized Hosmer-Lemeshow test is used as an important goodness of fit measure to assess whether or not the observed events match expected events, by sub grouping the probabilities estimated from the data [<xref ref-type="bibr" rid="scirp.77395-ref13">13</xref>] [<xref ref-type="bibr" rid="scirp.77395-ref14">14</xref>] . The data set, of size n, is sorted according to the probabilities estimated from the final fitted MNL model. Then the data set is partitioned into several (Hosmer and Lemeshow recommended 10) equal-sized groups. The first group corresponds to the n/10 observations having the highest estimated probabilities. The next group corresponds to the n/10 observations having the next highest estimated probabilities, etc. A Pearson-like chi square statistic is constructed based on the observed and expected group frequencies. In order to get the generalized test statistic (HL), we suppose that we have a sample of n independent observations,<inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-3500367x33.png" xlink:type="simple"/></inline-formula>. Recoding y<sub>i</sub> into binary indicator variables y<sub>ij</sub>, such that y<sub>ij</sub> = 1 when y<sub>i</sub>= j and y<sub>ij</sub>= 0, otherwise (<inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-3500367x33.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-3500367x34.png" xlink:type="simple"/></inline-formula>and<inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-3500367x33.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-3500367x34.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-3500367x35.png" xlink:type="simple"/></inline-formula>). After fitting the model, let π<sub>ij</sub> denote the estimated probabilities for each observation (<inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-3500367x33.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-3500367x34.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-3500367x35.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-3500367x36.png" xlink:type="simple"/></inline-formula>) for each possible outcome (<inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-3500367x33.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-3500367x34.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-3500367x35.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-3500367x36.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-3500367x37.png" xlink:type="simple"/></inline-formula>). By sorting the observations according to 1 − π<sub>i</sub><sub>0</sub>, the complement of the estimated probability of the reference outcome. We then form g groups, each containing approximately n/g observations. For each group, we calculate the sums of the observed and estimated frequencies for each outcome category as follows [<xref ref-type="bibr" rid="scirp.77395-ref15">15</xref>] :</p><disp-formula id="scirp.77395-formula71"><label>(12)</label><graphic position="anchor" xlink:href="http://html.scirp.org/file/4-3500367x38.png"  xlink:type="simple"/></disp-formula><disp-formula id="scirp.77395-formula72"><label>(13)</label><graphic position="anchor" xlink:href="http://html.scirp.org/file/4-3500367x39.png"  xlink:type="simple"/></disp-formula><p>where O<sub>kj</sub> is the observed frequency, E<sub>jk</sub> is the expected frequency,<inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-3500367x40.png" xlink:type="simple"/></inline-formula>;<inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-3500367x40.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-3500367x41.png" xlink:type="simple"/></inline-formula>; and Ω<sub>k</sub> denotes indices of the n/g observations in group k. The multinomial goodness-of-fit (HL) test statistic is the Pearson’s chi-squared statistic from the table of observed and estimated frequencies, and is given as [<xref ref-type="bibr" rid="scirp.77395-ref15">15</xref>] :</p><disp-formula id="scirp.77395-formula73"><label>(14)</label><graphic position="anchor" xlink:href="http://html.scirp.org/file/4-3500367x42.png"  xlink:type="simple"/></disp-formula><p>The distribution of C<sub>g</sub> is chi-squared and has <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-3500367x43.png" xlink:type="simple"/></inline-formula> degrees of freedom [<xref ref-type="bibr" rid="scirp.77395-ref16">16</xref>] . The null hypothesis is that the differences between the observed and predicted events are insignificant so the fitted model is correct, while the alternative hypothesis is that the differences are significant so the fitted model has deficiency and incorrect. If the test statistic HL is insignificant, then we will accept the null hypothesis, and conclude that the fitted model is a good fit. If the test statistic HL is significant, then we will reject the null hypothesis, and conclude that the data do not fit the hypothesized fitted MNL regression model.</p></sec><sec id="s10"><title>10. The Classification <xref ref-type="table" rid="table">Table </xref>of MNL</title><p>The classification table is another method to assess the goodness of fit of the MNL regression model. In this table the observed values for the dependent outcome and the predicted values (at a user defined cut-off value, for example p = 0.50) are cross-classified to indicate the correct % of predicted cases. This percent statistic assumes that if the estimated p is greater than or equal to 0.5 then the event is expected to occur and not occur otherwise. The bigger the % correct predictions, the better the model fit. We suppose for n observations that <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-3500367x44.png" xlink:type="simple"/></inline-formula> is the <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-3500367x44.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-3500367x45.png" xlink:type="simple"/></inline-formula> element of the classification table,<inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-3500367x44.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-3500367x45.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-3500367x46.png" xlink:type="simple"/></inline-formula>. <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-3500367x44.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-3500367x45.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-3500367x46.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-3500367x47.png" xlink:type="simple"/></inline-formula>is the sum of the frequencies for the observations whose actual response category is j (as row) and predicted response category is <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-3500367x44.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-3500367x45.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-3500367x46.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-3500367x47.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-3500367x48.png" xlink:type="simple"/></inline-formula> (as column) respectively. Then, the percentage of total correct predictions of the model is given by [<xref ref-type="bibr" rid="scirp.77395-ref4">4</xref>] [<xref ref-type="bibr" rid="scirp.77395-ref17">17</xref>] :</p><disp-formula id="scirp.77395-formula74"><label>(15)</label><graphic position="anchor" xlink:href="http://html.scirp.org/file/4-3500367x49.png"  xlink:type="simple"/></disp-formula><p>The percentage of correct predictions for response category j is given by:</p><disp-formula id="scirp.77395-formula75"><label>(16)</label><graphic position="anchor" xlink:href="http://html.scirp.org/file/4-3500367x50.png"  xlink:type="simple"/></disp-formula></sec><sec id="s11"><title>11. The Pseudo R-Squares</title><p>In ordinary least squared (OLS) regression there is a non-pseudo R-square, which is often generated as a goodness-of-fit measure, and is given by:</p><disp-formula id="scirp.77395-formula76"><label>(17)</label><graphic position="anchor" xlink:href="http://html.scirp.org/file/4-3500367x51.png"  xlink:type="simple"/></disp-formula><p>where n is the number of observations in the model, y is the dependent variable, y-bar is the mean of the y values, and y-hat is the value predicted by the model. The numerator of the ratio is the sum of the squared differences between the actual y values and the predicted y values. The denominator of the ratio is the sum of squared differences between the actual y values and their mean.</p><p>When analyzing data with a multinomial logistic regression, there is no an equivalent statistic to R-squared. The estimates from a logistic regression are found by the maximum likelihood estimation rather than the least squared estimation, so the OLS approach to goodness-of-fit does not apply. However, to evaluate the goodness-of-fit of logistic models, several pseudo R-squares have been developed. They are called “pseudo” R-squares because they are on a similar scale, ranging from 0 to 1 (though some pseudo R-squares never achieve 0 or 1) with higher values indicating better model fit, but they cannot be interpreted as one would interpret an OLS R-squared, and different pseudo R- squares can present different values [<xref ref-type="bibr" rid="scirp.77395-ref12">12</xref>] . Some of the popular pseudo R-squares are:</p><p>McFadden’s R-square, which is defined as [<xref ref-type="bibr" rid="scirp.77395-ref18">18</xref>] :</p><disp-formula id="scirp.77395-formula77"><label>(18)</label><graphic position="anchor" xlink:href="http://html.scirp.org/file/4-3500367x52.png"  xlink:type="simple"/></disp-formula><p>where L<sub>0</sub> is the value of the likelihood function for a model with no predictors (i.e. with intercept only), and L<sub>M</sub> is the likelihood function for the model being estimated. The ratio of the McFadden R-square indicates the level of improvement over the intercept model offered by the full model. Since a likelihood falls between 0.0 and 1.0, the log of a likelihood is less than or equal to zero. If a model has a very low likelihood, then the log of the likelihood will have a larger magnitude than the log of a more likely model. Thus, a small ratio of log likelihoods indicates that the full model is a far better fit than the intercept model. When comparing two models on the same data, McFadden’s would be higher for the model with the greater likelihood. Another pseudo R-square is the Cox and Snell R<sup>2</sup> which is defined as [<xref ref-type="bibr" rid="scirp.77395-ref19">19</xref>] :</p><disp-formula id="scirp.77395-formula78"><label>(19)</label><graphic position="anchor" xlink:href="http://html.scirp.org/file/4-3500367x53.png"  xlink:type="simple"/></disp-formula><p>where n is the sample size. The Cox and Snell R-square indicates the level of improvement of the full model over the intercept model. This pseudo R-squared has a maximum value that is less than 1.0 when the full model predicts the outcome perfectly and has a likelihood of 1.0. The Nagelkerke R-square adjusts Cox &amp; Snell’s so that the range of possible values extends to 1.0 by dividing by its maximum possible value,<inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-3500367x54.png" xlink:type="simple"/></inline-formula>. If the full model perfectly predicts the outcome and has a likelihood of 1.0, then the Nagelkerke R-square = 1.0, which is defined as [<xref ref-type="bibr" rid="scirp.77395-ref20">20</xref>] :</p><disp-formula id="scirp.77395-formula79"><label>(20)</label><graphic position="anchor" xlink:href="http://html.scirp.org/file/4-3500367x55.png"  xlink:type="simple"/></disp-formula><p>Pseudo R-squares are useful tools in evaluating multiple models predicting the same outcome on the same dataset, but they cannot be interpreted independently or compared across different datasets. In other words, a pseudo R-squared statistic without context has little meaning. A pseudo R-squared only has meaning when compared to another pseudo R-squared of the same type, on the same data, predicting the same outcome [<xref ref-type="bibr" rid="scirp.77395-ref12">12</xref>] [<xref ref-type="bibr" rid="scirp.77395-ref21">21</xref>] . In this case, the higher pseudo R-squared indicates which model better predicts the outcome.</p></sec><sec id="s12"><title>12. Estimation of Marginal Effects</title><p>Marginal effects are useful estimates of the impact of a one-unit change of an independent variable (predictor) on the dependent variable. The average marginal effects are interpreted as the effect of a one-unit change in an independent variable (keeping all other independent variables constant at their mean values) on dependent variable. It is common to use a single average marginal effect value for all observations of an independent variable. Elasticity analysis can also be used to interpret the effect of a specific independent variable on the dependent variable, but with a 1.0% change instead of a one-unit change. In MNL, the marginal effect of an explanatory variable (predictor) is the partial derivative of the event probability with respect to the predictor of interest (i.e. the change in the event probability for a unit change in the predictor). The marginal effect for a dummy independent variable is the difference of the predicted probability values at their different levels [<xref ref-type="bibr" rid="scirp.77395-ref17">17</xref>] . The values of the marginal effects reflect the slopes of lines tangent to each of the predictors that is drawn tangent to the fitted probability curve at the selected point. The slope of the tangent line is the change in event probability, p, measured at two points, one unit apart along this straight line. If the probability curve is linear (near p = 0.5) at the selected point, then the marginal effect will approximate the probability change when changing the predictor by one unit. If the probability curve is nonlinear (near the smallest and largest values of p), the marginal effect might deviate from the change [<xref ref-type="bibr" rid="scirp.77395-ref4">4</xref>] [<xref ref-type="bibr" rid="scirp.77395-ref17">17</xref>] . For multinomial logistic regression models, the possible response values are unordered with levels<inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-3500367x56.png" xlink:type="simple"/></inline-formula>. The probability of response level i is given by [<xref ref-type="bibr" rid="scirp.77395-ref22">22</xref>] :</p><disp-formula id="scirp.77395-formula80"><label>(20)</label><graphic position="anchor" xlink:href="http://html.scirp.org/file/4-3500367x57.png"  xlink:type="simple"/></disp-formula><p>where <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-3500367x58.png" xlink:type="simple"/></inline-formula> is the predictor of interest, and <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-3500367x58.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-3500367x59.png" xlink:type="simple"/></inline-formula> is the regression coefficient (i.e. log odd) of<inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-3500367x58.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-3500367x59.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-3500367x60.png" xlink:type="simple"/></inline-formula>. The marginal effect of the j<sup>th</sup> predictor, X<sub>j</sub>, on p<sub>i</sub> is given by:</p><disp-formula id="scirp.77395-formula81"><label>(21)</label><graphic position="anchor" xlink:href="http://html.scirp.org/file/4-3500367x61.png"  xlink:type="simple"/></disp-formula></sec><sec id="s13"><title>13. Testing the Effects of Independent Variables</title><p>Multinomial logistic regression (MNL) is usually conducted using maximum likelihood estimation, which is an iterative procedure. The first iteration (called iteration zero) is the log likelihood of the null or empty model; that is, a model with no predictors. At the next iteration, the predictors are included in the model. At each iteration, the log likelihood decreases as the goal is to minimize the log likelihood. When the difference between successive iterations is very small, the model is said to have converged, the iterating stops, and the final log likelihood (LR) statistic is computed. The log likelihood ration (LR) test statistic is obtained for the I-70 corridor for both the training and testing data, using the Stata 14 software package and reported in <xref ref-type="table" rid="table">Table </xref>1.</p><p>The effect of any independent variable on the outcome can be tested using the likelihood ratio (LR) statistic test. The null hypothesis of this test is that the independent variables do not affect the dependent variable. The null model is calculated by obtaining the log likelihood of the observations with just the response variable in the model from iteration zero (i.e. model with intercept alone). The final fitted model is calculated by obtaining the log likelihood of observations with all the independent variables in the model from the final iteration after convergence. The difference of these two yields a chi-squared LR statistic which is a measure of how well the independent variables affect the outcomes or dependent variable categories [<xref ref-type="bibr" rid="scirp.77395-ref1">1</xref>] . If the LR statistic for the overall model is significant, then there is evidence that the independent variables are effective and they have contributed to the prediction of the outcome. <xref ref-type="table" rid="table">Table </xref>1 shows that the Likelihood Ratio (LR) test statistic for the I-70 corridor is significant at the 95% confidence level with p-values less than 0.05 for the training and testing datasets, implying that all the independent variables included in the models are not equal to zero, and this indicates that they are effectively contributing to modeling the</p><table-wrap id="table1" ><label><xref ref-type="table" rid="table">Table </xref>1</label><caption><title> The LR statistic results</title></caption><table><tbody><thead><tr><th align="center" valign="middle" >Dataset</th><th align="center" valign="middle" ># Observations</th><th align="center" valign="middle" >LR statistic</th><th align="center" valign="middle" >p-value</th></tr></thead><tr><td align="center" valign="middle" >I-70 Training data</td><td align="center" valign="middle" >4108</td><td align="center" valign="middle" >339.12</td><td align="center" valign="middle" >0.0000</td></tr><tr><td align="center" valign="middle" >I-70 Testing data</td><td align="center" valign="middle" >1761</td><td align="center" valign="middle" >122.44</td><td align="center" valign="middle" >0.0000</td></tr></tbody></table></table-wrap><table-wrap id="table2" ><label><xref ref-type="table" rid="table">Table </xref>2</label><caption><title> The IIA assumption results</title></caption><table><tbody><thead><tr><th align="center" valign="middle"  rowspan="2"  >Dataset</th><th align="center" valign="middle"  colspan="2"  >Minor injury vs. disabled</th><th align="center" valign="middle"  colspan="2"  >Minor injury vs. fatal</th><th align="center" valign="middle"  colspan="2"  >Disabled vs. fatal</th></tr></thead><tr><td align="center" valign="middle" >H<sub>IIA</sub></td><td align="center" valign="middle" >p-value</td><td align="center" valign="middle" >H<sub>IIA</sub></td><td align="center" valign="middle" >p-value</td><td align="center" valign="middle" >H<sub>IIA</sub></td><td align="center" valign="middle" >p-value</td></tr><tr><td align="center" valign="middle" >I-70 training</td><td align="center" valign="middle" >1.46</td><td align="center" valign="middle" >0.5461</td><td align="center" valign="middle" >1.39</td><td align="center" valign="middle" >0.6725</td><td align="center" valign="middle" >1.73</td><td align="center" valign="middle" >0.7748</td></tr><tr><td align="center" valign="middle" >I-70 testing</td><td align="center" valign="middle" >1.08</td><td align="center" valign="middle" >0.6726</td><td align="center" valign="middle" >1.14</td><td align="center" valign="middle" >0.7453</td><td align="center" valign="middle" >1.24</td><td align="center" valign="middle" >0.6833</td></tr></tbody></table></table-wrap><p>crash severity for all categories. Thus, it can be concluded that the overall chosen models for the I-70 corridor data are good fits.</p></sec><sec id="s14"><title>14. Testing the IIA Assumption</title><p>The Independence of Irrelevant Alternatives (IIA) assumption in multinomial logistic regression means that adding or deleting alternative outcome categories does not affect the odd ratios among the remaining outcomes [<xref ref-type="bibr" rid="scirp.77395-ref2">2</xref>] [<xref ref-type="bibr" rid="scirp.77395-ref3">3</xref>] . The Hausman specification test is used to test the IIA assumption for the I-70 dataset (both training and testing datasets). The results of this test are shown in <xref ref-type="table" rid="table">Table </xref>2, as computed using the Stata 14 software package.</p><p>The null hypothesis of the test is that the IIA does not exist and under the alternative hypothesis the IIA does exist. The Hausman specification test statistic H<sub>IIA</sub> is asymptotically distributed as chi square, and significant values of H<sub>IIA</sub> indicate that the IIA assumption is violated [<xref ref-type="bibr" rid="scirp.77395-ref11">11</xref>] . The Hausman specification test was run on each outcome pair of the dependent variable (i.e. crash severity) separately, excluding the other category of the dependent variable. The base category was assumed to be the records were property damage was reported. First, the test was run on the second vs. the third categories (i.e. minor injury vs. disabled), second; it was run on the second vs. the fourth categories (i.e. minor injury vs. fatal), and lastly; it was run on the third vs. the fourth categories (i.e. disabled vs. fatal). <xref ref-type="table" rid="table">Table </xref>2 shows that for all cases the H<sub>IIA</sub> statistic was insignificant at the 95% confidence level with their p-values greater than 0.05 for the I-70 corridor datasets. Therefore, the null hypothesis can be accepted and it can be concluded that the IIA assumption has not been violated so that the odd ratios of any outcome pair of the dependent variable are determined without reference to the other category.</p></sec><sec id="s15"><title>15. Testing the Generalized Hosmer-Lemeshow Statistic</title><p>The generalized Hosmer-Lemeshow statistic assesses whether or not the observed events match the predicted events, by subgrouping the probabilities estimated from the data [<xref ref-type="bibr" rid="scirp.77395-ref13">13</xref>] [<xref ref-type="bibr" rid="scirp.77395-ref14">14</xref>] . This test works by sorting the data according to the probabilities estimated from the final fitted MNL model. Then the sorted dataset is partitioned into several equal-sized groups. Then, the HL test statistic that follows a chi-square distribution is constructed based on the observed and predicted group frequencies. The null hypothesis is that the differences between the observed and predicted events are insignificant so the fitted model is correct, while the alternative hypothesis is that the differences are significant so the fitted</p><table-wrap id="table3" ><label><xref ref-type="table" rid="table">Table </xref>3</label><caption><title> The Generalized Hosmer-Lemeshow test results</title></caption><table><tbody><thead><tr><th align="center" valign="middle" >Dataset</th><th align="center" valign="middle" ># Observations</th><th align="center" valign="middle" ># Groups</th><th align="center" valign="middle" >HL statistic</th><th align="center" valign="middle" >p-value</th></tr></thead><tr><td align="center" valign="middle" >I-70 training</td><td align="center" valign="middle" >4108</td><td align="center" valign="middle" >10</td><td align="center" valign="middle" >27.406</td><td align="center" valign="middle" >0.286</td></tr><tr><td align="center" valign="middle" >I-70 testing</td><td align="center" valign="middle" >1761</td><td align="center" valign="middle" >10</td><td align="center" valign="middle" >27.134</td><td align="center" valign="middle" >0.298</td></tr></tbody></table></table-wrap><p>model has deficiency and incorrect. If the test statistic HL is insignificant, then we will accept the null hypothesis, and conclude that the fitted model is a good fit. If the test statistic HL is significant, then we will reject the null hypothesis, and conclude that the data do not fit the hypothesized fitted MNL regression model. The generalized Hosmer-Lemeshow test is applied to the I-70 dataset (both training and testing datasets) with ten groups for each dataset. This test was again conducted using the Stata 14 software package and the results of this test are summarized in <xref ref-type="table" rid="table">Table </xref>3.</p><p><xref ref-type="table" rid="table">Table </xref>3 shows that the HL test statistic for the I-70 corridor is insignificant at the 95% confidence level with p-values larger than 0.05 for the training and testing datasets. Therefore, the null hypothesis cannot be rejected and it can be concluded that the overall models of I-70 corridor are good fit, and there is a good match between the predicted events and the observed events for all categories of the dependent variable.</p></sec><sec id="s16"><title>16. Testing the Multicollinearity</title><p>Multicollinearity occurs when two or more predictors in the model are highly correlated that can create inaccurate estimates of the regression coefficients, and inflate the standard errors. The MNL model requires that multicollinearity be low between predictors in the model. To test for this assumption, the variance inflation factor (VIF) is used to detect multicollinearity among all predictors in our MNL logistic regression models, as it is the most widely used test formulticollinearity [<xref ref-type="bibr" rid="scirp.77395-ref23">23</xref>] . The VIF measures how much the variance of the estimated regression coefficients is inflated as compared to when the predictors are not linearly related. The VIF may be calculated for each predictor by doing a linear regression of that predictor on all the other predictors. The VIFs obtained by the linear regression can still be used in logistic regression models, because the concern is with the relationship among the independent variables included in the model, not with the functional form of the model [<xref ref-type="bibr" rid="scirp.77395-ref12">12</xref>] . The VIF has a lower value of 1.0 but no upper bound. As a rule of thumb, if VIF is more than 10.0, then multicollinearity is considered a serious problem, and must be corrected [<xref ref-type="bibr" rid="scirp.77395-ref12">12</xref>] [<xref ref-type="bibr" rid="scirp.77395-ref23">23</xref>] . The VIF statistic is obtained for the I-70 corridor data using the Stata 14 and the results are reported in <xref ref-type="table" rid="table">Table </xref>4.</p><p>The VIFs of all the independent variables are considerably less than 10.0 for the I-70 datasets as can be seen from <xref ref-type="table" rid="table">Table </xref>4. The VIFs of the independent variables (Direction and Grade-Level) of the I-70 dataset are 6.397 and 6.457 respectively, but they are still less than 10.0. The VIFs of the other predictors are even less than 5.0. Based on this, it can be concluded that multicollinearity is not a serious problem in both datasets, and this implies that the assumption of low</p><table-wrap id="table4" ><label><xref ref-type="table" rid="table">Table </xref>4</label><caption><title> VIF results</title></caption><table><tbody><thead><tr><th align="center" valign="middle" >MONTH</th><th align="center" valign="middle" >1.023</th></tr></thead><tr><td align="center" valign="middle" >DAY_WEEK</td><td align="center" valign="middle" >1.013</td></tr><tr><td align="center" valign="middle" >HOUR</td><td align="center" valign="middle" >1.026</td></tr><tr><td align="center" valign="middle" >NO_VEHICLE</td><td align="center" valign="middle" >2.099</td></tr><tr><td align="center" valign="middle" >DIRECTION</td><td align="center" valign="middle" >6.397</td></tr><tr><td align="center" valign="middle" >LIGHT_COND</td><td align="center" valign="middle" >1.113</td></tr><tr><td align="center" valign="middle" >ACC_TYPE</td><td align="center" valign="middle" >2.264</td></tr><tr><td align="center" valign="middle" >DR_DRINK</td><td align="center" valign="middle" >1.046</td></tr><tr><td align="center" valign="middle" >SPEED</td><td align="center" valign="middle" >1.408</td></tr><tr><td align="center" valign="middle" >CZONE</td><td align="center" valign="middle" >1.072</td></tr><tr><td align="center" valign="middle" >DR_AGGRESSIVE</td><td align="center" valign="middle" >1.373</td></tr><tr><td align="center" valign="middle" >CELL_TEXT</td><td align="center" valign="middle" >1.008</td></tr><tr><td align="center" valign="middle" >DR_AGE</td><td align="center" valign="middle" >1.015</td></tr><tr><td align="center" valign="middle" >VEH_TYPE</td><td align="center" valign="middle" >1.044</td></tr><tr><td align="center" valign="middle" >RURAL_URBAN</td><td align="center" valign="middle" >2.455</td></tr><tr><td align="center" valign="middle" >NUMBER_LANES</td><td align="center" valign="middle" >3.504</td></tr><tr><td align="center" valign="middle" >AADT</td><td align="center" valign="middle" >4.896</td></tr><tr><td align="center" valign="middle" >GRADE_LEVEL</td><td align="center" valign="middle" >6.457</td></tr></tbody></table></table-wrap><p>multicollinearity is achieved in the MLN model.</p></sec><sec id="s17"><title>17. The Classification Table</title><p>The classification table is used to assess the goodness of fit of the MNL regression model. In this table the observed values for the dependent outcomes and the predicted values (at a user defined cut-off value) are cross-classified to indicate the correct % of predicted cases. This percent statistic assumes that if the predicted probability is greater than or equal to the (cut-off value) then the event is expected to occur and not occur otherwise. The bigger the % correct predictions, the better the model fit. The classification tables for the I-70 corridor dataset (for both training and testing data) are obtained using the SPSS 23 and the results are detailed in <xref ref-type="table" rid="table">Table </xref>5.</p><p><xref ref-type="table" rid="table">Table </xref>5 shows how many cases are correctly predicted for each category of the dependent variable. For example, for the I-70 training data, there are 3168 observed incidents involving property damage and the percent correctly predicted is 99.6%, 785 observed incidents involving minor injury with 65.4% correctly predicted, 114 observed incidents involving disabled with 72.8% correctly predicted, and 23 observed incidents involving fatal crashes and the percent correctly predicted is 77.1%. The overall percentage gives the overall percent of cases that are correctly predicted by the full model, which is 92.2% for the I-70 training data and 91.5% for testing data. This overall percentage is an important</p><table-wrap id="table5" ><label><xref ref-type="table" rid="table">Table </xref>5</label><caption><title> I-70 classification table results</title></caption><table><tbody><thead><tr><th align="center" valign="middle"  rowspan="2"  >Severity categories</th><th align="center" valign="middle"  colspan="3"  >I-70 training data</th><th align="center" valign="middle"  colspan="3"  >I-70 testing data</th></tr></thead><tr><td align="center" valign="middle" ># obs.</td><td align="center" valign="middle" >% correct</td><td align="center" valign="middle" >Overall % correct</td><td align="center" valign="middle" ># obs.</td><td align="center" valign="middle" >% correct</td><td align="center" valign="middle" >Overall % correct</td></tr><tr><td align="center" valign="middle" >Property damage</td><td align="center" valign="middle" >3186</td><td align="center" valign="middle" >99.6%</td><td align="center" valign="middle"  rowspan="4"  >92.2%</td><td align="center" valign="middle" >1372</td><td align="center" valign="middle" >97.3%</td><td align="center" valign="middle"  rowspan="4"  >91.5%</td></tr><tr><td align="center" valign="middle" >Minor injury</td><td align="center" valign="middle" >785</td><td align="center" valign="middle" >65.4%</td><td align="center" valign="middle" >323</td><td align="center" valign="middle" >69.8%</td></tr><tr><td align="center" valign="middle" >Disabled</td><td align="center" valign="middle" >114</td><td align="center" valign="middle" >72.8%</td><td align="center" valign="middle" >52</td><td align="center" valign="middle" >76.2%</td></tr><tr><td align="center" valign="middle" >Fatal</td><td align="center" valign="middle" >23</td><td align="center" valign="middle" >77.1%</td><td align="center" valign="middle" >14</td><td align="center" valign="middle" >83.6%</td></tr></tbody></table></table-wrap><table-wrap id="table6" ><label><xref ref-type="table" rid="table">Table </xref>6</label><caption><title> The pseudo R-squares results</title></caption><table><tbody><thead><tr><th align="center" valign="middle"  rowspan="2"  >Pseudo R-square</th><th align="center" valign="middle"  colspan="2"  >I-70 training</th><th align="center" valign="middle"  colspan="2"  >I-70 testing</th></tr></thead><tr><td align="center" valign="middle" >Intercept</td><td align="center" valign="middle" >Full</td><td align="center" valign="middle" >Intercept</td><td align="center" valign="middle" >Full</td></tr><tr><td align="center" valign="middle" >McFadden</td><td align="center" valign="middle" >0.025</td><td align="center" valign="middle" >0.118</td><td align="center" valign="middle" >0.028</td><td align="center" valign="middle" >0.138</td></tr><tr><td align="center" valign="middle" >Cox-Snell</td><td align="center" valign="middle" >0.031</td><td align="center" valign="middle" >0.123</td><td align="center" valign="middle" >0.047</td><td align="center" valign="middle" >0.147</td></tr><tr><td align="center" valign="middle" >Nagelkerke</td><td align="center" valign="middle" >0.046</td><td align="center" valign="middle" >0.132</td><td align="center" valign="middle" >0.054</td><td align="center" valign="middle" >0.166</td></tr></tbody></table></table-wrap><p>goodness-of-fit measure that indicates how well the data have fitted the full model. These overall percentages of correctly predicted cases demonstrate that our MNL models are good fit, confirming the results obtained by the generalized Hosmer-Lemeshow test statistic that there is a good match between the predicted events and the observed events for all categories of the dependent variable.</p></sec><sec id="s18"><title>18. The Pseudo R-Squares</title><p>Multinomial logistic regression does not have an equivalent to the R-squared that is found in ordinary least square regression; however, there are some pseudo-R-square statistics that have been developed for MNL. The McFadden R-square treats the log likelihood of the intercept model as a total sum of squares, and the log likelihood of the full model as the sum of squared errors, the Cox and Snell’s R-square reflects the improvement of the full model over the intercept model through the ratio of log likelihood, and the Nagelkerke R-square try to adjust the Cox and Snell’s so that the range of possible values extends to 1.0. Pseudo R-squares are generally useful tools in evaluating multiple models predicting the same outcome on the same dataset, but they cannot be interpreted independently or compared across different datasets [<xref ref-type="bibr" rid="scirp.77395-ref12">12</xref>] [<xref ref-type="bibr" rid="scirp.77395-ref21">21</xref>] . In this case, the higher pseudo R-squared indicates which model better predicts the outcome. Three types of pseudo R-squares (McFadden’s, Cox and Snell’s, and Nagelkerke’s) are obtained for the I-70 corridor (both training and testing datasets), using SPSS 23, as shown in <xref ref-type="table" rid="table">Table </xref>6. First, these pseudo R-squares are applied to the intercept only model for each dataset, and then they are applied to the full model with all predictors to capture any improvement in the fitted full model.</p><p>The improvement of the full model over the intercept model through the three types of pseudo R-squares is clear for both the training and testing datasets of I-70. For example, the McFadden R-square value for the I-70 training dataset is increased from 0.025 for the intercept to 0.118 for the full model, the Cox and Snell R-square value is increased from 0.031 for the intercept to 0.123 for the full model, and the Nagelkerke R-square is also increased from 0.046 for the intercept to 0.132 for the full mode. The higher pseudo R-squared values for the full models compared to the intercept models indicate that the fitted full models better predict the outcomes of the dependent variable, and the predictors are effective in modeling the different outcomes of the crash severity.</p></sec><sec id="s19"><title>19. Results of Multinomial Logistic Regression</title><p>The prediction results of the MNL are shown in the following sections:</p><sec id="s19_1"><title>19.1. Predicted Odd Ratios for I-70 Corridor</title><p>The odd ratios in MNL models present the probability of the event divided by the probability of the nonevent, and they can be obtained by exponentiating the multinomial logit coefficients (i.e. e<sup>(coef.)</sup>). The multinomial logistic regression model estimates (k − 1) models, where k is the number of outcome levels of the dependent variable, and the k<sup>th</sup> equation is relative to the referent group. In our model, the property damage is considered as the referent group (i.e. base level), because it is the most frequent outcome of crash severity, and the other outcome levels (i.e. minor injury, disabled, and fatal) are estimated relative to the property damage. The standard interpretation of the multinomial logistic regression is that for a unit change in the predictor variable, the odd ratio of outcome m relative to the referent group is expected to change by its respective parameter estimate given the other predictors in the model are held constant [<xref ref-type="bibr" rid="scirp.77395-ref1">1</xref>] [<xref ref-type="bibr" rid="scirp.77395-ref9">9</xref>] . The predicted odd ratios for the I-70 corridor (for both training and testing data) are obtained using Stata 14 and reported in <xref ref-type="table" rid="table">Table </xref>7. The odd ratios are significant when their related p-values at the 95% confidence level are less than 0.05. If the odd ratios are greater than 1.0, then the predictors are positively correlated with the dependent variable (i.e. crash severity), and if the odd ratios are smaller than 1.0, then the predictors are negatively correlated with the dependent variable. In other words, if the odd ratios are greater than 1.0, then the predictors would increase the likelihood of the crash severity occurrence at the specified level, indicating positive contribution to the crash severity occurrence at that level, and if the odd ratios are smaller than 1.0, then the predictors would decrease the likelihood of the crash severity occurrence at the specified level, indicating negative contribution to the crash occurrence at that level.</p><p>For example, when inspecting the MONTH predictor in the 1<sup>st</sup> case of crash severity (i.e. minor injury relative to property damage) in <xref ref-type="table" rid="table">Table </xref>7 for the training dataset, the odd ratio is greater than 1.0 (i.e. 1.015594), which indicates that this predictor is positively contributing to the crash severity at this level (i.e. minor injury), however, it is not significant at the 95% confidence as its p-value is greater than 0.05. In other words, the contribution of the predictor MONTH to the crash severity of the level of minor injury, would be expected to increase</p><table-wrap-group id="7"><label><xref ref-type="table" rid="table">Table </xref>7</label><caption><title> Predicted odd ratios for I-70, MO</title></caption><table-wrap id="7_1"><table><tbody><thead><tr><th align="center" valign="middle"  rowspan="2"  >Variable</th><th align="center" valign="middle"  colspan="3"  >I-70 training data</th><th align="center" valign="middle"  colspan="3"  >I-70 testing data</th></tr></thead><tr><td align="center" valign="middle" >Odd ratio</td><td align="center" valign="middle" >Std. error</td><td align="center" valign="middle" >p-value</td><td align="center" valign="middle" >Odd ratio</td><td align="center" valign="middle" >Std. error</td><td align="center" valign="middle" >p-value</td></tr><tr><td align="center" valign="middle"  colspan="7"  >Crash severity: Case 1: Minor Injury relative to base level (property damage)</td></tr><tr><td align="center" valign="middle" >MONTH</td><td align="center" valign="middle" >1.015594</td><td align="center" valign="middle" >0.0121626</td><td align="center" valign="middle" >0.196</td><td align="center" valign="middle" >1.098245</td><td align="center" valign="middle" >0.0187836</td><td align="center" valign="middle" >0.326</td></tr><tr><td align="center" valign="middle" >DAY_WEEK</td><td align="center" valign="middle" >0.9868066</td><td align="center" valign="middle" >0.0201894</td><td align="center" valign="middle" >0.516</td><td align="center" valign="middle" >0.9911457</td><td align="center" valign="middle" >0.0322842</td><td align="center" valign="middle" >0.421</td></tr><tr><td align="center" valign="middle" >HOUR</td><td align="center" valign="middle" >1.002493</td><td align="center" valign="middle" >0.0069472</td><td align="center" valign="middle" >0.719</td><td align="center" valign="middle" >1.017365</td><td align="center" valign="middle" >0.0112121</td><td align="center" valign="middle" >0.118</td></tr><tr><td align="center" valign="middle" >NO_VEHICLE</td><td align="center" valign="middle" >2.013444</td><td align="center" valign="middle" >0.1528305</td><td align="center" valign="middle" >0.000</td><td align="center" valign="middle" >1.603548</td><td align="center" valign="middle" >0.1673633</td><td align="center" valign="middle" >0.000</td></tr><tr><td align="center" valign="middle" >DIRECTION</td><td align="center" valign="middle" >1.001714</td><td align="center" valign="middle" >0.204671</td><td align="center" valign="middle" >0.993</td><td align="center" valign="middle" >1.299179</td><td align="center" valign="middle" >0.4009112</td><td align="center" valign="middle" >0.396</td></tr><tr><td align="center" valign="middle" >LIGHT_COND</td><td align="center" valign="middle" >1.018658</td><td align="center" valign="middle" >0.0539375</td><td align="center" valign="middle" >0.727</td><td align="center" valign="middle" >1.079072</td><td align="center" valign="middle" >0.0817907</td><td align="center" valign="middle" >0.800</td></tr><tr><td align="center" valign="middle" >ACC_TYPE</td><td align="center" valign="middle" >0.7646827</td><td align="center" valign="middle" >0.0322156</td><td align="center" valign="middle" >0.000</td><td align="center" valign="middle" >0.827777</td><td align="center" valign="middle" >0.0512309</td><td align="center" valign="middle" >0.002</td></tr><tr><td align="center" valign="middle" >DR_DRINK</td><td align="center" valign="middle" >0.4393219</td><td align="center" valign="middle" >0.0827939</td><td align="center" valign="middle" >0.000</td><td align="center" valign="middle" >0.4566945</td><td align="center" valign="middle" >0.1487597</td><td align="center" valign="middle" >0.016</td></tr><tr><td align="center" valign="middle" >SPEED</td><td align="center" valign="middle" >0.7628727</td><td align="center" valign="middle" >0.0832404</td><td align="center" valign="middle" >0.013</td><td align="center" valign="middle" >0.7331396</td><td align="center" valign="middle" >0.1258309</td><td align="center" valign="middle" >0.021</td></tr><tr><td align="center" valign="middle" >CZONE</td><td align="center" valign="middle" >0.8728007</td><td align="center" valign="middle" >0.1914342</td><td align="center" valign="middle" >0.882</td><td align="center" valign="middle" >0.8306115</td><td align="center" valign="middle" >0.4002926</td><td align="center" valign="middle" >0.384</td></tr><tr><td align="center" valign="middle" >DR_AGGRESSIVE</td><td align="center" valign="middle" >0.6820784</td><td align="center" valign="middle" >0.1231692</td><td align="center" valign="middle" >0.044</td><td align="center" valign="middle" >0.6812309</td><td align="center" valign="middle" >0.1762853</td><td align="center" valign="middle" >0.046</td></tr><tr><td align="center" valign="middle" >CELL_TEXT</td><td align="center" valign="middle" >0.5149235</td><td align="center" valign="middle" >0.1742725</td><td align="center" valign="middle" >0.049</td><td align="center" valign="middle" >0.3814188</td><td align="center" valign="middle" >0.2081773</td><td align="center" valign="middle" >0.047</td></tr><tr><td align="center" valign="middle" >DR_AGE</td><td align="center" valign="middle" >1.037926</td><td align="center" valign="middle" >0.3769126</td><td align="center" valign="middle" >0.158</td><td align="center" valign="middle" >1.078291</td><td align="center" valign="middle" >0.2189271</td><td align="center" valign="middle" >0.183</td></tr><tr><td align="center" valign="middle" >VEH_TYPE</td><td align="center" valign="middle" >0.8286522</td><td align="center" valign="middle" >0.1593428</td><td align="center" valign="middle" >0.462</td><td align="center" valign="middle" >0.857681</td><td align="center" valign="middle" >0.1783352</td><td align="center" valign="middle" >0.413</td></tr><tr><td align="center" valign="middle" >RURAL_URBAN</td><td align="center" valign="middle" >1.21414</td><td align="center" valign="middle" >0.1662723</td><td align="center" valign="middle" >0.157</td><td align="center" valign="middle" >1.194506</td><td align="center" valign="middle" >0.2581555</td><td align="center" valign="middle" >0.411</td></tr><tr><td align="center" valign="middle" >NUMBER_ LANES</td><td align="center" valign="middle" >1.043295</td><td align="center" valign="middle" >0.0714342</td><td align="center" valign="middle" >0.536</td><td align="center" valign="middle" >1.009117</td><td align="center" valign="middle" >0.1109496</td><td align="center" valign="middle" >0.081</td></tr><tr><td align="center" valign="middle" >AADT</td><td align="center" valign="middle" >1.000573</td><td align="center" valign="middle" >0.0018531</td><td align="center" valign="middle" >0.757</td><td align="center" valign="middle" >1.000707</td><td align="center" valign="middle" >0.0028542</td><td align="center" valign="middle" >0.804</td></tr><tr><td align="center" valign="middle" >GRADE_LEVEL</td><td align="center" valign="middle" >0.9969032</td><td align="center" valign="middle" >0.2049085</td><td align="center" valign="middle" >0.988</td><td align="center" valign="middle" >0.9728124</td><td align="center" valign="middle" >0.3983592</td><td align="center" valign="middle" >0.425</td></tr><tr><td align="center" valign="middle" >CONSTANT</td><td align="center" valign="middle" >0.504704</td><td align="center" valign="middle" >0.3106637</td><td align="center" valign="middle" >0.267</td><td align="center" valign="middle" >0.3146406</td><td align="center" valign="middle" >0.3145004</td><td align="center" valign="middle" >0.247</td></tr><tr><td align="center" valign="middle"  colspan="7"  >Crash severity: Case 2: Disabled relative to base level (property damage)</td></tr><tr><td align="center" valign="middle" >MONTH</td><td align="center" valign="middle" >1.04566</td><td align="center" valign="middle" >0.0294898</td><td align="center" valign="middle" >0.113</td><td align="center" valign="middle" >1.052662</td><td align="center" valign="middle" >0.044181</td><td align="center" valign="middle" >0.221</td></tr><tr><td align="center" valign="middle" >DAY_WEEK</td><td align="center" valign="middle" >0.9849045</td><td align="center" valign="middle" >0.055887</td><td align="center" valign="middle" >0.004</td><td align="center" valign="middle" >0.9713375</td><td align="center" valign="middle" >0.0714767</td><td align="center" valign="middle" >0.019</td></tr><tr><td align="center" valign="middle" >HOUR</td><td align="center" valign="middle" >1.0907501</td><td align="center" valign="middle" >0.0153366</td><td align="center" valign="middle" >0.548</td><td align="center" valign="middle" >1.0921144</td><td align="center" valign="middle" >0.0225957</td><td align="center" valign="middle" >0.067</td></tr><tr><td align="center" valign="middle" >NO_VEHICLE</td><td align="center" valign="middle" >2.325778</td><td align="center" valign="middle" >0.346116</td><td align="center" valign="middle" >0.000</td><td align="center" valign="middle" >1.303495</td><td align="center" valign="middle" >0.3296267</td><td align="center" valign="middle" >0.029</td></tr><tr><td align="center" valign="middle" >DIRECTION</td><td align="center" valign="middle" >1.0244691</td><td align="center" valign="middle" >0.102775</td><td align="center" valign="middle" >0.141</td><td align="center" valign="middle" >1.0231048</td><td align="center" valign="middle" >0.7614965</td><td align="center" valign="middle" >0.314</td></tr><tr><td align="center" valign="middle" >LIGHT_COND</td><td align="center" valign="middle" >1.0325387</td><td align="center" valign="middle" >0.1239202</td><td align="center" valign="middle" >0.836</td><td align="center" valign="middle" >1.0277047</td><td align="center" valign="middle" >0.2202536</td><td align="center" valign="middle" >0.156</td></tr><tr><td align="center" valign="middle" >ACC_TYPE</td><td align="center" valign="middle" >0.77145632</td><td align="center" valign="middle" >0.061105</td><td align="center" valign="middle" >0.000</td><td align="center" valign="middle" >0.79145609</td><td align="center" valign="middle" >0.1310232</td><td align="center" valign="middle" >0.006</td></tr><tr><td align="center" valign="middle" >DR_DRINK</td><td align="center" valign="middle" >0.1758408</td><td align="center" valign="middle" >0.0543585</td><td align="center" valign="middle" >0.000</td><td align="center" valign="middle" >0.2855924</td><td align="center" valign="middle" >0.1548372</td><td align="center" valign="middle" >0.021</td></tr><tr><td align="center" valign="middle" >SPEED</td><td align="center" valign="middle" >0.6718398</td><td align="center" valign="middle" >0.1729933</td><td align="center" valign="middle" >0.122</td><td align="center" valign="middle" >0.5888928</td><td align="center" valign="middle" >0.2284686</td><td align="center" valign="middle" >0.172</td></tr><tr><td align="center" valign="middle" >CZONE</td><td align="center" valign="middle" >0.8159377</td><td align="center" valign="middle" >0.5622705</td><td align="center" valign="middle" >0.760</td><td align="center" valign="middle" >0.81661143</td><td align="center" valign="middle" >0.3387404</td><td align="center" valign="middle" >0.375</td></tr><tr><td align="center" valign="middle" >DR_AGGRESSIVE</td><td align="center" valign="middle" >0.79283284</td><td align="center" valign="middle" >0.3286617</td><td align="center" valign="middle" >0.251</td><td align="center" valign="middle" >0.72908047</td><td align="center" valign="middle" >0.4614627</td><td align="center" valign="middle" >0.475</td></tr><tr><td align="center" valign="middle" >CELL_TEXT</td><td align="center" valign="middle" >0.6839739</td><td align="center" valign="middle" >0.518411</td><td align="center" valign="middle" >0.016</td><td align="center" valign="middle" >0.6161388</td><td align="center" valign="middle" >0.1346915</td><td align="center" valign="middle" >0.029</td></tr><tr><td align="center" valign="middle" >DR_AGE</td><td align="center" valign="middle" >1.098286</td><td align="center" valign="middle" >0.482946</td><td align="center" valign="middle" >0.243</td><td align="center" valign="middle" >1.08442</td><td align="center" valign="middle" >0.4398022</td><td align="center" valign="middle" >0.283</td></tr><tr><td align="center" valign="middle" >VEH_TYPE</td><td align="center" valign="middle" >0.7338291</td><td align="center" valign="middle" >0.172765</td><td align="center" valign="middle" >0.389</td><td align="center" valign="middle" >0.672993</td><td align="center" valign="middle" >0.1798307</td><td align="center" valign="middle" >0.317</td></tr><tr><td align="center" valign="middle" >RURAL_URBAN</td><td align="center" valign="middle" >1.154855</td><td align="center" valign="middle" >0.3503854</td><td align="center" valign="middle" >0.635</td><td align="center" valign="middle" >1.1281573</td><td align="center" valign="middle" >0.4360739</td><td align="center" valign="middle" >0.274</td></tr></tbody></table></table-wrap><table-wrap id="7_2"><table><tbody><thead><tr><th align="center" valign="middle" >NUMBER_ LANES</th><th align="center" valign="middle" >1.0623837</th><th align="center" valign="middle" >0.1297353</th><th align="center" valign="middle" >0.035</th><th align="center" valign="middle" >1.0729747</th><th align="center" valign="middle" >0.327797</th><th align="center" valign="middle" >0.041</th></tr></thead><tr><td align="center" valign="middle" >AADT</td><td align="center" valign="middle" >1.0900496</td><td align="center" valign="middle" >0.0048217</td><td align="center" valign="middle" >0.302</td><td align="center" valign="middle" >1.0993353</td><td align="center" valign="middle" >0.0064666</td><td align="center" valign="middle" >0.306</td></tr><tr><td align="center" valign="middle" >GRADE_LEVEL</td><td align="center" valign="middle" >0.99225575</td><td align="center" valign="middle" >0.095807</td><td align="center" valign="middle" >0.000</td><td align="center" valign="middle" >0.92474128</td><td align="center" valign="middle" >1.210746</td><td align="center" valign="middle" >0.015</td></tr><tr><td align="center" valign="middle" >CONSTANT</td><td align="center" valign="middle" >0.4430657</td><td align="center" valign="middle" >5.736671</td><td align="center" valign="middle" >0.273</td><td align="center" valign="middle" >0.42062742</td><td align="center" valign="middle" >0.40145</td><td align="center" valign="middle" >0.417</td></tr><tr><td align="center" valign="middle"  colspan="7"  >Crash severity: Case 3: Fatal relative to base level (property damage)</td></tr><tr><td align="center" valign="middle" >MONTH</td><td align="center" valign="middle" >1.204367</td><td align="center" valign="middle" >0.0806321</td><td align="center" valign="middle" >0.005</td><td align="center" valign="middle" >1.204406</td><td align="center" valign="middle" >0.0831499</td><td align="center" valign="middle" >0.014</td></tr><tr><td align="center" valign="middle" >DAY_WEEK</td><td align="center" valign="middle" >0.9863804</td><td align="center" valign="middle" >0.105922</td><td align="center" valign="middle" >0.737</td><td align="center" valign="middle" >0.9828217</td><td align="center" valign="middle" >0.1155401</td><td align="center" valign="middle" >0.177</td></tr><tr><td align="center" valign="middle" >HOUR</td><td align="center" valign="middle" >1.023859</td><td align="center" valign="middle" >0.0319797</td><td align="center" valign="middle" >0.450</td><td align="center" valign="middle" >1.036516</td><td align="center" valign="middle" >0.0377693</td><td align="center" valign="middle" >0.365</td></tr><tr><td align="center" valign="middle" >NO_VEHICLE</td><td align="center" valign="middle" >2.232134</td><td align="center" valign="middle" >0.5612323</td><td align="center" valign="middle" >0.001</td><td align="center" valign="middle" >1.707896</td><td align="center" valign="middle" >0.4912682</td><td align="center" valign="middle" >0.009</td></tr><tr><td align="center" valign="middle" >DIRECTION</td><td align="center" valign="middle" >1.099131</td><td align="center" valign="middle" >1.869167</td><td align="center" valign="middle" >0.515</td><td align="center" valign="middle" >1.042631</td><td align="center" valign="middle" >1.473025</td><td align="center" valign="middle" >0.231</td></tr><tr><td align="center" valign="middle" >LIGHT_COND</td><td align="center" valign="middle" >1.042018</td><td align="center" valign="middle" >0.6304126</td><td align="center" valign="middle" >0.001</td><td align="center" valign="middle" >1.038765</td><td align="center" valign="middle" >0.766612</td><td align="center" valign="middle" >0.007</td></tr><tr><td align="center" valign="middle" >ACC_TYPE</td><td align="center" valign="middle" >0.7563569</td><td align="center" valign="middle" >0.3752455</td><td align="center" valign="middle" >0.063</td><td align="center" valign="middle" >0.6287748</td><td align="center" valign="middle" >0.3629575</td><td align="center" valign="middle" >0.370</td></tr><tr><td align="center" valign="middle" >DR_DRINK</td><td align="center" valign="middle" >0.1747316</td><td align="center" valign="middle" >0.1104344</td><td align="center" valign="middle" >0.006</td><td align="center" valign="middle" >0.2648509</td><td align="center" valign="middle" >0.4530978</td><td align="center" valign="middle" >0.033</td></tr><tr><td align="center" valign="middle" >SPEED</td><td align="center" valign="middle" >0.3108948</td><td align="center" valign="middle" >0.2162619</td><td align="center" valign="middle" >0.093</td><td align="center" valign="middle" >0.3551321</td><td align="center" valign="middle" >0.334089</td><td align="center" valign="middle" >0.271</td></tr><tr><td align="center" valign="middle" >CZONE</td><td align="center" valign="middle" >0.82678563</td><td align="center" valign="middle" >0.1350873</td><td align="center" valign="middle" >0.081</td><td align="center" valign="middle" >0.8429472</td><td align="center" valign="middle" >0.244088</td><td align="center" valign="middle" >0.291</td></tr><tr><td align="center" valign="middle" >DR_AGGRESSIVE</td><td align="center" valign="middle" >0.8619844</td><td align="center" valign="middle" >3.320254</td><td align="center" valign="middle" >0.003</td><td align="center" valign="middle" >0.8827105</td><td align="center" valign="middle" >3.191887</td><td align="center" valign="middle" >0.008</td></tr><tr><td align="center" valign="middle" >CELL_TEXT</td><td align="center" valign="middle" >0.2562309</td><td align="center" valign="middle" >0.2849574</td><td align="center" valign="middle" >0.021</td><td align="center" valign="middle" >0.0714367</td><td align="center" valign="middle" >0.0850799</td><td align="center" valign="middle" >0.027</td></tr><tr><td align="center" valign="middle" >DR_AGE</td><td align="center" valign="middle" >1.0981655</td><td align="center" valign="middle" >0.7690331</td><td align="center" valign="middle" >0.295</td><td align="center" valign="middle" >1.0616548</td><td align="center" valign="middle" >0.2628931</td><td align="center" valign="middle" >0.319</td></tr><tr><td align="center" valign="middle" >VEH_TYPE</td><td align="center" valign="middle" >0.7822954</td><td align="center" valign="middle" >0.1692881</td><td align="center" valign="middle" >0.284</td><td align="center" valign="middle" >0.781194</td><td align="center" valign="middle" >0.1672393</td><td align="center" valign="middle" >0.342</td></tr><tr><td align="center" valign="middle" >RURAL_URBAN</td><td align="center" valign="middle" >1.3862095</td><td align="center" valign="middle" >0.4994118</td><td align="center" valign="middle" >0.605</td><td align="center" valign="middle" >1.4874849</td><td align="center" valign="middle" >0.4314113</td><td align="center" valign="middle" >0.217</td></tr><tr><td align="center" valign="middle" >NUMBER_ LANES</td><td align="center" valign="middle" >1.0718678</td><td align="center" valign="middle" >0.3198925</td><td align="center" valign="middle" >0.404</td><td align="center" valign="middle" >1.0565231</td><td align="center" valign="middle" >0.9127193</td><td align="center" valign="middle" >0.344</td></tr><tr><td align="center" valign="middle" >AADT</td><td align="center" valign="middle" >1.002445</td><td align="center" valign="middle" >0.0111129</td><td align="center" valign="middle" >0.226</td><td align="center" valign="middle" >1.0876658</td><td align="center" valign="middle" >0.0134131</td><td align="center" valign="middle" >0.361</td></tr><tr><td align="center" valign="middle" >GRADE_LEVEL</td><td align="center" valign="middle" >0.75853</td><td align="center" valign="middle" >0.7540079</td><td align="center" valign="middle" >0.781</td><td align="center" valign="middle" >0.82517107</td><td align="center" valign="middle" >2.630318</td><td align="center" valign="middle" >0.377</td></tr><tr><td align="center" valign="middle" >CONSTANT</td><td align="center" valign="middle" >0.68610</td><td align="center" valign="middle" >2.9507</td><td align="center" valign="middle" >0.916</td><td align="center" valign="middle" >0.677015</td><td align="center" valign="middle" >1.40901</td><td align="center" valign="middle" >0.487</td></tr></tbody></table></table-wrap></table-wrap-group><p>by a factor of 1.015594 given the other variables in the model are held constant. When inspecting the DAY_WEEK predictor in the 1<sup>st</sup> case of crash severity (i.e. minor injury relative to property damage) in <xref ref-type="table" rid="table">Table </xref>8 for the training dataset, the odd ratio is smaller than 1.0 (i.e. 0.9868066), which indicates that this predictor is negatively contributing to the crash severity at this level (i.e. minor injury), and it is not significant at the 95% confidence as its p-value is greater than 0.05. When inspecting the NO_VEHICLE predictor in the 1<sup>st</sup> case of crash severity (i.e. minor injury relative to property damage) in <xref ref-type="table" rid="table">Table </xref>8 for the training dataset, the odd ratio is greater than 1.0 (i.e. 2.013444), which indicates that this predictor is positively contributing to the crash severity at this level (i.e. minor injury), and it is significant at the 95% confidence as its p-value is less than 0.05. So, the contribution of the predictor NO_VEHICLE to the crash severity of the level of minor injury, would be expected to increase by a factor of 2.013444 given the other variables in the model are held constant. Likewise, when inspecting the MONTH predictor in the 2<sup>nd</sup> case of crash severity (i.e. disabled relative to</p><table-wrap id="table8" ><label><xref ref-type="table" rid="table">Table </xref>8</label><caption><title> Significant risk factors for I-70, MO</title></caption><table><tbody><thead><tr><th align="center" valign="middle"  rowspan="2"  >Crash severity level</th><th align="center" valign="middle"  colspan="2"  >INTERSTATE I-70, MO</th></tr></thead><tr><td align="center" valign="middle" >Significant risk factors</td><td align="center" valign="middle" >Significant group factors</td></tr><tr><td align="center" valign="middle" >Case 1: minor injury</td><td align="center" valign="middle" >1. NO_VEHICLE 2. ACC_TYPE 3. DR_DRINK 4. SPEED 5. DR_AGGRESSIVE 6. CELL_TEXT</td><td align="center" valign="middle" >1. Driver behavior 2. Accident type</td></tr><tr><td align="center" valign="middle" >Case 2: disabled</td><td align="center" valign="middle" >1. DAY_WEEK 2. NO_VEHICLE 3. ACC_TYPE 4. DR_DRINK 5. CELL_TEXT 6. NUMBER_LANES 7. GRADE_LEVEL</td><td align="center" valign="middle" >1. Time 2. Driver Behavior 3. Accident type 4. Road geometry</td></tr><tr><td align="center" valign="middle" >Case 3: fatal</td><td align="center" valign="middle" >1. MONTH 2. NO_VEHICLES 3. LIGHT_COND 4. DR_DRINK 5. DR_AGGRESSIVE 6. CELL_TEXT</td><td align="center" valign="middle" >1. Time 2. Driver behavior 3. Environment</td></tr></tbody></table></table-wrap><p>property damage) in <xref ref-type="table" rid="table">Table </xref>7 for the training dataset, the odd ratio is greater than 1.0 (i.e. 1.04566), which indicates that this predictor is positively contributing to the crash severity at this level (i.e. disabled), however it is not significant at the 95% confidence as its p-value is greater than 0.05. In other words, the contribution of the predictor MONTH to the crash severity of the level of “disabled”, would be expected to increase by a factor of 1.04566 given the other variables in the model are held constant. When inspecting the MONTH predictor in the 3<sup>rd</sup> case of crash severity (i.e. fatal relative to property damage) in <xref ref-type="table" rid="table">Table </xref>8 for the training dataset, the odd ratio is greater than 1.0 (i.e. 1.204367), which indicates that this predictor is positively contributing to the crash severity at this level (i.e. fatal), and it is significant at the 95% confidence as its p-value is less than 0.05. When inspecting the NO_VEHICLE predictor in the 2<sup>nd</sup> and 3<sup>rd</sup> cases of crash severity (i.e. disabled relative to property damage, and fatal relative to property damage) in <xref ref-type="table" rid="table">Table </xref>7 for the training dataset, the odd ratios are greater than 1.0 (i.e. 2.325778, 2.232134 respectively), which indicates that this predictor is positively contributing to the crash severity at these two levels (i.e. disabled, and fatal), and it is significant at the 95% confidence as its p-values are less than 0.05. So, the contribution of the predictor NO_VEHICLE to the crash severity of the levels of “disabled” and “fatal”, would be expected to increase by a factor of 2.325778 and 2.232134 respectively given the other variables in the model are held constant.</p></sec><sec id="s19_2"><title>19.2. Significant Risk Factors for I-70 Corridor</title><p>The statistically significant risk factors (i.e. predictors or independent variables) of the I-70 corridor in Missouri at the 95% confidence level are shown in <xref ref-type="table" rid="table">Table </xref>8.</p><p>For the 1<sup>st</sup> case of crash severity level (i.e. minor injury relative to property damage), the number of vehicles involved in the crashes, the accident type, the driver drink, the speed, the driver aggressiveness, and the cell-text, are significant at the 95% confidence level. For the 2<sup>nd</sup> case of crash severity level (i.e. disabled relative to property damage), the day of the week, the number of vehicles involved in the crashes, the accident type, the driver drink, the cell-text, the number of lanes, and the grade of the road are significant at the 95% confidence level. For the 3<sup>rd</sup> case of crash severity level (i.e. fatal relative to property damage), the month of the year, the number of vehicles involved in the crashes, the light condition, the driver drink, the driver aggressiveness, and the cell-text, are significant at the 95% confidence level. We can see that two risk factors (i.e. the number of vehicles involved in the crashes and using the cell phones or texts when driving) are significant at the three crash severity levels (i.e. minor injury, disabled, fatal), indicating the importance of these two risk factors in modeling the severity of crashes of the I-70 corridor in MO. Some other risk factors are significant at only two levels of crash severity, but not at the third level. These risk factors are the accident type, the driver drink, and the driver aggressiveness. The speed, the light condition, the number of lanes, the grade of the road, the day of the week, and the month of the year are significant at only one level of crash severity. In term of the significant group of factors, we can see that the driver’s behavior group is the most important one as it has been related to the three crash severity levels, whereas the accident type, the time, is the next in its importance.</p></sec><sec id="s19_3"><title>19.3. Marginal Effects for Crashes along I-70 Corridor</title><p>The marginal effect reflects the impact of a one-unit change of an independent variable (predictor) on the event probability of the dependent variable (keeping all other independent variables constant at their mean values). In MNL, the marginal effect of an explanatory variable (predictor) is the partial derivative of the event probability with respect to the predictor of interest (i.e. the change in the event probability of the dependent variable for a unit change in the predictor), and they could be positive or negative values. Positive values indicate that the predictor would positively contribute to crash severity (i.e. would increase the degree severity of crashes), and negative values indicate that the predictor would negatively contribute to crash severity (i.e. would decrease the degree severity of crashes). The marginal effect for a dummy or discrete independent variable is the difference of the predicted probability values at their different levels [<xref ref-type="bibr" rid="scirp.77395-ref17">17</xref>] . The marginal effects for the I-70 corridor (for both training and testing data) are obtained using Stata 14 and reported in <xref ref-type="table" rid="table">Table </xref>9. It can be seen from the table that some predictors have higher marginal effects than others. For instance, the driver drink predictor has a marginal effect of 15.56% for training data, and 16.07% for testing data. These values present the difference of the event probability of the crash severity when drivers using the road being drunk and not drunk.</p><p>In other words, if all the drivers that use the I-70 corridor in MO were not in intoxicated conditions, then the probability of crash severity at the I-70 corridor would decrease by 15.56% using training data and 16.07% using testing data. The speed predictor has a marginal effect of 8.04 % for training data, and 10.12% for testing data. These values present the difference of the event probability of the crash severity when drivers using the road are speeding and not speeding so that the crash severity would decrease by (8.04% using training data and 10.12% using testing data) if all drivers were not speeding. The cell-text predictor has a marginal effect of 12.54% for training data, and 14.17% for testing data. These values present the difference of the event probability of the crash severity when drivers are using the cell phones and/or texting during the driving and not using them so that the crash severity would decrease by 12.54% using training data and 14.17% using testing data if all drivers were not using cell-text when driving. The number of vehicles involved (assuming one vehicle) in the crash has a marginal effect of 9.58% for training data, and 10.62% for testing data. Meaning that if only one vehicle is involved in the crash, then it would increase the severity by 9.58% using training data and 10.62% using testing data. However, if the number of vehicles involved were increased to two vehicles, then this would increase the severity by 14.54% using training data and 15.87% using testing data. If the number of vehicles increased to three vehicles, then this would increase the severity by 13.17% using training data and 13.16% using testing data. If the number of vehicles further increased to four vehicles, then this would increase the severity by 14.39% using training data and 15.04% using testing data. The accident type predictor (ACC_TYPE) relative to an animal has a marginal effect of 1.78% for training data and 2.19% for testing data. Meaning if an animal would have caused the accident, then this would increase the severity by 1.78% using training data and 2.19% using testing data. However, the accident type predictor relative to a fixed object has a marginal effect of 7.06% for training data and 6.48% for testing data. Meaning if a fixed object (such as a tree or a traffic sign) would have caused the accident, then this would increase the severity by 7.06% using training data and 6.48% using testing data. However, the accident type predictor relative to an overturn has a marginal effect of 8.39% for training data and 7.79% for testing data. Meaning if an overturn was the accident type, then this would increase the severity by 8.39% using training data and 7.79% using testing data. Similarly, the accident type predictor relative to a pedestrian has a marginal effect of 7.17% for training data and 7.36% for testing data. Meaning if a pedestrian would have caused the accident, then this would increase the severity by 7.17% using training data and 7.36% using testing data. In similar manner, the accident type predictor relative to a vehicle in transport has a marginal effect of 7.38% for training data and 7.27% for testing data. Meaning if a vehicle in</p><table-wrap id="table9" ><label><xref ref-type="table" rid="table">Table </xref>9</label><caption><title> Marginal effects for crashes along I-70</title></caption><table><tbody><thead><tr><th align="center" valign="middle"  rowspan="2"  >Variable name</th><th align="center" valign="middle"  rowspan="2"  >Variable subgroup</th><th align="center" valign="middle"  colspan="2"  >% Marginal effect</th></tr></thead><tr><td align="center" valign="middle" >I-70 training</td><td align="center" valign="middle" >I-70 testing</td></tr><tr><td align="center" valign="middle" >GRADE_LEVEL</td><td align="center" valign="middle" >Grade Level</td><td align="center" valign="middle" >3.22 −1.58</td><td align="center" valign="middle" >3.62 −1.74</td></tr><tr><td align="center" valign="middle" >NUMBER_LANES</td><td align="center" valign="middle" >One lane Two lanes Three lanes Four lanes Five lanes Six lanes or more</td><td align="center" valign="middle" >1.06 2.05 −2.28 −2.94 1.31 0.42</td><td align="center" valign="middle" >1.23 2.16 −2.77 −2.49 1.53 0.22</td></tr><tr><td align="center" valign="middle" >RURAL_URBAN</td><td align="center" valign="middle" >Rural Urban</td><td align="center" valign="middle" >1.97 −1.56</td><td align="center" valign="middle" >2.31 −1.81</td></tr><tr><td align="center" valign="middle" >CZONE</td><td align="center" valign="middle" >n/a</td><td align="center" valign="middle" >1.71</td><td align="center" valign="middle" >2.33</td></tr><tr><td align="center" valign="middle" >AADT</td><td align="center" valign="middle" >n/a</td><td align="center" valign="middle" >1.92</td><td align="center" valign="middle" >1.72</td></tr><tr><td align="center" valign="middle" >HOUR</td><td align="center" valign="middle" >n/a</td><td align="center" valign="middle" >1.74</td><td align="center" valign="middle" >2.09</td></tr><tr><td align="center" valign="middle" >DAY_WEEK</td><td align="center" valign="middle" >Sun. Mon. Tues. Wed. Thurs. Fri. Sat.</td><td align="center" valign="middle" >−2.02 2.31 −2.09 −1.65 −1.38 3.15 2.88</td><td align="center" valign="middle" >−1.79 1.84 −1.98 −1.43 −1.17 3.37 2.49</td></tr><tr><td align="center" valign="middle" >MONTH</td><td align="center" valign="middle" >n/a</td><td align="center" valign="middle" >1.67</td><td align="center" valign="middle" >1.89</td></tr><tr><td align="center" valign="middle" >DIRECTION</td><td align="center" valign="middle" >East West</td><td align="center" valign="middle" >1.47 1.31</td><td align="center" valign="middle" >1.52 1.36</td></tr><tr><td align="center" valign="middle" >LIGHT_COND</td><td align="center" valign="middle" >Daylight Dark, lighted Dark, unlighted</td><td align="center" valign="middle" >−0.43 −0.79 0.59</td><td align="center" valign="middle" >−0.23 −0.62 0.44</td></tr><tr><td align="center" valign="middle" >DR_AGE</td><td align="center" valign="middle" >Less than 21 years From (21 - 64) years More than 64 years</td><td align="center" valign="middle" >2.58 −1.87 2.49</td><td align="center" valign="middle" >2.87 −1.63 2.61</td></tr><tr><td align="center" valign="middle" >VEH_TYPE</td><td align="center" valign="middle" >Passenger car Motorcycle Truck</td><td align="center" valign="middle" >−1.62 2.16 −1.79</td><td align="center" valign="middle" >−1.44 2.06 −1.48</td></tr><tr><td align="center" valign="middle" >NO_VEHICLE</td><td align="center" valign="middle" >One vehicle Two vehicles Three vehicles Four vehicles Five vehicles Six or more vehicles</td><td align="center" valign="middle" >9.58 14.54 13.17 14.39 13.33 15.17</td><td align="center" valign="middle" >10.62 15.87 13.16 15.04 13.94 14.81</td></tr><tr><td align="center" valign="middle" >ACC_TYPE</td><td align="center" valign="middle" >Animal Fixed object Overturn Pedestrian Vehicle in transport</td><td align="center" valign="middle" >1.78 7.06 8.39 7.17 7.38</td><td align="center" valign="middle" >2.19 6.48 7.79 7.36 7.27</td></tr><tr><td align="center" valign="middle" >DR_DRINK</td><td align="center" valign="middle" >n/a</td><td align="center" valign="middle" >−15.56</td><td align="center" valign="middle" >−16.07</td></tr><tr><td align="center" valign="middle" >SPEED</td><td align="center" valign="middle" >n/a</td><td align="center" valign="middle" >−8.04</td><td align="center" valign="middle" >−10.12</td></tr><tr><td align="center" valign="middle" >DR_AGGRESSIVE</td><td align="center" valign="middle" >n/a</td><td align="center" valign="middle" >−8.84</td><td align="center" valign="middle" >−8.41</td></tr><tr><td align="center" valign="middle" >CELL_TEXT</td><td align="center" valign="middle" >n/a</td><td align="center" valign="middle" >−12.54</td><td align="center" valign="middle" >−14.17</td></tr></tbody></table></table-wrap><p>transport would have caused the accident, then this would increase the severity by 7.38% using training data and 7.27% using testing data.</p></sec></sec><sec id="s20"><title>20. Conclusion</title><p>This paper applied multinomial logistic regression (MNL) to model the relationships of the crash severity categories with the independent variables. The I-70 corridor is tested under the assumptions of the MNL. The categories of the dependent variable (i.e. fatal, disabling injury, minor injury, property-damage- only) are considered nominal (i.e. cannot be ordered in any logical way). This paper investigated the use of a wider range of independent variables (i.e. risk factors) in crash severity modeling, given that past research has only made use of limited numbers/types of independent variables. In addition, this paper introduced a variety of new procedures in presenting the results of the MNL applications that have not been reported in other crash severity models, including: 1) the use of the odd ratios as regression estimates instead of using regression coefficients to interpret the results of prediction; 2) a focus on the assumption of the independence of irrelevant alternatives (IIA) that is very important in the MNL modeling, using the Hausman specification test; 3) consideration of the generalized Hosmer-Lemeshow test as an important goodness of fit measure to assess whether or not the observed incidents match the predicted incidents; 4) the use of the classification table as a measure of goodness of fit to determine the percent of corrected prediction cases; 5) testing for the multicollinearity among the independent variables as precondition assumption; 6) the use of the pseudo R squares as potential goodness of fits instead of classical measures of goodness of fit, such as the Deviance, the Akaike Information Criteria (AIC), and the Bayesian Information Criteria (BIC); and 7) presenting the marginal effects of all independent variables upon the dependent variable. Results showed the effectiveness of the MNL approach in crash severity modeling.</p></sec><sec id="s21"><title>Cite this paper</title><p>Abdulhafedh, A. (2017) Incorporating the Multinomial Logistic Regression in Vehicle Crash Severity Modeling: A Detailed Overview. Journal of Transportation Technologies, 7, 279-303. https://doi.org/10.4236/jtts.2017.73019</p></sec><sec id="s22"><title>NOTES</title></sec></body><back><ref-list><title>References</title><ref id="scirp.77395-ref1"><label>1</label><mixed-citation publication-type="other" xlink:type="simple">Greene, W. (2012) Econometric Analysis. 7th Edition, Prentice Hall, Upper Saddle River.</mixed-citation></ref><ref id="scirp.77395-ref2"><label>2</label><mixed-citation publication-type="other" xlink:type="simple">McFadden, D., Tye, W. and Train, K. (1976) An Application of Diagnostic Tests for the Independence from Irrelevant Alternatives Property of the Multinomial Logit Model. Transportation Research Record, 637, 39-45.</mixed-citation></ref><ref id="scirp.77395-ref3"><label>3</label><mixed-citation publication-type="other" xlink:type="simple">Hausman, J.A. (1978) Specification Tests in Econometrics. Econometrica, 46, 1251- 1271. https://doi.org/10.2307/1913827</mixed-citation></ref><ref id="scirp.77395-ref4"><label>4</label><mixed-citation publication-type="other" xlink:type="simple">Kleinbaum, D.G. and Klein, M. (2010) Logistic Regression: A Self-Learning Text. 3rd Edition, Springer, New York. https://doi.org/10.1007/978-1-4419-1742-3</mixed-citation></ref><ref id="scirp.77395-ref5"><label>5</label><mixed-citation publication-type="other" xlink:type="simple">Baltagi, B.H. (2011) Econometrics. 5th Edition, Springer, Berlin.  
https://doi.org/10.1007/978-3-642-20059-5</mixed-citation></ref><ref id="scirp.77395-ref6"><label>6</label><mixed-citation publication-type="other" xlink:type="simple">Abdel-Aty, M. (2003) Analysis of Driver Injury Severity Levels at Multiple Locations Using Ordered Probit Models. Journal of Safety Research, 34, 597-603.  
https://doi.org/10.1016/j.jsr.2003.05.009</mixed-citation></ref><ref id="scirp.77395-ref7"><label>7</label><mixed-citation publication-type="other" xlink:type="simple">Bham, G., Javvadi, B. and Manepalli, U. (2012) Multinomial Logistic Regression Model for Single-Vehicle and Multivehicle Collisions on Urban U.S. Highways in Arkansas. Journal of Transportation Engineering, 138, 786-797.  
https://doi.org/10.1061/(ASCE)TE.1943-5436.0000370</mixed-citation></ref><ref id="scirp.77395-ref8"><label>8</label><mixed-citation publication-type="other" xlink:type="simple">The Missouri State Highway Patrol (2016) Accident Investigation Reports.  
https://www.mshp.dps.missouri.gov/HP68/static/Official.html</mixed-citation></ref><ref id="scirp.77395-ref9"><label>9</label><mixed-citation publication-type="other" xlink:type="simple">Judge, G., Griffiths, W.E., Hill, R.C., Lutkepohl, H. and Lee, T.C. (1985) The Theory and Practice of Econometrics. 2nd Edition, Wiley, New York.</mixed-citation></ref><ref id="scirp.77395-ref10"><label>10</label><mixed-citation publication-type="other" xlink:type="simple">Long, S. (1996) Regression Models for Categorical and Limited Dependent Variables. Sage Publications, Thousand Oaks.</mixed-citation></ref><ref id="scirp.77395-ref11"><label>11</label><mixed-citation publication-type="other" xlink:type="simple">Hausman, J.A. and McFadden, D. (1984) Specification Tests for the Multinomial Logit Model. Econometrica, 52, 1219-1240. https://doi.org/10.2307/1910997</mixed-citation></ref><ref id="scirp.77395-ref12"><label>12</label><mixed-citation publication-type="other" xlink:type="simple">Menard, S. (2002) Applied Logistic Regression Analysis. Sage Publications, Thousand Oaks.</mixed-citation></ref><ref id="scirp.77395-ref13"><label>13</label><mixed-citation publication-type="other" xlink:type="simple">Lemeshow, S.A. and Hosmer, J.D.W. (1982) A Review of Goodness of Fit Statistics for the Use in the Development of Logistic Regression Models. American Journal of Epidemiology, 115, 92-106. https://doi.org/10.1093/oxfordjournals.aje.a113284</mixed-citation></ref><ref id="scirp.77395-ref14"><label>14</label><mixed-citation publication-type="other" xlink:type="simple">Hosmer, D.W., Lemeshow, S.A. and Sturdivant, R.X. (2013) Applied Logistic Regre- ssion. 3rd Edition, Wiley, Hoboken. https://doi.org/10.1002/9781118548387</mixed-citation></ref><ref id="scirp.77395-ref15"><label>15</label><mixed-citation publication-type="other" xlink:type="simple">Fagerland, M.W. and Hosmer, D.W.J. (2012) A Generalized Hosmer Leme Show Goodness-of-Fit Test for Multinomial Logistic Regression Models. Stata Journal, 12, 447-453.</mixed-citation></ref><ref id="scirp.77395-ref16"><label>16</label><mixed-citation publication-type="other" xlink:type="simple">Fagerland, M.W., Hosmer, D.W. and Bofin, A.M. (2008) Multinomial Goodness-of- 
Fit Tests for Logistic Regression Models. Statistics in Medicine, 27, 4238-4253.  
https://doi.org/10.1002/sim.3202</mixed-citation></ref><ref id="scirp.77395-ref17"><label>17</label><mixed-citation publication-type="other" xlink:type="simple">Long, S. and Freese, J. (2014) Regression Models for Categorical Dependent Variables Using Stata. 3rd Edition, Stata Press, College Station.</mixed-citation></ref><ref id="scirp.77395-ref18"><label>18</label><mixed-citation publication-type="other" xlink:type="simple">McFadden, D. (1974) Conditional Logit Analysis of Qualitative Choice Behavior. Frontiers in Econometrics, 105-142.</mixed-citation></ref><ref id="scirp.77395-ref19"><label>19</label><mixed-citation publication-type="other" xlink:type="simple">Cox, D.R. and Snell, E.J. (1989) Analysis of Binary Data. 2nd Edition, Chapman &amp; Hall, London.</mixed-citation></ref><ref id="scirp.77395-ref20"><label>20</label><mixed-citation publication-type="other" xlink:type="simple">Nagelkerke, N.J.D. (1991) A Note on a General Definition of the Coefficient of Determination. Biometrika, 78, 691-692. https://doi.org/10.1093/biomet/78.3.691</mixed-citation></ref><ref id="scirp.77395-ref21"><label>21</label><mixed-citation publication-type="other" xlink:type="simple">Tjur, T. (2009) Coefficients of Determination in Logistic Regression Models: A New Proposal: The Coefficient of Discrimination. The American Statistician, 63, 366- 372. https://doi.org/10.1198/tast.2009.08210</mixed-citation></ref><ref id="scirp.77395-ref22"><label>22</label><mixed-citation publication-type="other" xlink:type="simple">Freese, J. and Long, J.S. (2000) Tests for the Multinomial Logit Model. Stata Techni- cal Bulletin, 10, 247-255.</mixed-citation></ref><ref id="scirp.77395-ref23"><label>23</label><mixed-citation publication-type="other" xlink:type="simple">Greene, W. (2008) Econometric Analysis. 6th Edition, Prentice-Hall, Upper Saddle River.</mixed-citation></ref></ref-list></back></article>