<?xml version="1.0" encoding="UTF-8"?><!DOCTYPE article  PUBLIC "-//NLM//DTD Journal Publishing DTD v3.0 20080202//EN" "http://dtd.nlm.nih.gov/publishing/3.0/journalpublishing3.dtd"><article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" dtd-version="3.0" xml:lang="en" article-type="research article"><front><journal-meta><journal-id journal-id-type="publisher-id">OJS</journal-id><journal-title-group><journal-title>Open Journal of Statistics</journal-title></journal-title-group><issn pub-type="epub">2161-718X</issn><publisher><publisher-name>Scientific Research Publishing</publisher-name></publisher></journal-meta><article-meta><article-id pub-id-type="doi">10.4236/ojs.2021.113026</article-id><article-id pub-id-type="publisher-id">OJS-110121</article-id><article-categories><subj-group subj-group-type="heading"><subject>Articles</subject></subj-group><subj-group subj-group-type="Discipline-v2"><subject>Physics&amp;Mathematics</subject></subj-group></article-categories><title-group><article-title>
 
 
  Inference Procedures on the Generalized Poisson Distribution from Multiple Samples: Comparisons with Nonparametric Models for Analysis of Covariance (ANCOVA) of Count Data
 
</article-title></title-group><contrib-group><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Maha</surname><given-names>Al-Eid</given-names></name><xref ref-type="aff" rid="aff1"><sup>1</sup></xref></contrib><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Mohamed</surname><given-names>M. Shoukri</given-names></name><xref ref-type="aff" rid="aff2"><sup>2</sup></xref><xref ref-type="corresp" rid="cor1"><sup>*</sup></xref></contrib></contrib-group><aff id="aff1"><addr-line>Department of Biostatistics, Epidemiology and Scientific Computing, King Faisal Specialist Hospital and Research Center, Riyadh, Saudi Arabia</addr-line></aff><aff id="aff2"><addr-line>Department of Epidemiology and Biostatistics, Schulich School of Medicine and Dentistry, University of Western Ontario, London, Ontario Canada</addr-line></aff><pub-date pub-type="epub"><day>08</day><month>05</month><year>2021</year></pub-date><volume>11</volume><issue>03</issue><fpage>420</fpage><lpage>436</lpage><history><date date-type="received"><day>18,</day>	<month>May</month>	<year>2021</year></date><date date-type="rev-recd"><day>22,</day>	<month>June</month>	<year>2021</year>	</date><date date-type="accepted"><day>25,</day>	<month>June</month>	<year>2021</year></date></history><permissions><copyright-statement>&#169; Copyright  2014 by authors and Scientific Research Publishing Inc. </copyright-statement><copyright-year>2014</copyright-year><license><license-p>This work is licensed under the Creative Commons Attribution International License (CC BY). http://creativecommons.org/licenses/by/4.0/</license-p></license></permissions><abstract><p>
 
 
  Count data that exhibit over dispersion (variance of counts is larger than its mean) are commonly analyzed using discrete distributions such as negative binomial, Poisson inverse Gaussian and other models. The Poisson is characterized by the equality of mean and variance whereas the Negative Binomial and the Poisson inverse Gaussian have variance larger than the mean and therefore are more appropriate to model over-dispersed count data. As an alternative to these two models, we shall use the generalized Poisson distribution for group comparisons in the presence of multiple covariates. This problem is known as the ANCOVA and is solved for continuous data. Our objectives were to develop ANCOVA using the generalized Poisson distribution, and compare its goodness of fit to that of the nonparametric Generalized Additive Models. We used real life data to show that the model performs quite satisfactorily when compared to the nonparametric Generalized Additive Models.
 
</p></abstract><kwd-group><kwd>Count Regression</kwd><kwd> Over Dispersion</kwd><kwd> Generalized Linear Models</kwd><kwd> Analysis of Covariance</kwd><kwd> Generalized Additive Models</kwd></kwd-group></article-meta></front><body><sec id="s1"><title>1. Introduction</title><p>The Poisson distribution is commonly used to model count data. However, a restriction of this distribution is that the response variable must have a mean equal to the variance. This restriction does not often hold true for many biological and epidemiological data. In many applications the variance can be much larger than the mean, a phenomenon known as “over dispersion”. This over dispersion may occur due to population heterogeneity, or presence of outliers in the data [<xref ref-type="bibr" rid="scirp.110121-ref1">1</xref>]. An analysis of data with overly dispersed counts can lead to the underestimation of parameter standard error, if overdispersion is ignored. A review of the issue of overdispersion in both binary and count data was reviewed by Hinde and Demetrio [<xref ref-type="bibr" rid="scirp.110121-ref2">2</xref>], and in a more recent review by Hayat and Higgins [<xref ref-type="bibr" rid="scirp.110121-ref3">3</xref>]. Diagnosing and accounting for overdispersion is not a simple issue and should be appropriately dealt with to avoid bias in interpreting the results.</p><p>The Negative-Binomial (NB) distribution has been used as an alternative to the Poisson distribution for modeling data that exhibit overdispersion. The NB has two parameters and a variance that is a quadratic function of the mean. NB model has been the model of choice for the analysis of overly dispersed count data. The NB regression was reviewed by Hinde and Demetrio [<xref ref-type="bibr" rid="scirp.110121-ref2">2</xref>]. Joe and Zhu [<xref ref-type="bibr" rid="scirp.110121-ref4">4</xref>] drew a comparison between the NB and a mixture-based generalization of the Poisson distribution.</p><p>In this paper we discuss several inferential statistical issues related to a modified form of the Generalized Poisson Distribution (GPD). The GPD distribution was introduced to the statistical literature by Consul and Jain [<xref ref-type="bibr" rid="scirp.110121-ref5">5</xref>] and a detailed account of its properties was given by Consul [<xref ref-type="bibr" rid="scirp.110121-ref6">6</xref>]. The distribution has two parameters, and a variance that is cubic function of the mean. The distribution has been used to analyze data in the fields of genetics [<xref ref-type="bibr" rid="scirp.110121-ref7">7</xref>] as a queuing model [<xref ref-type="bibr" rid="scirp.110121-ref8">8</xref>] [<xref ref-type="bibr" rid="scirp.110121-ref9">9</xref>] [<xref ref-type="bibr" rid="scirp.110121-ref10">10</xref>] and genomics [<xref ref-type="bibr" rid="scirp.110121-ref11">11</xref>]. The modified form of the GPD, which we shall call “Modified Poisson Distribution” (MGPD) was first discussed in [<xref ref-type="bibr" rid="scirp.110121-ref12">12</xref>]. The modification was a double parametric transformation on the original parameters of the GPD. The main purpose of the transformation was to achieve parameters orthogonality [<xref ref-type="bibr" rid="scirp.110121-ref13">13</xref>] and make the MGPD a member of the class of “Generalized Linear Models” [<xref ref-type="bibr" rid="scirp.110121-ref14">14</xref>]. Recently Shoukri and Al-Eid investigated several inference procedures in the two samples situation [<xref ref-type="bibr" rid="scirp.110121-ref15">15</xref>].</p><p>This paper has three-fold objectives. In Section 2, we present the model. We then, assume that we have k independent samples and we demonstrate how to construct statistical testing procedures on the dispersion parameters. Specifically, we first validate the hypothesis of homogeneity of dispersion parameters, thereafter we test the significance of the common dispersion parameter. In Section 3 we test the hypothesis of equality of k-means in the presence of overdispersion. When covariates are measured, testing the equality of group means is therefore equivalent to the Analysis of covariance (ANCOVA) in the presence of overdispersion. In Section 4 we use the COVID-19 mortality data to draw a comparison between the MGPD, and the Generalized Additive Models (GAM). We demonstrate the differences between the two analytic strategies and highlight the superiority of the MGPD in the analysis of count data exhibiting overdispersion in Section 5. General discussion is presented in Section 6.</p></sec><sec id="s2"><title>2. The Model and Its Parameters Estimation</title><sec id="s2_1"><title>2.1. Modified Generalized Poisson Distribution</title><p>The GPD was introduced by Consul and Jain [<xref ref-type="bibr" rid="scirp.110121-ref5">5</xref>]</p><p>P ( Y = y ) = λ 1 ( λ 1 + λ 2 y ) y − 1 y ! exp [ − λ 1 − λ 2 y ] λ 1 &gt; 0 0 ≼ λ 2 &lt; 1 (2.1)</p><p>The GPD whose probability function is given in (2.1) reduces to the well-known Poisson distribution when λ 2 = 0 . Therefore the parameter λ 2 with the above restriction on its range, is considered the dispersion parameter. Shoukri and Mian [<xref ref-type="bibr" rid="scirp.110121-ref12">12</xref>] employed the parametric transformations:</p><p>λ 1 = μ / ( 1 + ϵ μ ) λ 2 = ϵ λ 1 (2.2)</p><p>Using the transformations in (2.2) we therefore have:</p><p>P ( X = x ) = ( 1 + ϵ x ) x − 1 x ! g x ( μ , ϵ ) exp [ − μ 1 + ϵ μ ] (2.3)</p><p>where g ( μ , ϵ ) = μ 1 + ϵ μ exp [ − ϵ μ 1 + ϵ μ ]</p><p>For fixed ϵ , the function g ( . ) in (2.3) is the natural parameter transformation which renders the GPD a member of the linear family of exponential class (see; [<xref ref-type="bibr" rid="scirp.110121-ref14">14</xref>] ), with a general structure:</p><p>f ( x ) = h ( x ) exp [ ϕ T ( x ) − A ( ϕ ) ] (2.4)</p><p>We call the transformed GPD, the “Modified Generalized Poisson Distribution” or MGPD.</p><p>Shoukri and Mian [<xref ref-type="bibr" rid="scirp.110121-ref12">12</xref>] showed that a recurrence relation among the r t h non-central moments ϑ ′ r is such that:</p><p>ϑ ′ r + 1 = σ 2 ( μ ) ∂ ϑ ′ r ∂ μ + μ ϑ ′ r (2.5)</p><p>From (2.5) we can show that:</p><p>ϑ ′ 0 ≡ 1 , ϑ ′ 1 ≡ μ = E ( Y ) , and σ 2 ( μ ) = μ ( 1 + ϵ μ ) 2 ≡ var ( Y ) (2.6)</p><p>That is the variance is a cubic function of the population mean. We shall deal with the situation when ϵ &gt; 0 .</p></sec><sec id="s2_2"><title>2.2. Point Estimators</title><p>Our approach for parametric estimation in this section will be for a single random sample. If Y 1 , Y 2 , ⋯ , Y n is a random sample from the GPD (2.3), Consul and Shoukri [<xref ref-type="bibr" rid="scirp.110121-ref16">16</xref>] showed that the unique maximum likelihood estimates of the parameters exist if and only if the sample variance is larger than the sample mean. Here we shall use the sample moments to obtain estimators for the model parameters ( μ , ϵ ) .</p><p>Equating the first two sample moments ( y &#175; , s 2 ) to their corresponding population moments</p><p>y &#175; = μ s 2 = 1 n ∑ i = 1 i = 1 ( y i − y &#175; ) 2 = μ ( 1 + ϵ μ ) 2</p><p>and solving for the parameters we get:</p><p>μ ˜ = y &#175; ϵ ˜ = ( s 2 ) 1 / 2 ( y &#175; ) − 3 / 2 − ( y &#175; ) − 1 (2.7)</p><p>The variance of the moment estimators of the mean and the dispersion parameter are respectively given by Shoukri and Al-Eid [<xref ref-type="bibr" rid="scirp.110121-ref15">15</xref>] as:</p><p>var ( μ ^ ) = μ ( 1 + ϵ μ ) 2 / n (2.8a)</p><p>v = var ( ϵ ^ ) = ( 1 + ϵ μ ) 2 2 n μ 2 [ 1 + 2 ϵ + 3 ϵ 2 μ ] (2.8b)</p><sec id="s2_2_1"><title>2.2.1. Homogeneity of Dispersion Parameters</title><p>Suppose that we have k independent random samples from (2.3), which we denote Y i j ~ G P D ( μ i , ϵ i ) with n i observations from the i t h population ( i = 1 , 2 , ⋯ , k ) .</p><p>Wedenote the variance of the estimator of ϵ given in Equation (2.8) by v i and, let w i = 1 / v i , v i is given in (2.8b).</p><p>Cochran [<xref ref-type="bibr" rid="scirp.110121-ref17">17</xref>] developed a general statistic Q which may be used to test the homogeneity of several population parameters. The Q statistics has asymptotically, chi-square distribution with k − 1 degrees of freedoms. It is defined as:</p><p>Q _ e s p = ∑ i = 1 k w i ( ϵ ^ i − ϵ &#175; ) 2 / ∑ i = 1 k w i (2.9)</p><p>where</p><p>ϵ &#175; = ∑ i = 1 k w i ϵ ^ i / ∑ i = 1 k w i (2.10)</p><p>The hypothesis H 0 : ϵ 1 = ϵ 2 = ϵ 3 = ⋯ ϵ k = ϵ of homogeneity of dispersion parameters is rejected whenever the statistic Q _ e s p exceeds Q α , k − 1 , the upper 5% quantile of a chi-square random variable with k − 1 degrees of freedom.</p></sec><sec id="s2_2_2"><title>2.2.2. Testing the Significance of the Common Dispersion Parameter: H 0 : ϵ = 0</title><p>Here we develop a test statistic on the null hypothesis of absence of overdispersion. For the case when μ i ’s are unknown, a uniformly most powerful test for H 0 : ϵ = 0 (Poisson) versus H 1 : ϵ &gt; 0 (GPD) cannot be obtained, however the locally powerful Neyman’s C ( α ) test can be constructed [<xref ref-type="bibr" rid="scirp.110121-ref18">18</xref>]. The log-likelihood function is given by</p><p>l = ∑ i = 1 k ∑ j = 1 n i ( y i j − 1 ) log ( 1 + ε y i j ) + ∑ i = 1 k n i y &#175; i . [ log μ i − log ( 1 + ε μ i ) ] − ∑ i = 1 k n i μ i ( 1 + ε y &#175; i . 1 + μ i ) (2.11)</p><p>where y &#175; i . = y i . n i = ∑ j = 1 n i y i j / n i .</p><p>The locally asymptotically most powerful C ( α ) test is to reject H 0 for large values of ( ∂ l / ∂ ϵ ) ϵ = 0 . From (2.11):</p><p>( ∂ l / ∂ ϵ ) ϵ = 0 = ∑ i = 1 k ∑ j = 1 n i y i j ( y i j − 1 ) − ∑ i = 1 k n i μ i y &#175; i . − ∑ i = 1 k n i μ i ( y &#175; i . − μ i ) (2.12)</p><p>Therefore, the locally asymptotically most powerful C ( α ) test is to reject H 0 for large values of T , where</p><p>T = ( ∂ l / ∂ ϵ ) ϵ = 0 = ∑ i = 1 k ∑ j = 1 n i [ ( y i j − y &#175; i . ) 2 − y i j ] (2.13)</p><p>The statistic (2.13) is obtained from (2.12) by replacing each μ i with root n i consistent estimator, μ ^ i . The simplest μ ^ i is the maximum likelihood estimator μ ^ i = y &#175; i . . Moran [<xref ref-type="bibr" rid="scirp.110121-ref18">18</xref>] pointed out that the C ( α ) test statistic T is asymptotically normal. It can be easily shown that:</p><p>E ( T ) = ∑ i = 1 k [ ( n i − 1 ) μ i ( 1 + ε μ i ) 2 − n i μ i ] (2.14)</p><p>and</p><p>var ( T ) = ∑ i = 1 k { 2 ( n i − 1 ) μ i 2 ( 1 + ε μ i ) 4 + 1 n i [ μ i ( 1 + ε μ i ) 4 ( 1 + 3 μ i + 10 ε μ i + 15 ε 2 μ i 2 ) − 3 μ i 2 ( 1 + ε μ i ) 4 + μ i ( 1 + ε μ i ) 2 − 2 μ i ( 1 + ε μ i ) 3 ] + μ i n i } (2.15)</p><p>Under H 0 : ϵ = 0 , (2.14) and (2.15) reduce respectively to E &#176; = − ∑ i = 1 k μ i and</p><p>v &#176; = ∑ i = 1 k ( 2 ( n i − 1 ) μ i 2 + μ i n i ) (2.16)</p><p>The hypothesis H 0 : ϵ = 0 is rejected whenever:</p><p>Q ( ϵ = 0 ) = ( T − E &#176; ) 2 / v &#176; exceeds Q α , 1 , the upper 5% quantile of a chi-square random variable with one degree of freedom.</p></sec></sec></sec><sec id="s3"><title>3. Testing Equality of Means</title><p>Based on the one-way layout data considered in the previous section, we would like to test the null hypothesis H 0 : μ 1 = ⋯ = μ k = μ against H a : at least two of the μ ′ i s are different, for all ϵ &gt; 0 . The log likelihood under the hypothesis H a : is given by (2.11), and will be denoted by l a , the log likelihood under H 0 will be denoted by l 0 and is obtained by replacing μ i = μ ( i = 1 , 2 , ⋯ , k ) in (2.11). Under H a , the maximum likelihood estimator of μ i is</p><p>μ ^ i = y &#175; i .</p><p>And the maximum likelihood estimator ϵ ^ a , of ϵ is the non-negative root of</p><p>∑ i = 1 k [ ∑ j = 1 n i ( y i j − 1 ) y i j 1 + ε ^ a y i j − n i ( y &#175; i . ) 2 ( 1 − ε ^ a y &#175; i . ) ] = 0 (3.1)</p><p>Under H 0 the maximum likelihood estimator of the common mean μ is μ ^ = y .. / N = y &#175; , where, N = ∑ i = 1 k n i .</p><p>The maximum likelihood estimator of ϵ ^ o and ϵ under H 0 is the positive root of</p><p>∑ i = 1 k ∑ j = 1 n i ( y i j − 1 ) y i j 1 + ϵ ^ o y i j − N ( y &#175; ) 2 ( 1 + ϵ ^ o Y &#175; ) = 0 (3.2)</p><p>Detailed discussion on the necessary and sufficient conditions that (3.1) and (3.2) to have a unique root is given in Consul and Shoukri [<xref ref-type="bibr" rid="scirp.110121-ref16">16</xref>].</p><p>Denoting the maximized log likelihood under H a by L a , and that under H 0 by L 0 , the likelihood ratio test, which has an asymptotic distribution of chi-squared with ( k − 1 ) degree of freedom is:</p><p>λ = 2 ( L a − L 0 ) = 2 [ ∑ i = 1 k ∑ j = 1 n i ( y i j − 1 ) log { 1 + ε ^ a y i j 1 + ε ^ 0 y i j } + ∑ i = 1 k n i y &#175; i log { y &#175; i . ( 1 + ε ^ 0 y &#175; ) y &#175; ( 1 + ε ^ a y &#175; i . ) } ] (3.3)</p><p>As an alternative to the likelihood ratio test (3.3), we present the Neyman’s C ( α ) statistic which has local optimal properties. Suppose that μ i can be written as μ i = μ + δ i with δ k = 0 . Then testing the null hypothesis H 0 : μ 1 = μ 2 = ⋯ = μ k is equivalent to testing H 0 : δ i = o ( i = 1 , 2 , 3 , ⋯ , k ) , where μ and ϵ are nuisance parameters. We reparametrize (11.2), and denote the resulting function by l * .</p><p>Define δ = ( δ 1 , δ 2 , ⋯ , δ k − 1 ) , τ = ( τ 1 , τ 2 ) ′ = ( μ , ϵ ) .</p><p>ϕ i ( τ ) = [ ∂ l * ∂ δ i ] δ &#175; = 0 i = 1 , 2 , ⋯ , k − 1</p><p>Δ i ( τ ) = [ ∂ l * ∂ τ j ] δ &#175; = 0 j = 1 , 2</p><p>Let τ ^ be any root-n consistent estimator of τ under the null hypothesis. Moran [<xref ref-type="bibr" rid="scirp.110121-ref18">18</xref>] showed that the C ( α ) test is based on F i ( τ ^ ) = ϕ i ( τ ^ ) − γ i 1 Δ i 1 ( τ ^ ) − γ i 2 Δ i 2 ( τ ^ ) , where γ i 1 and γ i 2 are the partial regression coefficients of ϕ i on Δ 1 and Δ 2 respectively. Define the following matrices:</p><p>P i j = − E [ ∂ 2 l * ∂ δ i ∂ δ j ] δ &#175; = 0 , Q i j = − E [ ∂ 2 l * ∂ δ i ∂ τ j ] δ &#175; = 0 ,</p><p>and R i j = − E [ ∂ 2 l * ∂ τ i ∂ τ j ] δ &#175; = 0 .</p><p>Here, we replace by its estimator τ ^ in F, P, Q, and R, the C ( α ) test statistic is given by</p><p>F ′ ( p − Q R − 1 Q ′ ) − 1 F (3.5)</p><p>The asymptotic distribution of the test statistic given in (3.5) will be that of a chi-square with k − 1 degrees of freedom.</p><p>Now, there are two possible root-n consistent estimators of τ , under H 0 :</p><p>The first is the maximum likelihood estimator τ ^ = ( y &#175; , ϵ ^ 0 ) ′ , which on substitution we get Δ j ( τ ^ ) = 0 ( j = 1 , 2 ) , and hence F j ( τ ^ ) = ϕ i ( τ ^ ) . Accordingly, (3.5) reduces to</p><p>C 2 = ∑ i = 1 k n i ( y &#175; i . − y &#175; ) 2 y &#175; ( 1 + ε ^ 0 y &#175; ) 2 (3.6)</p><p>The hypothesis of equality of population means is thus rejected whenever C 2 exceeds Q α , k − 1 , the upper 5% quantile of a chi-square random variable with k − 1 degrees of freedom. For more details we refer the reader to [<xref ref-type="bibr" rid="scirp.110121-ref12">12</xref>].</p></sec><sec id="s4"><title>4. ANCOVA: The Generalized Poisson Regression</title><p>It is well-known that ANOVA and regression are related techniques that are concerned with testing the differences in group means after adjusting for the confounding effects of potential risk factors and covariates. Since the MGPD is a member of the linear exponential family (for fixed ϵ ) Shoukri and Mian [<xref ref-type="bibr" rid="scirp.110121-ref12">12</xref>] expressed the expectation μ i of y i as:</p><p>η ( μ i ) = X i T β (4.1)</p><p>In Equation (4.1), X i is a set of measured ( P + 1 ) covariates, and a subset of theses covariates defines a set of indicators (dummy) variables to identify categorical effects. The transformation η ( ⋅ ) is a monotone, differentiable function named “the link function”. To estimate β 0 , β 1 , ⋯ , β p , and ϵ we construct the log-link so that:</p><p>μ i ( x ) = exp [ X i T β ]</p><p>The logarithm of the likelihood function will be proportional to</p><p>l = l ( β , ε ) = ∑ i = 1 k ( y i − 1 ) ln ( 1 + ε y i ) + ∑ i = 1 k y i ln μ i ( x ) − ∑ i = 1 k y i ln ( 1 + ε μ i ( x ) ) − ∑ i = 1 k μ i ( x ) ( 1 + ε y i ) 1 + ε μ i ( x ) (4.2)</p><p>The first and second partial derivatives are given by:</p><p>∂ l ∂ ε = l ˙ ε = ∑ i = 1 k y i ( y i − 1 ) 1 + ε y i − ∑ i = 1 k y i μ i ( x ) 1 + ε μ i ( x ) − ∑ i = 1 k μ i ( x ) ( y i − μ i ( x ) ) ( 1 + ε μ i ( x ) ) 2 (4.3)</p><p>∂ l ∂ β r = l ˙ r = ∑ i = 1 k ( y i − μ i ( x ) ) ( 1 + ε μ i ( x ) ) 2 x i r (4.4)</p><p>∂ 2 l ∂ ε ∂ β r = l &#168; ε r = − 2 ∑ i = 1 k μ i ( x ) ( y i − μ i ( x ) ) ( 1 + ε μ i ( x ) ) 3 x i r (4.5)</p><p>∂ 2 l ∂ β r ∂ β s = l &#168; r s = − ∑ i = 1 k [ μ i ( x ) ( 1 + ε μ i ( x ) ) 2 x i r x i s + 2 ε ( y i − μ i ( x ) ) μ i ( x ) ( 1 + ε μ i ( x ) ) 3 x i r x i s ] (4.6)</p><p>∂ 2 l ∂ ε 2 = l &#168; ε r − ∑ i = 1 k y i 2 ( y i − 1 ) ( 1 + ε y i ) 2 + ∑ i = 1 k y 1 μ i 2 ( x ) ( 1 + ε μ i ) 2 + 2 ∑ i = 1 k μ i 2 ( x ) ( y i − μ i ( x ) ) ( 1 + ε μ i ( x ) ) 3 (4.7)</p><p>Taking the expected value of the negative of the second partial derivatives we get the Fishers’ information matrix I, whose elements are:</p><p>− E [ l &#168; r s ] = I r s = ∑ i = 1 k μ i ( x ) ( 1 + ε μ i ( x ) ) 2 x i r x i s , r , s , = 1 , 2 , ⋯ , p (4.8)</p><p>− E [ l &#168; ε s ] = I ε r = 0 (4.9)</p><p>From Consul and Shoukri [<xref ref-type="bibr" rid="scirp.110121-ref16">16</xref>] we get:</p><p>− E [ l &#168; ε ε ] = i ε ε = 2 ( 1 + 2 ε ) − 1 ∑ i = 1 k μ i 2 ( x ) ( 1 + ε μ i ( x ) ) 2 (4.10)</p><p>The asymptotic distributions of the regression estimators can be established using the results in [<xref ref-type="bibr" rid="scirp.110121-ref12">12</xref>].</p><p>Our approach to the data analysis when the main interest is comparing group means in the presence of potential risk factors and confounders is summarized in three steps. In the first step we use the MLE to estimate the regression parameters using Equation (4.2), without including the groups as independent variable. In the second step, we extract the residuals (E) of the generalized Poisson regression model, defines as:</p><p>E = = Observed dependent variable − predicted val − ue of the dependent variable</p><p>In the final step we test the normality and variance homogeneity of E. Thereafter, we use nonparametric ANOVA with the residuals being the dependent variable, and the groups being the independent variables to complete the ANCOVA testing.</p></sec><sec id="s5"><title>5. Data Analyses</title><p>Al-Gahtani et al. [<xref ref-type="bibr" rid="scirp.110121-ref19">19</xref>] analyzed COVID-19 case fatality data collected retrospectively from the start of the of the epidemic to December 2-2020, the day the Pfizer vaccine was approved by the American Center for Disease Control (CDC). The data were collected from 120 countries grouped into 15 regions [<xref ref-type="bibr" rid="scirp.110121-ref19">19</xref>] as shown in <xref ref-type="table" rid="table1">Table 1</xref>. We will reanalyze the data such that:</p><p>The response variable is the aggregate number of COVID-19 deaths which we denote by “y”. We shall use different set of covariates, and these are:</p><p>&#183; Region: The factor variable which is the main effect.</p><p>&#183; The other covariates are:</p><p>1) X<sub>1</sub> =log (percentage of obese personsin a country reported in 2018) [<xref ref-type="bibr" rid="scirp.110121-ref21">21</xref>] [<xref ref-type="bibr" rid="scirp.110121-ref22">22</xref>];</p><p>2) X<sub>2</sub> = log (population density) [<xref ref-type="bibr" rid="scirp.110121-ref23">23</xref>];</p><p>3) X<sub>3</sub> = log (number of people with colorectal cancer in a country reported in 2017) [<xref ref-type="bibr" rid="scirp.110121-ref24">24</xref>];</p><p>4) X<sub>4</sub> = log (Chronic Kidney Disease—case fatality in a country as reported in 2017) [<xref ref-type="bibr" rid="scirp.110121-ref20">20</xref>].</p><p>In <xref ref-type="fig" rid="fig1">Figure 1</xref> we show the histogram of (y), the aggregate of COVID-19 deaths during the study period. The distribution is positively skewed with variance much larger than the mean.</p><p>Direct calculations from the summary statistics given in <xref ref-type="table" rid="table2">Table 2</xref> give:</p><p>Q _ e s p = 0.00004 , and the corresponding p-value = 0.999. Therefore, the hypothesis of homogeneity of dispersion parameters is supported by the data. Moreover, Q ( ϵ = 0 ) is quite large and the corresponding p-value = 0.00001.</p><table-wrap id="table1" ><label><xref ref-type="table" rid="table1">Table 1</xref></label><caption><title> Adaptedtable: Countries and the corresponding Regional classification as given in https://doi.org/10.1016/s0140-6736(20)30045-3 [<xref ref-type="bibr" rid="scirp.110121-ref20">20</xref>]. In the first column we have the countries, in the second column we have Region or group name followed by the number of countries within the group. In the last column we have Region code</title></caption><table><tbody><thead><tr><th align="center" valign="middle" >Countries</th><th align="center" valign="middle" >Region name</th><th align="center" valign="middle" >Region Code</th></tr></thead><tr><td align="center" valign="middle" >Peru, Ecuador, Bolivia</td><td align="center" valign="middle" >Andean. Latin (3)</td><td align="center" valign="middle" >10</td></tr><tr><td align="center" valign="middle" >Kazakhstan, Georgia, Armenia, Azerbaijan, Kyrgyzstan, Uzbekistan, Tajikistan</td><td align="center" valign="middle" >Central Asia (7)</td><td align="center" valign="middle" >2</td></tr><tr><td align="center" valign="middle" >Czechia, Romania, Hungary, Serbia, Bulgaria, Croatia, Slovakia, Bosnia, Slovenia, North-Macedonia, Albania, Montenegro</td><td align="center" valign="middle" >Central Europe (12)</td><td align="center" valign="middle" >5</td></tr><tr><td align="center" valign="middle" >Brazil, Columbia, Mexico, Panama, Costa Rica, Guatemala, Honduras, Venezuela Paraguay, El-Salvador</td><td align="center" valign="middle" >Central Latin America (10)</td><td align="center" valign="middle" >11</td></tr><tr><td align="center" valign="middle" >Dominican Republic, Puerto Rico, Jamaica</td><td align="center" valign="middle" >Caribbean (3)</td><td align="center" valign="middle" >9</td></tr><tr><td align="center" valign="middle" >Ethiopia, Kenya, Uganda, Zambia, Madagascar. Mozambique, Angola, French Guinea</td><td align="center" valign="middle" >CESSA (8)</td><td align="center" valign="middle" >13</td></tr><tr><td align="center" valign="middle" >Indonesia, Philippine, China, Myanmar, Malaysia, Sri Lanka, French Polynesia, Maldives</td><td align="center" valign="middle" >East Asia (8)</td><td align="center" valign="middle" >1</td></tr><tr><td align="center" valign="middle" >Russia, Poland, Ukraine Belarus, Lithuania, Latvia, Estonia</td><td align="center" valign="middle" >East Europe (7)</td><td align="center" valign="middle" >6</td></tr><tr><td align="center" valign="middle" >Japan, Singapore, Republic Korea, Australia</td><td align="center" valign="middle" >HIAP (4)</td><td align="center" valign="middle" >4</td></tr><tr><td align="center" valign="middle" >Iran, Iraq, Turkey, Morocco, Saudi Arabia, Israel, Jordan, United Arab, Kuwait, Qatar, Lebanon, Oman, Egypt, Occupied Palestine, Tunisia, Bahrain, Algeria, Libya, Afghanistan, Sudan</td><td align="center" valign="middle" >MENA (20)</td><td align="center" valign="middle" >12</td></tr><tr><td align="center" valign="middle" >India, Bangladesh, Pakistan, Nepal</td><td align="center" valign="middle" >South Asia (4)</td><td align="center" valign="middle" >3</td></tr><tr><td align="center" valign="middle" >Chile, Argentina, Uruguay</td><td align="center" valign="middle" >South Latin America (3)</td><td align="center" valign="middle" >8</td></tr><tr><td align="center" valign="middle" >South Africa, Namibia, Zimbabwe</td><td align="center" valign="middle" >SSSA (3)</td><td align="center" valign="middle" >15</td></tr><tr><td align="center" valign="middle" >USA, France, Spain, UK, Italy, Germany, Belgium, Netherland Canada, Switzerland, Portugal, Austria, Sweden, Greece, Denmark, Ireland, Norway, Luxemburg, Finland, Cyprus</td><td align="center" valign="middle" >West Europe and North America (20)</td><td align="center" valign="middle" >7</td></tr><tr><td align="center" valign="middle" >Mali, Nigeria, Ghana, Cameroon, Ivory Coast Senegal, Guinea, Cape Verde</td><td align="center" valign="middle" >WSSA (8)</td><td align="center" valign="middle" >14</td></tr></tbody></table></table-wrap><p>CSSA = Central Sub-Saharan-Africa; MENA = Middle East and North Africa; HIAP = High Income Asian Pacific; WSSA = Western Sub-Saharan-Africa; SSSA = Southern Sub-Saharan Africa.</p><table-wrap id="table2" ><label><xref ref-type="table" rid="table2">Table 2</xref></label><caption><title> Summary measures of COVID-19 deaths: group sizes (n), means (m), standard deviation (s) and the estimates of the dispersion parameter (eps) per group</title></caption><table><tbody><thead><tr><th align="center" valign="middle" >Region</th><th align="center" valign="middle" >n</th><th align="center" valign="middle" >m</th><th align="center" valign="middle" >s</th><th align="center" valign="middle" >eps</th></tr></thead><tr><td align="center" valign="middle" >1 An. Latin 2 C. Asia 3 C. EUROP 4 C. Latin 5 Caribbean 6 CESSA 7 E. ASIA 8 E. Europe 9 HIAP 10 MENA 11 S. Asia 12 S. Latin 13 SSSA 14 W. Eur 15 WSSA</td><td align="center" valign="middle" >3 7 11 10 3 7 6 7 4 19 4 2 3 20 7</td><td align="center" valign="middle" >19,496 1350 3515 33,158 976 640 5452 10,488 920 5986 38,606 27,080 7358 28,061 692</td><td align="center" valign="middle" >14,448 836 3577 59,252 1177 658 6495 15,213 934 11,027 66,403 16,476 12,372 59,609 806</td><td align="center" valign="middle" >0.005 0.016 0.018 0.01 0.037 0.039 0.016 0.014 0.032 0.024 0.009 0.004 0.019 0.013 0.043</td></tr></tbody></table></table-wrap><p>Therefore, the hypotheses that the common dispersion is not significantly different from zero is not supported by the data. The C<sup>2</sup>-statistic is quite large as well, and the corresponding p-value is near zero, therefore the hypothesis of equality of mean counts in all regions (aggregate COVID-19 deaths) is also not supported by the data.</p><p>We shall write a function using the R-program for the estimation of the regression parameters. The iteration process requires staring points. We obtain the staring points by first fitting the classical Poisson regression, which is done using the following code:</p><p>out1 = GLM (y~x<sub>1</sub> + x<sub>2</sub> + x<sub>3</sub> + x<sub>4</sub>, data = data2, family = Poisson).</p><p>Having obtained the parameter estimates from the Poisson regression, we use them to start the iteration process and obtain final estimates as shown in the Appendix.</p><p>The MGPD regression results are summarized in <xref ref-type="table" rid="table3">Table 3</xref>.</p><p>The correlation between the observed and predicted COVID-19 death counts is (0.758).</p><p><xref ref-type="fig" rid="fig2">Figure 2</xref> gives the Q-Q plot of the quantiles of model residuals exhibiting close agreement among with the quantiles of the standard normal distribution.</p><p>To complete the ANCOVA testing we use theKruskal-Wallis test whereby residuals of the MGPD regression model are used as dependent variables and the “Regions”, or groups as independent variables. The results are summarized as follows:</p><p>Kruskal-Wallis chi-squared = 14.936, p-value = 0.1344.</p><p>Therefore, after adjusting for the covariates, there is not sufficient evidence to</p><table-wrap id="table3" ><label><xref ref-type="table" rid="table3">Table 3</xref></label><caption><title> Maximum likelihood estimation of the MGPD regression using R</title></caption><table><tbody><thead><tr><th align="center" valign="middle" >Parameter</th><th align="center" valign="middle" >Estimate</th><th align="center" valign="middle" >SE</th><th align="center" valign="middle" >Z</th><th align="center" valign="middle" >p value</th></tr></thead><tr><td align="center" valign="middle" >1-Intercept 2-X<sub>1 </sub> 3-X<sub>2 </sub> 4-X<sub>3 </sub> 5-X<sub>4 </sub> 6-𝟄</td><td align="center" valign="middle" >−0.299 0.659 0.489 0.425 −0.385 0.028</td><td align="center" valign="middle" >1.138 0.186 0.131 0.107 0.090 0.002</td><td align="center" valign="middle" >−0.262 3.545 3.729 3.400 −4.280 13.537</td><td align="center" valign="middle" >0.7900 0.0004 0.0002 0.0002 0.00002 0.000001</td></tr></tbody></table></table-wrap><p>reject the hypothesis of equality of mean counts in COVID-19 deaths among the “Regions”.</p></sec><sec id="s6"><title>6. Nonparametric Regression Modeling: Generalized Additive Models</title><p>The Generalized Additive Models (GAM) are recent developments that are becoming popular as modeling techniques. It is nonparametric in nature and, even though less powerful, it is quite robust against departure from the assumptions required by classical GLM regression models. The GAM allow us to include non-linear smoothers into the modeling strategy. In mathematical terms GAM solve the following equation:</p><p>g ( μ i ) = f 1 ( x 1 ) + f 2 ( x 2 ) + f 3 ( x 3 ) + f 4 ( x 4 ) + f 5 ( x 5 ) (6.1)</p><p>The f j ( x j ) are smooth functions to be estimated. Equation (6.1) seems complex, but it is very simple to understand. The first thing to notice is that with GAM we are not necessarily estimating the response directly, i.e. we are not modelling y. In fact, as with GLM we have the possibility to use link functions to model non-normal response variables (and thus perform Poisson or logistic regression) [<xref ref-type="bibr" rid="scirp.110121-ref14">14</xref>]. Therefore, the term g ( μ ) is simply the transformation of y needed to “linearize” the model. When we are dealing with a normally distributed response this term is simply replaced by y. The second part of the equation, where we have two terms: the parametric and the non-parametric part. In GAM we can include all the parametric terms we can include in Linear Model or GLM, for example linear or polynomial terms. The second part is the non-parametric smoother that will be automatically fitted, and it is the key point of GAMs. A complete and lucid account of the GAM theory can be found in [<xref ref-type="bibr" rid="scirp.110121-ref25">25</xref>] [<xref ref-type="bibr" rid="scirp.110121-ref26">26</xref>] [<xref ref-type="bibr" rid="scirp.110121-ref27">27</xref>].</p><p>We fitted the GAM to the data using the R-package “GAM”, and the next two lines are the needed code:</p><p>library(gam);</p><p>agam=gam(y~Region+x1+x2+x3+x4,data=data2).</p><p>Call: gam(formula = y ~ code + x1 + x2 + x3 + x4, data = data2).</p><p>The following results are obtained from the GAM fitting to the data:</p><p>1) Null Deviance: 135512150629 on 112 degrees of freedom;</p><p>2) Residual Deviance: 74451674616 on 96 degrees of freedom.</p><p>From which: The correlation between observed counts and predicted counts is: (1-74451674616/135512150629)<sup>1/2</sup> = 0.671.</p><p>The GAM results are shown in <xref ref-type="table" rid="table4">Table 4</xref>, and the Q-Q plot of the model residuals is given in <xref ref-type="fig" rid="fig3">Figure 3</xref>, showing that the model residuals are not as close to normality as the residuals of the MGPD regression model.</p></sec><sec id="s7"><title>7. Discussion</title><p>In this paper we demonstrated the use of the MGPD as a model for the ANCOVA. We used a two-steps approach. In the first step we used the regression models to</p><table-wrap id="table4" ><label><xref ref-type="table" rid="table4">Table 4</xref></label><caption><title> The results of fitting the GAM: ANOVA for Parametric Effects</title></caption><table><tbody><thead><tr><th align="center" valign="middle" >Source</th><th align="center" valign="middle" >DF</th><th align="center" valign="middle" >Sum Sq</th><th align="center" valign="middle" >Mean Sq</th><th align="center" valign="middle" >F value</th><th align="center" valign="middle" >Pr. (&gt;F.)</th></tr></thead><tr><td align="center" valign="middle" >Region x<sub>1</sub> x<sub>2</sub> x<sub>3</sub> x<sub>4</sub> Residuals</td><td align="center" valign="middle" >12 1 1 1 1 96</td><td align="center" valign="middle" >1.5856e+10 1.9335e+09 2.3000e+10 1.1518e+10 8.7527e+09 7.4452e+10</td><td align="center" valign="middle" >1.3213e+09 1.9335e+09 2.3000e+10 1.1518e+10 8.7527e+09 7.7554e+08</td><td align="center" valign="middle" >1.7038 2.4932 29.6566 14.8522 11.2859</td><td align="center" valign="middle" >0.078 0.118 3.963e−07*** 0.0002 0.001</td></tr></tbody></table></table-wrap><p>***significant at level of significance less than 0.00001.</p><p>assess the influence of possible confounders and covariates on the outcome of interest. Thereafter we extracted the model residuals and used these residuals as a dependent variable of a nonparametric ANOVA, with the groups being the independent predictors. We note that while there was a significant difference among the group means in the univariate analysis, such difference was not significant in the second step of the ANCOVA. We note that the MGPD regression model showed high correlation (0.758) between the observed counts and the model based predicted counts, indicative of a good fit by the model to the given data. On using the Q-Q plot, model residuals are shown to have close agreement to the empirical quantiles of the standard normal distribution. This shows that the model is quite reliable as a predictive tool, and that the distribution of the estimated regression parameters is that of a multivariate normal.</p><p>For the sake of comparison, we fitted the data using the GAM, a nonparametric regression approach. This approach deals with the covariates as factors. The GAM model showed that after adjusting for the covariates within the same model, there are no significant differences among regions. These findings are in agreement with those based on the MGPD regression. The GAM did not produce estimate for the dispersion parameter 𝟄. The measure of goodness of fit of the GAM was (0.671), which is much lower than that of the MGPD. The MGPD model has several advantages when compared to the GAM. First, The GAM cannot be used as a predictive tool, while the MGPD model can be used to predict the mean of the response variable. Second, the residuals of the MGPD regression model have a distribution that is almost normal. This emphasizes the reliability of the likelihood based statistical estimation of the model parameters. Finally, our two-steps approach to data fitting makes helps avoiding both overfitting and possible multicollinearity.</p></sec><sec id="s8"><title>Conflicts of Interest</title><p>The authors declare no conflicts of interest regarding the publication of this paper.</p></sec><sec id="s9"><title>Cite this paper</title><p>Al-Eid, M. and Shoukri, M.M. (2021) Inference Procedures on the Generalized Poisson Distribution from Multiple Samples: Comparisons with Nonparametric Models for Analysis of Covariance (ANCOVA) of Count Data. Open Journal of Statistics, 11, 420-436. https://doi.org/10.4236/ojs.2021.113026</p></sec><sec id="s10"><title>Appendix A</title><p>R-Code fitting the Generalized Poisson using the method of maximum likelihood.</p><p>###Notations: b<sub>i</sub> are the regression parameters, mu is the mean function, “n” is the number of observations “ll” denotes the ###log-likelihood function, eta is the linear predictor ###</p><p>llik= function(y,par){</p><p>b0=par[<xref ref-type="bibr" rid="scirp.110121-ref1">1</xref>]</p><p>b1=par[<xref ref-type="bibr" rid="scirp.110121-ref2">2</xref>]</p><p>b2=par[<xref ref-type="bibr" rid="scirp.110121-ref3">3</xref>]</p><p>b3=par[<xref ref-type="bibr" rid="scirp.110121-ref4">4</xref>]</p><p>b4=par[<xref ref-type="bibr" rid="scirp.110121-ref5">5</xref>]</p><p>k=par[<xref ref-type="bibr" rid="scirp.110121-ref6">6</xref>]</p><p>n=length(y)</p><p>eta=b0+b1*x1+b2*x2+b3*x3+b4*x4</p><p>mu=exp(eta)</p><p>ll= sum(y*log(mu/(1+(k*mu))))+sum((y-1)*log(1+(k*y))</p><p>+((-mu*(1+(k*y)))/ (1+(k*mu)))-lgamma(y+1))</p><p>return(-ll)</p><p>}</p><p>res=optim(par=c(1.4,.84,.07,.95,-.37,.1),llik,y=y,method=“BFGS”,hessian=T)</p><p>theta=res$par</p><p>theta</p><p>#CALCULATING THE STANDARD ERRORS OF MLE</p><p>out2=nlm(llik,theta,y=y,hessian=TRUE)</p><p>summary(out2)</p><p>plot(data2$y,resid(out2))</p><p>data_new=data.rame(data2$y,resid(out2))</p><p>fish=out2$hessian</p><p>solve(fish)</p><p>element=diag((solve(fish)))</p><p>se=sqrt(element)</p><p>se</p><p>z=theta/se</p><p>out.GMPD=data.frame(theta,se,z)</p><p>out.GMPD</p><p>#### FINAL ESTIMATES -0.30007804 0.65884589 0.48926214 0.42536091 -0.38469265</p><p>data2$y_hat=exp(-.3+.66*data2$x1+.5*data2$x2+.425*data2$x3-.384*data2$x4)</p><p>data2$y</p><p>data2_error=data.frame(data2$y,data2$y_hat)</p><p>cor(data2$y,data2$y_hat) #####0.76</p><p>data2_error=data2$y-data2$y_hat</p><p>data2$response=sqrt((data2_error)^(1/3))</p><p>qqnorm(data2$response)</p><p>### SHAPITO WILK TEST OF NORMALITY####</p><p>shapiro.test(data2$response)</p><p>leveneTest(data2$response~data2$Region)</p><p>aov_result=aov(data2$response~data2$Region)</p><p>####ANOVA ON THE RESIDUALS WITH REGION BEING THE INDEPENDENT VARIABLEUSING KRUSKAL_WALLIS####</p><p>levels(data2$Region)</p><p>aov_result=aov(data2$response~data2$Region)</p><p>summary(aov_result)</p><p>boxplot(data2_error~data2$Region,xlab=“Region”,ylab=“GPD Residuals”,main=“CODID-19 Deaths”)</p><p>kruskal_result=kruskal.test(data2$response~data2$Region)</p><p>###END OF CODE####.</p></sec></body><back><ref-list><title>References</title><ref id="scirp.110121-ref1"><label>1</label><mixed-citation publication-type="other" xlink:type="simple">Cox, D.R. (1983) Some Remarks on Overdispersion. Biometrika, 70, 269-274.  
https://doi.org/10.1093/biomet/70.1.269</mixed-citation></ref><ref id="scirp.110121-ref2"><label>2</label><mixed-citation publication-type="other" xlink:type="simple">Hinde, J. and Demetrio, C.G.B. (1998) Overdispersion: Models and Estimation. Computational statistics and Data Analysis, 27, 151-170.  
https://doi.org/10.1016/S0167-9473(98)00007-3</mixed-citation></ref><ref id="scirp.110121-ref3"><label>3</label><mixed-citation publication-type="other" xlink:type="simple">Hayat, M.J. and Higgins, M. (2014) Understanding Poisson Regression. Journal of Nursing Education, 53, 207-215. https://doi.org/10.3928/01484834-20140325-04</mixed-citation></ref><ref id="scirp.110121-ref4"><label>4</label><mixed-citation publication-type="other" xlink:type="simple">Joe, H. and Zhu, R. (2005) Generalized Poisson Distribution: The Property of Mixture of Poisson and Comparison with the Negative Binomial Distribution. Biometrical Journal, 47, 219-229. https://doi.org/10.1002/bimj.200410102</mixed-citation></ref><ref id="scirp.110121-ref5"><label>5</label><mixed-citation publication-type="other" xlink:type="simple">Consul, P.C. and Jain, G.C. (1970) On the Generalization of Poisson Distribution. Annals of Mathematical Statistics, 41, 1387.</mixed-citation></ref><ref id="scirp.110121-ref6"><label>6</label><mixed-citation publication-type="other" xlink:type="simple">Consul, P.C. (1989) Generalized Poisson Distribution. Marcel Dekker Inc., New York.</mixed-citation></ref><ref id="scirp.110121-ref7"><label>7</label><mixed-citation publication-type="other" xlink:type="simple">Janardan, K.G. and Schaeffer, D.J. (1977) Models for the Analysis of Chromosomal Aberrations in Human Leukocytes. Biometrical Journal, 19, 599-612.  
https://doi.org/10.1002/bimj.4710190804</mixed-citation></ref><ref id="scirp.110121-ref8"><label>8</label><mixed-citation publication-type="other" xlink:type="simple">Tanner, J.C. (1961) A Derivation of Borel Distribution. Biometrika, 40, 222-224.  
https://doi.org/10.1093/biomet/48.1-2.222</mixed-citation></ref><ref id="scirp.110121-ref9"><label>9</label><mixed-citation publication-type="other" xlink:type="simple">Consul, P.C. and Shoukri, M.M. (1988) Some Chance Mechanisms Related to a Generalized Poisson Probability Model. American Journal of Mathematical and Management Sciences, 8, 181-202. https://doi.org/10.1080/01966324.1988.10737237</mixed-citation></ref><ref id="scirp.110121-ref10"><label>10</label><mixed-citation publication-type="other" xlink:type="simple">Jiang, H. and Wong, W.H. (2009) Statistical Inferences for Isoform Expression in RNA-Seq. Bioinformatics, 25, 1026-1032.  
https://doi.org/10.1093/bioinformatics/btp113</mixed-citation></ref><ref id="scirp.110121-ref11"><label>11</label><mixed-citation publication-type="other" xlink:type="simple">Srivastava, S. and Chen, L. (2010) A Two-Parameter Generalized Poisson Model to Improve the Analysis of RNA-seq Data. Nucleic Acids Research, 38, e170.  
https://doi.org/10.1093/nar/gkq670</mixed-citation></ref><ref id="scirp.110121-ref12"><label>12</label><mixed-citation publication-type="other" xlink:type="simple">Shoukri, M.M. and Mian, I.U.H. (1991) Some Aspects of Statistical Inference on the Lagrange (Generalized) Poisson Distribution. Communication in Statistics: Computations and Simulations, 20, 1115-1137.  
https://doi.org/10.1080/03610919108812999</mixed-citation></ref><ref id="scirp.110121-ref13"><label>13</label><mixed-citation publication-type="other" xlink:type="simple">Cox, D.R. and Reid, N. (1987) Parameter Orthogonality and Approximate Conditional Inference. Journal of the Royal Statistical Society. Series B (Methodological), 49, 1-39. https://doi.org/10.1111/j.2517-6161.1987.tb01422.x</mixed-citation></ref><ref id="scirp.110121-ref14"><label>14</label><mixed-citation publication-type="other" xlink:type="simple">McCullagh, P. and Nelder, J.A. (1989) Generalized Linear Models. Chapman and Hall, London. https://doi.org/10.1007/978-1-4899-3242-6</mixed-citation></ref><ref id="scirp.110121-ref15"><label>15</label><mixed-citation publication-type="other" xlink:type="simple">Shoukri, M.M. and Al-Eid, M. (2020) Inference Procedures on the Ratio of Modified Generalized Poisson Distribution Means: Applications to RNA_SEQ Data. International Journal of Statistics in Medical Research, 9, 41-49.  
https://doi.org/10.6000/1929-6029.2020.09.05</mixed-citation></ref><ref id="scirp.110121-ref16"><label>16</label><mixed-citation publication-type="other" xlink:type="simple">Consul, P.C. and Shoukri, M.M. (1984) Maximum Likelihood Estimation of the Generalized Poisson Distribution. Communications in Statistics, Theory and Methods, 13, 1533-1547. https://doi.org/10.1080/03610928408828776</mixed-citation></ref><ref id="scirp.110121-ref17"><label>17</label><mixed-citation publication-type="other" xlink:type="simple">Cochran, W.G. (1954) Some Methods for Strengthening the Common X2 Tests. Biometrics, 10, 417-451. https://doi.org/10.2307/3001616</mixed-citation></ref><ref id="scirp.110121-ref18"><label>18</label><mixed-citation publication-type="other" xlink:type="simple">Moran, P.A.P. (1970) On Asymptotically Optimal Test of Composite Hypotheses. Biometrika, 57, 47-55. https://doi.org/10.1093/biomet/57.1.47</mixed-citation></ref><ref id="scirp.110121-ref19"><label>19</label><mixed-citation publication-type="other" xlink:type="simple">Al-Gahtani, S., Shoukri, M. and Al-Eid, M. (2021) Predictors of the Aggregate of COVID-19 Cases and Its Case-Fatality: A Global Investigation Involving 120 Countries. Open Journal of Statistics, 11, 259-277. https://www.scirp.org/journal/ojs  
https://doi.org/10.4236/ojs.2021.112014</mixed-citation></ref><ref id="scirp.110121-ref20"><label>20</label><mixed-citation publication-type="other" xlink:type="simple">Cockwell, P. and Fisher, L.-A. (2017) Global, Regional, and National Burden of Chronic Kidney Disease, 1990-2017: A Systematic Analysis for the Global Burden of Disease Study. The Lancet, 395, 709-733.</mixed-citation></ref><ref id="scirp.110121-ref21"><label>21</label><mixed-citation publication-type="other" xlink:type="simple">Rottoli, M., Bernante, P., Garelli, S. and Gianella, M. (2020) How Important Is Obesity as a Risk Factor for Respiratory Failure, Intensive Care Admission and Death in Hospitalized COVID-19 Patients? Results from a Single Italian Center. European Journal of Endocrinology, 183, 389-397. https://doi.org/10.1530/EJE-20-0541</mixed-citation></ref><ref id="scirp.110121-ref22"><label>22</label><mixed-citation publication-type="other" xlink:type="simple">https://worldpopulationreview.com/en/country-rankings/obesity-rates-by-country</mixed-citation></ref><ref id="scirp.110121-ref23"><label>23</label><mixed-citation publication-type="other" xlink:type="simple">Rashed, E.A., Kodera, S., Gomez-Tames, J. and Hirata, A. (2020) Correlation between COVID-19 Morbidity and Mortality Rates in Japan and Local Population Density, Temperature, and Absolute Humidity. International Journal of Environmental Research and Public Health, 17, 5447.  
https://doi.org/10.3390/ijerph17155477</mixed-citation></ref><ref id="scirp.110121-ref24"><label>24</label><mixed-citation publication-type="other" xlink:type="simple">GBD Colorectal Cancer Collaborators (2019) The Global, Regional, and National Burden of Colorectal Cancer and Its Attributable Risk Factors in 195 Countries and Territories, 1990-2017: A Systematic Review for the Global Burden of Disease Study 2017. The Lancet Gastroenterology and Hepatology, 4, 913-933.</mixed-citation></ref><ref id="scirp.110121-ref25"><label>25</label><mixed-citation publication-type="other" xlink:type="simple">Hastie, T.J. and Tibshirani, R.J. (1990) Generalized Additive Models. Chapman &amp; Hall/CRC, London.</mixed-citation></ref><ref id="scirp.110121-ref26"><label>26</label><mixed-citation publication-type="other" xlink:type="simple">Ruppert, D., Wand, M.P. and Carroll, R.J. (2003) Semiparametric Regression. Cambridge University Press, Cambridge.  
https://doi.org/10.1017/CBO9780511755453</mixed-citation></ref><ref id="scirp.110121-ref27"><label>27</label><mixed-citation publication-type="other" xlink:type="simple">Wood, S.N. (2017) Generalized Additive Models: An Introduction with R. Second Edition, CRC Press, Boca Raton. https://doi.org/10.1201/9781315370279</mixed-citation></ref></ref-list></back></article>