<?xml version="1.0" encoding="UTF-8"?><!DOCTYPE article  PUBLIC "-//NLM//DTD Journal Publishing DTD v3.0 20080202//EN" "http://dtd.nlm.nih.gov/publishing/3.0/journalpublishing3.dtd"><article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" dtd-version="3.0" xml:lang="en" article-type="research article"><front><journal-meta><journal-id journal-id-type="publisher-id">JDAIP</journal-id><journal-title-group><journal-title>Journal of Data Analysis and Information Processing</journal-title></journal-title-group><issn pub-type="epub">2327-7211</issn><publisher><publisher-name>Scientific Research Publishing</publisher-name></publisher></journal-meta><article-meta><article-id pub-id-type="doi">10.4236/jdaip.2023.111005</article-id><article-id pub-id-type="publisher-id">JDAIP-122887</article-id><article-categories><subj-group subj-group-type="heading"><subject>Articles</subject></subj-group><subj-group subj-group-type="Discipline-v2"><subject>Computer Science&amp;Communications</subject><subject> Physics&amp;Mathematics</subject></subj-group></article-categories><title-group><article-title>
 
 
  Analysis of College Students’ Test Scores Based on Two-Component Mixed Generalized Normal Distribution
 
</article-title></title-group><contrib-group><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Luliang</surname><given-names>Wen</given-names></name><xref ref-type="aff" rid="aff1"><sup>1</sup></xref><xref ref-type="corresp" rid="cor1"><sup>*</sup></xref></contrib><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Haiwu</surname><given-names>Rong</given-names></name><xref ref-type="aff" rid="aff1"><sup>1</sup></xref><xref ref-type="corresp" rid="cor1"><sup>*</sup></xref></contrib><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Yanjun</surname><given-names>Qiu</given-names></name><xref ref-type="aff" rid="aff2"><sup>2</sup></xref></contrib></contrib-group><aff id="aff1"><addr-line>Foshan University, Foshan, China</addr-line></aff><aff id="aff2"><addr-line>Jinan University, Guangzhou, China</addr-line></aff><pub-date pub-type="epub"><day>18</day><month>01</month><year>2023</year></pub-date><volume>11</volume><issue>01</issue><fpage>69</fpage><lpage>80</lpage><history><date date-type="received"><day>19,</day>	<month>December</month>	<year>2022</year></date><date date-type="rev-recd"><day>5,</day>	<month>February</month>	<year>2023</year>	</date><date date-type="accepted"><day>8,</day>	<month>February</month>	<year>2023</year></date></history><permissions><copyright-statement>&#169; Copyright  2014 by authors and Scientific Research Publishing Inc. </copyright-statement><copyright-year>2014</copyright-year><license><license-p>This work is licensed under the Creative Commons Attribution International License (CC BY). http://creativecommons.org/licenses/by/4.0/</license-p></license></permissions><abstract><p>
 
 
  In order to improve the fitting accuracy of college students’ test scores, this paper proposes two-component mixed generalized normal distribution, uses maximum likelihood estimation method and Expectation Conditional Maxinnization (ECM) algorithm to estimate parameters and conduct numerical simulation, and performs fitting analysis on the test scores of Linear Algebra and Advanced Mathematics of F University. The empirical results show that the two-component mixed generalized normal distribution is better than the commonly used two-component mixed normal distribution in fitting college students’ test data, and has good application value.
 
</p></abstract><kwd-group><kwd>Two-Component Mixed Generalized Normal Distribution</kwd><kwd> Two-Component Mixed Normal Distribution</kwd><kwd> ECM Algorithm</kwd><kwd> Test Scores</kwd></kwd-group></article-meta></front><body><sec id="s1"><title>1. Introduction</title><p>With regard to the distribution of test scores, the traditional view is to use normal distribution for statistical analysis and inference. However, in reality, many test scores do not conform to the assumption of normal distribution. Li and Zhang (2021) [<xref ref-type="bibr" rid="scirp.122887-ref1">1</xref>] through theoretical reasoning and analysis of test data, suggest reducing or eliminating the requirements for normal distribution of scores in college course tests. In order to find a more accurate distribution to describe college students’ test scores, many scholars have carried out extensive research. Yin (2007) [<xref ref-type="bibr" rid="scirp.122887-ref2">2</xref>], Gu and Chi (2010) [<xref ref-type="bibr" rid="scirp.122887-ref3">3</xref>], Zhang and Ma (2021) [<xref ref-type="bibr" rid="scirp.122887-ref4">4</xref>] used the two-component mixed normal distribution to fit the distribution of college students’ test scores. Through numerical simulation and empirical analysis, it is more accurate and reasonable to use the two-component mixed normal distribution to fit the test scores than the normal distribution.</p><p>In this paper, the test scores of linear algebra (2417 samples) and advanced mathematics (2035 samples) of the students of relevant majors in F University in the second semester of 2019-2020 academic year are plotted as a distribution histogram. As shown in <xref ref-type="fig" rid="fig1">Figure 1</xref>, through observation, it can be found that the two score distributions show double peaks, and the two peaks are respectively located in the score interval [40, 50] and [70, 80], which indicates that the relevant literature has good applicability to select the two component mixed normal distribution for fitting. Wen et al. (2022) [<xref ref-type="bibr" rid="scirp.122887-ref5">5</xref>] proposed the mixed generalized normal distribution, and its degenerate distribution includes the mixed normal distribution. Therefore, this paper first introduces the two-component mixed generalized normal distribution into the analysis of college students’ test scores, compares the fitting effects of the plan and two-component mixed normal distribution, and tries to find a better fitting distribution of college students’ test scores than the two-component mixed normal distribution.</p><p>In terms of content arrangement, Section 2 gives the definition of two-component mixed generalized normal distribution; In Section 3, ECM algorithm is proposed to estimate and simulate the parameters of two-component mixed generalized normal distribution and two-component mixed normal distribution; Section 4 compares and analyzes the fitting effects of the two-component mixed generalized normal distribution and the two-component mixed normal distribution by using the higher mathematics and linear algebra test scores; Section 5 is the conclusion.</p></sec><sec id="s2"><title>2. Two-Component Mixed Generalized Normal Distribution</title><p>If variable X is subject to two-component generalized normal distribution, its probability density function is:</p><p>f ( x | λ , μ 1 , σ 1 , s 1 , μ 2 , σ 2 , s 2 ) = λ ( s 1 2 σ 1 Γ ( 1 / s 1 ) ) exp { − | x − μ 1 σ 1 | s 1 } + ( 1 − λ ) ( s 2 2 σ 2 Γ ( 1 / s 2 ) ) exp { − | x − μ 2 σ 2 | s 2 } . (1)</p><p>where θ = ( λ , μ 1 , σ 1 , s 1 , μ 2 , σ 2 , s 2 ) , Γ ( 1 / s 1 ) = ∫ 0 ∞ t 1 / s 1 − 1 e − t d t , Γ ( 1 / s 2 ) = ∫ 0 ∞ t 1 / s 2 − 1 e − t d t , 0   &lt; λ &lt; 1 , s 1 &gt; 0 , s 2 &gt; 0 , σ 1 &gt; 0 , σ 2 &gt; 0 , − ∞ &lt; μ 1 &lt; ∞ , − ∞ &lt; μ 2 &lt; ∞ , − ∞ &lt; x &lt; ∞ . μ 1 , μ 2 is called location parameter, σ 1 , σ 2 is called scale parameter, and s 1 , s 2 is called shape parameter. When s 1 = s 2 = 2 , it is a mixed normal distribution.</p><p>The expectation and variance of two-component mixed generalized normal distribution are:</p><p>E ( X ) = λ μ 1 + ( 1 − λ ) μ 2 , (2)</p><p>V a r ( X ) = λ σ 1 2 Γ ( 3 / s 1 ) Γ ( 1 / s 1 ) + ( 1 − λ ) σ 2 2 Γ ( 3 / s 2 ) Γ ( 1 / s 2 ) + λ ( 1 − λ ) ( μ 1 − μ 2 ) 2 . (3)</p><p>Given the value of the parameter, the probability density function of the two-component mixed generalized normal distribution and two-component mixed normal distribution can be drawn. It is found from <xref ref-type="fig" rid="fig2">Figure 2</xref> that it is a bimodal asymmetric graph, in which the thick tail of the control distribution is smaller, and the tail is thicker. By comparing the shapes in <xref ref-type="fig" rid="fig1">Figure 1</xref> and <xref ref-type="fig" rid="fig2">Figure 2</xref>, it can be preliminarily judged that it is feasible to use the two-component mixed generalized normal distribution and the two-component mixed normal distribution to fit college students’ test scores.</p></sec><sec id="s3"><title>3. Parameter Estimation</title><p>Expectation Maximization (EM) algorithm is an effective method to solve mixed distribution parameter estimation. Each iteration is divided into two steps: E-step and M-step (Dempster et al., 1977) [<xref ref-type="bibr" rid="scirp.122887-ref6">6</xref>].</p><p>E-step: According to the observed data and the estimated initial values of the current parameters, first calculate the log likelihood function log f ( θ | X , Z ) of the complete data, and then calculate the conditional expectation about the potential data Z:</p><p>Q ( θ | θ ( t ) , X ) = E Z [ log f ( θ | X , Z ) | | θ ( t ) , X ] = ∫ log f ( θ | X , Z ) f ( Z | θ ( t ) , X ) d Z</p><p>M-step: Maximize Q ( θ | θ ( t ) , X ) , solve θ ( t + 1 ) , make Q ( θ ( t + 1 ) | θ ( t ) , X ) = max θ Q ( θ | θ ( t ) , X ) , an iteration is completed θ ( t ) → θ ( t + 1 ) , and repeated until | Q ( θ ( t + 1 ) | θ ( t ) , X ) − Q ( θ ( t ) | θ ( t ) , X ) | is sufficiently small. This is the basic principle of EM algorithm.</p><p>Meng and Rubin (1993) [<xref ref-type="bibr" rid="scirp.122887-ref7">7</xref>] proposed a special EM algorithm called ECM or GEM algorithm. It decomposes the M-step in the original EM algorithm into the next k<sup>th</sup> conditional maximization: in the i + 1 iteration, remember that θ ( i ) = ( θ 1 ( i ) , θ 2 ( i ) , ⋯ , θ k ( i ) ) , after obtaining Q ( θ | θ ( t ) , Y ) , first, under the condition of θ 1 ( i ) , θ 2 ( i ) , ⋯ , θ k ( i ) keeping unchanged, Q ( θ | θ ( i ) , Y ) seek to maximize θ 1 ( i + 1 ) , and then under the conditions of θ 1 = θ 1 ( i + 1 ) , θ j = θ j ( i ) , j = 3 , ⋅ ⋅ ⋅ , k , Q ( θ | θ ( i ) , Y ) seek to maximize θ 2 ( i + 1 ) . Continue like this. After the k<sup>th</sup> condition maximized, we can get θ ( i + 1 ) and complete an iteration.</p><p>Chen et al.(2015) [<xref ref-type="bibr" rid="scirp.122887-ref8">8</xref>] used iterative Newton Raphson algorithm to solve the parameter estimation problem of generalized linear mixed model (GLMM). In this section, ECM algorithm is mainly used for parameter estimation and numerical simulation of two-component mixed generalized normal distribution.</p><sec id="s3_1"><title>3.1. ECM Algorithm</title><p>If the random sample obeys two-component mixed generalized normal distribution, its logarithmic likelihood function is:</p><p>log L ( θ ) = ∑ j = 1 n log { λ ( s 1 2 σ 1 Γ ( 1 / s 1 ) ) exp { − | x j − μ 1 σ 1 | s 1 }       + ( 1 − λ ) ( s 2 2 σ 2 Γ ( 1 / s 2 ) ) exp { − | x j − μ 2 σ 2 | s 2 } } (4)</p><p>with reference to Wen et al. (2022) [<xref ref-type="bibr" rid="scirp.122887-ref4">4</xref>], the maximum likelihood estimation of the two-component mixed generalized normal distribution ECM algorithm in the case of complete data is given below. Set the sample X 1 , X 2 , ⋯ , X n with the capacity of n from the two-component mixed generalized normal distribution, x 1 , x 2 , ⋯ , x n are the sample observation values:</p><p>f i ( x i | λ , μ 1 , μ 2 , σ 1 , σ 2 , s 1 , s 2 ) = f i ( x i | θ ) = λ f 1 i + ( 1 − λ ) f 2 i (5)</p><p>where f 1 i = f 1 i ( x i | μ 1 , σ 1 , s 1 ) = s 1 2 σ 1 Γ ( 1 / s 1 ) exp { − | x i − μ 1 σ 1 | s 1 } , f 2 i = f 2 i ( x i | μ 2 , σ 2 , s 2 ) = s 2 2 σ 2 Γ ( 1 / s 2 ) exp { − | x i − μ 2 σ 2 | s 2 } .</p><p>The indicator function I i is introduced. Suppose it follows the two-point distribution:</p><p>P ( I i = 1 ) = λ , P ( I i = 0 ) = 1 − λ ,</p><p>since the X i of generalized normal distribution population from f 1 i or f 2 i is unknown, the joint distribution of X i and I i is: g ( x i , I i , θ ) = ( λ f 1 i ) I i [ ( 1 − λ ) f 2 i ] 1 − I i , and in a given case X i , the conditional distribution of I i is:</p><p>P ( I i = 1 | x i , θ ) = α f 1 i f i , P ( I i = 0 | x i , θ ) = ( 1 − α ) f 2 i f i .</p><p>E-Step: seek expectations.</p><p>Q ( θ , θ ( m − 1 ) ) = ∑ i = 1 n z 1 i ( m − 1 ) ( log λ s 1 2 σ 1 Γ ( 1 / s 1 ) − | x i − μ 1 σ 1 | s 1 )       + ∑ i = 1 n z 2 i ( m − 1 ) ( log ( 1 − λ ) s 2 2 σ 2 Γ ( 1 / s 2 ) − | x i − μ 2 σ 2 | s 2 ) . (6)</p><p>where z 1 i m − 1 = λ ( m − 1 ) f 1 i ( m − 1 ) f i ( m − 1 ) , z 2 i m − 1 = ( 1 − λ ( m − 1 ) ) f 2 i ( m − 1 ) f i ( m − 1 ) .</p><p>CM-Step: maximizing conditions.</p><p>Ψ ( 1 υ ) represents the digamma function and Ψ ′ ( 1 υ ) represents the trigamma</p><p>function. The iterative formula for the seven parameters of the two-component mixed generalized normal distribution is derived as follows:</p><p>{ λ ( m ) = ∑ i = 1 n z 1 i ( m − 1 ) ∑ j = 1 2 ∑ i = 1 n z j i ( m − 1 ) = ∑ i = 1 n z 1 i ( m − 1 ) ∑ i = 1 n ( z 1 i ( m − 1 ) + z 2 i ( m − 1 ) ) μ j ( m ) = μ j ( m − 1 ) + ∑ x i ≥ μ j ( m − 1 ) z j i ( m − 1 ) ( x i − μ j ( m − 1 ) ) s j ( m − 1 ) − 1 − ∑ x i &lt; μ j ( m − 1 ) z j i ( m − 1 ) ( μ j ( m − 1 ) − x i ) s j ( m − 1 ) − 1 ( σ j ( m − 1 ) ) υ j ( m − 1 ) − 2 ( s j ( m − 1 ) − 1 ) ∑ i = 1 n z j i ( m − 1 ) | x i − μ j ( m − 1 ) σ j ( m − 1 ) | s j ( m − 1 ) − 2 ,       j = 1 , 2 σ j ( m ) = [ s j ( m − 1 ) ⋅ ∑ i = 1 n z i j ( m − 1 ) | x i − μ j ( m ) | s j ( m − 1 ) ∑ i = 1 n z i j ( m − 1 ) ] 1 s j ( m − 1 ) ,   j = 1 , 2 s j ( m ) = s j ( m − 1 ) − ∑ i = 1 n z i j ( m − 1 ) A 1 − ∑ i = 1 n z i j ( m − 1 ) A 2 ∑ i = 1 n z i j ( m − 1 ) A 3 − ∑ i = 1 n z i j ( m − 1 ) A 4 ,   j = 1 , 2 (7)</p><p>where</p><p>A 1 = 1 s j ( m − 1 ) ( 1 s j ( m − 1 ) Ψ ( 1 s j ( m − 1 ) ) + 1 ) , A 2 = | x i − μ j ( m ) σ j ( m ) | s j ( m - 1 ) log | x i − μ j ( m ) σ j ( m ) | , A 3 = − 1 s j 2 ( m − 1 ) [ 1 + 2 s j ( m − 1 ) Ψ ( 1 s j ( m − 1 ) ) + 1 s j 2 ( m − 1 ) Ψ ′ ( 1 s j ( m − 1 ) ) ] , A 4 = | x i − μ j ( m ) σ j ( m ) | s j ( m − 1 ) ( log | x i − μ j ( m ) σ j ( m ) | ) 2 .</p><p>On the basis of known observation data, numerical iterative method can be used to solve the above equations, because the transcendental equation is involved, the solution process is difficult. Two propositions of consistency and asymptotic normality for maximum likelihood estimation of two-component mixed generalized normal distribution are given below.</p><p>Proposition 1: (Consistency) For the two-component mixed generalized normal distribution, given arbitrarily θ 0 , its maximum likelihood estimates θ ^ are continuous in the interval, and then θ ^ converge to in probability θ 0 , that is θ ^ → p θ 0 .</p><p>It is proved that the parameters θ = ( α , μ 1 , σ 1 , s 1 , μ 2 , σ 2 , s 2 ) of the two-component mixed generalized normal distribution are assumed to be open sets:</p><p>Θ = ( 0 , 1 ) &#215; ( − ∞ , + ∞ ) &#215; ( 0 , + ∞ ) &#215; ( 0 , + ∞ ) &#215; ( − ∞ , + ∞ ) &#215; ( 0 , + ∞ ) &#215; ( 0 , + ∞ )</p><p>Let θ 0 = ( α 0 , μ 01 , σ 01 , s 01 , μ 02 , σ 02 , s 02 ) denote the true parameter value. Assume that for any θ 0 = ( α 0 , μ 01 , σ 01 , s 01 , μ 02 , σ 02 , s 02 ) , there is a compact set θ ⊂ Θ that satisfies:</p><p>1) θ 0 ∈ θ ,</p><p>2) ∀ ξ ≠ ξ 0 , ξ ∈ θ , f ( x i | θ ) ≠ f ( x i | θ 0 ) ,</p><p>3) ∀ ξ ∈ θ , log f ( x i | ξ ) is a continuous function,</p><p>4) E [ sup ξ | f ( x i | ξ ) | ] &lt; ∞ .</p><p>According to the relevant theorem content of Newey and McFadden (1994) [<xref ref-type="bibr" rid="scirp.122887-ref9">9</xref>], it can be proved that the two-component mixed generalized normal distribution satisfies the above four conditions.</p><p>With reference to some bounded conditions given by Redner and Walker (1984) [<xref ref-type="bibr" rid="scirp.122887-ref10">10</xref>], such as the set value range with s 1 , s 2 , it can be proved that the maximum likelihood estimator of the two-component mixed generalized normal distribution satisfies the asymptotic normality.</p><p>Proposition 2: (Asymptotic normality) If υ 1 &gt; 1 , υ 2 &gt; 1 , I ( θ 0 ) is an information matrix, then the maximum likelihood estimate θ ^ of θ 0 satisfies the asymptotic normality, that is</p><p>T ( θ ^ − θ 0 ) → d N ( 0 , I − 1 ( θ 0 ) ) .</p><p>To prove proposition 2, refer to Redner and Walker (1984) [<xref ref-type="bibr" rid="scirp.122887-ref10">10</xref>] and McLachlan and Peel (2004) [<xref ref-type="bibr" rid="scirp.122887-ref11">11</xref>] for the derivation process. It should be noted that the information matrix expression of the two component mixed generalized normal distribution is more complex. In the process of seeking expectations, it is a difficult problem how to effectively simplify an analytic expression.</p><p>I ( θ 0 ) ≡ E [ ∂ ln f ∂ θ i ⋅ ∂ ln f ∂ θ j ] = − E [ ∂ 2 ln f ∂ θ i ∂ θ j ] .</p></sec><sec id="s3_2"><title>3.2. Numerical Simulation</title><p>Before the analysis of real data, numerical simulation experiments are conducted. In this section, we will evaluate the performance of ECM algorithm for parameter estimation of two-component mixed generalized normal distribution. First, the formula of skewness and mean square error is given:</p><p>Bias ( θ ^ ) = | 1 N ∑ j = 1 N θ ^ j − θ | ,     MSE ( θ ^ ) = 1 N ∑ j = 1 N ( θ ^ j − θ ) 2 (8)</p><p>where θ represents the real parameter value and θ ^ j represents the θ estimated value for the j time. Skewness and mean square error are measures that reflect the difference between the estimator and the estimated. The smaller the skewness and mean square error, the better the estimation effect.</p><p>Based on the iterative formula given in Section 3.1, and referring to the rounding sampling method proposed by Tadikamalla (1980) [<xref ref-type="bibr" rid="scirp.122887-ref12">12</xref>] and Chiodi (1995) [<xref ref-type="bibr" rid="scirp.122887-ref13">13</xref>], random numbers of two-component mixed generalized normal distribution are generated. Since the sample size of higher mathematics and linear algebra examination scores selected in this paper is 2035 and 2417 respectively, in order to better simulate real data, we choose to generate 2000 and 2500 random numbers for each simulation.</p><p>For λ , we select the initial value λ ∈ ( 0 , 1 ) ; for s 1 , s 2 , we select the initial values by s 1 ∈ [ 1 , 4 ] , s 2 ∈ [ 1 , 4 ] ; for μ 1 , μ 2 and σ 1 , σ 2 , we use the initial values by μ 1 ∈ [ 1 , 4 ] , μ 2 ∈ [ 1 , 4 ] , σ 1 ∈ [ 1 , 3 ] , σ 2 ∈ [ 1 , 3 ] . The convergence criterion is set as | θ ^ ( m + 1 ) − θ ^ ( m ) | &lt; 10 − 4 , the numerical simulation experiment is carried out for 30 times, and the average value is calculated for analysis. Programming calculation with Python software.</p><p><xref ref-type="table" rid="table1">Table 1</xref> shows that the skewness of the seven parameters is within 0.25 and the mean square error is within 7% when estimating the parameters of the two-component mixed generalized normal distribution; <xref ref-type="table" rid="table2">Table 2</xref> shows that when estimating the five parameters of the two-component mixed normal distribution, the skewness of the parameters is within 0.09 and the mean square error is within 4%. Under the same sample size, it is found that the fewer parameters to be estimated, the better the estimation effect. In general, when the sample size is 2000 and 2500, the parameter estimates of the two distributions have reached convergence. Through the analysis of skewness and mean square error, the parameter estimates are also relatively ideal.</p><table-wrap id="table1" ><label><xref ref-type="table" rid="table1">Table 1</xref></label><caption><title> Parameter estimation results of two-components mixed generalized normal distribution ( e − 0 x = 10 − x )</title></caption><table><tbody><thead><tr><th align="center" valign="middle" >True value</th><th align="center" valign="middle" >λ = 0.65</th><th align="center" valign="middle" >μ 1 = 1.5</th><th align="center" valign="middle" >μ 2 = 3.5</th><th align="center" valign="middle" >σ 1 = 1.2</th><th align="center" valign="middle" >σ 2 = 2.6</th><th align="center" valign="middle" >s 1 = 3.2</th><th align="center" valign="middle" >s 2 = 1.5</th><th align="center" valign="middle" >sample size</th></tr></thead><tr><td align="center" valign="middle" >Est.</td><td align="center" valign="middle" >0.6418</td><td align="center" valign="middle" >1.4948</td><td align="center" valign="middle" >3.4443</td><td align="center" valign="middle" >1.2394</td><td align="center" valign="middle" >2.8493</td><td align="center" valign="middle" >3.0563</td><td align="center" valign="middle" >1.5973</td><td align="center" valign="middle"  rowspan="3"  >2000</td></tr><tr><td align="center" valign="middle" >Bias</td><td align="center" valign="middle" >0.0082</td><td align="center" valign="middle" >0.0052</td><td align="center" valign="middle" >0.0557</td><td align="center" valign="middle" >0.0394</td><td align="center" valign="middle" >0.2493</td><td align="center" valign="middle" >0.1437</td><td align="center" valign="middle" >0.0973</td></tr><tr><td align="center" valign="middle" >MSE</td><td align="center" valign="middle" >6.7e−05</td><td align="center" valign="middle" >2.7e−05</td><td align="center" valign="middle" >0.0031</td><td align="center" valign="middle" >0.0016</td><td align="center" valign="middle" >0.0621</td><td align="center" valign="middle" >0.0207</td><td align="center" valign="middle" >0.0095</td></tr><tr><td align="center" valign="middle" >Est.</td><td align="center" valign="middle" >0.6513</td><td align="center" valign="middle" >1.5053</td><td align="center" valign="middle" >3.4871</td><td align="center" valign="middle" >1.2714</td><td align="center" valign="middle" >2.8015</td><td align="center" valign="middle" >3.1837</td><td align="center" valign="middle" >1.5982</td><td align="center" valign="middle"  rowspan="3"  >2500</td></tr><tr><td align="center" valign="middle" >Bias</td><td align="center" valign="middle" >0.0013</td><td align="center" valign="middle" >0.0053</td><td align="center" valign="middle" >0.0129</td><td align="center" valign="middle" >0.0714</td><td align="center" valign="middle" >0.2015</td><td align="center" valign="middle" >0.0163</td><td align="center" valign="middle" >0.0982</td></tr><tr><td align="center" valign="middle" >MSE</td><td align="center" valign="middle" >1.6e−05</td><td align="center" valign="middle" >2.8e−05</td><td align="center" valign="middle" >0.0002</td><td align="center" valign="middle" >0.0051</td><td align="center" valign="middle" >0.0406</td><td align="center" valign="middle" >0.0003</td><td align="center" valign="middle" >0.0096</td></tr></tbody></table></table-wrap><table-wrap id="table2" ><label><xref ref-type="table" rid="table2">Table 2</xref></label><caption><title> Parameter estimation results of two-components mixed normal distribution ( s 1 = s 2 = 2 )</title></caption><table><tbody><thead><tr><th align="center" valign="middle" >True value</th><th align="center" valign="middle" >λ = 0.65</th><th align="center" valign="middle" >μ 1 = 1.5</th><th align="center" valign="middle" >μ 2 = 3.5</th><th align="center" valign="middle" >σ 1 = 1.2</th><th align="center" valign="middle" >σ 2 = 2.6</th><th align="center" valign="middle" >s 1 = 2</th><th align="center" valign="middle" >s 2 = 2</th><th align="center" valign="middle" >sample size</th></tr></thead><tr><td align="center" valign="middle" >Est.</td><td align="center" valign="middle" >0.6804</td><td align="center" valign="middle" >1.5046</td><td align="center" valign="middle" >3.6270</td><td align="center" valign="middle" >1.2811</td><td align="center" valign="middle" >2.5637</td><td align="center" valign="middle" >2.0</td><td align="center" valign="middle" >2.0</td><td align="center" valign="middle"  rowspan="3"  >2000</td></tr><tr><td align="center" valign="middle" >Bias</td><td align="center" valign="middle" >0.0304</td><td align="center" valign="middle" >0.0046</td><td align="center" valign="middle" >0.1270</td><td align="center" valign="middle" >0.0811</td><td align="center" valign="middle" >0.0363</td><td align="center" valign="middle" >0.0</td><td align="center" valign="middle" >0.0</td></tr><tr><td align="center" valign="middle" >MSE</td><td align="center" valign="middle" >0.0009</td><td align="center" valign="middle" >2.2e−05</td><td align="center" valign="middle" >0.0161</td><td align="center" valign="middle" >0.0066</td><td align="center" valign="middle" >0.0013</td><td align="center" valign="middle" >0.0</td><td align="center" valign="middle" >0.0</td></tr><tr><td align="center" valign="middle" >Est.</td><td align="center" valign="middle" >0.6922</td><td align="center" valign="middle" >1.5108</td><td align="center" valign="middle" >3.703</td><td align="center" valign="middle" >1.3008</td><td align="center" valign="middle" >2.5080</td><td align="center" valign="middle" >2.0</td><td align="center" valign="middle" >2.0</td><td align="center" valign="middle"  rowspan="3"  >2500</td></tr><tr><td align="center" valign="middle" >Bias</td><td align="center" valign="middle" >0.0422</td><td align="center" valign="middle" >0.0108</td><td align="center" valign="middle" >0.2031</td><td align="center" valign="middle" >0.1007</td><td align="center" valign="middle" >0.0920</td><td align="center" valign="middle" >0.0</td><td align="center" valign="middle" >0.0</td></tr><tr><td align="center" valign="middle" >MSE</td><td align="center" valign="middle" >0.0018</td><td align="center" valign="middle" >0.0001</td><td align="center" valign="middle" >0.0413</td><td align="center" valign="middle" >0.0102</td><td align="center" valign="middle" >0.0085</td><td align="center" valign="middle" >0.0</td><td align="center" valign="middle" >0.0</td></tr></tbody></table></table-wrap></sec></sec><sec id="s4"><title>4. Real Data Analysis</title><sec id="s4_1"><title>4.1. Descriptive Statistics</title><p>We choose the linear algebra (2417 samples) and advanced mathematics (2035 samples) test scores of students of relevant majors in F University in the second semester of 2019-2020 academic year, and the column chart of the distribution is shown in <xref ref-type="fig" rid="fig1">Figure 1</xref>. In order to facilitate data analysis, the scores of linear algebra test (abbreviated as XXDS) and advanced mathematics test (abbreviated as GDSX) are normalized. To generate new data for descriptive statistics, the scores</p><p>of each examinee are set as x i , y i = x i 100 ∈ [ 0 , 1 ] .</p><p>It can be seen from <xref ref-type="table" rid="table3">Table 3</xref> that linear algebra (XXDS) and advanced mathematics (GDSX) have common features, such as a small difference between their mean values; The skewness coefficients are all less than 0, showing the characteristics of left bias, and the kurtosis are all less than 3; At the 5% significance level, the results of J-B statistics are far greater than 5.99, and the assumption of normal distribution is rejected.</p></sec><sec id="s4_2"><title>4.2. Fitting Evaluation</title><p>AIC and BIC criteria are generally used to evaluate the model fitting effect:</p><p>{ AIC = 2 φ − 2 log L ( θ ^ | x ) , BIC = φ log ( n ) − 2 log L ( θ ^ | x ) , (9)</p><table-wrap id="table3" ><label><xref ref-type="table" rid="table3">Table 3</xref></label><caption><title> Descriptive statistics of linear algebra (XXDS) and advanced mathematics (GDSX)</title></caption><table><tbody><thead><tr><th align="center" valign="middle" ></th><th align="center" valign="middle" >sample size</th><th align="center" valign="middle" >Mean</th><th align="center" valign="middle" >Std.</th><th align="center" valign="middle" >Skewness</th><th align="center" valign="middle" >Kurtosis</th><th align="center" valign="middle" >J-B value</th><th align="center" valign="middle" >P-value</th></tr></thead><tr><td align="center" valign="middle" >XXDS</td><td align="center" valign="middle" >2417</td><td align="center" valign="middle" >0.685</td><td align="center" valign="middle" >0.175</td><td align="center" valign="middle" >−0.684</td><td align="center" valign="middle" >0.087</td><td align="center" valign="middle" >189.483</td><td align="center" valign="middle" >0</td></tr><tr><td align="center" valign="middle" >GDSX</td><td align="center" valign="middle" >2035</td><td align="center" valign="middle" >0.610</td><td align="center" valign="middle" >0.205</td><td align="center" valign="middle" >−0.682</td><td align="center" valign="middle" >0.179</td><td align="center" valign="middle" >160.465</td><td align="center" valign="middle" >0</td></tr></tbody></table></table-wrap><table-wrap id="table4" ><label><xref ref-type="table" rid="table4">Table 4</xref></label><caption><title> Estimates of two-component mixed generalized normal distribution (MGND) and two-component mixed normal distribution (MND) parameters using the ECM algorithm for XXDS and GDSX test score data</title></caption><table><tbody><thead><tr><th align="center" valign="middle"  rowspan="2"  >Parameter</th><th align="center" valign="middle"  colspan="2"  >MND</th><th align="center" valign="middle"  colspan="2"  >MGND</th><th align="center" valign="middle"  rowspan="2"  >Data Sources</th></tr></thead><tr><td align="center" valign="middle" >Est.</td><td align="center" valign="middle" >S.E.</td><td align="center" valign="middle" >Est.</td><td align="center" valign="middle" >S.E.</td></tr><tr><td align="center" valign="middle"  rowspan="2"  >λ</td><td align="center" valign="middle" >0.2674</td><td align="center" valign="middle" >3.4e−05</td><td align="center" valign="middle" >0.5166</td><td align="center" valign="middle" >0.0213</td><td align="center" valign="middle" >XXDS</td></tr><tr><td align="center" valign="middle" >0.4030</td><td align="center" valign="middle" >3.7e−06</td><td align="center" valign="middle" >0.4918</td><td align="center" valign="middle" >0.0533</td><td align="center" valign="middle" >GDSX</td></tr><tr><td align="center" valign="middle"  rowspan="2"  >μ 1</td><td align="center" valign="middle" >0.4493</td><td align="center" valign="middle" >0.0008</td><td align="center" valign="middle" >0.5195</td><td align="center" valign="middle" >0.0057</td><td align="center" valign="middle" >XXDS</td></tr><tr><td align="center" valign="middle" >0.4088</td><td align="center" valign="middle" >2.4e−06</td><td align="center" valign="middle" >0.4593</td><td align="center" valign="middle" >0.0299</td><td align="center" valign="middle" >GDSX</td></tr><tr><td align="center" valign="middle"  rowspan="2"  >μ 2</td><td align="center" valign="middle" >0.7710</td><td align="center" valign="middle" >3.7e−06</td><td align="center" valign="middle" >0.7925</td><td align="center" valign="middle" >0.0005</td><td align="center" valign="middle" >XXDS</td></tr><tr><td align="center" valign="middle" >0.7464</td><td align="center" valign="middle" >4e−07</td><td align="center" valign="middle" >0.7541</td><td align="center" valign="middle" >0.0056</td><td align="center" valign="middle" >GDSX</td></tr><tr><td align="center" valign="middle"  rowspan="2"  >σ 1</td><td align="center" valign="middle" >0.1573</td><td align="center" valign="middle" >2.2e−05</td><td align="center" valign="middle" >0.1626</td><td align="center" valign="middle" >0.0030</td><td align="center" valign="middle" >XXDS</td></tr><tr><td align="center" valign="middle" >0.2210</td><td align="center" valign="middle" >2.1e−06</td><td align="center" valign="middle" >0.2448</td><td align="center" valign="middle" >0.0503</td><td align="center" valign="middle" >GDSX</td></tr><tr><td align="center" valign="middle"  rowspan="2"  >σ 2</td><td align="center" valign="middle" >0.1380</td><td align="center" valign="middle" >2.6e−06</td><td align="center" valign="middle" >0.1617</td><td align="center" valign="middle" >0.0006</td><td align="center" valign="middle" >XXDS</td></tr><tr><td align="center" valign="middle" >0.1285</td><td align="center" valign="middle" >2.4e−07</td><td align="center" valign="middle" >0.1442</td><td align="center" valign="middle" >0.0024</td><td align="center" valign="middle" >GDSX</td></tr><tr><td align="center" valign="middle"  rowspan="2"  >s 1</td><td align="center" valign="middle" >2.0</td><td align="center" valign="middle" >--</td><td align="center" valign="middle" >1.2350</td><td align="center" valign="middle" >0.0069</td><td align="center" valign="middle" >XXDS</td></tr><tr><td align="center" valign="middle" >2.0</td><td align="center" valign="middle" >--</td><td align="center" valign="middle" >2.0822</td><td align="center" valign="middle" >0.1034</td><td align="center" valign="middle" >GDSX</td></tr><tr><td align="center" valign="middle"  rowspan="2"  >s 2</td><td align="center" valign="middle" >2.0</td><td align="center" valign="middle" >--</td><td align="center" valign="middle" >5.7705</td><td align="center" valign="middle" >0.2971</td><td align="center" valign="middle" >XXDS</td></tr><tr><td align="center" valign="middle" >2.0</td><td align="center" valign="middle" >--</td><td align="center" valign="middle" >2.4348</td><td align="center" valign="middle" >0.3449</td><td align="center" valign="middle" >GDSX</td></tr><tr><td align="center" valign="middle"  rowspan="2"  >L ( θ ^ | x )</td><td align="center" valign="middle"  colspan="2"  >−297.2428</td><td align="center" valign="middle"  colspan="2"  >−226.0337</td><td align="center" valign="middle" >XXDS</td></tr><tr><td align="center" valign="middle"  colspan="2"  >−312.0936</td><td align="center" valign="middle"  colspan="2"  >−307.7131</td><td align="center" valign="middle" >GDSX</td></tr><tr><td align="center" valign="middle"  rowspan="2"  >AIC</td><td align="center" valign="middle"  colspan="2"  >604.4856</td><td align="center" valign="middle"  colspan="2"  >466.0673</td><td align="center" valign="middle" >XXDS</td></tr><tr><td align="center" valign="middle"  colspan="2"  >634.1871</td><td align="center" valign="middle"  colspan="2"  >629.4261</td><td align="center" valign="middle" >GDSX</td></tr><tr><td align="center" valign="middle"  rowspan="2"  >BIC</td><td align="center" valign="middle"  colspan="2"  >645.0176</td><td align="center" valign="middle"  colspan="2"  >506.5993</td><td align="center" valign="middle" >XXDS</td></tr><tr><td align="center" valign="middle"  colspan="2"  >673.5149</td><td align="center" valign="middle"  colspan="2"  >668.7539</td><td align="center" valign="middle" >GDSX</td></tr></tbody></table></table-wrap><p>where φ represents the number of parameters and the sample size. The smaller the AIC and BIC values, the better the fitting effect. To achieve more accurate results, the parameter estimation is repeated for 30 times to get the final result.</p><p><xref ref-type="table" rid="table4">Table 4</xref> shows the parameter estimation and AIC and BIC calculation results of the two-component mixed normal distribution and two-component generalized normal distribution using ECM algorithm. By analyzing the calculation results of AIC and BIC, it can be seen that the fitting effect of the two-component mixed generalized normal distribution is obviously better than the two-component mixed normal distribution for the test results of linear algebra (XXDS), and for the test results of advanced mathematics (GDSX), the fitting effect of the two-component mixed generalized normal distribution is slightly better than the two-component mixed normal distribution, and there is little difference between the two distributions.</p><p>Through the analysis of the parameter estimation s 1 , s 2 results, we can see that for the linear algebra test scores, the estimation s 1 results are significantly less than 2, and the s 2 results are significantly greater than 2. For the estimation of the higher mathematics scores, we find that the s 1 , s 2 results are all near 2, which also explains the reason why the AIC and BIC results of the two-component mixed generalized normal distribution and the two-component mixed normal distribution are very close.</p><p><xref ref-type="fig" rid="fig3">Figure 3</xref> is the histogram of the two-component mixed normal distribution and the two-component generalized normal distribution. It can be found that the two-component generalized mixed normal distribution better fits the peak on the left, and the two-component mixed normal distribution better fits the peak on the right. In general, the fitting effect of the two-component generalized mixed normal distribution is better than the two-component mixed normal distribution, which is consistent with the conclusion drawn from the analysis of AIC and BIC calculation results.</p></sec></sec><sec id="s5"><title>5. Conclusion</title><p>In this paper, a two-component mixed generalized normal distribution is proposed, and the results of linear algebra and higher mathematics examinations are fitted and analyzed. The following conclusions are drawn: 1) When studying the maximum likelihood estimation of two-component mixed generalized normal distribution, ECM algorithm is proposed to estimate parameters, which is verified to be an effective method by numerical simulation. 2) Through the comparative analysis of the fitting effect of two-component mixed generalized normal distribution and two-component mixed normal distribution on college students’ math test scores, the empirical results show that the overall fitting effect of two-component mixed generalized normal distribution is better than that of two-component mixed normal distribution, especially in characterizing low or high scoring groups, it avoids the problem of too much or too little tail fitting of two-component mixed normal distribution, It is the innovation of research methods of Zhang and Ma (2021) [<xref ref-type="bibr" rid="scirp.122887-ref4">4</xref>] and other scholars. 3) For the “bimodal distribution” of college students’ test scores, there are mainly two types of students, one is the students who fail the test, and the other is the students who pass the test. Huang et al. (2019) [<xref ref-type="bibr" rid="scirp.122887-ref14">14</xref>] gave a statistical analysis of the influencing factors for students who failed in the exam. The two-component mixed generalized normal distribution proposed in this paper has a good application value for accurately analyzing the test scores of different types of college students and optimizing the teaching methods of different types of students.</p></sec><sec id="s6"><title>Acknowledgements</title><p>We thank the reviewers for their valuable suggestions, which helped us to improve the manuscript. This research is supported by the 13th Five-Year Plan of Philosophy and Social Sciences of Guangdong Province (No. GD20XYJ19).</p></sec><sec id="s7"><title>Conflicts of Interest</title><p>The authors declare no conflicts of interest regarding the publication of this paper.</p></sec><sec id="s8"><title>Cite this paper</title><p>Wen, L.L., Rong, H.W. and Qiu, Y.J. (2023) Analysis of College Students’ Test Scores Based on Two-Component Mixed Generalized Normal Distribution. Journal of Data Analysis and Information Processing, 11, 69-80. https://doi.org/10.4236/jdaip.2023.111005</p></sec></body><back><ref-list><title>References</title><ref id="scirp.122887-ref1"><label>1</label><mixed-citation publication-type="other" xlink:type="simple">Li, L. and Zhang, W.H. (2021) Analysis on Normal Distribution of College Course Examination Results. China Exam, 4, 86-93.</mixed-citation></ref><ref id="scirp.122887-ref2"><label>2</label><mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Yin</surname><given-names> X.F. </given-names></name>,<etal>et al</etal>. (<year>2007</year>)<article-title>Fitting the Distribution of College Students’ Test Scores Based on Mixed Normal Distribution</article-title><source> Statistics and Decision Making</source><volume> 8</volume>,<fpage> 133</fpage>-<lpage>135</lpage>.<pub-id pub-id-type="doi"></pub-id></mixed-citation></ref><ref id="scirp.122887-ref3"><label>3</label><mixed-citation publication-type="other" xlink:type="simple">Gu, C.C. and Chi, Z.Y. (2010) Research on the Distribution Law of Students’ Achievements. Journal of Anyang Institute of Technology, 9, 88-90.</mixed-citation></ref><ref id="scirp.122887-ref4"><label>4</label><mixed-citation publication-type="other" xlink:type="simple">Zhang, J.J. and Ma, D.J. (2021) Analysis of Mixed Normal Distribution of Test Results. Mathematical Statistics and Management, 40, 815-821.</mixed-citation></ref><ref id="scirp.122887-ref5"><label>5</label><mixed-citation publication-type="other" xlink:type="simple">Wen, L.L., Qiu, Y.J., Wang, M.H., Yin, J.L. and Chen, P.Y. (2022) Numerical Characteristics and Parameter Estimation of Finite Mixed Generalized Normal Distribution. Communications in Statistics—Simulation and Computation, 51, 3596-3620. https://doi.org/10.1080/03610918.2020.1720733</mixed-citation></ref><ref id="scirp.122887-ref6"><label>6</label><mixed-citation publication-type="other" xlink:type="simple">Dempster, A.P., Laird, N.M. and Rubin, D.B. (1977) Maximum Likelihood from Incomplete Data via the EM Algorithm. Journal of the Royal Statistical Society: Series B (Methodological), 39, 1-22. https://doi.org/10.1111/j.2517-6161.1977.tb01600.x</mixed-citation></ref><ref id="scirp.122887-ref7"><label>7</label><mixed-citation publication-type="other" xlink:type="simple">Meng, X.L. and Rubin, D.B. (1993) Maximum Likelihood Estimation via the ECM Algorithm: A General Framework. Biometrika, 80, 267-278. https://doi.org/10.1093/biomet/80.2.267</mixed-citation></ref><ref id="scirp.122887-ref8"><label>8</label><mixed-citation publication-type="other" xlink:type="simple">Chen, Y., Fei, Y. and Pan, J.X. (2015) Statistical Inference in Generalized Linear Mixed Models by Joint Modelling Mean and Covariance of Non-Normal Random Effects. Open Journal of Statistics, 5, 568-584. https://doi.org/10.4236/ojs.2015.56059</mixed-citation></ref><ref id="scirp.122887-ref9"><label>9</label><mixed-citation publication-type="other" xlink:type="simple">Newey, W.K. and McFadden, D. (1994) Large Sample Estimation and Hypothesis Testing. Handbook of Econometrics, 4, 2111-2245. https://doi.org/10.1016/S1573-4412(05)80005-4</mixed-citation></ref><ref id="scirp.122887-ref10"><label>10</label><mixed-citation publication-type="other" xlink:type="simple">Redner, R.A. and Walker, H.F. (1984) Mixture Densities, Maximum Likelihood and the EM Algorithm. SIAM Review, 26, 195-239. https://doi.org/10.1137/1026034</mixed-citation></ref><ref id="scirp.122887-ref11"><label>11</label><mixed-citation publication-type="other" xlink:type="simple">McLachlan, G. J. and Peel, D. (2004) Finite Mixture Models. John Wiley &amp; Sons, New York.</mixed-citation></ref><ref id="scirp.122887-ref12"><label>12</label><mixed-citation publication-type="other" xlink:type="simple">Tadikamalla, P. (1980) Random Sampling from the Exponential Power Distribution. Journal of the American Statistical Association, 75, 683-686. https://doi.org/10.1080/01621459.1980.10477533</mixed-citation></ref><ref id="scirp.122887-ref13"><label>13</label><mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Chiodi</surname><given-names> M. </given-names></name>,<etal>et al</etal>. (<year>1995</year>)<article-title>Generation of Pseudo-Random Variates from a Normal Distribution of Order P</article-title><source> Italian Journal of Applied Statistics</source><volume> 7</volume>,<fpage> 401</fpage>-<lpage>416</lpage>.<pub-id pub-id-type="doi"></pub-id></mixed-citation></ref><ref id="scirp.122887-ref14"><label>14</label><mixed-citation publication-type="other" xlink:type="simple">Huang, G.J., Ou, S.D. and Li, Q. (2019) Estimation of Failure Rate of College Mathematics Examination and Statistical Analysis of Its Influencing Factors. Journal of Guangxi University: Natural Science Edition, 44, 1835-1841.</mixed-citation></ref></ref-list></back></article>