<?xml version="1.0" encoding="UTF-8"?><!DOCTYPE article  PUBLIC "-//NLM//DTD Journal Publishing DTD v3.0 20080202//EN" "http://dtd.nlm.nih.gov/publishing/3.0/journalpublishing3.dtd"><article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" dtd-version="3.0" xml:lang="en" article-type="research article"><front><journal-meta><journal-id journal-id-type="publisher-id">OJS</journal-id><journal-title-group><journal-title>Open Journal of Statistics</journal-title></journal-title-group><issn pub-type="epub">2161-718X</issn><publisher><publisher-name>Scientific Research Publishing</publisher-name></publisher></journal-meta><article-meta><article-id pub-id-type="doi">10.4236/ojs.2019.94031</article-id><article-id pub-id-type="publisher-id">OJS-94295</article-id><article-categories><subj-group subj-group-type="heading"><subject>Articles</subject></subj-group><subj-group subj-group-type="Discipline-v2"><subject>Physics&amp;Mathematics</subject></subj-group></article-categories><title-group><article-title>
 
 
  Using Excel to Explore the Effects of Assumption Violations on One-Way Analysis of Variance (ANOVA) Statistical Procedures
 
</article-title></title-group><contrib-group><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>William</surname><given-names>Laverty</given-names></name><xref ref-type="aff" rid="aff1"><sup>1</sup></xref><xref ref-type="corresp" rid="cor1"><sup>*</sup></xref></contrib><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Ivan</surname><given-names>Kelly</given-names></name><xref ref-type="aff" rid="aff2"><sup>2</sup></xref></contrib></contrib-group><aff id="aff1"><addr-line>Department of Mathematics and Statistics, University of Saskatchewan, Saskatoon, Canada</addr-line></aff><aff id="aff2"><addr-line>Professor Emeritus, Department of Educational Psychology &amp;amp; Special Education, University of Saskatchewan, Saskatoon, Canada</addr-line></aff><pub-date pub-type="epub"><day>05</day><month>08</month><year>2019</year></pub-date><volume>09</volume><issue>04</issue><fpage>458</fpage><lpage>469</lpage><history><date date-type="received"><day>25,</day>	<month>June</month>	<year>2019</year></date><date date-type="rev-recd"><day>10,</day>	<month>August</month>	<year>2019</year>	</date><date date-type="accepted"><day>13,</day>	<month>August</month>	<year>2019</year></date></history><permissions><copyright-statement>&#169; Copyright  2014 by authors and Scientific Research Publishing Inc. </copyright-statement><copyright-year>2014</copyright-year><license><license-p>This work is licensed under the Creative Commons Attribution International License (CC BY). http://creativecommons.org/licenses/by/4.0/</license-p></license></permissions><abstract><p>
 
 
  To understand any statistical tool requires not only an understanding of the relevant computational procedures but also an awareness of the assumptions upon which the procedures are based, and the effects of violations of these assumptions.
   
  In our earlier articles (Laverty, Miket, &amp; Kell
  y [1]
  ) and (Laverty &amp; Kelly,
   
  [2] [3]
  ) we used Microsoft Excel to simulate both a Hidden Markov model and heteroskedastic models showing different realizations of these models and the performance of the techniques for identifying the underlying hidden states using simulated data. The advantage of using Excel is that the simulations are regenerated when the spreadsheet is recalculated allowing the user to observe the performance of the statistical technique under different realizations of the data. In this article we will show how to use Excel to generate data from a one-way ANOVA (Analysis of Variance) model and how the statistical methods behave both when the fundamental assumptions of the model hold and when these assumptions are violated. The purpose of this article is to provide tools for individuals to gain an intuitive understanding of these violations using this readily available program.
 
</p></abstract><kwd-group><kwd>Excel</kwd><kwd> One-Way ANOVA</kwd><kwd> Assumption Violations</kwd><kwd> t-Distribution</kwd><kwd> Cauchy Distribution</kwd></kwd-group></article-meta></front><body><sec id="s1"><title>1. Introduction</title><p>An important aspect of any statistical procedure is the assumptions that the procedure is based on. For example, using the t-distribution to calculate a 95% confidence interval for the centre of the population that is being sampled requires that the population being sampled is a normal distribution and that the observations in the sample are independent. If these underlying assumptions do not hold, the desired performance of the statistical procedure may no longer hold true. Sometimes the effect of an invalid assumption on a property of the procedure is minimal, sometimes not so. If the population is non-normal but has a finite mean and variance (such that the Law of Large Numbers and the Central Limit theorem applies), the departure from normality will have little effect on the properties of confidence intervals computed assuming normality when the sample size is adequately large. The reason for this is that it is a consequence of the Central Limit Theorem. The purpose of this paper is to show how to use the program Excel to simulate data for which the statistical technique of one-way Analysis of Variance (ANOVA) is used. The advantage of the using the program Excel is that when you press the recalculate button, under the Formulas menu, the data that is generated at random will be regenerated, statistical calculations will be recalculated and relevant graphs will be redrawn. This allows the user to observe the variation in these procedures for different realizations of the data. See <xref ref-type="fig" rid="fig1">Figure 1</xref>.</p></sec><sec id="s2"><title>2. A Model for Non-Normality (The Cauchy Distribution, the t-Distribution)</title><p>For most cases when one-way ANOVA is applicable the normality assumption is appropriate, i.e. the departures of individual observations from their central value are normally distributed. There are however, many examples where this is not the case and extreme departures are more prevalent than predicted by the Normal distribution. This would be dependent on the measurements being collected. For example, if the measurements were measurements of blood pressure, IQ, performance of a political leader one may expect the presence of extreme measurements. In such cases an appropriate model of the departures from the central value would be the t-distribution (a heavy tailed distribution). In this article the reader can use the technique provided to explore the effects of sampling from heavy tailed distributions on ANOVA calculations that assume normality.</p><p>The probability density function of the standard Normal, Students t-distribution with ν degrees of freedom and the standard Cauchy distribution is given in (1).</p><p>f Normal ( z ) = 1 2 π e − z 2 f t ( t : ν ) = Γ ( ν + 1 2 ) ν π Γ ( ν 2 ) ( 1 + t 2 ν ) − ν + 1 2 f Cauchy ( x : 0 , 1 ) = 1 π ( 1 + x 2 ) } (1)</p><p>The Standard Cauchy distribution is equivalent to the t distribution with 1 degree of freedom. A graph of the standard normal distribution, the t-distribution with 5 degrees of freedom, and the Cauchy distribution is in <xref ref-type="fig" rid="fig2">Figure 2</xref>.</p><p>The Cauchy Distribution is an example of a distribution where the Law of Large numbers and the Central limit Theorem do not apply [<xref ref-type="bibr" rid="scirp.94295-ref4">4</xref>] . In order for these two Laws to hold both the mean and higher moments have to exist and be finite. This is not the case for the Cauchy distribution. There is no convergence of the distribution of the sample mean to the central value. In fact the distribution of the sample mean is the Cauchy distribution for any sample size (i.e. the distribution of the sample mean is the same as that of any individual observation when the data comes from the Cauchy distribution). The Cauchy distribution is a heavy-tailed distribution. The t-distribution is also a heavy-tailed distribution (but not as extreme) when the degrees of freedom ν is small. As the degrees of freedom increases the t distribution approaches the standard normal distribution. Tsay [<xref ref-type="bibr" rid="scirp.94295-ref5">5</xref>] uses the t-distribution with 5 degrees to model random disturbances that appear in various time series models of financial data. This accounts for the sometimes extreme changes that appear in financial data. The Cauchy distribution is appropriate if extreme values are prevalent in the data (the t-distribution with degrees of freedom higher than 1 in the less extreme case). This could occur in surveys where individuals were asked to make a continuous measurement of some quantity and extreme values were prevalent in the populations. For example, measurements of blood pressure, IQ, and performance of a political leader, could result in non-normal data with extreme values at either end. In such cases alternatives to ANOVA are appropriate.1 We haven’t considered these alternatives in this paper.</p><p>The t-distribution with ν degrees of freedom can also be shown to be mixture of Normal distributions with mean 0 and variance W, where the weighting distribution for W is the inverse gamma distribution with α = ν/2 and β = ν/2 (Cook [<xref ref-type="bibr" rid="scirp.94295-ref6">6</xref>] ). This implies that a random variable T will have the t-distribution with ν degrees of freedom if W is selected from the inverse gamma distribution with α = ν/2 and β = ν/2 and then T is selected from Normal distributions with mean 0 and variance W. <sub> </sub></p></sec><sec id="s3"><title>3. Simulation of Data from a Continuous Distribution in Excel</title><p>Uniform random variates on [0,1] can be generated in Excel with the function “RAND()”. The generation of random variates from a continuous distribution with measure of central location μ and measure of scale σ, can be carried out using the inverse-transform method (Fishman [<xref ref-type="bibr" rid="scirp.94295-ref7">7</xref>] ). Namely Y = F<sup>−1(</sup>U) where F(u) is the desired cumulative distribution of Y and U has a uniform distribution on [0,1] (see<xref ref-type="fig" rid="fig3">Figure 3</xref>). In Excel this is achieved for the Normal distribution (mean μ, standard deviation σ) with the function “μ+σ*NORMSINV(RAND())” and for the Cauchy (t with 1 d.f.) location parameter, μ, and scale parameter, σ, “μ+σ*TINV(2*(1-RAND()),1)” (<xref ref-type="fig" rid="fig3">Figure 3</xref>).</p><p>Comment:The Excel function TINV(U,df) does not calculate F<sup>−1</sup>(U) for the t-distn with degrees of freedom df, however the excel function TINV(2*(1-U),df) does achieve the desired calculation.</p></sec><sec id="s4"><title>4. Setting Up the Excel Worksheet to Simulate Anova Data</title><p>The data simulated will come from 3 populations (this can easily be generalized to more than 3 populations). The parameters of the populations</p><p>1) mean(central location), stored in cells C2:E2</p><p>2) standarddeviation (scale parameter), stored in cells C3:E3</p><p>3) samplesize), stored in cells C4:E4</p><p>4) aparameter that determines normality of the data versus non-normality. stored in cells C1:E1. This parameter is set to zero if the desired data is normal. If this parameter is set to an integer, ν, greater than 0 the data will come from a t -distribution with ν degrees of freedom. The t -distribution is a non-normal heavy-tailed, centered and symmetric about zero.</p><p>5) A final parameter (precision), located in cell A2 specifies the of decimal places that the raw data is rounded to (<xref ref-type="table" rid="table1">Table 1</xref> below)</p><table-wrap id="table1" ><label><xref ref-type="table" rid="table1">Table 1</xref></label><caption><title> Excel worksheet</title></caption><table><tbody><thead><tr><th align="center" valign="middle" >A</th><th align="center" valign="middle" >B</th><th align="center" valign="middle" >C</th><th align="center" valign="middle" >D</th><th align="center" valign="middle" >E</th></tr></thead><tr><td align="center" valign="middle" >precision</td><td align="center" valign="middle" >normality</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >0</td></tr><tr><td align="center" valign="middle" >2</td><td align="center" valign="middle" >loc. par.</td><td align="center" valign="middle" >10</td><td align="center" valign="middle" >10</td><td align="center" valign="middle" >15</td></tr><tr><td align="center" valign="middle" ></td><td align="center" valign="middle" >scale par.</td><td align="center" valign="middle" >3</td><td align="center" valign="middle" >3</td><td align="center" valign="middle" >3</td></tr><tr><td align="center" valign="middle" ></td><td align="center" valign="middle" >n</td><td align="center" valign="middle" >10</td><td align="center" valign="middle" >10</td><td align="center" valign="middle" >10</td></tr></tbody></table></table-wrap></sec><sec id="s5"><title>5. Generating Simulated Data</title><p>Copy the observation numbers (1 to 10) in Cells B7:B:16</p><p>Paste in cell C7 the formula =IF($B7&gt;C$4,&quot;&quot;,ROUND(C$2+C$3*IF(C$1=0,NORMSINV(RAND()),TINV(2*(1-RAND()),C$1)),$A$2))”</p><p>Copy this formula to cells C7:E16. If the normality parameter is 0, the data generated will be from the normal distribution with mean = “loc. Par.” And standard deviation = “scale par.”). If the normality parameter is an integer greater than 0, the data will be a random number with a t-distribution scaled by the “scale par.” and location shifted by the “loc. par.” The data will be rounded to the number of decimals specified by “precision”.</p><p>For each population compute T<sub>i</sub> = Sxand Sx<sup>2</sup>. Paste formula “=SUM(C7:C16)” and formula “=SUMSQ(C7:C16)” in cells C18 and C19. Copy these formulae to cells C18:E19.</p></sec><sec id="s6"><title>6. Computation of Statistics Required for One-Way ANOVA</title><p>Suppose we have data from k Normal populations with means μ 1 , μ 2 , μ 3 , ⋯ , μ k and common standard deviation σ.Let { x i j , i = 1 , 2 , ⋯ , k ; j = 1 , 2 , ⋯ , n i } denote data from these populations. Let x<sub>ij</sub> = the j<sup>th</sup> observation from the i<sup>th</sup> population, n<sub>i</sub> = the sample size from the i<sup>th</sup> population.</p><p>Let</p><p>x &#175; i = ∑ j = 1 n i x i j n i and s i = ∑ j = 1 n i ( x i j − x &#175; i ) 2 n i − 1 (2)</p><p>denotethe sample mean and standard deviation from the i<sup>th</sup> population. To compute the sample mean and sample Standard deviation for each population, paste the formulae “=AVERAGE(C7:C16)”and “=STDEV(C7:C16)”in cells C21 and C22. Copy these formulae to cells C21:E22.</p><p>To test the null hypothesis H<sub>0</sub>: μ 1 = μ 2 = ⋯ = μ k against H<sub>A</sub>: μ i ≠ μ j for at least one pair i, j we use the test statistic</p><p>F = ∑ i = 1 k ( x &#175; i − x &#175; . ) 2 / ( k − 1 ) ∑ i = 1 k ∑ j = 1 n i ( x i j − x &#175; i ) 2 / ( N − k ) = SS Between / ( k − 1 ) SS Within / ( N − k ) . (3)</p><p>where</p><p>SS Between = ∑ i = 1 k ( x &#175; i − x &#175; . ) 2 and SS Within = ∑ i = 1 k ∑ j = 1 n i ( x i j − x &#175; i ) 2 (4)</p><p>This statistic has an F-distribution with ν<sub>1</sub> = k – 1 degrees of freedom in the numerator and ν<sub>2</sub> = N – k degrees of freedom in the denominator.</p><p>The computing formulae for</p><p>SS Between = ∑ i T i 2 n i − G 2 N and SS Within = ∑ i ∑ j x i j 2 − ∑ i T i 2 n i (5)</p><p>where</p><p>T i = ∑ i x i j = ∑ ∑ x i j and G = ∑ i T i = ∑ ∑ x i j (6)</p><p>The testing for One-way ANOVA is carried out using the Analysis of Variance table (<xref ref-type="table" rid="table2">Table 2</xref>).</p><p>Place the formula “=SUM(C18:E18)” in cell G18 to compute the grand total, G = ∑ i T i = ∑ ∑ x i j and the formula “=SUM(C19:E19)” in cell G19 to compute ∑ ∑ x i j 2 .</p><p>Place the formula “=C18<sup>2</sup>/C4” in cell C24 and copy to E24 to compute T i 2 n i for each sample. Then place the formula “=SUM(C24:E24)” in cell G24 to compute ∑ i T i 2 n i .</p><p>To compute SS Between = ∑ i T i 2 n i − G 2 N place the formula “=G24-G18<sup>2</sup>/F4” in cell J22 and to compute SS Within = ∑ i ∑ j x i j 2 − ∑ i T i 2 n i place the formula = G19-G24” in J23.</p><table-wrap id="table2" ><label><xref ref-type="table" rid="table2">Table 2</xref></label><caption><title> One-way Anova format</title></caption><table><tbody><thead><tr><th align="center" valign="middle" >Source</th><th align="center" valign="middle" >d.f.</th><th align="center" valign="middle" >Sum of Squares</th><th align="center" valign="middle" >Mean Square</th><th align="center" valign="middle" >F</th><th align="center" valign="middle" >Significance</th></tr></thead><tr><td align="center" valign="middle" >Between</td><td align="center" valign="middle" >k− 1</td><td align="center" valign="middle" >SS<sub>Between</sub></td><td align="center" valign="middle" >MS<sub>Between</sub></td><td align="center" valign="middle" >MS<sub>Between</sub>/MS<sub>Within</sub></td><td align="center" valign="middle" >p-value</td></tr><tr><td align="center" valign="middle" >Within</td><td align="center" valign="middle" >N− k</td><td align="center" valign="middle" >SS<sub>Within</sub></td><td align="center" valign="middle" >MS<sub>Within</sub></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td></tr><tr><td align="center" valign="middle" >Total</td><td align="center" valign="middle" >N− 1</td><td align="center" valign="middle" >SS<sub>Total</sub></td><td align="center" valign="middle" >MS<sub>Total</sub></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td></tr></tbody></table></table-wrap><p>The formulae for degrees of freedom, Mean Square can be placed in the appropriate cells L22:L23 and K22:K23.</p><p>The formula for the F statistic “=L22/L23” can be placed in cell M22. The formula for the p-value of the observed F value “=FDIST(M22, K22,K23)” can be placed in cell N22.</p><p>The formula for a (1−α)100% confidence interval for the mean of the ith sample is:</p><p>x &#175; i &#177; t α / 2 ( d f Error ) MS Error n i (7)</p><p>This formula “=C$21-TINV(0.05,$K$23)*(SQRT($L$23)/$C$4)” can be placed in Cell I28 for the lower limit and in cell I29 “=C$21+TINV(0.05,$K$23)*(SQRT($L$23)/$C$4)” for the upper limit. These formulae can be copied to cells I28:K29 to do the computation for all samples.</p><p>The spreadsheet should now look like <xref ref-type="fig" rid="fig4">Figure 4</xref>.</p><p>To construct Box-whisker plots of the data</p><p>1) Select a range containing the data C6:E16 for 10 observations from each sample from the 3 Populations.</p><p>2) The menu item for Box-plots can be found under the histogram item (<xref ref-type="fig" rid="fig5">Figure 5</xref>).</p><p>Comment: There is a problem with Excel’s method of drawing box-plots. If in the data range there is a blank cell, when drawing a box-plot Excel treats that cell as containing a zero rather than treating the observation as non-existent.</p></sec><sec id="s7"><title>7. Exercises That Can Be Performed to Illustrate the Effects of Assumption Violations on ANOVA</title><p>In these exercises we generate samples using different ANOVA assumptions to examine the violations of these assumptions on the ANOVA calculations.</p><p>1) Equal means, Equal Standard deviations, Equal sample size, Normality:</p><p>μ<sub>1</sub>= 10, σ<sub>1</sub> = 2, n<sub>1</sub> = 10,μ<sub>2</sub>= 10, σ<sub>2</sub> = 2, n<sub>2</sub> = 10, μ<sub>3</sub> = 10, σ<sub>3</sub> = 2, n<sub>3</sub> = 10; normality = 0 (normal distribution)</p><disp-formula id="scirp.94295-formula3"><graphic  xlink:href="//html.scirp.org/file/4-1241238x29.png"  xlink:type="simple"/></disp-formula><disp-formula id="scirp.94295-formula4"><graphic  xlink:href="//html.scirp.org/file/4-1241238x30.png"  xlink:type="simple"/></disp-formula><disp-formula id="scirp.94295-formula5"><graphic  xlink:href="//html.scirp.org/file/4-1241238x31.png"  xlink:type="simple"/></disp-formula><p>Comment: When the population means are all equal and the assumptions are satisfied the p-values come from a uniform distribution from 0 to 1. Thus 5% of the time the p-value will be less than or equal to 0.05 resulting in a type I error.</p><p>2) Unequal means (H<sub>0</sub> false), Equal Standard deviations, Equal sample size, Normality:</p><p>μ<sub>1</sub>= 15, σ<sub>1</sub> = 2, n<sub>1</sub> = 10,μ<sub>2</sub>= 10, σ<sub>2</sub> = 2, n<sub>2</sub> = 10, μ<sub>3</sub> = 5, σ<sub>3</sub> = 2, n<sub>3</sub> = 10; normality = 0 (normal distribution)</p><disp-formula id="scirp.94295-formula6"><graphic  xlink:href="//html.scirp.org/file/4-1241238x32.png"  xlink:type="simple"/></disp-formula><disp-formula id="scirp.94295-formula7"><graphic  xlink:href="//html.scirp.org/file/4-1241238x33.png"  xlink:type="simple"/></disp-formula><disp-formula id="scirp.94295-formula8"><graphic  xlink:href="//html.scirp.org/file/4-1241238x34.png"  xlink:type="simple"/></disp-formula><p>Comment: The ability to detect differences among the means will depend on the non-centralityparameter δ = ∑ i n i ( μ i − μ ) 2 σ 2 where μ = ∑ i n i μ i ∑ i n i .</p><p>(Kirk, [<xref ref-type="bibr" rid="scirp.94295-ref8">8</xref>] ) The larger the value of the non-centrality parameter, δ, the greater the power of the F-test. (i.e. the greater the probability of picking out existent differences.)</p><p>3) Unequal means (H<sub>0</sub> false), Equal Standard deviations, Equal sample size, Normality (low non-centrality parameter):</p><p>μ<sub>1</sub>= 11, σ<sub>1</sub> = 5, n<sub>1</sub> = 10,μ<sub>2</sub>= 10, σ<sub>2</sub> = 5, n<sub>2</sub> = 10, μ<sub>3</sub> = 9, σ<sub>3</sub> = 5, n<sub>3</sub> = 10; normality = 0 (normal distribution)</p><disp-formula id="scirp.94295-formula9"><graphic  xlink:href="//html.scirp.org/file/4-1241238x37.png"  xlink:type="simple"/></disp-formula><disp-formula id="scirp.94295-formula10"><graphic  xlink:href="//html.scirp.org/file/4-1241238x38.png"  xlink:type="simple"/></disp-formula><p>Comment: In this case the non-centrality parameter is smaller than the previous example. The p-value of the F-test is considerably higher resulting in an inability to detect a difference in the means.</p><p>4) Equal means, Unequal Standard deviations, Equal sample size, Normality:</p><p>μ<sub>1</sub>= 10, σ<sub>1</sub> = 2, n<sub>1</sub> = 10,μ<sub>2</sub>= 10, σ<sub>2</sub> = 5, n<sub>2</sub> = 10, μ<sub>3</sub> = 10, σ<sub>3</sub> = 10, n<sub>3</sub> = 10; normality = 0 (normal distribution))</p><disp-formula id="scirp.94295-formula11"><graphic  xlink:href="//html.scirp.org/file/4-1241238x39.png"  xlink:type="simple"/></disp-formula><disp-formula id="scirp.94295-formula12"><graphic  xlink:href="//html.scirp.org/file/4-1241238x40.png"  xlink:type="simple"/></disp-formula><disp-formula id="scirp.94295-formula13"><graphic  xlink:href="//html.scirp.org/file/4-1241238x41.png"  xlink:type="simple"/></disp-formula><p>Comment: The anova F-test is to some extent robust against the violation of the assumption of the homogeneity of variance (Bathke [<xref ref-type="bibr" rid="scirp.94295-ref9">9</xref>] ).</p><p>5) Equal means, Equal Standard deviations, Equal sample size, non-Normality:</p><p>μ<sub>1</sub>= 10, σ<sub>1</sub> = 2, n<sub>1</sub> = 10,μ<sub>2</sub>= 10, σ<sub>2</sub> = 2, n<sub>2</sub> = 10, μ<sub>3</sub> = 10, σ<sub>3</sub> = 2, n<sub>3</sub> = 10; normality = 1 (Cauchy distribution)</p><disp-formula id="scirp.94295-formula14"><graphic  xlink:href="//html.scirp.org/file/4-1241238x42.png"  xlink:type="simple"/></disp-formula><disp-formula id="scirp.94295-formula15"><graphic  xlink:href="//html.scirp.org/file/4-1241238x43.png"  xlink:type="simple"/></disp-formula><disp-formula id="scirp.94295-formula16"><graphic  xlink:href="//html.scirp.org/file/4-1241238x44.png"  xlink:type="simple"/></disp-formula><p>Comment: Recall when the data comes from the Cauchy distribution (t-distribution 1 d.f.) neither the law of large numbers or the Central Limit Theorem are applicable. In fact, the distribution of the sample mean for n observations is the same as a single observation. This is illustrated in this example.</p></sec><sec id="s8"><title>8. Discussion</title><p>In applying any statistical procedure it is important understanding the assumptions on which it is based. It is also important to understand the effects on these procedures of the violations of these assumptions. Sometimes the effects of the violations can be extreme, sometimes minimal. The purpose of this article is to provide tools for individuals to gain an intuitive understanding of these violations using the readily available program Microsoft Excel. The advantage of the using the program Excel is that when you press the recalculate button, under the Formulas menu, the data that is generated at random will be regenerated, statistical calculations will be recalculated and relevant graphs will be redrawn. The statistical procedure that we have chosen to illustrate these tools is one-way ANOVA. This procedure is an important component of introductory statistical courses and textbooks. The tools can be easily extended to other and more advanced univariate procedures.</p></sec><sec id="s9"><title>9. Conclusion</title><p>Excel is a very useful tool for examining the performance of One-Way Anova of variance both when the assumptions hold and more importantly when the assumptions are violated.</p></sec><sec id="s10"><title>Conflicts of Interest</title><p>The authors declare no conflicts of interest regarding the publication of this paper.</p></sec><sec id="s11"><title>Cite this paper</title><p>Laverty, W. and Kelly, I. (2019) Using Excel to Explore the Effects of Assumption Violations on One-Way Analysis of Variance (ANOVA) Statistical Procedures. Open Journal of Statistics, 9, 458-469. https://doi.org/10.4236/ojs.2019.94031</p></sec><sec id="s12"><title>NOTES</title></sec></body><back><ref-list><title>References</title><ref id="scirp.94295-ref1"><label>1</label><mixed-citation publication-type="other" xlink:type="simple">Laverty, W.H., Miket, M.J. and Kelly, I.W. (2002) Simulation of Hidden Markov Models with EXCEL. Journal of Royal Statistical Society: Series D, 51, 31-40.  
https://doi.org/10.1111/1467-9884.00296</mixed-citation></ref><ref id="scirp.94295-ref2"><label>2</label><mixed-citation publication-type="other" xlink:type="simple">Laverty, W.H. and Kelly, I.W. (2018) Using Excel to Simulate and Visualize Conditional Heteroskedastic Models. American Journal of Theoretical and Applied Statistics, 7, 242-246.</mixed-citation></ref><ref id="scirp.94295-ref3"><label>3</label><mixed-citation publication-type="other" xlink:type="simple">Laverty, W.H. and Kelly, I.W. (2019) Using Excel to Visualize State Identification in Hidden Markov Models Using the Forward and Backward Algorithms. Applied Mathematical Sciences, 13, 151-162. https://doi.org/10.12988/ams.2019.812195</mixed-citation></ref><ref id="scirp.94295-ref4"><label>4</label><mixed-citation publication-type="other" xlink:type="simple">Feller, W. (1971) An Introduction to Probability Theory and Its Applications, Volume II. 2nd Edition, John Wiley &amp; Sons Inc., New York.</mixed-citation></ref><ref id="scirp.94295-ref5"><label>5</label><mixed-citation publication-type="other" xlink:type="simple">Tsay, R.S. (2010) Analysis of Financial Time Series. 3rd Edition, Wiley, Hoboken.  
https://doi.org/10.1002/9780470644560</mixed-citation></ref><ref id="scirp.94295-ref6"><label>6</label><mixed-citation publication-type="other" xlink:type="simple">Cook, J.D. (2018) Statistical Odds and Ends Blog.  
https://statisticaloddsandends.wordpress.com/2018/03/03/t-distribution-as-a-mixture-of-normals/</mixed-citation></ref><ref id="scirp.94295-ref7"><label>7</label><mixed-citation publication-type="other" xlink:type="simple">Fishman, G.S. (1995) Monte Carlo, Concepts, Algorithms and Applications. Springer, Berlin.</mixed-citation></ref><ref id="scirp.94295-ref8"><label>8</label><mixed-citation publication-type="other" xlink:type="simple">Kirk, R. (2012) Experimental Design: Procedures for Behavioral Sciences. Sage Publications, Thousand Oaks. https://doi.org/10.4135/9781483384733</mixed-citation></ref><ref id="scirp.94295-ref9"><label>9</label><mixed-citation publication-type="other" xlink:type="simple">Bathke, A. (2004) The Anova F Test Can Still Be Used in Some Balanced Designs with Unequal Variances and Non-normal Data. Journal of Statistical Planning and Inference, 126, 413-422. https://doi.org/10.1016/j.jspi.2003.09.010</mixed-citation></ref></ref-list></back></article>