<?xml version="1.0" encoding="UTF-8"?><!DOCTYPE article  PUBLIC "-//NLM//DTD Journal Publishing DTD v3.0 20080202//EN" "http://dtd.nlm.nih.gov/publishing/3.0/journalpublishing3.dtd"><article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" dtd-version="3.0" xml:lang="en" article-type="research article"><front><journal-meta><journal-id journal-id-type="publisher-id">OJS</journal-id><journal-title-group><journal-title>Open Journal of Statistics</journal-title></journal-title-group><issn pub-type="epub">2161-718X</issn><publisher><publisher-name>Scientific Research Publishing</publisher-name></publisher></journal-meta><article-meta><article-id pub-id-type="doi">10.4236/ojs.2022.122020</article-id><article-id pub-id-type="publisher-id">OJS-116834</article-id><article-categories><subj-group subj-group-type="heading"><subject>Articles</subject></subj-group><subj-group subj-group-type="Discipline-v2"><subject>Physics&amp;Mathematics</subject></subj-group></article-categories><title-group><article-title>
 
 
  On Sample Size Determination When Comparing Two Independent Spearman or Kendall Coefficients
 
</article-title></title-group><contrib-group><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Justine</surname><given-names>O. May</given-names></name><xref ref-type="aff" rid="aff1"><sup>1</sup></xref></contrib><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Stephen</surname><given-names>W. Looney</given-names></name><xref ref-type="aff" rid="aff2"><sup>2</sup></xref></contrib></contrib-group><aff id="aff2"><addr-line>Department of Population Health Sciences, Augusta University, Augusta, USA</addr-line></aff><aff id="aff1"><addr-line>Health System Information Technology, Augusta University, Augusta, USA</addr-line></aff><pub-date pub-type="epub"><day>14</day><month>03</month><year>2022</year></pub-date><volume>12</volume><issue>02</issue><fpage>291</fpage><lpage>302</lpage><history><date date-type="received"><day>18,</day>	<month>March</month>	<year>2022</year></date><date date-type="rev-recd"><day>24,</day>	<month>April</month>	<year>2022</year>	</date><date date-type="accepted"><day>27,</day>	<month>April</month>	<year>2022</year></date></history><permissions><copyright-statement>&#169; Copyright  2014 by authors and Scientific Research Publishing Inc. </copyright-statement><copyright-year>2014</copyright-year><license><license-p>This work is licensed under the Creative Commons Attribution International License (CC BY). http://creativecommons.org/licenses/by/4.0/</license-p></license></permissions><abstract><p>
 
 
  One of the most commonly used statistical methods is bivariate correlation analysis. However, it is usually the case that little or no attention is given to power and sample size considerations when planning a study in which correlation will be the primary analysis. In fact, when we reviewed studies published in clinical research journals in 2014, we found that none of the 111 articles that presented results of correlation analyses included a sample size justification. It is sometimes of interest to compare two correlation coefficients between independent groups. For example, one may wish to compare diabetics and non-diabetics in terms of the correlation of systolic blood pressure with age. Tools for performing power and sample size calculations for the comparison of two independent Pearson correlation coefficients are widely available; however, we were unable to identify any easily accessible tools for 
  power and sample size calculations when comparing two independent
   Spearman rank correlation coefficients or two independent Kendall coefficients of concordance. In this article, we provide formulas and charts that can be used to calculate the sample size that is needed when testing the hypothesis that two independent Spearman or Kendall coefficients are equal.
 
</p></abstract><kwd-group><kwd>Fisher z-Transform</kwd><kwd> Hypothesis Testing</kwd><kwd> Power</kwd><kwd> Significance Level</kwd></kwd-group></article-meta></front><body><sec id="s1"><title>1. Introduction</title><p>One of the most commonly used statistical methods is bivariate correlation analysis. Sometimes it is of interest to compare the correlation of two variables X and Y that have been calculated using two independent samples. For example, in an unpublished Master’s thesis [<xref ref-type="bibr" rid="scirp.116834-ref1">1</xref>], Stuart evaluated several potential biomarkers for severity of symptoms of dry mouth (xerostomia). As part of her assessment of these biomarkers, she examined the association of one of the potential biomarkers, p21, with another potential biomarker, PCNA. Stuart wished to compare two independent groups in terms of the association between p21 and PCNA: 1) healthy subjects and 2) dental patients with xerostomia. An important part of the planning of this study is to ask the question “How large a sample is needed in the two groups to achieve adequate power?”</p><p>Tools are widely available for performing sample size and power calculations when the analysis involves the comparison of two Pearson correlation coefficients (PCCs). These include, for example, tables [<xref ref-type="bibr" rid="scirp.116834-ref2">2</xref>], software packages (PASS, nQuery, G*Power), and internet-based tools (e.g., https://www.unistat.com/guide/sample-size-and-power-two-correlations/). However, as best we can determine, there are no easily accessible tools for sample size calculation when the planned analysis will be a comparison of either two Spearman rank correlation coefficients (SCCs) or two Kendall coefficients of concordance (KCCs). The PCC is the most commonly used measure of bivariate association; however, the SCC and KCC are also widely used. For example, Brough et al. [<xref ref-type="bibr" rid="scirp.116834-ref3">3</xref>] used the SCC to measure the associations between peanut protein levels found in various household environments, including dust, surfaces, bedding, furnishings and air. Heist et al. [<xref ref-type="bibr" rid="scirp.116834-ref4">4</xref>] used the KCC as the measure of association in their study of the use of bevacizumab as a chemotherapeutic agent for the treatment of advanced non-small cell lung cancer. The KCC is also frequently used in the analysis of environmental data. For example, Helsel ( [<xref ref-type="bibr" rid="scirp.116834-ref5">5</xref>], pp. 227-228) used the KCC in his examination of concentrations of dissolved iron in water samples.</p><p>In May and Looney [<xref ref-type="bibr" rid="scirp.116834-ref6">6</xref>], we provided formulas and charts that can be used to determine the sample size needed for a hypothesis test of a single Spearman or Kendall coefficient. In this article, we extend those results to the situation in which it is desired to compare two independent Spearman or Kendall coefficients.</p><p>In Section 2, we briefly describe our previously published methods for sample size determination for a single Spearman or Kendall coefficient [<xref ref-type="bibr" rid="scirp.116834-ref6">6</xref>]. This will facilitate the description of our methods for determining the sample size when comparing two independent coefficients since the results in the two situations are similar in many ways. In Section 3, we present the formulas and charts for comparing two independent Spearman or Kendall coefficients; in Section 4, we provide some notes on the use of the sample size charts; and, in Section 5, we discuss our results.</p></sec><sec id="s2"><title>2. Methods for Sample Size Determination for a Single Measure of Association</title><p>Let ξ denote the population value of the measure of association to be used (either the PCC, SCC or KCC). Suppose that we wish to test a hypothesis involving ξ. In the two-sided case, for example, the hypotheses to be tested can be stated as:</p><p>H 0 : ξ = ξ 0     vs .     H a : ξ ≠ ξ 0 , (1)</p><p>where ξ<sub>0</sub> is the pre-specified null value of the desired measure of association and − 1 &lt; ξ 0 &lt; 1 . For the PCC, the most common approach for testing the hypotheses in (1) is to use a test statistic based on the Fisher z-transform of the sample value of the PCC, denoted by r. For any − 1 &lt; r &lt; 1 , the Fisher z-transform of r is given by [<xref ref-type="bibr" rid="scirp.116834-ref7">7</xref>]:</p><p>z ( r ) = tanh − 1 r = ln ( ( 1 + r ) / ( 1 − r ) ) / 2 , (2)</p><p>which is asymptotically distributed as N ( tanh − 1 ρ , σ z 2 ) , where ρ denotes the population value of the PCC and σ z 2 denotes the asymptotic variance of z(r). The transformation in (2) can be applied to the sample SCC or KCC; this also yields an approximately normally distributed transformed coefficient. For the PCC, σ z 2 = 1 / ( n − 3 ) [<xref ref-type="bibr" rid="scirp.116834-ref7">7</xref>] and, for the KCC, σ z 2 = 0.437 / ( n − 4 ) [<xref ref-type="bibr" rid="scirp.116834-ref8">8</xref>]. For the SCC, the preferred value of σ z 2 depends on ρ<sub>s</sub>, the population value of the Spearman coefficient: for | ρ s | &lt; 0.95 , σ z 2 = ( 1 + ρ s 2 / 2 ) / ( n − 3 ) [<xref ref-type="bibr" rid="scirp.116834-ref9">9</xref>]; for | ρ s | ≥ 0.95 , σ z 2 = 1.06 / ( n − 3 ) [<xref ref-type="bibr" rid="scirp.116834-ref8">8</xref>]. The value σ z 2 = 1.06 / ( n − 3 ) is more commonly used for the SCC than the improved value proposed by Bonett and Wright [<xref ref-type="bibr" rid="scirp.116834-ref9">9</xref>] when | ρ s | &lt; 0.95 ; however, the results in this article are based on the Bonett and Wright values.</p><p>A test statistic for testing H 0 : ξ = ξ 0 based on the Fisher z-transform is given by:</p><p>z ξ 0 = z ( ξ ^ ) − z ( ξ 0 ) c 2 / ( n − b ) , (3)</p><p>where ξ ^ denotes the sample estimate of ξ, ξ<sub>0</sub> denotes the hypothesized value of ξ, n denotes the sample size, z ( &#183; ) denotes the Fisher z-transform, and b and c<sup>2</sup> are obtained from <xref ref-type="table" rid="table1">Table 1</xref>.</p><p>An approximate p-value is obtained using the appropriate tail probability for z<sub>ξ</sub><sub>0</sub> given in (3) using the standard normal distribution.</p><p>Suppose we wish to test H 0 : ξ = ξ 0 using the test statistic in (3). Let ξ<sub>1</sub> denote the alternative value of the measure of association that we wish to detect with our planned hypothesis test. The required sample size for detecting the value ξ<sub>1</sub> with power φ using a two-tailed test of H 0 : ξ = ξ 0 with significance level α is given by:</p><p>n = b + c 2 [ z α / 2 + z 1 − φ z ( ξ 1 ) − z ( ξ 0 ) ] 2 , (4)</p><table-wrap id="table1" ><label><xref ref-type="table" rid="table1">Table 1</xref></label><caption><title> Constants needed to apply the Fisher z-transform to measures of association</title></caption><table><tbody><thead><tr><th align="center" valign="middle"  colspan="2"  >Measure of Association</th><th align="center" valign="middle" >b</th><th align="center" valign="middle" >c<sup>2 </sup></th><th align="center" valign="middle" >Source</th></tr></thead><tr><td align="center" valign="middle" >Pearson</td><td align="center" valign="middle"  colspan="2"  >3</td><td align="center" valign="middle" >1</td><td align="center" valign="middle" >[<xref ref-type="bibr" rid="scirp.116834-ref7">7</xref>]</td></tr><tr><td align="center" valign="middle"  colspan="2"  >Spearman<sup>a</sup><sup> </sup></td><td align="center" valign="middle" >3</td><td align="center" valign="middle" >1 + ( ρ s 0 2 / 2 )</td><td align="center" valign="middle" >[<xref ref-type="bibr" rid="scirp.116834-ref9">9</xref>]</td></tr><tr><td align="center" valign="middle"  colspan="2"  >Spearman<sup>a</sup></td><td align="center" valign="middle" >3</td><td align="center" valign="middle" >1.06</td><td align="center" valign="middle" >[<xref ref-type="bibr" rid="scirp.116834-ref8">8</xref>]</td></tr><tr><td align="center" valign="middle"  colspan="2"  >Kendall</td><td align="center" valign="middle" >4</td><td align="center" valign="middle" >0.437</td><td align="center" valign="middle" >[<xref ref-type="bibr" rid="scirp.116834-ref8">8</xref>]</td></tr><tr><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td></tr></tbody></table></table-wrap><p><sup>a#</sup><sup> ρ s 0 </sup> denotes the null value of the SCC. c 2 = 1 + ( ρ s 0 2 / 2 ) is used in the sample size calculation if | ρ s 0 | &lt; 0.95 ; otherwise 1.06 is used.</p><p>where z<sub>γ</sub> = upper γ-percentage point of the standard normal distribution, z ( &#183; ) denotes the Fisher z-transform, and b and c<sup>2</sup> are obtained from <xref ref-type="table" rid="table1">Table 1</xref>.</p><p>For a one-tailed test of H 0 : ξ = ξ 0 , replace z<sub>α/</sub><sub>2</sub> in (4) by z<sub>α</sub>. If the formula in (4) does not yield an integer value, round up to the next largest integer.</p><p>In our previous article [<xref ref-type="bibr" rid="scirp.116834-ref6">6</xref>], we provided charts based on (4) for finding the sample size needed to achieve 80% power for tests of the SCC and KCC using significance level 0.05.</p></sec><sec id="s3"><title>3. Methods for Sample Size Determination for Two Independent Coefficients</title><p>Suppose that we wish to test the equality of two coefficients ξ<sub>1</sub> and ξ<sub>2</sub>, where ξ denotes either the SCC or the KCC, and that ξ<sub>1</sub> will be estimated using a sample that is independent of the sample used to estimate ξ<sub>2</sub>. In the two-sided case, for example, the hypotheses to be tested are given by:</p><p>H 0 : ξ 1 = ξ 2     vs .     H a : ξ 1 ≠ ξ 2 . (5)</p><p>Let ξ<sub>0</sub> denote the common value of ξ<sub>1</sub> and ξ<sub>2</sub> in H<sub>0</sub>.</p><p>We can use the approximate distributional results given in Section 2 for the Fisher z-transform of the estimators of the SCC and KCC to derive a test statistic for testing the hypotheses in (5). Let z ( ξ ^ 1 ) and z ( ξ ^ 2 ) denote the Fisher z-transformed estimators of ξ<sub>1</sub> and ξ<sub>2</sub>, respectively. Assume that ξ<sub>1</sub> is estimated using a sample that is independent of the sample used to estimate ξ<sub>2</sub>; hence, ξ ^ 1 and ξ ^ 2 are independent. Suppose that the same sample size n will be used to estimate ξ<sub>1</sub> and ξ<sub>2</sub>. (The assumption of equal sample sizes will be relaxed in Section 4.5.) Under the null hypothesis H<sub>0</sub> in (5), the random variable z ( ξ ^ 1 ) − z ( ξ ^ 2 ) is asymptotically normally distributed with mean 0 and variance c 2 ( 2 n − b ) , where the appropriate values of b and c<sup>2</sup> are given in <xref ref-type="table" rid="table1">Table 1</xref>. A test statistic for testing the hypotheses in (5) is given by:</p><p>z ξ 0 = z ( ξ ^ 1 ) − z ( ξ ^ 2 ) 2 c 2 / ( n − b ) , (6)</p><p>where ξ ^ 1 and ξ ^ 2 denote the sample estimates of ξ<sub>1</sub> and ξ<sub>2</sub>, respectively, n denotes the sample size in each group, z ( &#183; ) denotes the Fisher z-transform, and b and c<sup>2</sup> are given in <xref ref-type="table" rid="table1">Table 1</xref>.</p><p>An approximate p-value is then obtained by calculating the appropriate tail probability for z<sub>ξ</sub><sub>0</sub> in (6) using the standard normal distribution.</p><p>For the SCC, the hypotheses in (5) would be written as</p><p>H 0 : ρ s 1 = ρ s 2     vs .     H a : ρ s 1 ≠ ρ s 2 . (7)</p><p>Let ρ<sub>s</sub><sub>0</sub> denote the common value of the two Spearman coefficients under the null hypothesis in (7); namely, ρ s 0 = ρ s 1 = ρ s 2 . The value of c<sup>2</sup> in the test statistic in (6) when testing two SCCs is therefore c 2 = 1 + ( ρ s 0 2 / 2 ) if | ρ s 0 | &lt; 0.95 ; otherwise c<sup>2</sup> = 1.06 is used.</p><p>Assume that the two independent measures of association ξ<sub>1</sub> and ξ<sub>2</sub> (either two PCCs, two SCCs, or two KCCs) will be estimated using samples of equal size. Let ξ<sub>0</sub> denote the common value of ξ<sub>1</sub> and ξ<sub>2</sub> in the null hypothesis in (5) and let ξ<sub>21</sub> denote the alternative value of ξ<sub>2</sub> that one wishes to detect, assuming that ξ<sub>1</sub> = ξ<sub>0</sub>. For a two-tailed test of the null hypothesis in (5), the required sample size for detecting the alternative value ξ<sub>21</sub> with power φ using a test based on (6) with significance level α is given by:</p><p>n = b + 2 c 2 ( z α / 2 + z 1 − φ ) 2 [ z ( ξ 0 ) − z ( ξ 21 ) ] 2 , (8)</p><p>where z<sub>γ</sub> = upper γ-percentage point of the standard normal distribution, and z ( &#183; ) denotes the Fisher z-transform.</p><p>The values of b and c<sup>2</sup> in (8) are given in <xref ref-type="table" rid="table1">Table 1</xref>. For a one-tailed test of H 0 : ξ 1 = ξ 2 , replace z<sub>α/</sub><sub>2</sub> in (8) by z<sub>α</sub>. If the formula in (8) does not yield an integer value, round up to the next largest integer.</p><p>A chart for finding the required per-group sample size that will yield 80% power for comparing two independent Spearman coefficients based on samples of equal size using significance level 0.05 is provided in <xref ref-type="fig" rid="fig1">Figure 1</xref> for a two-tailed test and in <xref ref-type="fig" rid="fig2">Figure 2</xref> for a one-tailed test. The corresponding charts for two independent Kendall coefficients are given in <xref ref-type="fig" rid="fig3">Figure 3</xref> and <xref ref-type="fig" rid="fig4">Figure 4</xref>, respectively.</p><p>To use the chart in <xref ref-type="fig" rid="fig1">Figure 1</xref>, first locate the larger of the two SCCs associated with the alternative hypothesis along the horizontal axis. Next, draw a vertical line that intersects with the curve corresponding to the smaller value of the SCC associated with the alternative hypothesis. Finally, draw a horizontal line from the curve to the vertical axis. The point of intersection is the required sample size in each group.</p><p>For example, suppose one wishes to test</p><p>H 0 : ρ s 1 = ρ s 2     vs .     H a : ρ s 1 ≠ ρ s 2</p><p>and that the common null value under H<sub>0</sub> is ρ s 0 = ρ s 1 = ρ s 2 = 0.6 and the alternative value of the SCC that one wishes to detect is 0.4. In other words, the alternative difference that one wishes to detect is ρ s 1 − ρ s 2 = 0.6 − 0.4 = 0.2 . To use <xref ref-type="fig" rid="fig1">Figure 1</xref> to determine the sample size needed in each group to detect this difference with 80% power using a significance level of 0.05, locate max ( | ρ s 1 | , | ρ s 2 | ) = 0.6 on the horizontal axis and locate the curve corresponding to min ( | ρ s 1 | , | ρ s 2 | ) = 0.4 . After drawing the horizontal and vertical lines as described above, we find n = 258 (<xref ref-type="fig" rid="fig5">Figure 5</xref>). For ρ<sub>s</sub><sub>1</sub> = 0.4 and ρ<sub>s</sub><sub>2</sub> = 0.2, locate max ( | ρ s 1 | , | ρ s 2 | ) = 0.4 on the horizontal axis and locate the curve corresponding to min ( | ρ s 1 | , | ρ s 2 | ) = 0.2 ; the resulting sample size is n = 351. For a one-tailed test, the sample sizes obtained using <xref ref-type="fig" rid="fig2">Figure 2</xref> are 204 and 277, respectively.</p><p>If we use the KCC instead of the SCC as the measure of association, the required sample size for a two-tailed test, assuming τ<sub>1</sub> = 0.6 and τ<sub>2</sub> = 0.4 is n = 99, according to <xref ref-type="fig" rid="fig3">Figure 3</xref>. For τ<sub>1</sub> = 0.4 and τ<sub>2</sub> = 0.2, the sample size is n = 145. For a one-tailed test, the required sample sizes obtained from <xref ref-type="fig" rid="fig4">Figure 4</xref> are 79 and 115, respectively. Note that, in the KCC charts, we use the maximum and minimum values associated with the alternative hypothesis differently than in the SCC charts. For the KCC charts, we first locate min ( | τ 1 | , | τ 2 | ) (rather than the max) along the horizontal axis, and then draw a vertical line that intersects with the curve corresponding to max ( | τ 1 | , | τ 2 | ) (rather than the min).</p></sec><sec id="s4"><title>4. Notes on Using the Charts</title><sec id="s4_1"><title>4.1. Specification of Planning Values</title><p>It may not be possible for the analyst to specify the relevant planning values ρ<sub>s</sub><sub>0</sub>, ρ<sub>s</sub><sub>1</sub>, and ρ<sub>s</sub><sub>2</sub> for the SCC (or τ<sub>0</sub>, τ<sub>1</sub>, and τ<sub>2</sub> for the KCC). Since the null hypothesis for the SCC in (7) can also be written as H 0 : ρ s 1 − ρ s 2 = 0 , it is not necessary to specify the common null value ρ s 0 = ρ s 1 = ρ s 2 in H<sub>0</sub>. However, the required sample sizes in the two independent groups depend on the alternative values of ρ<sub>s</sub><sub>1</sub> and ρ<sub>s</sub><sub>2</sub> to be detected, as illustrated in <xref ref-type="table" rid="table2">Table 2</xref> for ρ s 1 − ρ s 2 = 0.2 (assuming that ρ s 1 &gt; ρ s 2 &gt; 0 ).</p><table-wrap id="table2" ><label><xref ref-type="table" rid="table2">Table 2</xref></label><caption><title> Required sample sizes to detect ρ<sub>s</sub><sub>1</sub> − ρ<sub>s</sub><sub>2</sub>= 0.2, Power = 80%, Significance Level = 0.05</title></caption><table><tbody><thead><tr><th align="center" valign="middle" >Common Value of ρ<sub>s</sub><sub>1</sub> and ρ<sub>s</sub><sub>2</sub> under H<sub>0</sub></th><th align="center" valign="middle" >Value of ρ<sub>s</sub><sub>1</sub> − ρ<sub>s</sub><sub>2</sub> <sub> </sub>to Be Detected</th><th align="center" valign="middle" >ρ<sub>s</sub><sub>1</sub></th><th align="center" valign="middle" >ρ<sub>s</sub><sub>2</sub></th><th align="center" valign="middle" >Required Sample Size in Each Group</th></tr></thead><tr><td align="center" valign="middle" >0.1</td><td align="center" valign="middle" >0.2</td><td align="center" valign="middle" >0.3</td><td align="center" valign="middle" >0.1</td><td align="center" valign="middle" >378</td></tr><tr><td align="center" valign="middle" >0.2</td><td align="center" valign="middle" >0.2</td><td align="center" valign="middle" >0.4</td><td align="center" valign="middle" >0.2</td><td align="center" valign="middle" >351</td></tr><tr><td align="center" valign="middle" >0.3</td><td align="center" valign="middle" >0.2</td><td align="center" valign="middle" >0.5</td><td align="center" valign="middle" >0.3</td><td align="center" valign="middle" >311</td></tr><tr><td align="center" valign="middle" >0.4</td><td align="center" valign="middle" >0.2</td><td align="center" valign="middle" >0.6</td><td align="center" valign="middle" >0.4</td><td align="center" valign="middle" >258</td></tr><tr><td align="center" valign="middle" >0.5</td><td align="center" valign="middle" >0.2</td><td align="center" valign="middle" >0.7</td><td align="center" valign="middle" >0.5</td><td align="center" valign="middle" >197</td></tr><tr><td align="center" valign="middle" >0.6</td><td align="center" valign="middle" >0.2</td><td align="center" valign="middle" >0.8</td><td align="center" valign="middle" >0.6</td><td align="center" valign="middle" >129</td></tr><tr><td align="center" valign="middle" >0.7</td><td align="center" valign="middle" >0.2</td><td align="center" valign="middle" >0.9</td><td align="center" valign="middle" >0.7</td><td align="center" valign="middle" >64</td></tr><tr><td align="center" valign="middle" >0.75</td><td align="center" valign="middle" >0.2</td><td align="center" valign="middle" >0.95</td><td align="center" valign="middle" >0.75</td><td align="center" valign="middle" >34</td></tr></tbody></table></table-wrap><p>As can be seen from <xref ref-type="table" rid="table2">Table 2</xref>, for a given alternative difference ρ<sub>s</sub><sub>1</sub> − ρ<sub>s</sub><sub>2</sub> to be detected, the required sample size in each group depends very heavily on the values of ρ<sub>s</sub><sub>1</sub> and ρ<sub>s</sub><sub>2</sub>. If there is no information available to help specify reasonable planning values, we recommend performing the type of sensitivity analysis illustrated in <xref ref-type="table" rid="table2">Table 2</xref> to assist in selecting values of ρ<sub>s</sub><sub>0</sub>, ρ<sub>s</sub><sub>1</sub>, and ρ<sub>s</sub><sub>2</sub> for the SCC (or τ<sub>0</sub>, τ<sub>1</sub>, and τ<sub>2</sub> for the KCC).</p></sec><sec id="s4_2"><title>4.2. Minimum Detectable Difference</title><p>In addition to finding the sample size needed for planning statistical inference for the SCC and the KCC, the charts can also be used to find the minimum detectable difference for a given sample size. For example, suppose that one wishes to compare two independent Kendall coefficients using a two-tailed test, and that the largest possible sample size available to the investigators is n = 100 in each group. Assuming that the smaller of the two alternative values is τ<sub>1</sub> = 0.4, one could determine the minimum detectable difference by first drawing a horizontal line from n = 100 on the vertical axis, and then drawing a vertical line from 0.4 on the horizontal axis. One can then simply read off the alternative values larger than 0.4 from this vertical line until it intersects with the horizontal line drawn from the vertical axis. Using <xref ref-type="fig" rid="fig3">Figure 3</xref>, we see that with n = 100 in each group, we could detect any alternative value τ<sub>2</sub> greater than or equal to 0.6 with 80% power using a two-tailed significance level of 0.05.</p></sec><sec id="s4_3"><title>4.3. Negative Values</title><p>The sample size charts in this article can be used only when ξ<sub>1</sub> and ξ<sub>2</sub> have the same sign. For negative values of ξ<sub>1</sub> and ξ<sub>2</sub>, one simply enters the appropriate chart with |ξ<sub>1</sub>| and |ξ<sub>2</sub>|. If ξ<sub>1</sub> and ξ<sub>2</sub> are of opposite signs, the formula in (8) is still valid; however, the charts provided in this article do not apply.</p></sec><sec id="s4_4"><title>4.4. Interpolation Errors</title><p>As with any graphical method, these charts are subject to error. For example, one must interpolate graphically if the horizontal line drawn from the appropriate curve in the charts intersects the vertical axis at a value between the tick marks. This is more of a problem for larger sample sizes since the distances between the tick marks are much larger. However, the error for the interpolated value of n will be no larger than the difference between the values of n corresponding to the two relevant tick marks, and an adjustment can be made by slightly inflating the value read from the chart. Our experience has been that increasing the n value obtained from the chart by 5 is usually adequate. We have found that an 6-inch ruler marked off in millimeters is particularly useful for carrying out any necessary graphical interpolation in the charts.</p></sec><sec id="s4_5"><title>4.5. Unequal Sample Sizes</title><p>In Section 3, the assumption was made that the sample sizes were equal in the two independent samples used to estimate ξ<sub>1</sub> and ξ<sub>2</sub>. If it is desirable that the sample sizes be different in the two groups, this can be accomplished as follows. Let n<sub>1</sub> and n<sub>2</sub> denote the sizes of the samples on which the estimates of ξ<sub>1</sub> and ξ<sub>2</sub> will be based, respectively. Let a = n<sub>2</sub>/n<sub>1</sub> denote the desired allocation ratio of the sample sizes in the two groups. Let n denote the per-group sample size obtained from the charts in Figures 1-4. Then allocating n 1 = 2 n / ( 1 + a ) to the sample used to estimate ξ<sub>1</sub> and allocating 2n – n<sub>1</sub> to the sample used to estimate ξ<sub>2</sub> will yield the desired sample sizes n<sub>1</sub> and n<sub>2</sub>. A non-integer value n<sub>1</sub> obtained from the above formula should be rounded up to the next largest integer. To illustrate, consider the example in Section 3, in which the alternative values to be detected were ρ<sub>s</sub><sub>1</sub> = 0.6 and ρ<sub>s</sub><sub>2</sub> = 0.4. The per-group sample size obtained from <xref ref-type="fig" rid="fig1">Figure 1</xref> was n = 258. Assume that the desired allocation ratio is a = 1/2. Then, n 1 = 2 n / ( 1 + a ) = 2 ( 258 ) / ( 1 + 0.5 ) = 344 and n 2 = 2 ( 258 ) − 344 = 172 . Equal sample sizes in the two independent groups will maximize power for the test of H 0 : ρ s 1 = ρ s 2 or H 0 : τ 1 = τ 2 , so allocating unequal sample sizes to the two groups will result in a loss of power and should not be done unless absolutely necessary.</p></sec><sec id="s4_6"><title>4.6. Choice of Values of ρ<sub>s</sub><sub>1</sub> and ρ<sub>s</sub><sub>2</sub> in the Charts</title><p>The values of ρ<sub>s</sub><sub>1</sub> and ρ<sub>s</sub><sub>2</sub> presented in our charts were chosen to make the charts as easy to use as possible and to avoid cluttering the graphs. In particular, for the Spearman charts (<xref ref-type="fig" rid="fig1">Figure 1</xref>, <xref ref-type="fig" rid="fig2">Figure 2</xref>, <xref ref-type="fig" rid="fig5">Figure 5</xref>), we presented results for the following choices of max ( | ρ s 1 | , | ρ s 2 | ) , which was plotted on the horizontal axis: 0.05 (0.05) 0.90. We presented results for sample size curves corresponding to the following choices of min ( | ρ s 1 | , | ρ s 2 | ) : 0.0 (0.1) 0.8. If we had included a curve for min ( | ρ s 1 | , | ρ s 2 | ) = 0.9 , for example, this would have consisted of a single point, which we felt would have detracted from the overall visual appeal and interpretability of the charts. In general, for any values of ρ<sub>s</sub><sub>1</sub> and ρ<sub>s</sub><sub>2</sub> that are not available on the sample size curves for either the Spearman or Kendall coefficients, the simple formula for n given in Equation (8) can be used to determine the required sample size.</p></sec></sec><sec id="s5"><title>5. Discussion</title><p>In this article, we have presented charts that can be used for sample size determination when planning a study in which hypothesis testing will be used to compare two independent Spearman or Kendall coefficients. In addition to the charts, we have provided simple sample size formulas that can be used for more accurate calculations. We have found Microsoft Excel<sup>&#169;</sup> to be particularly useful for performing these calculations, and a spreadsheet that accomplishes this is available from the second author.</p><p>Despite the widespread use of correlation analysis, it is usually the case that, when planning a study in which correlation will be the primary analysis, little or no attention is given to sample size determination. This general impression was confirmed by our review of studies that used correlation as the primary analysis; none of the 111 studies published in clinical research journals in 2014 provided a power analysis or sample size calculation.</p><p>While it is true that two of the key references in this article ( [<xref ref-type="bibr" rid="scirp.116834-ref7">7</xref>] [<xref ref-type="bibr" rid="scirp.116834-ref9">9</xref>] ) are rather old, they provide valid results that are directly relevant to the present article. As stated in the Introduction, we were unable to locate any modern tools (software, tables, graphs, etc.) that can be used to determine the sample size needed for comparing two independent Spearman or Kendall coefficients. Hence, we made use of the classical results from these two articles to develop our new tools for addressing this problem. The present article represents an application of the results in [<xref ref-type="bibr" rid="scirp.116834-ref7">7</xref>] [<xref ref-type="bibr" rid="scirp.116834-ref8">8</xref>], and [<xref ref-type="bibr" rid="scirp.116834-ref9">9</xref>] to extend our results of our recently published article [<xref ref-type="bibr" rid="scirp.116834-ref6">6</xref>], which considered only a single Spearman or Kendall coefficient.</p><p>We hope that, by making available the easy-to-use tools presented in this article, analysts will be encouraged to perform sample size calculations for correlation coefficient inference. Given the adverse consequences that can occur when studies are either under- or over-powered, it is extremely important that such calculations be made prior to beginning a research study. Our future research efforts will focus on extending our results in [<xref ref-type="bibr" rid="scirp.116834-ref6">6</xref>] to sample size estimation for the intra-class correlation.</p></sec><sec id="s6"><title>Conflicts of Interest</title><p>The authors declare no conflicts of interest regarding the publication of this paper.</p></sec><sec id="s7"><title>Cite this paper</title><p>May, J.O. and Looney, S.W. (2022) On Sample Size Determination When Comparing Two Independent Spearman or Kendall Coefficients. Open Journal of Statistics, 12, 291-302. https://doi.org/10.4236/ojs.2022.122020</p></sec></body><back><ref-list><title>References</title><ref id="scirp.116834-ref1"><label>1</label><mixed-citation publication-type="other" xlink:type="simple">Stuart, M. (2013) Identification of Novel Molecular Biomarkers for Diagnosis of Salivary Dysfunction. Master’s Thesis, Georgia Regents University, Augusta.</mixed-citation></ref><ref id="scirp.116834-ref2"><label>2</label><mixed-citation publication-type="other" xlink:type="simple">Cohen, J. (1988) Statistical Power Analysis for the Behavioral Sciences. 2nd Edition, Lawrence Erlbaum Associates, Hillsdale, New Jersey.</mixed-citation></ref><ref id="scirp.116834-ref3"><label>3</label><mixed-citation publication-type="other" xlink:type="simple">Brough, H.A., Makinson, K., Penagos, M., Maleki, S.J., Cheng, H., Douiri, A., Stephens, A.C., Turcanu, V. and Lack, G. (2013) Distribution of Peanut Protein in the Home Environment. Journal of Allergy and Clinical Immunology, 132, 623-629.  
https://doi.org/10.1016/j.jaci.2013.02.035</mixed-citation></ref><ref id="scirp.116834-ref4"><label>4</label><mixed-citation publication-type="other" xlink:type="simple">Heist, R.S., Duda, G.D., Sahani, D., Pennell, N.A., Neal, J.W., Ancukiewicz, M., Engelman, J.A., Lynch, T.J. and Jain, R.K. (2010) In Vivo Assessment of the Effects of Bevacizumab in Advanced Non-Small Cell Lung Cancer (NSCLC). Journal of Clinical Oncology, 28, 7612. https://doi.org/10.1200/jco.2010.28.15_suppl.7612</mixed-citation></ref><ref id="scirp.116834-ref5"><label>5</label><mixed-citation publication-type="other" xlink:type="simple">Helsel, D.R. (2012) Statistics for Censored Environmental Data Using Minitab and R. 2nd Edition, John Wiley &amp; Sons, Hoboken, New Jersey.  
https://doi.org/10.1002/9781118162729</mixed-citation></ref><ref id="scirp.116834-ref6"><label>6</label><mixed-citation publication-type="other" xlink:type="simple">May, J.O. and Looney, S.W. (2020) Sample Size Charts for Spearman and Kendall Coefficients. Journal of Biometrics &amp; Biostatistics, 11, 7 p.</mixed-citation></ref><ref id="scirp.116834-ref7"><label>7</label><mixed-citation publication-type="other" xlink:type="simple">Fisher, R.A. (1925) Statistical Methods for Research Workers. Hafner Press, London.</mixed-citation></ref><ref id="scirp.116834-ref8"><label>8</label><mixed-citation publication-type="other" xlink:type="simple">Fieller, E.C., Hartley, H.O. and Pearson, E.S. (1957) Tests for Rank Correlation Coefficients. Biometrika, 44, 470-481. https://doi.org/10.1093/biomet/44.3-4.470</mixed-citation></ref><ref id="scirp.116834-ref9"><label>9</label><mixed-citation publication-type="other" xlink:type="simple">Bonett, D.G. and Wright, T.A. (2000) Sample Size Requirements for Estimating Pearson, Kendall, and Spearman Correlations. Psychometrika, 65, 23-28.  
https://doi.org/10.1007/BF02294183</mixed-citation></ref></ref-list></back></article>