<?xml version="1.0" encoding="UTF-8"?><!DOCTYPE article  PUBLIC "-//NLM//DTD Journal Publishing DTD v3.0 20080202//EN" "http://dtd.nlm.nih.gov/publishing/3.0/journalpublishing3.dtd"><article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" dtd-version="3.0" xml:lang="en" article-type="research article"><front><journal-meta><journal-id journal-id-type="publisher-id">OJS</journal-id><journal-title-group><journal-title>Open Journal of Statistics</journal-title></journal-title-group><issn pub-type="epub">2161-718X</issn><publisher><publisher-name>Scientific Research Publishing</publisher-name></publisher></journal-meta><article-meta><article-id pub-id-type="doi">10.4236/ojs.2019.93024</article-id><article-id pub-id-type="publisher-id">OJS-93080</article-id><article-categories><subj-group subj-group-type="heading"><subject>Articles</subject></subj-group><subj-group subj-group-type="Discipline-v2"><subject>Physics&amp;Mathematics</subject></subj-group></article-categories><title-group><article-title>
 
 
  Test for Homogeneity of Odds Ratios Using &lt;i&gt;U&lt;/i&gt;-Statistics
 
</article-title></title-group><contrib-group><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Qi</surname><given-names>Wei</given-names></name><xref ref-type="aff" rid="aff1"><sup>1</sup></xref></contrib><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Dejian</surname><given-names>Lai</given-names></name><xref ref-type="aff" rid="aff2"><sup>2</sup></xref><xref ref-type="corresp" rid="cor1"><sup>*</sup></xref></contrib></contrib-group><aff id="aff2"><addr-line>Department of Biostatistics and Data Science, School of Public Health, The University of Texas Health Science Center at Houston, Houston, USA</addr-line></aff><aff id="aff1"><addr-line>Surgery Core Research, Baylor College of Medicine, Houston, USA</addr-line></aff><pub-date pub-type="epub"><day>17</day><month>06</month><year>2019</year></pub-date><volume>09</volume><issue>03</issue><fpage>347</fpage><lpage>360</lpage><history><date date-type="received"><day>9,</day>	<month>May</month>	<year>2019</year></date><date date-type="rev-recd"><day>15,</day>	<month>June</month>	<year>2019</year>	</date><date date-type="accepted"><day>18,</day>	<month>June</month>	<year>2019</year></date></history><permissions><copyright-statement>&#169; Copyright  2014 by authors and Scientific Research Publishing Inc. </copyright-statement><copyright-year>2014</copyright-year><license><license-p>This work is licensed under the Creative Commons Attribution International License (CC BY). http://creativecommons.org/licenses/by/4.0/</license-p></license></permissions><abstract><p>
 
 
  There are a few statistics testing the homogeneity of odds ra
  t
  ios across strata. Asymptotic statistics los
  e
   their power in the “sparse-data” setting. Both asymptotic statistics and exact tests have low power when the sample sizes are small. We created a set of U statistics and compared them with some existing statistics in testing homogeneity of OR at different data settings. We evaluated their performance in terms of the empirical size and power via Monto Carlo simulations. Our results showed that two of the U-statistics under our study had higher power for testing homogeneity of odds ratios for 2 by 2 contingency tables. The application of the tests was illustrated in two real examples.
 
</p></abstract><kwd-group><kwd>Homogeneity Test</kwd><kwd> Odds Ratio</kwd><kwd> &lt;i&gt;U&lt;/i&gt;-Statistics</kwd></kwd-group></article-meta></front><body><sec id="s1"><title>1. Introduction</title><p>Odds ratio is commonly used in the analysis of association of two factors that both have two categories. In epidemiological studies and clinical trials, these two factors usually refer to the exposure (treatment/intervention/risk) factor X and the outcome factor Y respectively. The association between X and Y, however, could be modified or confounded by a third factor Z. For example, in a multi-center clinical trial, factor Z could be the center. Each center is corresponding to a stratum of Z. Because the presence of the heterogeneity of odds ratios may lead to different methods of analysis, researchers usually want to test whether the odds ratios are homogeneous across the strata of the factor Z or not. This type of tests is called the tests for homogeneity of the odds ratios, or the tests for null interaction.</p><p>A few procedures have been developed for testing the homogeneity of odds ratios. They are usually categorized into two classes: exact tests and asymptotic tests. Most of the asymptotic statistics were derived for “large-stratum” settings, where the sample size is large, and the number of strata is small. Liang and Self [<xref ref-type="bibr" rid="scirp.93080-ref1">1</xref>] developed two asymptotic statistics—score statistics for the “sparse-data” setting, where there are many cells with small counts and/or zeros. The asymptotic tests could be poor if there were some cells with small cell counts in the contingency tables even when the sample size was quite large. The empirical sizes of the asymptotic tests were shown to be conservative when the data were “small-stratum” setting [<xref ref-type="bibr" rid="scirp.93080-ref2">2</xref>] . Jones et al. [<xref ref-type="bibr" rid="scirp.93080-ref3">3</xref>] showed in their simulation study that, in the “sparse-data” setting, five asymptotic tests suffered in both empirical size and power. Zelen [<xref ref-type="bibr" rid="scirp.93080-ref4">4</xref>] constructed an exact test for homogeneity that employs the ordering principle for a single 2 by 2 table, which was reexamined by Halperin et al. [<xref ref-type="bibr" rid="scirp.93080-ref5">5</xref>] . Algorithms were developed for the Zelen exact test later [<xref ref-type="bibr" rid="scirp.93080-ref6">6</xref>] [<xref ref-type="bibr" rid="scirp.93080-ref7">7</xref>] [<xref ref-type="bibr" rid="scirp.93080-ref8">8</xref>] [<xref ref-type="bibr" rid="scirp.93080-ref9">9</xref>] . Hirji et al. [<xref ref-type="bibr" rid="scirp.93080-ref10">10</xref>] showed that their algorithm for the Zelen statistic was more versatile and efficient in terms of the accuracy and the usage of memory in the computer computation. Hirji et al. [<xref ref-type="bibr" rid="scirp.93080-ref10">10</xref>] also constructed five other exact statistics from their asymptotic counterparts, which are score, likelihood ratio, Pearson χ<sup>2</sup> and mixture model χ<sup>2</sup> tests. The exact tests are computationally intensive, particularly when the cell counts are large. Reis, Hirji and Afifi [<xref ref-type="bibr" rid="scirp.93080-ref11">11</xref>] examined the performance of the empirical power of six asymptotic statistics and six exact tests mentioned above; the results of the study showed that the power of both exact and asymptotic tests was low when sample sizes were small even with relatively large heterogeneity among odds ratios. Baghri et al. [<xref ref-type="bibr" rid="scirp.93080-ref12">12</xref>] compared three tests of homogeneity of odds ratios. A recent study investigated profile likelihood tests for common odds ratios in meta-analysis [<xref ref-type="bibr" rid="scirp.93080-ref13">13</xref>] .</p><p>Our study presented in this paper compared a class of U-statistics with the established tests such as the Zelen test and the Breslow-Day test.</p></sec><sec id="s2"><title>2. U-Statistics</title><p>The statistics of the form</p><p>U = [ l ! ( n − l ) ! / n ! ] ∑ ( n , l ) h ( X i 1 , ⋯ , X i l ) (1)</p><p>are known as U-statistics, where { X i 1 , ⋯ , X i l } is a set of l-subset of { 1 , ⋯ , n } ;</p><p>the sum ∑ ( n , l ) is taken over all subsets 1 ≤ i 1 &lt; ⋯ &lt; i l ≤ n of { 1 , 2 , ⋯ , n } ; h is the</p><p>kernel function and symmetric in its arguments. In our study, we investigated the applicability of this class of statistics with l = 2 in testing homogeneity of odds ratios among 2 by 2 tables. U-statistics were first identified as a minimum-variance unbiased estimator by Halmos [<xref ref-type="bibr" rid="scirp.93080-ref14">14</xref>] and were named by Hoeffding [<xref ref-type="bibr" rid="scirp.93080-ref15">15</xref>] , who demonstrated the asymptotic normality of this class of statistics in his seminal paper [<xref ref-type="bibr" rid="scirp.93080-ref15">15</xref>] , Many statistics in common use are members of this class, including the sample mean, the sample variance and the sample covariance [<xref ref-type="bibr" rid="scirp.93080-ref16">16</xref>] [<xref ref-type="bibr" rid="scirp.93080-ref17">17</xref>] .</p><p><xref ref-type="table" rid="table1">Table 1</xref> shows a 2 by 2 table of a stratum k.</p><p>Assume that a<sub>k</sub> and b<sub>k</sub> are counts of independent binomial outcomes from number n<sub>k</sub> and m<sub>k</sub> of trials with or without exposure at the stratum k respectively; N<sub>k</sub> is the total sample size of the stratum k; and the tables are independent among strata. The commonly used estimate of the odds ratio of the kth stratum is expressed as: ψ ^ k = ( a k / ( n k − a k ) ) / ( b k / ( m k − b k ) ) . We want to test whether the odds ratios are homogeneous among all K strata (or all levels of the third variable Z), that is to test H<sub>0</sub>: ψ 1 = ψ 2 = ⋯ = ψ K = ψ against Ha: ψ i ≠ ψ j for at least one pair of (i, j), where i , j = 1 , ⋯ , K , i ≠ j .</p><p>Our study evaluated a class of weighted U-statistics ∑ i &lt; j w i j h ( ψ ^ i , ψ ^ j ) as well as</p><p>a class of unweighted U-statistics ( 2 ! ( K − 2 ) ! / K ! ) ∑ i &lt; j h ( ψ ^ i , ψ ^ j ) , where ψ ^ i and ψ ^ j are the estimates of the odds ratios in the ith and jth stratum of Z, h ( ψ ^ i , ψ ^ j ) is a function of ψ ^ i and ψ ^ j , and w i j is the weight associated with h ( ψ ^ i , ψ ^ j ) . Based on the simulation results, we only focus our attention on the following two statistics in this paper:</p><p>U3: ( 2 ! ( K − 2 ) ! / K ! ) ∑ i &lt; j | log ψ ^ i − log ψ ^ j | , (2)</p><p>WU3: ∑ i &lt; j K w i j | log ψ ^ i − log ψ ^ j | . (3)</p><p>The base of log in all formula is e. The sample distribution of the estimated odds ratio is highly skewed when the sample size is small or moderate. Because of this, we used the natural logarithm of ψ ^ in U3 and WU3 to reduce the skewness. Consider that a large stratum offers more accurate estimate for the odds ratio, a weight was selected for each h ( ψ ^ i , ψ ^ j ) in expression (3), which is proportional to the stratum’s size. It has the following form:</p><p>w i j = ( w i w j ) / ∑ w k (4)</p><p>where w i = 1 / var ( log ( ψ ^ i ) ) . The log transform of the sample odds ratio has an asymptotic variance in a simple form, which is, var ( log ψ ^ k ) ≈ 1 / a k + 1 / ( n k − a k ) + 1 / b k + 1 / ( m k − b k ) . If there was any cell count that equals to zero, the odds ratio estimate ψ ^ k and var ( log ψ ^ k ) from the above formula would be undefined. We added 0.5 to each cell count of that table in the calculations to get the amended estimators [<xref ref-type="bibr" rid="scirp.93080-ref18">18</xref>] . In our simulation, we studied other forms of U-statistics. Due to the relatively poor performance of the others, we only report the results of the U3 and the WU3 and their comparisons to the Zelen’s test and the Breslow-Day test in this paper. Detailed results are available upon request.</p><table-wrap id="table1" ><label><xref ref-type="table" rid="table1">Table 1</xref></label><caption><title> 2 &#215; 2 table of X and Y at Z = k</title></caption><table><tbody><thead><tr><th align="center" valign="middle" ></th><th align="center" valign="middle" >Y = 1</th><th align="center" valign="middle" >Y = 0</th><th align="center" valign="middle" ></th></tr></thead><tr><td align="center" valign="middle" >X = 1</td><td align="center" valign="middle" >a<sub>k</sub></td><td align="center" valign="middle" >n<sub>k</sub> - a<sub>k</sub></td><td align="center" valign="middle" >n<sub>k</sub></td></tr><tr><td align="center" valign="middle" >X = 0</td><td align="center" valign="middle" >b<sub>k</sub></td><td align="center" valign="middle" >m<sub>k</sub> - b<sub>k</sub></td><td align="center" valign="middle" >m<sub>k</sub></td></tr><tr><td align="center" valign="middle" ></td><td align="center" valign="middle" >t<sub>k</sub></td><td align="center" valign="middle" >N<sub>k</sub> - t<sub>k</sub></td><td align="center" valign="middle" >N<sub>k</sub></td></tr></tbody></table></table-wrap></sec><sec id="s3"><title>3. Simulation Study</title><p>A total of 10,000 data sets were simulated using the SAS subroutine RANBIN. Each data set contains pre-specified K sets of 2 by 2 tables. The cell counts a<sub>k</sub> and b<sub>k</sub> were independently generated from binomial distributions (n<sub>k</sub>, p<sub>1k</sub>) and (m<sub>k</sub>, p<sub>0k</sub>), where p x k = P ( Y = 1 | X = x , Z = k ) is the probability of Y = 1 when X = x (x = 1, 0) in the kth stratum. Each set of the tables was simulated with a given n<sub>k</sub>, m<sub>k</sub>, the number of the strata K and the odds ratios. For a given odd ratio ψ<sub>k</sub> and a binomial proportion p<sub>0k</sub>, p<sub>1k</sub> was calculated by solving:</p><p>ψ k = [ p 1 k / ( 1 − p 1 k ) ] / [ p 0 k / ( 1 − p 0 k ) ] .</p><p>Following the previous simulation study by Reis, Hirji and Afifi [<xref ref-type="bibr" rid="scirp.93080-ref11">11</xref>] , five factors that might influence the sizes and power of the test statistics were studied. These five factors are: 1) the sample size for X = 1 (n<sub>k</sub>); 2) the ratio of the sample sizes of X = 0 and X = 1 (m<sub>k</sub>:n<sub>k</sub>); 3) the probability of Y = 1 among X = 0 (p<sub>0k</sub>); 4) the odds ratio (OR) ψ k ; 5) the number of 2 &#215; 2 tables, K (strata). In order to study the performance of the statistics, different values of the above five factors (parameters) were chosen to simulate different data sets. In choosing the values of n<sub>k</sub>, m<sub>k</sub>:n<sub>k</sub>, p<sub>0k</sub>, ψ k , and K, we took the following factors into account: the situations in practice, the characteristics of U-statistics and the design of the simulation study by Reis, Hirji and Afifi [<xref ref-type="bibr" rid="scirp.93080-ref11">11</xref>] . When the effect of one factor was studied, the other four factors’ values were fixed.</p><p>We compared the performance of the U-statistics with the Breslow-Day statistic and Zelen’s exact test in our simulation study. A C++ program was written to calculate these statistics’ exact P-values, empirical sizes and power. The empirical size was calculated as the percentage of times that the test rejected the null hypothesis of a common odds ratio at a pre-specified α level among 10,000 tests that were simulated with same odds ratios among K tables. The empirical power was calculated as the percentage of times that a test rejected the null hypothesis of a common odds ratio at a prescribed α level when data were simulated under alternative hypotheses. Because the U-statistics studied here are functions of the sums across all the absolute distances between all possible pairs of the estimated odds ratios in log scale, a large value of U statistics indicates the heterogeneity of the odds ratio.</p><p>Theoretically, under suitable conditions, [ T − E ( T ) ] / var ( T ) will be asymptotically following N ( 0 , 1 ) as K → ∞ , where T represents a U-statistic. In our study, the sample mean and the sample variance of 10,000 statistics under the null hypothesis were used to estimate the E(T) and the var(T). In an actual application, the sample mean and variance may be estimated as suggested in our application section. In the simulation, one would use the result from expression [<xref ref-type="bibr" rid="scirp.93080-ref9">9</xref>] for the variance. In so doing, one would get a different numerical value of the variance for each realization, which was inconsistent with our assumption of common variance. The effect of unequal variance is still under investigation.</p></sec><sec id="s4"><title>4. Results of the Simulation Study</title><sec id="s4_1"><title>4.1. Empirical Sizes of the Tests</title><p>The five factors affected the test statistics differently. The empirical size of the Breslow-Day test was improved (moved closer to the predefined α level) as the values of n<sub>k</sub>, m<sub>k</sub>:n<sub>k</sub>, p<sub>0k</sub> and odds ratios increased but diverged from the pre-specified α level when the number of stratum increased. A weak trend was observed that the empirical size of U3 and WU3 moved closer and then diverged from the pre-specified α level when the sample size n<sub>k</sub> increased (<xref ref-type="fig" rid="fig1">Figure 1</xref>). Similar results were observed in studying the effect of ratio m<sub>k</sub>:n<sub>k</sub>; their empirical sizes were improved when m<sub>k</sub>:n<sub>k</sub> increased to moderate value (m<sub>k</sub>:n<sub>k</sub> equal to 2 or 3), they then diverged from the predefined significant level 0.05 if the value of m<sub>k</sub>:n<sub>k</sub> was higher (<xref ref-type="fig" rid="fig2">Figure 2</xref>). The empirical sizes of U-statistics increased as the value of the odds ratio increased. The probability p<sub>0k</sub> had the least impact on the empirical sizes of the two U-statistics.</p><p>The number of strata K had an apparent effect on the sizes of U-statistics; their empirical sizes were improved as K increased (<xref ref-type="fig" rid="fig3">Figure 3</xref>), even when the total sample size remained the same (<xref ref-type="fig" rid="fig4">Figure 4</xref>). The empirical size of Breslow-Day test diverged from the predefined nominal size of 5% as the value of K increased. The empirical size of Zelen test improved as K and total sample size increased but diverged from nominal size of 5% when the number of stratums increased without increasing the total sample size.</p></sec><sec id="s4_2"><title>4.2. Empirical Power of the Statistics</title><p>Seven settings of heterogeneous odds ratios were evaluated as alternative hypotheses in our study. However, in this article, we only reported the empirical powers from the scenario that the alterative odds ratios were generated following the pattern of 1, 2, 3, 7. That is, 25% of the generated tables under Ha have odds ratios being 1, 2, 3 and 7, respectively. In order to show the effects of different factors on the test statistics, we also simulated the critical values based on these factors (Figures 5-11).</p><p>All the statistics’ empirical power increased as n<sub>k</sub> increased (<xref ref-type="fig" rid="fig5">Figure 5</xref>, <xref ref-type="fig" rid="fig6">Figure 6</xref>). The value of the ratio m<sub>k</sub>:n<sub>k</sub> had very little effect on the empirical power of U3; the empirical power of the weighted U-statistic and Breslow-Day test were improved as the value of m<sub>k</sub>:n<sub>k</sub> increased (<xref ref-type="fig" rid="fig7">Figure 7</xref>). Increasing the value of the odds ratio under the null hypothesis decreased the empirical power of all statistics studied (<xref ref-type="fig" rid="fig8">Figure 8</xref>). The empirical power increased as the value of p<sub>0k</sub> increased (<xref ref-type="fig" rid="fig9">Figure 9</xref>). When the number of strata K increased, together with total sample size increased, the power of the U-statistics and Zelen’s exact test were improved (<xref ref-type="fig" rid="fig1">Figure 1</xref>0); If the total sample size remained the same when K increased, the empirical sizes of the U-statistics were improved; the empirical size of the Breslow-Day and Zelen’s statistics diverged from 5%; the empirical power of U3 almost remained unchanged, the others’ power decreased (<xref ref-type="fig" rid="fig1">Figure 1</xref>1).</p></sec></sec><sec id="s5"><title>5. Application to Published Data</title><p>Generally, U3 and WU3 performed well in terms of both size and power. Their empirical sizes were stable under various situations and had relatively high power.</p><p>With the assumption that the odds ratios of K 2 &#215; 2 tables are the same; and log ( ψ ^ i ) is normally distributed with variance equal to σ 2 , we can derive the estimated expected value of U3 and WU3. They are:</p><p>E ( U 3 ) = ( 2 ! ( K − 2 ) ! / K ! ) E ( ∑ i &lt; j | log ψ ^ i − log ψ ^ j | ) = 2 ( 1 / ( 2 σ π ) ) ∫ 0 ∞ x exp ( − x 2 / ( 4 σ 2 ) ) d x = 2 σ π (5)</p><p>E ( W U 3 ) = E ( ∑ i &lt; j K w i j | log ψ ^ i − log ψ ^ j | ) = E ( | log ψ ^ i − log ψ ^ j | ) ∑ i &lt; j K w i j = ∑ i &lt; j K w i j 2 σ π (6)</p><p>The variance of U3 can be also expressed as:</p><p>var ( U 3 ) = [ 2 ! ( K − 2 ) ! / K ! ] [ 2 ( K − 2 ) 1.436 σ 2 + ( K − 2 ) 2 σ 2 ] . (7)</p><p>The variance of WU3 can be also expressed as:</p><p>var ( W U 3 ) = [ 1.436 σ 2 ∑ j ≠ k w i j w i k + 2 σ 2 ∑ i &lt; j w i j 2 ] . (8)</p><p>To estimate the σ 2 , consider the Mantel-Haenzel estimator of common odds ratio ψ ^ M H as a weighted average of odds ratios. Given the weight, we can solve the σ 2 as a function of var ( log ( ψ ^ M H ) ) , which is:</p><p>σ 2 = var ( log ( ψ ^ M H ) ) [ ∑ ( b k c k / N k ) ] 2 / ∑ ( b k c k / N k ) 2 (9)</p><p>And the value of var ( log ( ψ ^ M H ) ) can be calculated by the following formula:</p><p>var ( log ( ψ ^ M H ) ) = ( ∑ G i P i ) / [ 2 [ ∑ G i ] 2 ] + ∑ ( G i Q i + H i P i ) / [ 2 ∑ G i ∑ H i ]     + ∑ H i Q i / 2 [ ∑ H i ] 2</p><p>where G i = a i d i / N i , H i = b i c i / N i , P i = ( a i + d i ) / N i , Q i = ( b i + c i ) / N i .</p><p>To illustrate the application of these two U-statistics, we applied and compared them to the Breslow-Day statistic and the Zelen statistic in two published data sets: 1) Alcohol assumption data (<xref ref-type="table" rid="table2">Table 2</xref>) are from Statistical Methods in Cancer Research [<xref ref-type="bibr" rid="scirp.93080-ref19">19</xref>] , volume 1 page 137, which is an example for introducing the use of Breslow-Day statistic; 2) new drug data (<xref ref-type="table" rid="table4">Table 4</xref>) are from a study of comparing a new drug with a controlled drug among 22 hospital sites [<xref ref-type="bibr" rid="scirp.93080-ref20">20</xref>] . Because of the sparseness of the data, asymptotic tests (Breslow-Day statistic) might not be able to yield an accurate p-value. The test results for the homogeneity of odds ratios are presented in <xref ref-type="table" rid="table2">Table 2</xref> and <xref ref-type="table" rid="table4">Table 4</xref>. For data set 1), according to the p-values of the test statistics, U3 rejected the null hypothesis; the WU3, Breslow-Day, and Zelen statistics accepted the null hypothesis at 0.05 level. For data set 2), Zelen and WU3 rejected the null hypothesis, yet Breslow-Day and U3 accepted the null hypothesis of no difference among odds ratios at 0.05 level (Tables 2-5).</p></sec><sec id="s6"><title>6. Discussion and Conclusions</title><p>To summarize the simulation study that we conducted, the following are some remarks: When the number of strata is not very small, (K ≥ 6), the empirical size of U3 and WU3 were very stable under various situations and stay very close to the nominal of 0.05. In terms of size and power, U3 and WU3 performed better than the Breslow-Day statistic and the Zelen’s exact test. Therefore, U3 and WU3 are considered as better statistics for testing the homogeneity of odds ratios in this situation. The test statistic U3 is recommended when the sample size is the same in each stratum, the number of strata is large and the sample size in each stratum is not large. Otherwise, WU3 is recommended.</p><p>Breslow-Day test is conservative in most situations; its empirical size is close</p><table-wrap id="table2" ><label><xref ref-type="table" rid="table2">Table 2</xref></label><caption><title> Alcohol consumption data</title></caption><table><tbody><thead><tr><th align="center" valign="middle" >Age</th><th align="center" valign="middle" ></th><th align="center" valign="middle"  colspan="2"  >Daily Alcohol Consumption</th><th align="center" valign="middle" >Odds Ratio</th></tr></thead><tr><td align="center" valign="middle" >(Years)</td><td align="center" valign="middle" ></td><td align="center" valign="middle" >80 + g</td><td align="center" valign="middle" >0 - 79 g</td><td align="center" valign="middle" ></td></tr><tr><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td></tr><tr><td align="center" valign="middle" >25 - 34</td><td align="center" valign="middle" >Case</td><td align="center" valign="middle" >1</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >33.63</td></tr><tr><td align="center" valign="middle" ></td><td align="center" valign="middle" >Control</td><td align="center" valign="middle" >9</td><td align="center" valign="middle" >106</td><td align="center" valign="middle" ></td></tr><tr><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td></tr><tr><td align="center" valign="middle" >35 - 44</td><td align="center" valign="middle" >Case</td><td align="center" valign="middle" >4</td><td align="center" valign="middle" >5</td><td align="center" valign="middle" >5.05</td></tr><tr><td align="center" valign="middle" ></td><td align="center" valign="middle" >Control</td><td align="center" valign="middle" >26</td><td align="center" valign="middle" >164</td><td align="center" valign="middle" ></td></tr><tr><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td></tr><tr><td align="center" valign="middle" >45 - 54</td><td align="center" valign="middle" >Case</td><td align="center" valign="middle" >25</td><td align="center" valign="middle" >21</td><td align="center" valign="middle" >5.67</td></tr><tr><td align="center" valign="middle" ></td><td align="center" valign="middle" >Control</td><td align="center" valign="middle" >29</td><td align="center" valign="middle" >138</td><td align="center" valign="middle" ></td></tr><tr><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td></tr><tr><td align="center" valign="middle" >55 - 64</td><td align="center" valign="middle" >Case</td><td align="center" valign="middle" >42</td><td align="center" valign="middle" >34</td><td align="center" valign="middle" >6.36</td></tr><tr><td align="center" valign="middle" ></td><td align="center" valign="middle" >Control</td><td align="center" valign="middle" >27</td><td align="center" valign="middle" >139</td><td align="center" valign="middle" ></td></tr><tr><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td></tr><tr><td align="center" valign="middle" >65 - 74</td><td align="center" valign="middle" >Case</td><td align="center" valign="middle" >19</td><td align="center" valign="middle" >26</td><td align="center" valign="middle" >2.58</td></tr><tr><td align="center" valign="middle" ></td><td align="center" valign="middle" >Control</td><td align="center" valign="middle" >18</td><td align="center" valign="middle" >88</td><td align="center" valign="middle" ></td></tr><tr><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td></tr><tr><td align="center" valign="middle" >75+</td><td align="center" valign="middle" >Case</td><td align="center" valign="middle" >5</td><td align="center" valign="middle" >8</td><td align="center" valign="middle" >40.76</td></tr><tr><td align="center" valign="middle" ></td><td align="center" valign="middle" >Control</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >31</td><td align="center" valign="middle" ></td></tr></tbody></table></table-wrap><p>Source: Statistical Methods in Cancer Research, volume 1, page 137.</p><table-wrap id="table3" ><label><xref ref-type="table" rid="table3">Table 3</xref></label><caption><title> Homogeneity odds ratio test results for alcohol consumption data</title></caption><table><tbody><thead><tr><th align="center" valign="middle" >Statistics</th><th align="center" valign="middle" >Observed value</th><th align="center" valign="middle" >Expected value</th><th align="center" valign="middle" >Variance</th><th align="center" valign="middle" >z value</th><th align="center" valign="middle" >p-value</th></tr></thead><tr><td align="center" valign="middle" >U3</td><td align="center" valign="middle" >2.24525</td><td align="center" valign="middle" >0.452702</td><td align="center" valign="middle" >0.144733</td><td align="center" valign="middle" >2.24525</td><td align="center" valign="middle" >0.000001228</td></tr><tr><td align="center" valign="middle" >WU3</td><td align="center" valign="middle" >6.25017</td><td align="center" valign="middle" >4.47819</td><td align="center" valign="middle" >8.0402</td><td align="center" valign="middle" >0.624919</td><td align="center" valign="middle" >0.26601</td></tr><tr><td align="center" valign="middle" >Breslow_Day</td><td align="center" valign="middle" >9.38159</td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" >0.0968</td></tr><tr><td align="center" valign="middle" >Zelen exact</td><td align="center" valign="middle" >0.0968552</td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" >0.0969</td></tr></tbody></table></table-wrap><table-wrap-group id="4"><label><xref ref-type="table" rid="table4">Table 4</xref></label><caption><title> New drug data</title></caption><table-wrap id="4_1"><table><tbody><thead><tr><th align="center" valign="middle" >Test Site</th><th align="center" valign="middle"  colspan="2"  >New Drug</th><th align="center" valign="middle"  colspan="2"  >Control Drug</th><th align="center" valign="middle" >Odds Ratio</th></tr></thead><tr><td align="center" valign="middle" ></td><td align="center" valign="middle" >Response</td><td align="center" valign="middle" >No</td><td align="center" valign="middle" >Response</td><td align="center" valign="middle" >No</td><td align="center" valign="middle" ></td></tr><tr><td align="center" valign="middle" >1</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >15</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >15</td><td align="center" valign="middle" >1</td></tr><tr><td align="center" valign="middle" >2</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >39</td><td align="center" valign="middle" >6</td><td align="center" valign="middle" >32</td><td align="center" valign="middle" >0.06</td></tr><tr><td align="center" valign="middle" >3</td><td align="center" valign="middle" >1</td><td align="center" valign="middle" >20</td><td align="center" valign="middle" >3</td><td align="center" valign="middle" >18</td><td align="center" valign="middle" >0.3</td></tr><tr><td align="center" valign="middle" >4</td><td align="center" valign="middle" >1</td><td align="center" valign="middle" >14</td><td align="center" valign="middle" >2</td><td align="center" valign="middle" >15</td><td align="center" valign="middle" >0.54</td></tr><tr><td align="center" valign="middle" >5</td><td align="center" valign="middle" >1</td><td align="center" valign="middle" >20</td><td align="center" valign="middle" >2</td><td align="center" valign="middle" >19</td><td align="center" valign="middle" >0.48</td></tr><tr><td align="center" valign="middle" >6</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >12</td><td align="center" valign="middle" >2</td><td align="center" valign="middle" >10</td><td align="center" valign="middle" >0.17</td></tr></tbody></table></table-wrap><table-wrap id="4_2"><table><tbody><thead><tr><th align="center" valign="middle" >7</th><th align="center" valign="middle" >3</th><th align="center" valign="middle" >49</th><th align="center" valign="middle" >10</th><th align="center" valign="middle" >42</th><th align="center" valign="middle" >0.26</th></tr></thead><tr><td align="center" valign="middle" >8</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >19</td><td align="center" valign="middle" >2</td><td align="center" valign="middle" >17</td><td align="center" valign="middle" >0.18</td></tr><tr><td align="center" valign="middle" >9</td><td align="center" valign="middle" >1</td><td align="center" valign="middle" >14</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >15</td><td align="center" valign="middle" >1.07</td></tr><tr><td align="center" valign="middle" >10</td><td align="center" valign="middle" >2</td><td align="center" valign="middle" >26</td><td align="center" valign="middle" >2</td><td align="center" valign="middle" >27</td><td align="center" valign="middle" >1.04</td></tr><tr><td align="center" valign="middle" >11</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >19</td><td align="center" valign="middle" >2</td><td align="center" valign="middle" >18</td><td align="center" valign="middle" >0.19</td></tr><tr><td align="center" valign="middle" >12</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >12</td><td align="center" valign="middle" >1</td><td align="center" valign="middle" >11</td><td align="center" valign="middle" >0.31</td></tr><tr><td align="center" valign="middle" >13</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >24</td><td align="center" valign="middle" >5</td><td align="center" valign="middle" >19</td><td align="center" valign="middle" >0.07</td></tr><tr><td align="center" valign="middle" >14</td><td align="center" valign="middle" >2</td><td align="center" valign="middle" >10</td><td align="center" valign="middle" >2</td><td align="center" valign="middle" >11</td><td align="center" valign="middle" >1.1</td></tr><tr><td align="center" valign="middle" >15</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >14</td><td align="center" valign="middle" >11</td><td align="center" valign="middle" >3</td><td align="center" valign="middle" >0.01</td></tr><tr><td align="center" valign="middle" >16</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >53</td><td align="center" valign="middle" >4</td><td align="center" valign="middle" >48</td><td align="center" valign="middle" >0.1</td></tr><tr><td align="center" valign="middle" >17</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >20</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >20</td><td align="center" valign="middle" >1</td></tr><tr><td align="center" valign="middle" >18</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >21</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >21</td><td align="center" valign="middle" >1</td></tr><tr><td align="center" valign="middle" >19</td><td align="center" valign="middle" >1</td><td align="center" valign="middle" >50</td><td align="center" valign="middle" >1</td><td align="center" valign="middle" >48</td><td align="center" valign="middle" >0.96</td></tr><tr><td align="center" valign="middle" >20</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >13</td><td align="center" valign="middle" >1</td><td align="center" valign="middle" >13</td><td align="center" valign="middle" >0.32</td></tr><tr><td align="center" valign="middle" >21</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >13</td><td align="center" valign="middle" >1</td><td align="center" valign="middle" >13</td><td align="center" valign="middle" >0.32</td></tr><tr><td align="center" valign="middle" >22</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >21</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >21</td><td align="center" valign="middle" >1</td></tr></tbody></table></table-wrap></table-wrap-group><table-wrap id="table5" ><label><xref ref-type="table" rid="table5">Table 5</xref></label><caption><title> Homogeneity odds ratio test results for new drug data</title></caption><table><tbody><thead><tr><th align="center" valign="middle" >Statistics</th><th align="center" valign="middle" >Observed value</th><th align="center" valign="middle" >Expected value</th><th align="center" valign="middle" >Variance</th><th align="center" valign="middle" >z value</th><th align="center" valign="middle" >p-value</th></tr></thead><tr><td align="center" valign="middle" >U3</td><td align="center" valign="middle" >1.3056</td><td align="center" valign="middle" >0.762194</td><td align="center" valign="middle" >0.117405</td><td align="center" valign="middle" >1.58593</td><td align="center" valign="middle" >0.0564</td></tr><tr><td align="center" valign="middle" >WU3</td><td align="center" valign="middle" >6.56955</td><td align="center" valign="middle" >4.11497</td><td align="center" valign="middle" >1.09492</td><td align="center" valign="middle" >2.34578</td><td align="center" valign="middle" >0.0095</td></tr><tr><td align="center" valign="middle" >Breslow-Day</td><td align="center" valign="middle" >25.7844</td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" >0.0785</td></tr><tr><td align="center" valign="middle" >Zelen exact</td><td align="center" valign="middle" >0.0292015</td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" >0.0292</td></tr></tbody></table></table-wrap><p>to 0.05 when the sample size is large; when sample size is small, Breslow-Day test is not recommended. Breslow-Day test is never recommended for sparse data;</p><p>When the sample size is small and the number of strata is small, say less than 5, Zelen’s exact test is recommended;</p><p>In our application, the sample mean and the variance were estimated based on certain assumptions. The empirical power and size of U3 and WU3 would be highly dependent on how well the estimator of σ 2 would be.</p></sec><sec id="s7"><title>Acknowledgements</title><p>This work was partially supported by Cancer Prevention Research Institute of Texas (RP170668).</p></sec><sec id="s8"><title>Conflicts of Interest</title><p>The authors declare no conflicts of interest regarding the publication of this paper.</p></sec><sec id="s9"><title>Cite this paper</title><p>Wei, Q. and Lai, D.J. (2019) Test for Homogeneity of Odds Ratios Using U-Statistics. Open Journal of Statistics, 9, 347-360. https://doi.org/10.4236/ojs.2019.93024</p></sec></body><back><ref-list><title>References</title><ref id="scirp.93080-ref1"><label>1</label><mixed-citation publication-type="other" xlink:type="simple">Liang, K.Y. and Self, S.G. (1985) Tests of Homogeneity of Odds Ratio When the Data Are Sparse. Biometrika, 72, 353-358. https://doi.org/10.1093/biomet/72.2.353</mixed-citation></ref><ref id="scirp.93080-ref2"><label>2</label><mixed-citation publication-type="other" xlink:type="simple">Paul, S.R. and Donner, A. (1992) Small Sample Performance of Tests of Homogeneity of Odds Ratios in k 2 × 2 Tables. Statistics in Medicine, 11, 159-165. 
https://doi.org/10.1002/sim.4780110203</mixed-citation></ref><ref id="scirp.93080-ref3"><label>3</label><mixed-citation publication-type="other" xlink:type="simple">Jones, M.P., O’Gorman, T.W., Lemke, J.H. and Woolson, R.F. (1989) A Monte Carlo Investigation of Homogeneity Tests of the Odds Ratio under Various Sample Size Configurations. Biometrics, 45, 171-181. https://doi.org/10.2307/2532043</mixed-citation></ref><ref id="scirp.93080-ref4"><label>4</label><mixed-citation publication-type="other" xlink:type="simple">Zelen, M. (1971) The Analysis of Several 2 × 2 Contingency Tables. Biometrika, 58, 129-137. https://doi.org/10.1093/biomet/58.1.129</mixed-citation></ref><ref id="scirp.93080-ref5"><label>5</label><mixed-citation publication-type="other" xlink:type="simple">Halperin, M., Ware, J.H., Byar, D.P., Mantel, M., Brown, C.C., Koziol, J., Jail, M. and Green, S.B. (1977) Testing for Interaction in an I × J × K Contingency Table. Biometrika, 64, 271-275. https://doi.org/10.1093/biomet/64.2.271</mixed-citation></ref><ref id="scirp.93080-ref6"><label>6</label><mixed-citation publication-type="other" xlink:type="simple">Thomas, D.G. (1975) Exact and Asymptotic Methods for the Combination of 2 × 2 Tables. Computers and Biomedical Research, 8, 423-446.  
https://doi.org/10.1016/0010-4809(75)90048-8</mixed-citation></ref><ref id="scirp.93080-ref7"><label>7</label><mixed-citation publication-type="other" xlink:type="simple">Pagano, M. and Tritchler, D. (1983). Algorithms for the Analysis of Several 2 × 2 Contingency Tables. SIAM Journal of Scientific and Statistical Computing, 4, 302-309. https://doi.org/10.1137/0904024</mixed-citation></ref><ref id="scirp.93080-ref8"><label>8</label><mixed-citation publication-type="other" xlink:type="simple">Metha, C.R., Patel, N.R. and Wei, L.J. (1988) Computing Exact Significance Tests with Restricted Randomization Rules. Biometrika, 75, 295-302.  
https://doi.org/10.1093/biomet/75.2.295</mixed-citation></ref><ref id="scirp.93080-ref9"><label>9</label><mixed-citation publication-type="other" xlink:type="simple">Thomas, D.G. and Gart, J.J. (1992) Improved and Extended Exact and Asymptotic Methods for the Combination of 2 × 2 Tables. Computers and Biometrical Research, 25, 75-84. https://doi.org/10.1016/0010-4809(92)90036-A</mixed-citation></ref><ref id="scirp.93080-ref10"><label>10</label><mixed-citation publication-type="other" xlink:type="simple">Hirji, K.F., Vollset, S.E., Reis, I.M. and Afifi, A.A. (1996) Exact Tests for Interaction in Several 2 × 2 Tables. Journal of Computational and Graphical Statistics, 5, 209-224. https://doi.org/10.1080/10618600.1996.10474706</mixed-citation></ref><ref id="scirp.93080-ref11"><label>11</label><mixed-citation publication-type="other" xlink:type="simple">Reis, I. M., Hirji, K.F. and Afifi, A.A. (1999) Exact and Asymptotic Tests for Homogeneity in Several 2 × 2 Tables. Statistics in Medicine, 18, 893-906.  
https://doi.org/10.1002/(SICI)1097-0258(19990430)18:8&lt;893::AID-SIM84&gt;3.0.CO;2-5</mixed-citation></ref><ref id="scirp.93080-ref12"><label>12</label><mixed-citation publication-type="other" xlink:type="simple">Bagheri, Z., Ayatollahi, S.M.T. and Jafari, P. (2011) Comparison of Three Tests of Homogeneity of Odds Ratios in Multicenter Trials with Unequal Sample Sizes within and among Centers. BMC Medical Research Methodology, 11, 58. 
https://doi.org/10.1186/1471-2288-11-58</mixed-citation></ref><ref id="scirp.93080-ref13"><label>13</label><mixed-citation publication-type="other" xlink:type="simple">Viwatwongkasem, C., Donjdee, K. and Poodphraw, T. (2018) Profile Likelihood Tests for Common Risk Ratios in Meta-Analysis Studies. Open Journal of Statistics, 8, 915-930. https://doi.org/10.4236/ojs.2018.86061</mixed-citation></ref><ref id="scirp.93080-ref14"><label>14</label><mixed-citation publication-type="other" xlink:type="simple">Halmos, P.R. (1946) The Theory of Unbiased Estimation. The Annals of Mathematical Statistics, 17, 34-43. https://doi.org/10.1214/aoms/1177731020</mixed-citation></ref><ref id="scirp.93080-ref15"><label>15</label><mixed-citation publication-type="other" xlink:type="simple">Hoeffding, W. (1948) A Class of Statistics with Asymptotically Normal Distribution. The Annals of Mathematical Statistics, 19, 293-325.  
https://doi.org/10.1214/aoms/1177730196</mixed-citation></ref><ref id="scirp.93080-ref16"><label>16</label><mixed-citation publication-type="other" xlink:type="simple">Serfling, R.J. (1980) Approximation Theorems of Mathematical Statistics. Wiley, New York. https://doi.org/10.1002/9780470316481</mixed-citation></ref><ref id="scirp.93080-ref17"><label>17</label><mixed-citation publication-type="other" xlink:type="simple">Lee, A.J. (1990) U-Statistics: Theory and Practice. Dekker, New York.</mixed-citation></ref><ref id="scirp.93080-ref18"><label>18</label><mixed-citation publication-type="other" xlink:type="simple">Gart, J.J. and Zweifel, J.R. (1967) On the Bias of Various Estimators of the Logit and Its Variance, with Application to Quantal Bioassay. Biometrika, 54, 181-187.  
https://doi.org/10.1093/biomet/54.1-2.181</mixed-citation></ref><ref id="scirp.93080-ref19"><label>19</label><mixed-citation publication-type="other" xlink:type="simple">Breslow, N.E. and Day, N.E. (1980) Statistical Methods in Cancer Research. Vol. 1—The Analysis of Case-Control Studies. International Agency for Research on Cancer (IARC). Lyon, France, 136-146.</mixed-citation></ref><ref id="scirp.93080-ref20"><label>20</label><mixed-citation publication-type="other" xlink:type="simple">Mehta, C. and Patel, N. (2002) StatXact 4 for Windows User Manual. Cytel Software, Cambridge, MA, 486-490.</mixed-citation></ref></ref-list></back></article>