<?xml version="1.0" encoding="UTF-8"?><!DOCTYPE article  PUBLIC "-//NLM//DTD Journal Publishing DTD v3.0 20080202//EN" "http://dtd.nlm.nih.gov/publishing/3.0/journalpublishing3.dtd"><article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" dtd-version="3.0" xml:lang="en" article-type="research article"><front><journal-meta><journal-id journal-id-type="publisher-id">OJS</journal-id><journal-title-group><journal-title>Open Journal of Statistics</journal-title></journal-title-group><issn pub-type="epub">2161-718X</issn><publisher><publisher-name>Scientific Research Publishing</publisher-name></publisher></journal-meta><article-meta><article-id pub-id-type="doi">10.4236/ojs.2017.75053</article-id><article-id pub-id-type="publisher-id">OJS-79361</article-id><article-categories><subj-group subj-group-type="heading"><subject>Articles</subject></subj-group><subj-group subj-group-type="Discipline-v2"><subject>Physics&amp;Mathematics</subject></subj-group></article-categories><title-group><article-title>
 
 
  Estimating the Empirical Null Distribution of Maxmean Statistics in Gene Set Analysis
 
</article-title></title-group><contrib-group><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Xing</surname><given-names>Ren</given-names></name><xref ref-type="aff" rid="aff1"><sup>1</sup></xref></contrib><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Jianmin</surname><given-names>Wang</given-names></name><xref ref-type="aff" rid="aff2"><sup>2</sup></xref></contrib><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Song</surname><given-names>Liu</given-names></name><xref ref-type="aff" rid="aff2"><sup>2</sup></xref></contrib><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Jeffrey</surname><given-names>C. Miecznikowski</given-names></name><xref ref-type="aff" rid="aff1"><sup>1</sup></xref><xref ref-type="corresp" rid="cor1"><sup>*</sup></xref></contrib></contrib-group><aff id="aff1"><addr-line>Department of Biostatistics, University at Buffalo, Buffalo, USA</addr-line></aff><aff id="aff2"><addr-line>Department of Biostatistics and Bioinformatics, Roswell Park Cancer Institute, Buffalo, USA</addr-line></aff><author-notes><corresp id="cor1">* E-mail:<email>jcm38@buffalo.edu(JCM)</email>;</corresp></author-notes><pub-date pub-type="epub"><day>26</day><month>09</month><year>2017</year></pub-date><volume>07</volume><issue>05</issue><fpage>761</fpage><lpage>767</lpage><history><date date-type="received"><day>22,</day>	<month>May</month>	<year>2017</year></date><date date-type="rev-recd"><day>22,</day>	<month>September</month>	<year>2017</year>	</date><date date-type="accepted"><day>27,</day>	<month>September</month>	<year>2017</year></date></history><permissions><copyright-statement>&#169; Copyright  2014 by authors and Scientific Research Publishing Inc. </copyright-statement><copyright-year>2014</copyright-year><license><license-p>This work is licensed under the Creative Commons Attribution International License (CC BY). http://creativecommons.org/licenses/by/4.0/</license-p></license></permissions><abstract><p>
 
 
  Gene Set Analysis (GSA) is a framework for testing the association of a set of genes and the outcome, e.g. disease status or treatment group. The method replies on computing a maxmean statistic and estimating the null distribution of the maxmean statistics via a restandardization procedure. In practice, the pre-determined gene sets have stronger intra-correlation than genes across sets. This may result in biases in the estimated null distribution. We derive an asymptotic null distribution of the maxmean statistics based on sparsity assumption. We propose a flexible two group mixture model for the maxmean statistics. The mixture model allows us to estimate the null parameters empirically via maximum likelihood approach. Our empirical method is compared with the restandardization procedure of GSA in simulations. We show that our method is more accurate in null density estimation when the genes are strongly correlated within gene sets.
 
</p></abstract><kwd-group><kwd>Gene Set Analysis</kwd><kwd> Maxmean</kwd><kwd> Empirical Null</kwd><kwd> Mixture Model</kwd></kwd-group></article-meta></front><body><sec id="s1"><title>1. Introduction</title><p>A gene pathway commonly refers to a set of genes that share a particular property, carry out a biological function or lead to a certain product in cells/tissues. Performing differential expression (DE) analysis on such gene sets aggregates the signal of individual genes and potentially increases the power of a hypothesis test. Gene set analysis also provides comprehensive understanding of the biological activities associated with the outcome phenotype and may shed light on treatment of disease.</p><p>A variety of tools are available for gene set analysis. These methods can be roughly classified into two broad categories, self-contained and competitive [<xref ref-type="bibr" rid="scirp.79361-ref1">1</xref>] . The self-contained methods test the association between the phenotype and the gene set while ignoring the other genes. Competitive methods overcome this limitation by taking into account other genes when evaluating the association. A common approach is to compute the gene level statistics, and then aggregate them into a gene-set level summary statistic. Of the many competitive methods gene set enrichment analysis (GSEA) [<xref ref-type="bibr" rid="scirp.79361-ref2">2</xref>] and (GSA) [<xref ref-type="bibr" rid="scirp.79361-ref3">3</xref>] are two representative algorithms. In GSEA the Kolmogorov-Smirnov (KS) statistic is employed. GSA improves the power of GSEA by using a more powerful maxmean statistic. First the gene level z statistics are calculated, z i , i = 1 , ⋯ , n . Let S denote the indices of the gene set and n<sub>S</sub> be the size of S. The maxmean statistic S is defined as,</p><p>S + = 1 n S ∑ i ∈ S z i I { z i &gt; 0 } , S − = − 1 n S ∑ i ∈ S z i I { z i &lt; 0 } , S = max ( S + , S − ) , (1.1)</p><p>where I {   } is the indicator function.</p><p>The GSA method estimates the null distribution of S through a restandardization procedure, which is a combination of row (gene) randomization and column (sample label) permutation. It compares the gene set against its permutations and also takes into account the overall distribution of randomly selected null sets. The permutation maintains the correlation structure in the gene set, while row randomization rescales and shifts the permutations to include the competition of genes outside the set.</p><p>In practice, the gene sets for analysis are obtained from a public database e.g. Kyoto Encyclopedia of Genes and Genomes (KEGG) [<xref ref-type="bibr" rid="scirp.79361-ref4">4</xref>] , MSigDB [<xref ref-type="bibr" rid="scirp.79361-ref2">2</xref>] or gene ontology (GO) [<xref ref-type="bibr" rid="scirp.79361-ref5">5</xref>] . It is reasonable to expect that genes in the same set have stronger correlation than genes across different sets. In such cases there may be a mismatch from the null distribution estimated from restandardization and the true null, which leads to biased inference.</p><p>We propose a two group mixture model for the gene set maxmean statistics. Our model assumes that only a small proportion of the gene sets are truly significantly DE. Based on the mixture model, we apply the maximum likelihood method to estimate the empirical null. The empirical method improves the accuracy of GSA in large scale hypothesis testing. It also reduces the computational burden of the permutation steps in restandardization procedure of GSA. The analysis is demonstrated in simulation studies and a data set from the MSigDB database [<xref ref-type="bibr" rid="scirp.79361-ref2">2</xref>] .</p></sec><sec id="s2"><title>2. Methods</title><p>The maxmean statistic in GSA is essentially a maximum statistic of two correlated sample means, S<sub>+</sub> and S<sub>−</sub> defined in (1.1), which asymptotically follow bivariate normal distribution for adequately large n<sub>S</sub>,</p><p>( S + S − ) ~ N 2 { ( μ + μ − ) , ( σ + 2 ρ σ + σ − ρ σ + σ − σ − 2 ) } , (2.1)</p><p>where μ + = E ( S + ) , σ + 2 = Var ( S + ) , μ − = E ( S − ) , σ − 2 = Var ( S − ) , and ρ = corr ( S + , S − ) . Based on the work in [<xref ref-type="bibr" rid="scirp.79361-ref6">6</xref>] [<xref ref-type="bibr" rid="scirp.79361-ref7">7</xref>] , an asymptotic distribution of the maxmean statistics under the Lindeberg condition (see [<xref ref-type="bibr" rid="scirp.79361-ref8">8</xref>] ) is</p><p>f 0 ( s ) = 1 σ + ϕ ( − s + μ + σ + ) &#215; Φ ( ρ ( − s + μ + ) σ + 1 − ρ 2 − − s + μ − σ − 1 − ρ 2 )     + 1 σ − ϕ ( − s + μ − σ − ) &#215; Φ ( ρ ( − s + μ − ) σ − 1 − ρ 2 − − s + μ + σ + 1 − ρ 2 ) (2.2)</p><p>where ϕ and Φ are the probability density function (pdf) and cumulative density function (cdf) of the standard normal distribution.</p><p>We can estimate the parameters in f<sub>0</sub> by fitting the null gene sets. A special case would be that z i ∼ N ( 0 , σ 2 ) independently. We can easily compute the parameters:</p><p>μ + = μ − = 0.40 σ ,   σ + 2 = σ − 2 = 0.34 σ 2 ,   ρ = − 0.467.</p><p>However, genes in the same set are often correlated. Therefore, the theoretically computed parameters may not match the actual null distribution well. We propose an empirical method to estimate the null distribution of S. Let f be the density function of the maxmean statistics for all the gene sets. Adopting the two group mixture model [<xref ref-type="bibr" rid="scirp.79361-ref9">9</xref>] [<xref ref-type="bibr" rid="scirp.79361-ref10">10</xref>] , we specify a similar model that f is comprised by a large proportion (p<sub>0</sub>) of the null density f<sub>0</sub> and a small proportion of the non-null density f<sub>1</sub>,</p><p>f ( s ) = p 0 f 0 ( s ) + p 1 f 1 ( s ) . (2.3)</p><p>The null density f<sub>0</sub> is assumed to have the form in (2.2), in which the parameters (&#181;<sub>+</sub>, &#181;<sub>−</sub>, σ<sub>+</sub>, σ<sub>−</sub>, ρ) need to be estimated. For identifiability of f<sub>0</sub> we further assume the non-null density f 1 ( s ) ≈ 0   ∀ s ∈ A 0 for some interval A<sub>0</sub>. Under this assumption, S follows a truncated distribution f<sub>T</sub>(s) for s ∈ A<sub>0</sub>,</p><p>f T ( s ) ≈ p 0 f 0 ( s ) ∫ A 0 p 0 f 0 ( s ) = f 0 ( s ) ∫ A 0 f 0 ( s ) . (2.4)</p><p>Fitting the maxmean statistics in A<sub>0</sub> to (2.4) by maximum likelihood yields f ^ 0 . In addition, p<sub>0</sub> can be estimated by</p><p>p ^ 0 = # { s i ∈ A 0 } n ∫ A 0 f ^ 0 ( s ) . (2.5)</p><p>In this paper we let A<sub>0</sub> be the interval (0, q<sub>S</sub>) where q<sub>S</sub> is 90% quantile of the maxmean statistics of all gene sets.</p></sec><sec id="s3"><title>3. Simulations</title><p>We simulate 2000 gene sets, 5% of them are DE sets and the other 95% are null sets. All sets contain n<sub>S</sub> = 50 genes. Let C<sub>1</sub> and C<sub>2</sub> be the sample indices of two conditions and each condition has m = 10 samples. For gene i in sample j, the expression data x<sub>ij</sub> is generated in a hierarchical fashion,</p><p>x 0 k j ~ N ( δ k I { j ∈ C 1 } , τ k 2 ) , x i j ~ N ( α i x 0 k j , σ i 2 )   ∀ i ∈ S k ,</p><p>where τ k = σ i = 1 , δ k = 1 for DE sets and δ<sub>k</sub> = 0 for null nets. x<sub>0kj</sub> is the expression of the hub gene [<xref ref-type="bibr" rid="scirp.79361-ref11">11</xref>] of set S<sub>k</sub> and all genes in the same set are correlated with the hub gene. The DE genes of set S<sub>k</sub> are jointly controlled by δ<sub>k</sub> and α<sub>i</sub>. In particular, δ<sub>k</sub> controls the differential expression of the hub gene between two conditions. The parameter α<sub>i</sub> controls the inter-correlation within the set. When α<sub>i</sub> =0 genes are independent within gene sets. For null sets we let α<sub>i</sub> = 0 (independent) or &#177;0.2 (correlated). For DE sets α<sub>i</sub> = &#177;1 so that the correlation is stronger than the null sets. <xref ref-type="fig" rid="fig1">Figure 1</xref> shows the f<sub>0</sub> density estimate by our empirical method and the restandardization procedure in GSA.</p><p>When genes within the gene sets are independent, the null distribution obtained by the two methods are very close (<xref ref-type="fig" rid="fig1">Figure 1</xref> left). As 95% of the gene sets are null sets, the estimated distributions match the majority of the sets. When genes are correlated, the null distribution obtained by GSA shows a mismatch from the null gene sets, while our empirical method maintains a good fit to the data (<xref ref-type="fig" rid="fig1">Figure 1</xref> right).</p></sec><sec id="s4"><title>4. Application</title><p>We evaluate our empirical method and GSA with the 4722 curated gene sets in MSigDB database using the gender dataset included in the GSEA package [<xref ref-type="bibr" rid="scirp.79361-ref2">2</xref>] . The minimum size of the gene sets is 5, the maximum size is 1972 and median size is 39. The dataset contains transcriptional profiles of 15,056 genes in 17 female and 15 male lymphoblastoid cell line samples. There are 302 gene sets with unadjusted p-value &lt; 0.05 by our empirical method and 507 gene sets by GSA. Controlling the false discovery rate (FDR) at 0.1 with Benjamini-Hochberg pro-</p><p>cedure [<xref ref-type="bibr" rid="scirp.79361-ref12">12</xref>] , the empirical method identifies 39 significant sets while the restandardization identifies 8 significant sets. <xref ref-type="fig" rid="fig2">Figure 2</xref> shows the null distribution obtained by the empirical method and GSA.</p></sec><sec id="s5"><title>5. Discussion</title><p>GSA is a representative tool for gene set DE analysis. Compared with the state- of-the-art method GSEA, the maxmean statistic used in GSA is more powerful than the KS statistic in GSEA [<xref ref-type="bibr" rid="scirp.79361-ref3">3</xref>] . GSA also establishes a restandardization algorithm to assess the maxmean statistic. Due to the correlation of the genes in pre-defined gene sets or pathways, the null distribution obtained by the restandardization procedure in GSA may not match the true null distribution well.</p><p>In studying the GSA method, we propose a new method to estimate the null distribution of the gene set maxmean statistics. Unlike the permutation test of GSA in which every gene set is compared against its own permutations, the fundamental difference of our method is that it estimates an overall null distribution f<sub>0</sub> parametrically and all gene sets are compared against f<sub>0</sub>. The possibility of parametric estimation of f<sub>0</sub> is rendered by the large number of gene sets in the hypothesis testing. A similar idea is proposed in [<xref ref-type="bibr" rid="scirp.79361-ref9">9</xref>] for large-scale test of DE genes. We extend the idea to the test of gene sets when a large number of sets are available.</p><p>Our method is based on the sparsity assumption that only a small proportion of the gene sets are truly DE. Further, we adopt a two group parametric mixture distribution to model the gene set maxmean statistics. The parameters of the null distribution is estimated under the sparsity assumption. We show that our new method provides more accurate estimation of the null distribution. It also avoids the computational intensity in the permutation steps of GSA.</p><p>In the simulations, we compare the two methods under independence and correlation. When the intra-set correlation is greater than the cross-set correlation, the GSA shows a mismatch to the true null. The reason is that the randomization step samples genes in the entire dataset, which has a different correlation structure from gene sets. As a result the mean and standard deviation obtained in the randomization step is inaccurate.</p><p>In application to the gender data set, our method has fewer gene sets with unadjusted p-value &lt; 0.05, but identifies more gene sets after controlling for FDR. A reason is that in GSA, gene-set p-values are limited by the number of permutations. Due to the large number of genes and gene sets, with 10,000 permutations, GSA takes 64 minutes on an Intel i7 3770K 3.5 GHz CPU. In contrast, our empirical method takes less than 1 minute.</p><p>An important aspect of our empirical estimation method is specifying an appropriate range for the zero assumption region A<sub>0</sub>. In this paper we arbitrarily choose A<sub>0</sub> to be the interval (0, q<sub>S</sub>), where q<sub>S</sub> is some quantile of the observed maxmean statistics. Determination of q<sub>S</sub> is a trade-off between variance and bias. With small q<sub>S</sub>, the bias of the null estimate is small but the variance is large and vice-versa. How to determine an optimal A<sub>0</sub> interval requires further exploration.</p></sec><sec id="s6"><title>Cite this paper</title><p>Ren, X., Wang, J.M., Liu, S. and Miecznikowski, J.C. (2017) Estimating the Empirical Null Distribution of Maxmean Statistics in Gene Set Analysis. Open Journal of Statistics, 7, 761-767. https://doi.org/10.4236/ojs.2017.75053</p></sec></body><back><ref-list><title>References</title><ref id="scirp.79361-ref1"><label>1</label><mixed-citation publication-type="other" xlink:type="simple">Nam, D. and Kim, S.-Y. (2008) Gene-Set Approach for Expression Pattern Analysis. Briefings in Bioinformatics, 9, 189-197. https://doi.org/10.1093/bib/bbn001</mixed-citation></ref><ref id="scirp.79361-ref2"><label>2</label><mixed-citation publication-type="other" xlink:type="simple">Subramanian, A., Tamayo, P., Mootha, V.K., Mukherjee, S., Ebert, B.L., Gillette, M.A., Paulovich, A., Pomeroy, S.L., Golub, T.R., Lander, E.S., et al. (2005) Gene Set Enrichment Analysis: A Knowledge-Based Approach for Interpreting Genome-Wide Expression Profiles. Proceedings of the National Academy of Sciences of the United States of America, 102, 15545-15550.  
https://doi.org/10.1073/pnas.0506580102</mixed-citation></ref><ref id="scirp.79361-ref3"><label>3</label><mixed-citation publication-type="other" xlink:type="simple">Efron, B. and Tibshirani, R. (2007) On Testing the Significance of Sets of Genes. The Annals of Applied Statistics, 1, 107-129. https://doi.org/10.1214/07-AOAS101</mixed-citation></ref><ref id="scirp.79361-ref4"><label>4</label><mixed-citation publication-type="other" xlink:type="simple">Kanehisa, M. and Goto, S. (2000) KEGG: Kyoto Encyclopedia of Genes and Genomes. Nucleic Acids Research, 28, 27-30. https://doi.org/10.1093/nar/28.1.27</mixed-citation></ref><ref id="scirp.79361-ref5"><label>5</label><mixed-citation publication-type="other" xlink:type="simple">Ashburner, M., Ball, C.A., Blake, J.A., Botstein, D., Butler, H., Cherry, J.M., Davis, A.P., Dolinski, K., Dwight, S.S., Eppig, J.T., et al. (2000) Gene Ontology: Tool for the Unification of Biology. Nature Genetics, 25, 25-29.  
https://doi.org/10.1038/75556</mixed-citation></ref><ref id="scirp.79361-ref6"><label>6</label><mixed-citation publication-type="other" xlink:type="simple">Basu, A. and Ghosh, J. (1978) Identifiability of the Multinormal and Other Distributions under Competing Risks Model. Journal of Multivariate Analysis, 8, 413-429. https://doi.org/10.1016/0047-259X(78)90064-7</mixed-citation></ref><ref id="scirp.79361-ref7"><label>7</label><mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Cain</surname><given-names> M. </given-names></name>,<etal>et al</etal>. (<year>1994</year>)<article-title>The Moment-Generating Function of the Minimum of Bivariate Normal Random Variables</article-title><source> The American Statistician</source><volume> 48</volume>,<fpage> 124</fpage>-<lpage>125</lpage>.<pub-id pub-id-type="doi"></pub-id></mixed-citation></ref><ref id="scirp.79361-ref8"><label>8</label><mixed-citation publication-type="other" xlink:type="simple">Billingsley, P. (1995) Probability and Measure. 3rd Edition, Wiley Series in Probability and Mathematical Statistics.</mixed-citation></ref><ref id="scirp.79361-ref9"><label>9</label><mixed-citation publication-type="other" xlink:type="simple">Efron, B. (2004) Large-Scale Simultaneous Hypothesis Testing. Journal of the American Statistical Association, 99, 96-104.  
https://doi.org/10.1198/016214504000000089</mixed-citation></ref><ref id="scirp.79361-ref10"><label>10</label><mixed-citation publication-type="other" xlink:type="simple">Efron, B. (2012) Large-Scale Inference: Empirical Bayes Methods for Estimation, Testing, and Prediction, Volume 1. Cambridge University Press.</mixed-citation></ref><ref id="scirp.79361-ref11"><label>11</label><mixed-citation publication-type="other" xlink:type="simple">Langfelder, P. and Horvath, S. (2008) WGCNA: An R Package for Weighted Correlation Network Analysis. BMC Bioinformatics, 9, 1-13.  
https://doi.org/10.1186/1471-2105-9-559</mixed-citation></ref><ref id="scirp.79361-ref12"><label>12</label><mixed-citation publication-type="other" xlink:type="simple">Benjamini, Y. and Hochberg, Y. (1995) Controlling the False Discovery Rate: A Practical and Powerful Approach to Multiple Testing. Journal of the Royal Statistical Society. Series B (Methodological), 57, 289-300.</mixed-citation></ref></ref-list></back></article>