<?xml version="1.0" encoding="UTF-8"?><!DOCTYPE article  PUBLIC "-//NLM//DTD Journal Publishing DTD v3.0 20080202//EN" "http://dtd.nlm.nih.gov/publishing/3.0/journalpublishing3.dtd"><article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" dtd-version="3.0" xml:lang="en" article-type="research article"><front><journal-meta><journal-id journal-id-type="publisher-id">OJS</journal-id><journal-title-group><journal-title>Open Journal of Statistics</journal-title></journal-title-group><issn pub-type="epub">2161-718X</issn><publisher><publisher-name>Scientific Research Publishing</publisher-name></publisher></journal-meta><article-meta><article-id pub-id-type="doi">10.4236/ojs.2017.71005</article-id><article-id pub-id-type="publisher-id">OJS-74217</article-id><article-categories><subj-group subj-group-type="heading"><subject>Articles</subject></subj-group><subj-group subj-group-type="Discipline-v2"><subject>Physics&amp;Mathematics</subject></subj-group></article-categories><title-group><article-title>
 
 
  Microarray Analysis Using Rank Order Statistics for ARCH Residual Empirical Process
 
</article-title></title-group><contrib-group><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Hiroko</surname><given-names>Kato Solvang</given-names></name><xref ref-type="aff" rid="aff1"><sup>1</sup></xref><xref ref-type="corresp" rid="cor1"><sup>*</sup></xref></contrib><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Masanobu</surname><given-names>Taniguchi</given-names></name><xref ref-type="aff" rid="aff2"><sup>2</sup></xref></contrib></contrib-group><aff id="aff2"><addr-line>Deparment of Applied Mathematics, Waseda University, Tokyo, Japan</addr-line></aff><aff id="aff1"><addr-line>Marine Mammals Research Group, Institute of Marine Research, Bergen, Norway</addr-line></aff><author-notes><corresp id="cor1">* E-mail:<email>hiroko.solvang@imr.no(HKS)</email>;</corresp></author-notes><pub-date pub-type="epub"><day>09</day><month>02</month><year>2017</year></pub-date><volume>07</volume><issue>01</issue><fpage>54</fpage><lpage>71</lpage><history><date date-type="received"><day>20,</day>	<month>October</month>	<year>2016</year></date><date date-type="rev-recd"><day>17,</day>	<month>February</month>	<year>2017</year>	</date><date date-type="accepted"><day>20,</day>	<month>February</month>	<year>2017</year></date></history><permissions><copyright-statement>&#169; Copyright  2014 by authors and Scientific Research Publishing Inc. </copyright-statement><copyright-year>2014</copyright-year><license><license-p>This work is licensed under the Creative Commons Attribution International License (CC BY). http://creativecommons.org/licenses/by/4.0/</license-p></license></permissions><abstract><p>
 
 
  Statistical two-group comparisons are widely used to identify the significant differentially expressed (DE) signatures against a therapy response for microarray data analysis. We applied a rank order statistics based on an Autoregressive Conditional Heteroskedasticity (ARCH) residual empirical process to DE analysis. This approach was considered for simulation data and publicly available datasets, and was compared with two-group comparison by original data and Auto-regressive (AR) residual. The significant DE genes by the ARCH and AR residuals were reduced by about 20% - 30% to these genes by the original data. Almost 100% of the genes by ARCH are covered by the genes by the original data unlike the genes by AR residuals. GO enrichment and Pathway analyses indicate the consistent biological characteristics between genes by ARCH residuals and original data. ARCH residuals array data might contribute to refining the number of significant DE genes to detect the biological feature as well as ordinal microarray data.
 
</p></abstract><kwd-group><kwd>Time Series Model</kwd><kwd> ARCH</kwd><kwd> Wilcoxon Statistic</kwd><kwd> Volatility</kwd><kwd> Deferentially Expressed Gene Signatures</kwd><kwd> Two-Group Comparison</kwd><kwd> Breast Cancer GEO</kwd><kwd> Genome-Wide Expression Profiling</kwd><kwd> GO Analysis</kwd></kwd-group></article-meta></front><body><sec id="s1"><title>1. Introduction</title><p>Microarray technology provides a high-throughput way to simultaneously investigate gene expression information in a whole genome level. In the field of cancer research, the genome-wide expression profiling of tumors has become an important tool to identify gene sets and signatures that can be used as clinical endpoints, such as survival and therapy response [<xref ref-type="bibr" rid="scirp.74217-ref1">1</xref>] . When we are contrasting expressions between different groups or conditions (i.e., the response ispolytumous), such important genes are described as differentially expressed (DE) [<xref ref-type="bibr" rid="scirp.74217-ref2">2</xref>] . To identify important genes, a statistical scheme is required that measures and captures evidence for a DE per gene. If the response consists of binary data, the DE is measured using a two-group comparison for which such statistical methods as t-statistics, the statistical analysis of microarray (SAM) [<xref ref-type="bibr" rid="scirp.74217-ref3">3</xref>] ), fold change, and B statistics have been proposed [<xref ref-type="bibr" rid="scirp.74217-ref4">4</xref>] . The p-value of the statistics is calculated to assess the significance of the DE genes. The p-value per gene is ranked in ascending order; however, selecting significant genes must be considered by multiple testing corrections, e.g., false discovery rate (FDR) [<xref ref-type="bibr" rid="scirp.74217-ref5">5</xref>] , to avoid type I errors. Even if significant DE genes are identified by the FDR procedure, the gene list may still include too many to apply a statistical test for a substantial number of probes through whole genomic locations. Such a long list of significant DE genes complicates capturing gene signatures that should provide the availability of robust clinical and pathological prognostic and predictive factors to guide patient decision-making and the selection of treatment options.</p><p>As one approach for this challenge, based on the residuals from the Autoregressive Conditional Heteroskedasticity (ARCH) models, the proposed rank order statistic for two-sample problems pertaining to empirical processes refines the significant DE gene list. The ARCH process was proposed by Engle [<xref ref-type="bibr" rid="scirp.74217-ref6">6</xref>] , and the model was developed in much research to investigate a daily return series from finance domains. The series indicate time-inhomogeneous fluctuations and sudden changes of variance called volatility in finance. Financial analysts have attempted more suitable time series modeling for estimating this volatility. Chandra and Taniguchi [<xref ref-type="bibr" rid="scirp.74217-ref7">7</xref>] proposed a rank-order statistics and the theory provided an idea for applying residuals from two classes of ARCH models to test the innovation distributions of two financial returns generated by such varied mechanisms as different countries and/or industries. Empirical residuals called “innovation” generally perturb systems behind data. Theories of innovation approaches to time series analysis have historically been closely related to the idea of predicting dynamic phenomena from time series observations. Wiener’s theory is a well-known example that deems prediction error to be a source of information for improving the predictions of future phenomena. In a sense, innovation is a more positive label than prediction error [<xref ref-type="bibr" rid="scirp.74217-ref8">8</xref>] . As we see in innovation distribution for ARCH processes, it resembles the sequential expression level based on the whole genomic location. For applying time indices of ARCH model to the genomic location, the time series mining has been practically used to DNA sequence data analysis [<xref ref-type="bibr" rid="scirp.74217-ref9">9</xref>] and microarray data analysis [<xref ref-type="bibr" rid="scirp.74217-ref10">10</xref>] . To investigate the data’s properties, we believe that innovation analysis is more effective than analysis just based on the original data. While the original idea in Chandra and Taniguchi [<xref ref-type="bibr" rid="scirp.74217-ref7">7</xref>] was based on squared residuals from an ARCH model, not-squared empirical residuals are also theoretically applicable, as introduced in Lee and Taniguchi [<xref ref-type="bibr" rid="scirp.74217-ref11">11</xref>] . In this article, we apply this idea to test DEs between two sample groups in microarray datasets that we assume to be generated by different biological conditions.</p><p>To investigate whether ARCH residuals can consistently refine a list of significant DE genes, we apply publicly available datasets called Affy947 [<xref ref-type="bibr" rid="scirp.74217-ref12">12</xref>] for breast cancer research to compare significant gene signatures. As a statistical test for two-group comparisons, the estrogen receptor (ER) is applied in clinical outcomes to identify prognostic gene expression signatures. Estrogen is an important regulator of the development, the growth, and the differentiation of normal mammary glands. It is well documented that endogenous estrogen plays a major role in the development and progression of breast cancer. ER expression in breast tumors is frequently used to group breast cancer patients in clinical settings, both as a prognostic indicator and to predict the likelihood of response to treatment with antiestrogen [<xref ref-type="bibr" rid="scirp.74217-ref13">13</xref>] . If the cancer is ER+, hormone therapy using medication slows or stops the growth of breast cancer cells. If the cancer is ER-, then hormonal therapy is unlikely to succeed. Based on these two categorical factors for ER status, we applied our proposed statistical test to the expression levels for each genomic location. After identifying significant DE genes, biological enrichment analyses use the gene list and seek biological processes and interconnected pathways. These analyses support the consistency for refined gene lists obtained by ARCH residuals.</p></sec><sec id="s2"><title>2. Method</title><p>Denote the sample and the genomic location by i and j in microarray data x i j . The samples for the microarray data are divided by two biological different groups, one group is for breast cancer tumors driven by ER + and another group is for breast cancer tumors driven by ER−. We apply the two-group comparison testing to identify significant different expression level between two groups of ER+ and ER− samples for each gene (genomic location). As the statistical test, we propose the rank order statistics for ARCH residual empirical process introduced in 2.1. For comparisons with the ARCH model’s performance, we consider applying the two-group comparison testing to original array data and applying the test to the residuals obtained by ordinal AR (autoregressive) model. The details about both methods are summarized in 2.2. For the obtained significant DE gene lists, biologists or medical scientists require further analysis for their biological interpretation to investigate the biological process or biological network. In this article, we apply GO (gene ontology) analysis shown in 2.3 and Pathway analysis shown in 2.4, which methods are generally used to investigate specific genes or relationships among gene groups.</p><sec id="s2_1"><title>2.1. The Rank Order Statistic for ARCH Residual Empirical Process</title><p>Suppose that a classes of ARCH (p) models is generated by the following equations</p><p>X t = { σ t ( θ X ) ε t , σ t 2 ( θ X ) = θ X 0 + ∑ i = 1 p X θ X i X t − i 2 for t = 1 , ⋯ , m 0 , for t = − p X + 1 , ⋯ , 0 (2.1.1)</p><p>where { ε t } is a sequence of i.i.d.(0,1) random variables with fourth-order cumulant κ 4 X , θ X = ( θ X 0 , θ X 1 , ⋯ , θ X p X ) ′ ∈ Θ X ⊂ ℝ p X + 1 is an unknown parameter vector satisfying θ X 0 &gt; 0 , θ X i ≥ 0 , i = 1 , ⋯ , p X − 1 , θ X p X &gt; 0 , and ε t is independent of X s , s &lt; t . Denote by F ( x ) the distribution function of ε t 2 and we assume that f ( x ) = F ′ ( x ) exists and is continuous on ( 0 , ∞ ) .</p><p>Suppose that another class of ARCH(p) models, independent of { X t } , is generated similarly by the equations</p><p>Y t = { σ t ( θ Y ) ξ t , σ t 2 ( θ Y ) = θ Y 0 + ∑ i = 1 p Y θ Y i Y t − i 2 for t = 1 , ⋯ , m 0 , for t = − p Y + 1 , ⋯ , 0 (2.1.2)</p><p>where { ξ t } is a sequence of i.i.d. (0,1) random variables with fourth-order cumulant κ 4 Y , θ Y = ( θ Y 0 , θ Y 1 , ⋯ , θ Y p Y ) ′ ∈ Θ Y ⊂ ℝ p Y + 1 is an unknown parameter vector satisfying θ Y 0 &gt; 0 , θ Y i ≥ 0 , i = 1 , ⋯ , p Y − 1 , θ Y p Y &gt; 0 , and ξ t is independent of Y s , s &lt; t . The distribution function of ξ t 2 is denoted by G ( x ) and we assume that g ( x ) = G ′ ( x ) exists and is continuous on ( 0 , ∞ ) . For (2.1.1) and (2.1.2), we assume that θ X 1 + ⋯ + θ X p X &lt; 1 and θ Y 1 + ⋯ + θ Y p Y &lt; 1 for stationarity (see [<xref ref-type="bibr" rid="scirp.74217-ref14">14</xref>] ).</p><p>Now we are interested in the two-sample problem of testing</p><p>H 0 : F ( x ) = G ( x ) forall x agains t H A : F ( x ) ≠ G ( x ) forsome x .</p><p>In this article, F ( x ) and G ( x ) correspond to the distribution for the expression data of samples driven by ER+ and ER−, individually.</p><p>For this testing problem, we consider a class of rank order statistics including, such as Wilcoxon’s two-sample test. The form is derived from the empirical residuals ε ^ t 2 = X t 2 / σ t 2 ( θ ^ X ) , t = 1 , ⋯ , n and ξ ^ t 2 = Y t 2 / σ t 2 ( θ ^ Y ) , t = 1 , ⋯ , n . Because Lee and Taniguchi [<xref ref-type="bibr" rid="scirp.74217-ref11">11</xref>] developed the asymptotic theory for not squared empirical residuals, we may apply the results to ε ^ t and ξ ^ t .</p></sec><sec id="s2_2"><title>2.2. Two-Group Comparison for Microarray Data</title><p>To obtain the empirical residuals as mentioned in 2.1, the ARCH model is applied to a vector { x i 1 , x i 2 , ⋯ , x i L } for the i th sample, where L is the total number of genomic locations in the microarray data. Assuming that the ER+ and ER− samples correspond to distributions F ( x ) and G ( x ) as shown in 2.1, orders p X and p Y of the ARCH model are identified by model selection using the Akaike Information Criterion (AIC), where the model with the minimum AIC is defined as the best fit model [<xref ref-type="bibr" rid="scirp.74217-ref15">15</xref>] (see 1. in <xref ref-type="fig" rid="fig1">Figure 1</xref>). According to those responses, the empirical residuals are grouped as ε i j + and ξ i j − . Wilcoxon statistic is applied as order statistic to those two groups for each genomic location j , and the p-value is calculated (see 2. in <xref ref-type="fig" rid="fig1">Figure 1</xref>). The p-values obtained for all genes are adjusted for multiple testing corrections using false discovery rate (FDR) [<xref ref-type="bibr" rid="scirp.74217-ref5">5</xref>] (see 3. in <xref ref-type="fig" rid="fig1">Figure 1</xref>).</p><p>For comparisons with the ARCH model’s performance, the two-group comparison testing to original array data and applying the test to the residuals obtained by ordinal AR (autoregressive) model.AR model represents the current</p><p>value using the weighted average of the past values as x i j = ∑ k = 1 K β i x i j − k + w i j , where β i , k , and w i j are the AR coefficient, the AR order, and the error terms. The AR model is widely applied in time series analysis and the signal processing of economics, engineering, and science. In this article, we apply it to the expression data for two ER+ and ER− groups. The AR order for the best fit model is identified by AIC. Empirical residuals w i j + for ER+ and w i j − for ER− are subtracted from the data by predictions.</p><p>These procedures are finally summarized as follows: 1) take the original microarray and the clinical data for ER from one study cohort; 2) apply the ARCH and AR models to the original data for each sample and identify the best fit model among the model candidates within 1 - 10 time lags; 3) subtract the residuals from the data by the prediction for the best fit model; 4) apply Wilcoxon statistic to the original data and to the empirical residuals by ARCH and AR; 5) list the p-values and identify the significant FDR (5%) corrected genes. 6) apply Gene Ontology analysis and pathway analysis (see the details in 2.3 and 2.4) for biological interpretation to the obtained gene list (see 4. in <xref ref-type="fig" rid="fig1">Figure 1</xref>).</p><p>The computational programs were done by the garchFit function (in “fGARCH”) for ARCH fitting, by the ar.ols function for AR fitting, the wilcox.test as a rank- sum test, and fdr.R for the FDR adjustment in the R package.</p></sec><sec id="s2_3"><title>2.3. GO Analysis</title><p>To investigate the gene product attributes from the gene list, Gene Ontology (GO) analysis was performed to find specific gene sets that are statistically associated among several biological categories. GO is designed as a functional annotation database to capture the known relationships between biological terms and all the genes that are instances of those terms. It is widely used by many functional enrichment tools and is highly regarded both for its comprehensiveness and its unified approach for annotating genes in different species to the same basic set of underlying functions [<xref ref-type="bibr" rid="scirp.74217-ref16">16</xref>] . Many tools have been developed to explore, filter, and search the GO database. In our study, Gorilla [<xref ref-type="bibr" rid="scirp.74217-ref17">17</xref>] was used as a GO analysis tool. GOrilla is an efficient web-based interactive user interface that is based on a statistical framework called minimum hypergenometric (mHG) for enrichment analysis in ranked gene lists, which are naturally represented as functional genomic information. For each GO term, the method independently identifies the threshold at which the most significant enrichment is obtained. The significant mHG scores are accurately and tightly corrected for threshold multiple testing without time-consuming simulations [<xref ref-type="bibr" rid="scirp.74217-ref17">17</xref>] . The tool identifies enriched GO terms in ranked gene lists for background gene sets which are obtained by the whole genomic location of microarray data. GO consists of three hierarchically structured vocabularies (ontologies) that describe gene products in terms of their associated biological processes, cellular components, and molecular functions. The building blocks of GO are called terms, and the relationship among them can be described by a directed acyclic graph (DAG), which is a hierarchy where each gene product may be annotated to one or more terms in each ontology [<xref ref-type="bibr" rid="scirp.74217-ref16">16</xref>] . GOrilla requires a list of gene symbols as input data. The obtained significant Etrez gene lists by FDR correction are converted into gene symbols using a web-based database called SOURCE [<xref ref-type="bibr" rid="scirp.74217-ref18">18</xref>] , which was developed by the Genetics Department of Stanford University.</p></sec><sec id="s2_4"><title>2.4. Pathway Analysis</title><p>As well as for GO analysis, the identified genes are mapped to the well-defined- biological pathways. Pathway analysis determines which pathways are overrepresented among genes that present significant variations. The difference from GO analysis is that pathway analysis includes interactions among a given set of genes. Several tools for pathway analysis have been published. In this study, we used a web-based analysis tool called REACTOME, which is a manually curated open-source open-data resource of human pathways and reactions [<xref ref-type="bibr" rid="scirp.74217-ref19">19</xref>] . REAC- TOME is a recent fast and sophisticated tool that has grown to include annotations for 7088 of the 20,774 protein-coding genes in the current Ensembl human genome assembly, 15,107 literature references, and 1421 small molecules organized into 6744 reactions collected in 1481 pathways [<xref ref-type="bibr" rid="scirp.74217-ref19">19</xref>] .</p></sec></sec><sec id="s3"><title>3. Simulation Study</title><p>To investigate the performance of our proposed algorithm, we performed a simulation study. We first prepared the clinical indicator like ER+ and ER−. The artificial indicator includes “1” for 50 samples and “0” for 50 samples. Next, we considered two types of artificial 1000-array and 100 samples: one array data (A) generating by normal distributions was set. The mean and variance values of the distribution were set as 1.0 to generate overall array data at once. In addition, the array data for the 201 - 400 array and the 601 - 600 array were replaced with the data generating different normal distribution with 1.8 mean and 10 variance; another array data (B) was generated by ARCH model. The model was applied to real array (DES data, see the detail in Section 4) and the parameters (mu: the intercept, omega: the constant coefficient of the variance equation, alpha: the coefficients of the variance equation, skew: the skewness of the data, shape: the shape parameter of the conditional distribution setting as 3) for the model was estimated for ER+ and ER−, respectively. We used these parameters and random number to generate the simulation data. For the computational programs, we conducted normrnd of Matlab<sup>&#210;</sup> command to generate random variables by normal distributions for array data A, and conducted garchSim of the R package fGARCH for array data B. We iterated 100 times to generate the two array data sets. To 100 data sets for A and B, we applied two-group comparison for the original simulation data and the ARCH residuals of them and identify 5% FDR significant parts.</p></sec><sec id="s4"><title>4. Material</title><p>Due to the extensive usage of microarray technology, in recent years publicly available datasets have exploded [<xref ref-type="bibr" rid="scirp.74217-ref4">4</xref>] , including the Gene Expression Omnibus (GEO, http://www.ncbi.nlm.nih.gov/geo/) [<xref ref-type="bibr" rid="scirp.74217-ref20">20</xref>] and Array Express (https://www.ebi.ac.uk/arrayexpress/). In this study for breast cancer research, we used five different expression datasets, collectively called the Affy947 expression dataset [<xref ref-type="bibr" rid="scirp.74217-ref12">12</xref>] . These datasets, which all measure the Human Genome HG U133A Affymetrix arrays, are normalized using the same protocol and are assessable from GEO with the following identifiers: GSE6532 for the Loi et al. dataset [<xref ref-type="bibr" rid="scirp.74217-ref21">21</xref>] (Loi), GSE3494 for the Miller et al. dataset [<xref ref-type="bibr" rid="scirp.74217-ref22">22</xref>] (Mil), GSE7390 for the Desmedt et al. dataset [<xref ref-type="bibr" rid="scirp.74217-ref23">23</xref>] (Des), and GSE5327 for the Minn et al. dataset [<xref ref-type="bibr" rid="scirp.74217-ref24">24</xref>] (Min). The Chine et al. dataset [<xref ref-type="bibr" rid="scirp.74217-ref25">25</xref>] (Chin) is available from ArrayExpress. This pooled dataset was preprocessed and normalized, as described in Zhao et al. [<xref ref-type="bibr" rid="scirp.74217-ref26">26</xref>] . Microarray quality-quality-control assessment was carried out using the R AffyPLM package from the Bioconductor web site (http://www.bioconductor.org [<xref ref-type="bibr" rid="scirp.74217-ref27">27</xref>] ). The Relative Log Expression (RLE) and Normalized Unscaled Standard Errors (NUSE) tests were applied. Chip pseudo-images were produced to assess artifacts on the arrays that did not pass the preceding quality control tests. The selected arrays were normalized by a three- step procedure using the RMA expression measure algorithm (http://www.bioconductor.org [<xref ref-type="bibr" rid="scirp.74217-ref28">28</xref>] ): RMA background correction convolution, the median centering of each gene separately across arrays for each dataset, and the quantile normalization of all arrays. Gene mean centering effectively removes many dataset specific biases, allowing for effective integration of multiple datasets [<xref ref-type="bibr" rid="scirp.74217-ref29">29</xref>] . 22,268 is the total number of probes for these microarray data.</p><p>Against all probes that covered the whole genome, we use the probes that correspond to the intrinsic signatures that were obtained by classifying breast tumors into five molecular subtypes [<xref ref-type="bibr" rid="scirp.74217-ref30">30</xref>] . We extracted 777 probes from the whole 22 K probes for the microarray data-sets using the intrinsic annotation included in the R codes in Zhao et al. [<xref ref-type="bibr" rid="scirp.74217-ref26">26</xref>] . As the response contrasting expression between two groups, we used a hormone receptor called ER, which indicates whether a hormone drug works as well for treatment as a progesterone receptor and is critical to determine the prognosis and predictive factors. ERsare used for classifying breast tumors into ER-positive (ER+) and ER-negative (ER−) diseases. The two upper figures in <xref ref-type="fig" rid="fig2">Figure 2</xref> present the mean of the microarray data by averaging all of the previously obtained samples [<xref ref-type="bibr" rid="scirp.74217-ref23">23</xref>] . The left and right plots correspond to a sample indicating ER+ and ER−. The data for ER− show more fluctuation than for ER+. The two lower figures illustrate histograms of the averaged data for ER+ and ER− and present sharper peakedness and heavier tails than the shape of an ordinary Gaussian distribution.</p></sec><sec id="s5"><title>5. Results and Discussion</title><sec id="s5_1"><title>5.1. Simulation Data</title><p>For the simulation data and ARCH residuals, we summarized the average of the number of the identified 5% FDR significant parts and the number of the overlapped parts in <xref ref-type="table" rid="table1">Table 1</xref>. In the case of the simulation data generated by normal</p><table-wrap id="table1" ><label><xref ref-type="table" rid="table1">Table 1</xref></label><caption><title> Summary of the average for the identified significant parts</title></caption><table><tbody><thead><tr><th align="center" valign="middle" ></th><th align="center" valign="middle" >1. Original series</th><th align="center" valign="middle" >2. ARCH residuals</th><th align="center" valign="middle" >Overlapped # of 2 with 1</th><th align="center" valign="middle" >Ratio for the overlapped #</th></tr></thead><tr><td align="center" valign="middle" >Array set A</td><td align="center" valign="middle" >30.2</td><td align="center" valign="middle" >29.8</td><td align="center" valign="middle" >29.3</td><td align="center" valign="middle" >98.2 %</td></tr><tr><td align="center" valign="middle" >Array set B</td><td align="center" valign="middle" >512.2</td><td align="center" valign="middle" >338.9</td><td align="center" valign="middle" >212.1</td><td align="center" valign="middle" >62.6 %</td></tr></tbody></table></table-wrap><p>distribution, the significant number for original series and ARCH residuals in array sets A and B was not differ. The parts identified in A were mostly same as ones in B. On the other hand, in the simulation data generated by ARCH model, the ARCH residuals identified more significant parts from the data than the original series. The number of significant parts for ARCH residuals was about 30% less than the number of significant parts for the original series. The overlapped number was less than the case A, however over 50% parts were covered.</p></sec><sec id="s5_2"><title>5.2. Affy947 Expression Dataset</title><p>Based on the method explained in 3.2, the best fit AR and GARCH models were selected by AIC for each sample. The estimated orders of all the best fit models of all the studies are summarized in Supplementary <xref ref-type="table" rid="table1">Table 1</xref>. <xref ref-type="fig" rid="fig3">Figure 3</xref> summarizes the ratio of the sample numbers for each selected order against the total number of samples. These figures suggest that the most often selected orders were one while ER+ samples tended to take more complicated models than for the ER− samples.</p><p>Using residuals obtained by the best fit ARCH and AR models and the original data, we applied Wilcoxon statistic to compare DEs between two groups divided by ER+ and ER−. The significant genomic locations were assessed by a FDR. The locations were mapped on Entrez gene IDs according to the Affy probes presented in the original microarray data and converted into gene symbols by SOURCE. The identified genes in the original data and the ARCH residual analyses are listed in Supplementary <xref ref-type="table" rid="table2">Table 2</xref>. Based on these gene lists, we investigated the overlapped significant genes for the original data with significant genes for the ARCH and AR residuals and summarized the results in <xref ref-type="table" rid="table2">Table 2</xref>. About 200 - 280 significant DE genes in the studies of Des, Mil, Min, and Chin were identified with FDR correction. For Loi, the significant genes were fewer than the other studies in all cases. Except for Loi, the number of significant genes for the ARCH residuals was reduced by about 20% - 35% less than the number of genes for the original data. The estimated genes for all the datasets (except for Loi’s case) overlapped 100% with the estimated genes for the original data. The number of significant genes for the AR residuals in the cases of Mil, Min, and Loi was less than the number of genes by ARCH and resulted in about a 20% - 35% reduction of significant genes for the original data except for Loi. The reduction rate was similar to the rate shown in the case B of our simulation study. Furthermore, these genes for the AR residuals did not completely 100% overlap with the genes for the original data unlike the case of the ARCH residuals. The results suggest that the real array data might be generated by a</p><p>similar structure as the ARCH process and empirical ARCH residuals might be more effective to specify important genes from a list of long genes than AR residuals.</p><p>To investigate the overlapping genes by the ARCH residuals with genes by the original data, the corresponding cytoband and gene symbols are summarized in <xref ref-type="table" rid="table3">Table 3</xref>. The total numbers of common genes by the original and ARCH residuals in four studies were 132 and 99. If we take into account Loi’s case, the total numbers of common genes across all studies for the original and ARCH residuals are 12 and 9. The genes obtained by the ARCH residuals were completely covered by the genes obtained by the original data. The results by the ARCH residuals covered several important genes for breast cancer, such as TP53 in the</p><table-wrap id="table2" ><label><xref ref-type="table" rid="table2">Table 2</xref></label><caption><title> Summary for FDR 5% adjusted Entrez genes of five datasets. Values in parentheses indicate number of unique genes to avoid duplicate and multiple genes from obtained gene list. Percentages for overlapped with original indicate ratios for overlapped significant genes for ARCH or AR residuals with significant genes in original data</title></caption><table><tbody><thead><tr><th align="center" valign="middle" >Model</th><th align="center" valign="middle" >Data</th><th align="center" valign="middle" >Des</th><th align="center" valign="middle" >Mil</th><th align="center" valign="middle" >Min</th><th align="center" valign="middle" >Loi</th><th align="center" valign="middle" >Chin</th></tr></thead><tr><td align="center" valign="middle" >-</td><td align="center" valign="middle" >Original #EntrezID (unique)</td><td align="center" valign="middle" >245 (186)</td><td align="center" valign="middle" >238 (176)</td><td align="center" valign="middle" >274 (195)</td><td align="center" valign="middle" >53 (47)</td><td align="center" valign="middle" >277 (201)</td></tr><tr><td align="center" valign="middle"  rowspan="2"  >ARCH</td><td align="center" valign="middle" >Residual #EntrezID (unique)</td><td align="center" valign="middle" >193 (152)</td><td align="center" valign="middle" >175 (139)</td><td align="center" valign="middle" >207 (154)</td><td align="center" valign="middle" >46 (41)</td><td align="center" valign="middle" >177 (133)</td></tr><tr><td align="center" valign="middle" >Overlapped with original [%]</td><td align="center" valign="middle" >100</td><td align="center" valign="middle" >100</td><td align="center" valign="middle" >100</td><td align="center" valign="middle" >98</td><td align="center" valign="middle" >100</td></tr><tr><td align="center" valign="middle"  rowspan="2"  >AR</td><td align="center" valign="middle" >Residual #EntrezID (unique)</td><td align="center" valign="middle" >194 (152)</td><td align="center" valign="middle" >161 (131)</td><td align="center" valign="middle" >183 (141)</td><td align="center" valign="middle" >37 (34)</td><td align="center" valign="middle" >178 (139)</td></tr><tr><td align="center" valign="middle" >Overlapped with original [%]</td><td align="center" valign="middle" >95</td><td align="center" valign="middle" >99</td><td align="center" valign="middle" >87</td><td align="center" valign="middle" >92</td><td align="center" valign="middle" >98</td></tr></tbody></table></table-wrap><p>chromosome 1q region, ERBB2 in the chromosome 17q region, and ESR1 in the chromosome 6q region, even if the number of identified Entrez genes was less than the number of identified genes from the original data.</p><p>Next, we performed GO enrichment analysis using significant DE gene lists for the original data and ARCH’s residual analyses in all studies. To correctly find the enriched GO terms for the associated genes, a background list was prepared of all the probes included in the original microarray data. The Entrez genes in the background list were converted into 13,177 gene symbols without any duplication by SOURCE. As the input gene lists to GOrilla, the numbers of summarized unique genes are shown in the parentheses of <xref ref-type="table" rid="table2">Table 2</xref>. All the associated GO terms for the original and ARCH residuals in all the studies are summarized in Supplementary <xref ref-type="table" rid="table3">Table 3</xref>. Since the estimated gene symbols in Loi’s case were less than half of the amount taken in other studies, few associated GO terms were identified in the biological process and cellular component and no GO terms in the molecular function. Also, significant DE genes for the ARCH residuals contributed to finding additional associated GO terms that did not appear in the GO terms for the original data, e.g., mammary gland epithelial cell proliferation for Des, a single-organism metabolic process for Des, an organonitrogen compound metabolic process for Mil and Min, and a single-organism developmental process for Min and Chin, all of which are related to meaningful biological associations like cellular differentiation, proliferation, and metabolic pathways in cancer cells [<xref ref-type="bibr" rid="scirp.74217-ref16">16</xref>] . <xref ref-type="table" rid="table3">Table 3</xref> summarizes the common GO terms of the biological processes for Des, Mil, Min, and Chin and presented 13 terms for the original data. The terms for the ARCH residuals mostly overlapped with them except for Mil’s case. As shown in Supplementary <xref ref-type="table" rid="table3">Table 3</xref>, two terms in the molecular function and eight in the cellular components were commonly identified by the original data. The GO terms for the ARCH residuals covered them, and more terms were shown in the molecular function.</p><p>Furthermore, to investigate the consistency of the refined significant gene</p><table-wrap-group id="3"><label><xref ref-type="table" rid="table3">Table 3</xref></label><caption><title> Identified differentially expressed genes (FDR 5%) and cytobands for ER status in original data and empirical ARCH residuals</title></caption><table-wrap id="3_1"><table><tbody><thead><tr><th align="center" valign="middle"  rowspan="2"  >Studies</th><th align="center" valign="middle"  colspan="2"  >Original</th><th align="center" valign="middle"  colspan="2"  >ARCH residuals</th></tr></thead><tr><td align="center" valign="middle" >cytoband</td><td align="center" valign="middle" >genes</td><td align="center" valign="middle" >cytoband</td><td align="center" valign="middle" >genes</td></tr><tr><td align="center" valign="middle" >Des Mil Min Chin</td><td align="center" valign="middle" >1p13.3 1p32.3 1p34.1 1p35 1p35.3 1p35.3-p33 1q21 1q21.1 1q21.3 1q23.2 1q24-q25 1q32.2 1q41 1q42.11 1q42.13 2p11.2 2q35 2q37.3 3p14.3 3p21 3q13.1 3q23-q25 3q24-q25.1 3q25 4q12 4q21.1 4q28.3 4q32.1 4q35.1 5q13.1 5q14-q21 5q22-q23 5q31.1 5q33.2 5q33.3 5q35.2 5q35.3 6p12 6p21.3 6p22.3 6q22.31 6q22.33 6q22-q23 6q23.3 6q25.1 7p13 7p15 7q21 7q21-q31 7q31.1 7q36 8p12 8p21 8p22</td><td align="center" valign="middle" >VAV3, GSTM3, CHI3L2 ECHDC2 CTPS1 IFI6 ATPIF1 MEAF6 S100A11, S100A1 PEA15 CRABP2 COPA CACYBP ELF3 TP53BP2 DEGS1 ADCK3 TMSB10 IGFBP5, IGFBP2 LRRFIP1, SNED1 ACOX2 MST1 ALCAM CP GYG1 SIAH2 KIT, PDGFRA USO1 MGST2 GRIA2 ACSL1 PIK3R1 PAM REEP5 JADE2 GALNT10 CYFIP2 MSX2 GNB2L1 MCM3, TFAP2B HIST1H1C ID4 ASF1A ECHDC1 FABP7 CITED2 ESR1 BLVRA GARS FZD1 SEMA3C IFRD1 PTPRN2 NRG1, PLAT EPHX2 ASAH1, TUSC3</td><td align="center" valign="middle" >1p13.3 1p35 1p35.3 1q21 1q21.1 1q21.3 1q23.2 1q24-q25 1q41 1q42.13 2q35 2q37.3 3q23-q25 3q24-q25.1 3q25 4q12 4q21.1 4q35.1 5q13.1 5q14-q21 5q22-q23 5q33.2 5q35.2 5q35.3 6p12 6p21.3 6p22.3 6q22.31 6q22-q23 6q23.3 6q25.1 7p13 7p15 7q21 7q21-q31 7q31.1 7q36 8p12 8p21 8p22</td><td align="center" valign="middle" >VAV3, GSTM3 IFI6 ATPIF1 S100A1, S100A11 PEA15 CRABP2 COPA CACYBP TP53BP2 ADCK3 IGFBP5, IGFBP2 LRRFIP1, SNED1 CP GYG1 SIAH2 KIT, PDGFRA USO1 ACSL1 PIK3R1 PAM REEP5 GALNT10 MSX2 GNB2L1 MCM3 HIST1H1C ID4 ASF1A FABP7 CITED2 ESR1 BLVRA GARS FZD1 SEMA3C IFRD1 PTPRN2 NRG1 EPHX2 ASAH1, TUSC3</td></tr></tbody></table></table-wrap><table-wrap id="3_2"><table><tbody><thead><tr><th align="center" valign="middle" >Des Mil Min Chin</th><th align="center" valign="middle" >8q21.1 8q22 8q22.1 8q24.1 8q24.12 9q33.3 9q34.1 9q34.11 10p15 10q24 11p12-p11 11p15 11q11-q12 11q12.3 11q13 11q14.1 12p13 12q12 12q13 12q13.12 12q14 12q14.1 12q24.21 13q12 13q21.1-q32 13q22.2 13q31.2-q32.3 13q33 14q11.2 15q24 15q24.2 15q26.3 16p12.2 16p13.3 16q13 16q22.1 16q24.3 17p11.2 17q11.2 17q11.2-q12 17q11-q12 17q12 17q21.2 17q21.31 17q24-q25 18p11.3 18q21.1 18q22-q23 18q23 19p13.3 19p13.3-p13.2 19q13.2 19q13.3 19q13.4 20p11.21 21q21.1 22q11.2 22q13.1 Xp21.3</th><th align="center" valign="middle" >PEX2 CA2 LAPTM4B SQLE TRPS1 RALGPS1 CRAT SPTAN1 GATA3 PDCD4, MYOF, PAPSS2 EXT2 RPL27A TCN1 PLA2G16 NUMA1 RSF1 PTMS, SCNN1A TWF1 STAT6, SLC11A2 FKBP11 GNS PPM1H MED13L FLT1 CLN5 LMO7 STK24 EFNB2 MMP14 CIB2 COMMD4 IGF1R POLR3E HCFC1R1 ARL2BP CDH1 PIEZO1 ALDH3A2, PEMT FAM222B, CCL18 LIG3 FLOT2 ERBB2 KRT17 ACBD4 CDC42EP4 RAB31 ACAA2 ZNF236 CYB5A KDM4B EPOR CYP2A6 CA11, ARHGAP35 PEG3 ENTPD6 BTG3 IGL POLR2F, H1F0 ZFX</th><th align="center" valign="middle" >8q22 8q22.1 8q24.12 9q34.1 10p15 10q24 11p12-p11 11p15 11q11-q12 11q12.3 11q13 11q14.1 12p13 12q12 12q13 12q13.12 12q14 12q14.1 12q24.21 13q12 13q21.1-q32 13q22.2 13q33 15q26.3 16p12.2 16p13.3 16q13 16q24.3 17p11.2 17q11.2 17q11.2-q12 17q11-q12 17q12 17q24-q25 18q21.1 18q22-q23 18q23 19p13.3 19q13.2 19q13.3 19q13.4 20p11.21 21q21.1 22q11.2 22q13.1</th><th align="center" valign="middle" >CA2 LAPTM4B TRPS1 CRAT GATA3 PDCD4, MYOF, PAPSS2 EXT2 RPL27A TCN1 PLA2G16 NUMA1 RSF1 PTMS TWF1 STAT6 FKBP11 GNS PPM1H MED13L FLT1 CLN5 LMO7 EFNB2 IGF1R POLR3E HCFC1R1 ARL2BP PIEZO1 ALDH3A2 FAM222B, CCL18 LIG3 FLOT2 ERBB2 CDC42EP4 ACAA2 ZNF236 CYB5A KDM4B CYP2A6 CA11,ARHGAP35 PEG3 ENTPD6 BTG3 IGL POLR2F, H1F0</th></tr></thead></tbody></table></table-wrap><table-wrap id="3_3"><table><tbody><thead><tr><th align="center" valign="middle"  rowspan="2"  >Des Mil Min Chin</th><th align="center" valign="middle" >Xp22.1</th><th align="center" valign="middle" >SAT1</th><th align="center" valign="middle" ></th><th align="center" valign="middle" ></th></tr></thead><tr><td align="center" valign="middle" >Xq26.3</td><td align="center" valign="middle" >VGLL1</td><td align="center" valign="middle" >Xq26.3</td><td align="center" valign="middle" >VGLL1</td></tr><tr><td align="center" valign="middle"  rowspan="13"  >+Loi</td><td align="center" valign="middle" >1p13.3</td><td align="center" valign="middle" >CHI3L2</td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td></tr><tr><td align="center" valign="middle" >1q24-q25</td><td align="center" valign="middle" >CACYBP</td><td align="center" valign="middle" >1q24-q25</td><td align="center" valign="middle" >CACYBP</td></tr><tr><td align="center" valign="middle" >2q35</td><td align="center" valign="middle" >IGFBP5</td><td align="center" valign="middle" >2q35</td><td align="center" valign="middle" >IGFBP5</td></tr><tr><td align="center" valign="middle" >3p21</td><td align="center" valign="middle" >3p21</td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td></tr><tr><td align="center" valign="middle" >5q35.2</td><td align="center" valign="middle" >MSX2</td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td></tr><tr><td align="center" valign="middle" >7p13</td><td align="center" valign="middle" >BLVRA</td><td align="center" valign="middle" >7p13</td><td align="center" valign="middle" >BLVRA</td></tr><tr><td align="center" valign="middle" >10q24</td><td align="center" valign="middle" >PDCD4, MYOF</td><td align="center" valign="middle" >10q24</td><td align="center" valign="middle" >PDCD4, MYOF</td></tr><tr><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" >17q11.2</td><td align="center" valign="middle" >CCL18</td></tr><tr><td align="center" valign="middle" >17q24-q25</td><td align="center" valign="middle" ></td><td align="center" valign="middle" >17q24-q25</td><td align="center" valign="middle" >CDC42EP4</td></tr><tr><td align="center" valign="middle" >19p13.3-p13.2</td><td align="center" valign="middle" >EPOR</td><td align="center" valign="middle" >19p13.3-p13.2</td><td align="center" valign="middle" >EPOR</td></tr><tr><td align="center" valign="middle" >21q21.1</td><td align="center" valign="middle" >BTG3</td><td align="center" valign="middle" ></td><td align="center" valign="middle" >BTG3</td></tr><tr><td align="center" valign="middle" >22q11.2</td><td align="center" valign="middle" >IGL</td><td align="center" valign="middle" >21q21.1</td><td align="center" valign="middle" >IGL</td></tr><tr><td align="center" valign="middle" >Xp22.1</td><td align="center" valign="middle" >SAT1</td><td align="center" valign="middle" >22q11.2</td><td align="center" valign="middle" ></td></tr></tbody></table></table-wrap></table-wrap-group><p>lists, we applied pathway analysis to the significant DE genes for the original and ARCH residuals listed in <xref ref-type="table" rid="table4">Table 4</xref>. All the identified pathways with Entities FDR (&lt;1.0) and associated genes are summarized in Supplementary <xref ref-type="table" rid="table4">Table 4</xref>. In the pathway components shown in Supplementary <xref ref-type="table" rid="table3">Table 3</xref>, ERBB2 signaling, EGFR, cell-cycle, immune system, metabolic pathway, AKT signaling and Wnt pathway are well- known important breast cancer-signaling pathways [<xref ref-type="bibr" rid="scirp.74217-ref31">31</xref>] . We took them to be representative of important pathways and counted the number of identified pathways related to these components in the case of the original and ARCH residuals. The number and associated gene symbols are summarized in <xref ref-type="table" rid="table5">Table 5</xref>. The representative pathways were mostly covered by the significant DE genes for the ARCH residuals. This result supports that the refined gene lists obtained by the ARCH residuals generally captured the differentiating breast tumors based on ER status and did not overlook any important biological information by the limited DE gene lists for the ARCH residuals.</p></sec></sec><sec id="s6"><title>6. Conclusion</title><p>We applied a rank order statistic for an ARCH residual empirical process to refine significant DE genes by two-group comparison in microarray analysis. Our approach considered publicly available gene expression datasets and the clinical output for ER in addition to the simulation study. We compared the analysis performances by the ARCH residuals with the AR residuals and the ordinal original microarray data. While the genes for the AR residuals did not cover 100% of the genes for the original data analysis, the genes by the ARCH residuals were mostly 100% overlapped with the original data, and the gene lists were reduced about 30% from the gene lists obtained by the original data analysis. We confirmed the similar property for the 30% reduction in the simulation study. In GO enrichment and pathway analyses, the result by the ARCH residuals was mostly covered with associated biological terms obtained by the original data</p><table-wrap id="table4" ><label><xref ref-type="table" rid="table4">Table 4</xref></label><caption><title> Common associated biological processes among Des, Mil, Min, and Chin for original and ARCH residuals</title></caption><table><tbody><thead><tr><th align="center" valign="middle"  colspan="2"   rowspan="2"  >Associated GO terms</th><th align="center" valign="middle"  colspan="2"  >Des</th><th align="center" valign="middle"  colspan="2"  >Mil</th><th align="center" valign="middle"  colspan="2"  >Min</th><th align="center" valign="middle"  colspan="2"  >Chin</th></tr></thead><tr><td align="center" valign="middle" >Orig</td><td align="center" valign="middle" >Arch</td><td align="center" valign="middle" >Orig</td><td align="center" valign="middle" >Arch</td><td align="center" valign="middle" >Orig</td><td align="center" valign="middle" >Arch</td><td align="center" valign="middle" >Orig</td><td align="center" valign="middle" >Arch</td></tr><tr><td align="center" valign="middle"  rowspan="13"  >Biological Process</td><td align="center" valign="middle" >epithelial cell proliferation</td><td align="center" valign="middle" >+</td><td align="center" valign="middle" >+</td><td align="center" valign="middle" >+</td><td align="center" valign="middle" >+</td><td align="center" valign="middle" >+</td><td align="center" valign="middle" >+</td><td align="center" valign="middle" >+</td><td align="center" valign="middle" >+</td></tr><tr><td align="center" valign="middle" >response to estrogen</td><td align="center" valign="middle" >+</td><td align="center" valign="middle" >+</td><td align="center" valign="middle" >+</td><td align="center" valign="middle" >+</td><td align="center" valign="middle" >+</td><td align="center" valign="middle" >+</td><td align="center" valign="middle" >+</td><td align="center" valign="middle" >+</td></tr><tr><td align="center" valign="middle" >epidermis development</td><td align="center" valign="middle" >+</td><td align="center" valign="middle" >+</td><td align="center" valign="middle" >+</td><td align="center" valign="middle" ></td><td align="center" valign="middle" >+</td><td align="center" valign="middle" >+</td><td align="center" valign="middle" >+</td><td align="center" valign="middle" ></td></tr><tr><td align="center" valign="middle" >regulation of phosphatidylinositol 3-kinase activity</td><td align="center" valign="middle" >+</td><td align="center" valign="middle" >+</td><td align="center" valign="middle" >+</td><td align="center" valign="middle" >+</td><td align="center" valign="middle" >+</td><td align="center" valign="middle" ></td><td align="center" valign="middle" >+</td><td align="center" valign="middle" >+</td></tr><tr><td align="center" valign="middle" >erythropoietin-mediated signaling pathway</td><td align="center" valign="middle" >+</td><td align="center" valign="middle" >+</td><td align="center" valign="middle" >+</td><td align="center" valign="middle" ></td><td align="center" valign="middle" >+</td><td align="center" valign="middle" ></td><td align="center" valign="middle" >+</td><td align="center" valign="middle" >+</td></tr><tr><td align="center" valign="middle" >regulation of lipid kinase activity</td><td align="center" valign="middle" >+</td><td align="center" valign="middle" >+</td><td align="center" valign="middle" >+</td><td align="center" valign="middle" >+</td><td align="center" valign="middle" >+</td><td align="center" valign="middle" >+</td><td align="center" valign="middle" >+</td><td align="center" valign="middle" >+</td></tr><tr><td align="center" valign="middle" >phosphatidylinositol 3-kinase signaling</td><td align="center" valign="middle" >+</td><td align="center" valign="middle" >+</td><td align="center" valign="middle" >+</td><td align="center" valign="middle" >+</td><td align="center" valign="middle" >+</td><td align="center" valign="middle" >+</td><td align="center" valign="middle" >+</td><td align="center" valign="middle" >+</td></tr><tr><td align="center" valign="middle" >positive regulation of phosphatidylinositol 3-kinase activity</td><td align="center" valign="middle" >+</td><td align="center" valign="middle" >+</td><td align="center" valign="middle" >+</td><td align="center" valign="middle" >+</td><td align="center" valign="middle" >+</td><td align="center" valign="middle" >+</td><td align="center" valign="middle" >+</td><td align="center" valign="middle" >+</td></tr><tr><td align="center" valign="middle" >phenylpropanoid catabolic process</td><td align="center" valign="middle" >+</td><td align="center" valign="middle" >+</td><td align="center" valign="middle" >+</td><td align="center" valign="middle" >+</td><td align="center" valign="middle" >+</td><td align="center" valign="middle" >+</td><td align="center" valign="middle" >+</td><td align="center" valign="middle" >+</td></tr><tr><td align="center" valign="middle" >mast cell differentiation</td><td align="center" valign="middle" >+</td><td align="center" valign="middle" >+</td><td align="center" valign="middle" >+</td><td align="center" valign="middle" >+</td><td align="center" valign="middle" >+</td><td align="center" valign="middle" >+</td><td align="center" valign="middle" >+</td><td align="center" valign="middle" >+</td></tr><tr><td align="center" valign="middle" >extracellular vesicle</td><td align="center" valign="middle" >+</td><td align="center" valign="middle" >+</td><td align="center" valign="middle" >+</td><td align="center" valign="middle" ></td><td align="center" valign="middle" >+</td><td align="center" valign="middle" >+</td><td align="center" valign="middle" >+</td><td align="center" valign="middle" >+</td></tr><tr><td align="center" valign="middle" >extracellular vesicular exosome</td><td align="center" valign="middle" >+</td><td align="center" valign="middle" >+</td><td align="center" valign="middle" >+</td><td align="center" valign="middle" ></td><td align="center" valign="middle" >+</td><td align="center" valign="middle" >+</td><td align="center" valign="middle" >+</td><td align="center" valign="middle" >+</td></tr><tr><td align="center" valign="middle" >extracellular region part</td><td align="center" valign="middle" >+</td><td align="center" valign="middle" >+</td><td align="center" valign="middle" >+</td><td align="center" valign="middle" ></td><td align="center" valign="middle" >+</td><td align="center" valign="middle" >+</td><td align="center" valign="middle" >+</td><td align="center" valign="middle" >+</td></tr></tbody></table></table-wrap><table-wrap id="table5" ><label><xref ref-type="table" rid="table5">Table 5</xref></label><caption><title> Identified important breast cancer-signaling pathways and associated gene symbols obtained from original data and ARCH residuals</title></caption><table><tbody><thead><tr><th align="center" valign="middle"  rowspan="2"  >Pathways</th><th align="center" valign="middle"  colspan="2"  >Original</th><th align="center" valign="middle"  colspan="2"  >ARCH residuals</th></tr></thead><tr><td align="center" valign="middle" >Number</td><td align="center" valign="middle" >Gene symbol</td><td align="center" valign="middle" >Number</td><td align="center" valign="middle" >Gene symbol</td></tr><tr><td align="center" valign="middle" >ERBB2 signaling</td><td align="center" valign="middle" >7</td><td align="center" valign="middle" >ERBB2, KIT, NRG1</td><td align="center" valign="middle" >6</td><td align="center" valign="middle" >ERBB2</td></tr><tr><td align="center" valign="middle" >EGFR pathways</td><td align="center" valign="middle" >11</td><td align="center" valign="middle" >FLT1, KIT, PIK3R1, VAV3</td><td align="center" valign="middle" >10</td><td align="center" valign="middle" >ERBB2, FLT1, PIK3R1</td></tr><tr><td align="center" valign="middle" >Cell cycle</td><td align="center" valign="middle" >5</td><td align="center" valign="middle" >CDH1, MCM3</td><td align="center" valign="middle" >3</td><td align="center" valign="middle" >MCM3, NUMA1</td></tr><tr><td align="center" valign="middle" >Immune system</td><td align="center" valign="middle" >5</td><td align="center" valign="middle" >CDH1, KIT, PIK3R1, STAT6</td><td align="center" valign="middle" >5</td><td align="center" valign="middle" >ERBB2, IFI6, STAT6</td></tr><tr><td align="center" valign="middle" >Metabolic disorder</td><td align="center" valign="middle" >1</td><td align="center" valign="middle" >SAT1</td><td align="center" valign="middle" >1</td><td align="center" valign="middle" >FZD1</td></tr><tr><td align="center" valign="middle" >PI3K/AKT signaling</td><td align="center" valign="middle" >5</td><td align="center" valign="middle" >KIT, PIK3R1</td><td align="center" valign="middle" >5</td><td align="center" valign="middle" >ERBB2, PIK3R1</td></tr><tr><td align="center" valign="middle" >Wnt pathway</td><td align="center" valign="middle" >5</td><td align="center" valign="middle" >FZD1</td><td align="center" valign="middle" >5</td><td align="center" valign="middle" >FZD1</td></tr></tbody></table></table-wrap><p>analysis and presented additional important GO terms in biological processes. These results suggest that data processing using ARCH residuals array data could contribute to refining significant DE genes that follow the required gene signatures and provide prognostic accuracy and guide clinical decisions.</p></sec><sec id="s7"><title>Acknowledgements and Funding</title><p>The research by the second author was supported by Japanese Grant-in-Aids A23244011 (Taniguchi, M., Waseda Univ.).</p></sec><sec id="s8"><title>Competing Interest</title><p>The authors declare no competing interests.</p></sec><sec id="s9"><title>Cite this paper</title><p>Solvang, H.K. and Taniguchi, M. (2017) Microarray Analysis Using Rank Order Statistics for ARCH Residual Empirical Process. Open Journal of Statistics, 7, 54-71. https://doi.org/10.4236/ojs.2017.71005</p></sec><sec id="s10"><title>Supplementary (see  https://www.dropbox.com/sh/sotz8jufje73eg6/AACKHD-tXqB02h_rxlvVlXqsa?dl=0 )</title><p>Supplementary <xref ref-type="table" rid="table1">Table 1</xref>. Estimated order of all best fit models for each sample.</p><p>Supplementary <xref ref-type="table" rid="table2">Table 2</xref>. Significant DE Entrez genes and gene symbols identified by original data and ARCH residuals.</p><p>Supplementary <xref ref-type="table" rid="table3">Table 3</xref>. Associated GO terms for original and ARCH residuals.</p><p>Supplementary <xref ref-type="table" rid="table4">Table 4</xref>. Identified pathways and associated genes.</p></sec></body><back><ref-list><title>References</title><ref id="scirp.74217-ref1"><label>1</label><mixed-citation publication-type="other" xlink:type="simple">Zhao, X., Rodland, E.A., Sorlie, T., Naume, B., Langerod, A., Frigessi, A., Kristensen, V.N., Borresen-Dale, A.L. and Lingjaerde, O.C. (2011) Combining Gene Signatures Improves Prediction of Breast Cancer Survival. PLoS ONE, 6.</mixed-citation></ref><ref id="scirp.74217-ref2"><label>2</label><mixed-citation publication-type="other" xlink:type="simple">Yang, Y.H., Xiao, Y. and Segal, M.R. (2005) Identifying Differentially Expressed Genes from Microarray Experiments via Statistics Synthesis. Bioinformatics, 21, 1084-1093. https://doi.org/10.1093/bioinformatics/bti108</mixed-citation></ref><ref id="scirp.74217-ref3"><label>3</label><mixed-citation publication-type="other" xlink:type="simple">Tusher, V.G., Tibshirani, R. and Chu, G. (2001) Significance Analysis of Microarrays Applied to the Ionizing Radiation Response. Proceedings of the National Academy of Sciences, 98, 5116-5121. https://doi.org/10.1073/pnas.091062498</mixed-citation></ref><ref id="scirp.74217-ref4"><label>4</label><mixed-citation publication-type="other" xlink:type="simple">Campain, A. and Yang, Y.H. (2010) Comparison Study of Microarray Meta-Analysis Methods. BMC Bioinformatics, 11, 408. https://doi.org/10.1186/1471-2105-11-408</mixed-citation></ref><ref id="scirp.74217-ref5"><label>5</label><mixed-citation publication-type="other" xlink:type="simple">Benjamini, Y. and Hockberg, Y. (1995) Controlling the False Discovery Rate: A Practical and Powerful Approach to Multiple Testing. Journal of the Royal Statistical Society B, 57, 289-300.</mixed-citation></ref><ref id="scirp.74217-ref6"><label>6</label><mixed-citation publication-type="other" xlink:type="simple">Engle, R.F. (1982) Autoregressive Conditional Heteroskedasticity with Estimates of the Variance of UK Inflation. Econometrica, 50, 987-1008.  
https://doi.org/10.2307/1912773</mixed-citation></ref><ref id="scirp.74217-ref7"><label>7</label><mixed-citation publication-type="other" xlink:type="simple">Chandra, S.A. and Taniguchi, M. (2003) Asymptotics of Rank Order Statistics for ARCH Residual Empirical Processes. Stochastic Processes and Their Applications, 104, 301-324. https://doi.org/10.1016/S0304-4149(02)00239-9</mixed-citation></ref><ref id="scirp.74217-ref8"><label>8</label><mixed-citation publication-type="other" xlink:type="simple">Ozaki, T. and Iino, M. (2001) An Innovation Approach to Non-Gaussian Time Series Analysis. Journal of Applied Probability, 38A, 78-92.  
https://doi.org/10.1017/S0021900200112690</mixed-citation></ref><ref id="scirp.74217-ref9"><label>9</label><mixed-citation publication-type="other" xlink:type="simple">Stoffer, D.S., Tyler, D.E. and Wendt, D.A. (2000) The Spectral Envelope and Its Applications. Statistical Science, 15, 224-253. https://doi.org/10.1214/ss/1009212816</mixed-citation></ref><ref id="scirp.74217-ref10"><label>10</label><mixed-citation publication-type="other" xlink:type="simple">Koren, A., Tirosh I. and Barki, N. (2007) Autocorrelation Analysis Reveals Widespread Spatial Biases in Microarray Experiments. BMC Genomics, 8, 164.  
https://doi.org/10.1186/1471-2164-8-164</mixed-citation></ref><ref id="scirp.74217-ref11"><label>11</label><mixed-citation publication-type="other" xlink:type="simple">Lee, S. and Taniguchi, M. (2005) Asymptotic Theory for ARCH-SM Models: LAN and Residual Empirical Processes. Statistica Sinica, 15, 215-234.</mixed-citation></ref><ref id="scirp.74217-ref12"><label>12</label><mixed-citation publication-type="other" xlink:type="simple">Van Vliet, M.H., Reyal, F., Horlings, H.M., van de Vijver, M.J., Reinders, M.J.T. and Wessels, L.F.A. (2010) Pooling Breast Cancer Datasets Has a Synergetic Effect on Classification Performance and Improves Signature Stability. BMC Genomics, 9, 375. https://doi.org/10.1186/1471-2164-9-375</mixed-citation></ref><ref id="scirp.74217-ref13"><label>13</label><mixed-citation publication-type="other" xlink:type="simple">Rezaul, K., Thumar, J.K., Lundgren, D.H., Eng, J.K., Claffey, K.P., Wilson, L. and Han, D.K. (2010) Differential Protein Expression Profiles in Estrogen Receptor-Positive and -Negative Breast Cancer Tissues Using Label-Free Quantitative Proteomics. Genes Cancer, 1, 251-271. https://doi.org/10.1177/1947601910365896</mixed-citation></ref><ref id="scirp.74217-ref14"><label>14</label><mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Milhoj</surname><given-names> A. </given-names></name>,<etal>et al</etal>. (<year>1985</year>)<article-title>The Moment Structure of ARCH Processes</article-title><source> Scandinavian Journal of Statistics</source><volume> 12</volume>,<fpage> 281</fpage>-<lpage>292</lpage>.<pub-id pub-id-type="doi"></pub-id></mixed-citation></ref><ref id="scirp.74217-ref15"><label>15</label><mixed-citation publication-type="other" xlink:type="simple">Akaike, H. (1974) A New Look at the Statistical Model Identification. IEEE Transactions on Automatic Control, 19, 716-723.  
https://doi.org/10.1109/TAC.1974.1100705</mixed-citation></ref><ref id="scirp.74217-ref16"><label>16</label><mixed-citation publication-type="other" xlink:type="simple">Glass, K. and Girvan, M. (2014) Annotation Enrichment Analysis: An Alternative Method for Evaluating the Functional Properties of Gene Sets. Scientific Reports, 4, 4191. https://doi.org/10.1038/srep04191</mixed-citation></ref><ref id="scirp.74217-ref17"><label>17</label><mixed-citation publication-type="other" xlink:type="simple">Eden, E., Navon, R., Steinfeld, I., Lipson, D. and Yakhini, Z. (2009) GOrilla: A Tool for Discovery and Visualization of Enriched GO Terms in Ranked Gene Lists. BMC Bioinformatics, 10, 48. https://doi.org/10.1186/1471-2105-10-48</mixed-citation></ref><ref id="scirp.74217-ref18"><label>18</label><mixed-citation publication-type="other" xlink:type="simple">Deihn, M., Sherlock, G., Binkley, G., Jin, H., Matese, J.C., Hernandez-Boussard, T., Rees, C.A., Cherry, J.M., Botstein, D., Brown, P.O. and Alizadeh, A.A. (2003) SOURCE: A Unified Genomic Resource of Functional Annotations, Ontologies, and Gene Exrpression Data. Nucleic Acids Research, 31, 219-223.  
http://source-search.princeton.edu/cgi-bin/source/sourceSearch  
https://doi.org/10.1093/nar/gkg014</mixed-citation></ref><ref id="scirp.74217-ref19"><label>19</label><mixed-citation publication-type="other" xlink:type="simple">Joshi-Tope, G., Gillespie, M., Vastrik, I., D’Eustachio, P., Schmidt, E., de Bono, B., Jassal, B., Gopinah, G.R., Wu, G.R., Matthews, L., Lewis, S., Birney, E. and Stein, L. (2005) Reactome: A Knowledgebase of Biological Pathways. Nucleic Acids Research, 1, D428-D432.</mixed-citation></ref><ref id="scirp.74217-ref20"><label>20</label><mixed-citation publication-type="other" xlink:type="simple">Edgar, R., Domrachev, M. and Lash, A.E. (2002) Gene Expression Omnibus: NCBI Gene Expression and Hybridization Array Data Repository. Nucleic Acids Research, 30, 207-210. https://doi.org/10.1093/nar/30.1.207</mixed-citation></ref><ref id="scirp.74217-ref21"><label>21</label><mixed-citation publication-type="other" xlink:type="simple">Loi, S., Haibe-Kains, B., Desmedt, C., Lallemand, F., Tutt, A.M., Gillet, C., Ellis, P., Harris, A., Bergh, J., Foekens, J.A., Klijn, J.G., Larsimont, D., Buyse, M., Botempi, G., Delorenzi, M., Piccart, M.J. and Sotiriou, C. (2007) Definition of Clinically Distinct Molecular Subtypes in Estrogen Receptor-Positive Breast Carcinomas through Genomic Grade. Journal of Clinical Oncology, 25, 1239-1246.  
https://doi.org/10.1200/JCO.2006.07.1522</mixed-citation></ref><ref id="scirp.74217-ref22"><label>22</label><mixed-citation publication-type="other" xlink:type="simple">MillerL, D., Smeds, J., George, J., Vega, V.B., Vergara, L., Ploner, A., Pawitan, Y., Hall, P., Klaar, S., Liu, E.T. and Bergh, J. (2005) An Expression Signature for p53 Status in Human Breast Cancer Predicts Mutation Status, Transcriptional Effects, and Patient Survival. Proceedings of the National Academy of Sciences of the United States of America, 102, 13550-13555. https://doi.org/10.1073/pnas.0506230102</mixed-citation></ref><ref id="scirp.74217-ref23"><label>23</label><mixed-citation publication-type="other" xlink:type="simple">Desmedt, C., Piette, F., Loi, S., Wang, Y., Lallemand, F., Haibe-Kains, B., Delorenzi, M., d’Assignies, M.S., Bergh, J., Lidereau, R., Ellis, P., Harris, A.L., Klijn, J.G., Foekens, J.A., Cardoso, F., Piccart, M.J., Buyse, M. and Sotiriou, C. (2007) Strong Time Dependence of the 76-Gene Prognostic Signature for Node-Negative Breast Cancer Patients in the TRANSBIG Multicenter Independent Validation Series. Clinical Cancer Research, 13, 3207-3214. https://doi.org/10.1158/1078-0432.CCR-06-2765</mixed-citation></ref><ref id="scirp.74217-ref24"><label>24</label><mixed-citation publication-type="other" xlink:type="simple">Minn, A.J., Gupta, G.P., Siegel, P.M., Bos, P.D., Shu, W., Giri, D.D., Viale, A., Olshen, A.B., Gerald, W.L. and Massaqué, J. (2005) Genes That Mediate Breast Cancer Metastasis to Lung. Nature, 436, 518-524. https://doi.org/10.1038/nature03799</mixed-citation></ref><ref id="scirp.74217-ref25"><label>25</label><mixed-citation publication-type="other" xlink:type="simple">Chin, K., DeVries, S., Fridlyand, J., Spellman, P.T., Roydasgupta, R., Kuo, W.L., Lapuk, A., Neve, R.M., Qian, Z., Ryder, T., Chen, F., Feiler, H., Tokuyasu, T., Kingsley, C., Dairkee, S., Meng, Z., Chew, K., Pinkel, D., Jain, A., Ljung, B.M., Esseman, L., Albertson, D.G., Waldman, F.M. and Gray, J.W. (2006) Genomic and Transcriptional Aberrations Linked to Breast Cancer Pathophysiologies. Cancer Cell, 10, 529-541. https://doi.org/10.1016/j.ccr.2006.10.009</mixed-citation></ref><ref id="scirp.74217-ref26"><label>26</label><mixed-citation publication-type="other" xlink:type="simple">Zhao, X., Rodland, E.A., Sorlie, T., Vollan, H.K.M., Russnes, H.G., Kristensen, V.N., Lingjorde, O.C. and Borresen-Dale, A.L. (2014) Systematic Assessment of Prognstic Gene Signatures for Breast Cancer Shows Distinct Influence of Time and ER Status. BMC Cancer, 14, 211. https://doi.org/10.1186/1471-2407-14-211</mixed-citation></ref><ref id="scirp.74217-ref27"><label>27</label><mixed-citation publication-type="book" xlink:type="simple">Bolstad, B.M., Collin, F., Brettschneider, J., Simpson, K., Cope, L., Irizarry, R.A. and Speed, T.P. (2005) Quality Assessment of Affymetrix Gene Chip Data. In: Gentleman, R., Carey, V., Huber, W., Irizarry, R. and Dudoit, S., Eds., Bioinformatics and Computational Biology Solutions Using R and Bioconductor Statistics for Biology and Health, Springer, Berlin, 33-47. https://doi.org/10.1007/0-387-29362-0_3</mixed-citation></ref><ref id="scirp.74217-ref28"><label>28</label><mixed-citation publication-type="other" xlink:type="simple">Irizarry, R.A., Bolstad, B.M., Collin, F., Cope, L.M., Hobbs, B. and Speed, T.P. (2003) Summaries of Affymetrix Gene Chip Probe Level Data. Nucleic Acids Research, 31, e15. https://doi.org/10.1093/nar/gng015</mixed-citation></ref><ref id="scirp.74217-ref29"><label>29</label><mixed-citation publication-type="other" xlink:type="simple">Sims, A.H., Smethurst, G.J., Hey, Y., Okoniewski, M.J., Pepper, S.D., Howell, A., Miller, C.J. and Clarke, R.B. (2008) The Removal of Multiplicative, Systematic Bias Allows Integration of Breast Cancer Gene Expression Datasets—Improving Meta-Analysis and Prediction of Prognosis. BMC Medical Genomics, 1, 42.  
https://doi.org/10.1186/1755-8794-1-42</mixed-citation></ref><ref id="scirp.74217-ref30"><label>30</label><mixed-citation publication-type="other" xlink:type="simple">Perou, C.M., Sorlie, T., Eisen, M.B., van de Rijn, M., Jeffrey, S.S., Rees, C.A., Pollack, J.R., Ross, D.T., Johnsen, H., Akslen, L.A., Fluge, O., Pergamenschikov, A., Williams, C., Zhu, S.X., Lonning, P.E., Borresen-Dale, A.L., Brown, P.O. and Botstein, D. (2000) Molecular Portraits of Human Breast Tumours. Nature, 406, 747-752.  
https://doi.org/10.1038/35021093</mixed-citation></ref><ref id="scirp.74217-ref31"><label>31</label><mixed-citation publication-type="other" xlink:type="simple">Teschendorff, A.E., Journée, M., Absil, P.A., Sepulchre, R. and Caldas, C. (2007) Elucidating the Altered Transcriptional Programs in Breast Cancer Using Independent Component Analysis. PLoS Computational Biology, 3, e161.  
https://doi.org/10.1371/journal.pcbi.0030161</mixed-citation></ref></ref-list></back></article>