<?xml version="1.0" encoding="UTF-8"?><!DOCTYPE article  PUBLIC "-//NLM//DTD Journal Publishing DTD v3.0 20080202//EN" "http://dtd.nlm.nih.gov/publishing/3.0/journalpublishing3.dtd"><article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" dtd-version="3.0" xml:lang="en" article-type="research article"><front><journal-meta><journal-id journal-id-type="publisher-id">OJS</journal-id><journal-title-group><journal-title>Open Journal of Statistics</journal-title></journal-title-group><issn pub-type="epub">2161-718X</issn><publisher><publisher-name>Scientific Research Publishing</publisher-name></publisher></journal-meta><article-meta><article-id pub-id-type="doi">10.4236/ojs.2016.64053</article-id><article-id pub-id-type="publisher-id">OJS-69891</article-id><article-categories><subj-group subj-group-type="heading"><subject>Articles</subject></subj-group><subj-group subj-group-type="Discipline-v2"><subject>Physics&amp;Mathematics</subject></subj-group></article-categories><title-group><article-title>
 
 
  Challenges Analyzing RNA-Seq Gene Expression Data
 
</article-title></title-group><contrib-group><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Liliana</surname><given-names>López-Kleine</given-names></name><xref ref-type="aff" rid="aff1"><sup>1</sup></xref></contrib><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Cristian</surname><given-names>González-Prieto</given-names></name><xref ref-type="aff" rid="aff1"><sup>1</sup></xref></contrib></contrib-group><aff id="aff1"><addr-line>Department of Statistics, Universidad Nacional de Colombia—Sede Bogotá, Bogotá, Colombia</addr-line></aff><pub-date pub-type="epub"><day>22</day><month>07</month><year>2016</year></pub-date><volume>06</volume><issue>04</issue><fpage>628</fpage><lpage>636</lpage><history><date date-type="received"><day>25</day>	<month>June</month>	<year>2016</year></date><date date-type="rev-recd"><day>accepted</day>	<month>16</month>	<year>August</year>	</date><date date-type="accepted"><day>19</day>	<month>August</month>	<year>2016</year></date></history><permissions><copyright-statement>&#169; Copyright  2014 by authors and Scientific Research Publishing Inc. </copyright-statement><copyright-year>2014</copyright-year><license><license-p>This work is licensed under the Creative Commons Attribution International License (CC BY). http://creativecommons.org/licenses/by/4.0/</license-p></license></permissions><abstract><p>
 
 
  The analysis of messenger Ribonucleic acid obtained through sequencing techniques (RNA-se- quencing) data is very challenging. Once technical difficulties have been sorted, an important choice has to be made during pre-processing: Two different paths can be chosen: Transform RNA- sequencing count data to a continuous variable or continue to work with count data. For each data type, analysis tools have been developed and seem appropriate at first sight, but a deeper analysis of data distribution and structure, are a discussion worth. In this review, open questions regarding RNA-sequencing data nature are discussed and highlighted, indicating important future research topics in statistics that should be addressed for a better analysis of already available and new appearing gene expression data. Moreover, a comparative analysis of RNAseq count and transformed data is presented. This comparison indicates that transforming RNA-seq count data seems appropriate, at least for differential expression detection.
 
</p></abstract><kwd-group><kwd>RNA-Seq Analysis</kwd><kwd> Count Data</kwd><kwd> Preprocessing</kwd><kwd> Differential Expression</kwd><kwd> Gene Co-Expression  Network</kwd></kwd-group></article-meta></front><body><sec id="s1"><title>1. Introduction</title><p>This sequencing of messenger RNA transcripts (RNA-seq) is a recently developed approach to gene expression or transcriptome profiling that uses deep-sequencing technologies. Studies using this method have allowed assessing the complexity of transcriptomes. RNA-seq also provides more precise measurement of levels of transcripts and their isoforms than other methods based on hybridization (such as microarrays), that were used previously, but poses also new challenges [<xref ref-type="bibr" rid="scirp.69891-ref1">1</xref>] . Great issues concerning the identification of the real number of RNA fragments taking into account isoforms, mitochondrial and ribosomal RNA have appear but are beyond the interest of this review. Several satisfactory developments assure a good characterization of RNA-seq transcripts [<xref ref-type="bibr" rid="scirp.69891-ref2">2</xref>] to be used for increasing comprehension of biological knowledge. Here, statistical challenges that arise once RNA counts are obtained (after mapping), are discussed.</p><sec id="s1_1"><title>1.1. RNA-Seq Statistical Challenges</title><p>During the last 15 years, statistical research has been done, driven by the need to analyze properly data from high-throughput genomic assays, in particular microarrays. In the last five years, high-throughput sequencing technology has been changing the face of biological research, replacing the old microarray technology. As mentioned by Datta and Net-tleton (2014), with any new high-throughput technology come new data analytic challenges that have been solved in proposing new analytical methods based on novel and older concepts of error rate control for testing multiple hypotheses, various adaptations of existing [<xref ref-type="bibr" rid="scirp.69891-ref3">3</xref>] .</p><p>Analyzing mapped reads is a major challenge than continuous microarray data, because count data has to be modeled using discrete distributions that had not been used so far for gene expression data analysis. Moreover, an issue concerning dimensionality appears, because often less replicate samples are available than were for microarray data. Even though, data produced using these technologies are proving to be the most informative of any thus far, very little attention has been paid to fundamental design aspects of data collection and analysis, namely sampling, randomization, replication, and blocking [<xref ref-type="bibr" rid="scirp.69891-ref4">4</xref>] (Auer and Doerge, 2010).</p><p>RNA-seq data and its proper analysis has an enormous potential to promote genomic research and enhance understanding of biological processes, but a detailed comprehension of this technology and the type of data produced is needed in order to obtain confident results. In this review open questions regarding RNA-sequenc- ing data nature are discussed and highlighted, indicating important future research topics that should be addressed for a better analysis of already available and new appearing gene expression data. A comparative analysis of RNAseq count and transformed data is also presented allowing interesting results that make count data transformations generally applicable.</p></sec><sec id="s1_2"><title>1.2. RNA-Seq Data Preprocessing</title><p>As for microarray data, several similar steps of preprocessing need to be achieved before RNA-seq data can be used for analysis. Nevertheless, two main paths can be chosen for RNA-seq data. The first one is to transform the count data to a continuous variable using RPKM (reads per kilobase per million mapped reads) as originally introduced by [<xref ref-type="bibr" rid="scirp.69891-ref5">5</xref>] and the second path is to continue statistical analysis with count data as it is. Each path requires different analytic tools because each type of data need to be treated in a different way.</p><sec id="s1_2_1"><title>1.2.1. Transformation of RNA-Seq Count Data into a Continuous Variable</title><p>In the case of transformation to RPKM, the preprocessing begins by equalizing sequencing depths, to compare the ex-pression measures across different genes and samples. These “normalization” is made by dividing counts by gene length (a variable) and the total amount of reads in each experiment (a constant). Then, analysis conceived for continuous microarray data are applied without apparent concern about the distribution differences between these transformations on count data compared to transformations done on continuous microarray data. Several authors have discussed inconsistencies but no deep discussion on distribution of RPKM has been done [<xref ref-type="bibr" rid="scirp.69891-ref6">6</xref>] - [<xref ref-type="bibr" rid="scirp.69891-ref8">8</xref>] .</p><p>More realistic models than RPKM addressed the case for multiple isoforms [<xref ref-type="bibr" rid="scirp.69891-ref9">9</xref>] proposing a Poisson distribution for counts and create a continuous variable during the mapping process. The major assumption was that the number of reads coming from an exon of a certain length is Poisson where the mean is a normalized function of the exon length. The first insert length model extended the approach of [<xref ref-type="bibr" rid="scirp.69891-ref9">9</xref>] to paired-end reads [<xref ref-type="bibr" rid="scirp.69891-ref10">10</xref>] . Their algorithm was made available through the software called Cuffllinks in early 2011 and uses FPKM (fragments instead of reads). FPKM is based on a probabilistic assignment method indicating the probability that a fragment selected at random originates from a given transcript [<xref ref-type="bibr" rid="scirp.69891-ref11">11</xref>] . Similarities between both types of estimations have been reported, but again no known discussion on data distribution of FPKM has been undertaken. FPKM is analyzed with Cufflinks developed especially for this data transformation of RNA-seq counts. Therefore, use of FPKM seems a more restricted transformation than RPKM.</p></sec><sec id="s1_2_2"><title>1.2.2. Preprocessing or RNA-Seq Count Data without Transformation</title><p>Raw read counts from different experiments are not directly comparable without adjustment for technical variation due to sequencing depth (a process also called normalization). Complex normalization schemes for RNA- seq data have been proposed by [<xref ref-type="bibr" rid="scirp.69891-ref12">12</xref>] - [<xref ref-type="bibr" rid="scirp.69891-ref14">14</xref>] Bullard et al., 2010 Robinson and Oshlack, 2010. Sample specific normalizations are combined with library sizes in these methods. Trimmed mean of M-values normalization (TMM) [<xref ref-type="bibr" rid="scirp.69891-ref14">14</xref>] and the normalization scheme provided by [<xref ref-type="bibr" rid="scirp.69891-ref13">13</xref>] are among the most efficient and easy to use. When these path is chosen, new methods developed for counts are used for analysis.</p></sec></sec><sec id="s1_3"><title>1.3. Maintaining the Integrity of the Specifications</title><p>Depending on the type of normalization or transformation that has been undertaken to raw RNA-seq data, several tools are available for subsequent analysis.</p><sec id="s1_3_1"><title>1.3.1. Differential Expression for RPKM and FPKM</title><p>Generally, methods developed for continuous microarray data are applied on RPKM [<xref ref-type="bibr" rid="scirp.69891-ref5">5</xref>] transformed data. For FPKM, [<xref ref-type="bibr" rid="scirp.69891-ref11">11</xref>] have developed the algorithm Cuffdiff for differential expression. It estimates the expression transcript-level resolution and controls for variability across replicates. Following the number of citations of these two articles during 2016 (google academics consulted on march 20th 2016), both types of transformations are almost equally used (157 citations for [<xref ref-type="bibr" rid="scirp.69891-ref5">5</xref>] and 140 for [<xref ref-type="bibr" rid="scirp.69891-ref11">11</xref>] ). No methods developed based on distributional properties have been proposed but are also rare for microarray data [<xref ref-type="bibr" rid="scirp.69891-ref15">15</xref>] .</p></sec><sec id="s1_3_2"><title>1.3.2. Count Data</title><p>Several statistical methods and related R packages for differential gene expression analysis based on RNA-seq data have been developed over the years. The packages DESeq [<xref ref-type="bibr" rid="scirp.69891-ref13">13</xref>] and EdgeR [<xref ref-type="bibr" rid="scirp.69891-ref16">16</xref>] are a popular choice amongst users of RNA-seq. BaySeq [<xref ref-type="bibr" rid="scirp.69891-ref17">17</xref>] is a Bioconductor package that identifies differential expression using high throughput sequencing data via empirical Bayesian methods. Another method, called TSPM, is based on a two-stage Poisson model [<xref ref-type="bibr" rid="scirp.69891-ref18">18</xref>] . These four methods are compared by [<xref ref-type="bibr" rid="scirp.69891-ref19">19</xref>] . The results suggest that baySeq performs best in terms of ranking genes according to their significance to be declared differentially expressed. Both edgeR and DESeq perform similarly and close to baySeq. The results from TSPM are most variable and often the poorest when the number of replicates is small [<xref ref-type="bibr" rid="scirp.69891-ref19">19</xref>] .</p><p>One year later, six methods where compared by [<xref ref-type="bibr" rid="scirp.69891-ref20">20</xref>] : DESeq, DEGseq, edgeR, NBPSeq, TSPM and baySeq using both real and simulated data with the result that all six methods produce similar fold changes and reasonable overlapping of differentially expressed genes based on p-value, edgeR being little bit superior. However, all six methods suffer from over-sensitivity as reported by the authors.</p><p>A recent and not yet popular method based on a hierarchical negative binomial model (borrowing information across gene-variety means and across gene-specific over dispersion parameters) and using a computationally tractable empirical Bayes approach to inference has been proposed by [<xref ref-type="bibr" rid="scirp.69891-ref20">20</xref>] .</p><p>Also, threshold-independent methods for particular cases, for example the detection of marginal expression changes in cognitively stratified patients at different disease stages, have been developed [<xref ref-type="bibr" rid="scirp.69891-ref21">21</xref>] . This approach is based on the comparison of the distribution of changes in a well-defined gene group with the global distribution of the experiment [<xref ref-type="bibr" rid="scirp.69891-ref21">21</xref>] .</p><p>The question of preprocessing and making samples comparable is not yet completely solved for microarrays and very far from being solved for RNA-seq data. RPKM seems like a general solution for prokaryotes because the nature of transformed data allows using methods developed for microarrays, but limitations have been shown, especially for eukaryotes [<xref ref-type="bibr" rid="scirp.69891-ref10">10</xref>] . Algorithms for all other particular cases are being developed. Although RNA-seq data has been analyzed and biological conclusions have been drawn, a main question is still unanswered: How general is RPKM, can it really be analyzed with algorithms developed for microarray data? An answer to this question, based on empirical observations, is given in Section 4.</p><p>A second open question is: Can we combine microarray data and RNA-seq data to increase biological knowledge or do we need to repeat all microarray experiments? Assess this comparison is tricky because influence of experiment type, data type and method would confounded. Nevertheless, even a partial answer to this question would be of general interest for biologists and applied statisticians.</p></sec></sec></sec><sec id="s2"><title>2. Comparison of Differential Gene Expression between RPKM and Count Data</title><sec id="s2_1"><title>2.1. Empirical Characterization of RPKM Data Distribution</title><sec id="s2_1_1"><title>2.1.1. Simulation of Count Data</title><p>The function make Example Count Data Set from the DeSeq Bioconductor package [<xref ref-type="bibr" rid="scirp.69891-ref13">13</xref>] was used to simulate ten RNAseq count tables with 20,000 genes (rows) and 100 samples (50 treatments and 50 controls) as follows:</p><p>1) Mean expression values were sampled from an exponential distribution with parameter 1/250.</p><p>2) Once graphically verified, one table of count data with 50 treatments and 50 controls, and a proportion of differentially expressed genes of 0.3 was constructed.</p><p>3) Mean values of gene expression are divided in two conditions (treatment and control) such that the log2 fold-change is still centered in 0 and with a standard deviation of 2.</p><p>4) Counts were sampled from a negative binomial distribution with mean values of gene expression mentioned before and multiplied by size factors between 0.1 and 2.55. The size parameter of the binomial distribution was set at 1/0.2.</p></sec><sec id="s2_1_2"><title>2.1.2. Simulations of Gene Length</title><p>Using the real gene length of three organisms obtained from NCBI (Escherichia coli, Homo sapiens and Arabidobsis thaliana) the package fitdistrplus [<xref ref-type="bibr" rid="scirp.69891-ref22">22</xref>] was used to adjust a distribution to these data. We concluded that the gene length distribution is gamma with outliers.</p><p>In order to simulate gene lengths of the 20.000 genes present in the count data table, the mean of the gamma distribution was estimated based on mean gene length of the three organisms we had taken as example. Thousand extreme values were sampled from a distribution with different mean (third quartile of 19.000 simulated gene lengths) to end up with 20.000 gene lengths.</p></sec><sec id="s2_1_3"><title>2.1.3. RPKM Transformed Data</title><p>One count tables and gene lengths were generated as explained above. Those counts were then transformed to continuous RPKM data [<xref ref-type="bibr" rid="scirp.69891-ref5">5</xref>] dividing each count value by the simulated gene length and the total sum of counts for each column (sample). Finally, RPKM data were log2 transformed. The result of the density adjusted to 100 samples of one of the simulations (treatments and controls) can be observed in <xref ref-type="fig" rid="fig1">Figure 1</xref>.</p></sec><sec id="s2_1_4"><title>2.1.4. Verification of RPKM Distribution</title><p>With the estimated mean and variance parameters of each of the 100 simulated samples (50 treatments, 50 controls with 0.3 proportion of differentially expressed genes), 100 samples of normal distribution were generated and adjustment was tested using Kolmogorov-Smirnov test. As shown in <xref ref-type="fig" rid="fig2">Figure 2</xref>, most p-values are not significant, indicating that the transformed RPKM data can be considered having a normal distribution.</p><fig id="fig1"  position="float"><label><xref ref-type="fig" rid="fig1">Figure 1</xref></label><caption><title> Densities of 100 simulated samples of RPKM transformed count data</title></caption><graphic mimetype="image"   position="float"  xlink:type="simple"  xlink:href="http://html.scirp.org/file/7-1240742x6.png"/></fig><fig id="fig2"  position="float"><label><xref ref-type="fig" rid="fig2">Figure 2</xref></label><caption><title> Boxplot of 100 p-values obtained from Kolmogorov-Smirnov goodness to fit test comparing transformed RPKM data to a theoretical normal distribution. Red line indicates a p-value of 0.05</title></caption><graphic mimetype="image"   position="float"  xlink:type="simple"  xlink:href="http://html.scirp.org/file/7-1240742x7.png"/></fig></sec></sec><sec id="s2_2"><title>2.2. Analysis of RNA-Seq Count Data without Transformation</title><sec id="s2_2_1"><title>2.2.1. Real Data</title><p>Three available data sets (NCBI, Genome Expression Omnibus) comparing two conditions (controls vs. treatments) were analyzed using DeSeq standard normalization and detection of differentially expressed genes [<xref ref-type="bibr" rid="scirp.69891-ref13">13</xref>] :</p><p>- GSE67402: Controlled measurement and comparative analysis of cellular components in E. coli reveals broad regulatory changes under long-term starvation [<xref ref-type="bibr" rid="scirp.69891-ref23">23</xref>] .</p><p>- GSE76268: Integration of ATAC-seq and RNA-seq Identies Human Alpha Cell and Beta Cell Signature Genes [<xref ref-type="bibr" rid="scirp.69891-ref24">24</xref>] .</p><p>- GSE72548: RNA-seq analysis of Arabidopsis thaliana wild-type roots and type-A arr3, 4, 5, 6, 7, 8, 9, 15 mutant roots non-infected and infected with Heterodera schachtii nematodes [<xref ref-type="bibr" rid="scirp.69891-ref25">25</xref>] . Results are shown in <xref ref-type="table" rid="table1"><xref ref-type="table" rid="table">Table </xref>1</xref>.</p></sec><sec id="s2_2_2"><title>2.2.2. Simulated Data</title><p>Deseq results on 10 simulated count data tables are shown in <xref ref-type="table" rid="table2"><xref ref-type="table" rid="table">Table </xref>2</xref> and indicate (as expected) approximately 30% of genes with differential expression.</p></sec></sec><sec id="s2_3"><title>2.3. Analysis of RPKM Transformed Data</title><sec id="s2_3_1"><title>2.3.1. Real Data</title><p>Count tables of the three experiments used above, were RPKM and log2 transformed to be analyzed using Significance Analysis of Microarray (SAM) [<xref ref-type="bibr" rid="scirp.69891-ref26">26</xref>] , a standard method for microarray data analysis. Moreover, we retained the propor-tion of genes common to both analysis (<xref ref-type="table" rid="table3"><xref ref-type="table" rid="table">Table </xref>3</xref> and <xref ref-type="table" rid="table4"><xref ref-type="table" rid="table">Table </xref>4</xref>). The results indicate that the proportion and identity of genes identified on count data using Deseq or on RPKM transformed data are equivalent.</p></sec><sec id="s2_3_2"><title>2.3.2. Conclusion on RPKM vs. Count Data Comparison</title><p>The most important conclusion of the above comparison is that RPKM transformation with posterior log2 normalization conduces to a data distribution which is very similar to continuous microarray data and that tools</p><table-wrap id="table1" ><label><xref ref-type="table" rid="table1"><xref ref-type="table" rid="table">Table </xref>1</xref></label><caption><title> <xref ref-type="table" rid="table">Table </xref>indicating number of differentially expressed genes (DEG), proportion of the total (POT) using Deseq on real data. FDR for these results is lower than 0.0001. COUNT-RPKM indicates the proportion of genes that were detected by both methods</title></caption><table><tbody><thead><tr><th align="center" valign="middle" >Experiment</th><th align="center" valign="middle" >DEG</th><th align="center" valign="middle" >POT</th><th align="center" valign="middle" >COUNT-RPKM</th></tr></thead><tr><td align="center" valign="middle" >GSE67402</td><td align="center" valign="middle" >776</td><td align="center" valign="middle" >0.173</td><td align="center" valign="middle" >0.834</td></tr><tr><td align="center" valign="middle" >GSE76268</td><td align="center" valign="middle" >1350</td><td align="center" valign="middle" >0.067</td><td align="center" valign="middle" >0.951</td></tr><tr><td align="center" valign="middle" >GSE72548</td><td align="center" valign="middle" >9955</td><td align="center" valign="middle" >0.296</td><td align="center" valign="middle" >0.798</td></tr></tbody></table></table-wrap><table-wrap id="table2" ><label><xref ref-type="table" rid="table2"><xref ref-type="table" rid="table">Table </xref>2</xref></label><caption><title> <xref ref-type="table" rid="table">Table </xref>indicating number of differentially expressed genes (DEG), proportion of the total (POT) Deseq on simulations. FDR for these results is lower than 0.0001. COUNT-RPKM indicates the proportion of genes that were detected by both methods</title></caption><table><tbody><thead><tr><th align="center" valign="middle" >Simulation</th><th align="center" valign="middle" >DEG</th><th align="center" valign="middle" >POT</th><th align="center" valign="middle" >COUNT-RPKM</th></tr></thead><tr><td align="center" valign="middle" >1</td><td align="center" valign="middle" >5458</td><td align="center" valign="middle" >0.273</td><td align="center" valign="middle" >0.972</td></tr><tr><td align="center" valign="middle" >2</td><td align="center" valign="middle" >5969</td><td align="center" valign="middle" >0.298</td><td align="center" valign="middle" >0.903</td></tr><tr><td align="center" valign="middle" >3</td><td align="center" valign="middle" >5952</td><td align="center" valign="middle" >0.297</td><td align="center" valign="middle" >0.902</td></tr><tr><td align="center" valign="middle" >4</td><td align="center" valign="middle" >6002</td><td align="center" valign="middle" >0.300</td><td align="center" valign="middle" >0.885</td></tr><tr><td align="center" valign="middle" >5</td><td align="center" valign="middle" >5714</td><td align="center" valign="middle" >0.286</td><td align="center" valign="middle" >0.917</td></tr><tr><td align="center" valign="middle" >6</td><td align="center" valign="middle" >6323</td><td align="center" valign="middle" >0.316</td><td align="center" valign="middle" >0.850</td></tr><tr><td align="center" valign="middle" >7</td><td align="center" valign="middle" >5486</td><td align="center" valign="middle" >0.274</td><td align="center" valign="middle" >0.966</td></tr><tr><td align="center" valign="middle" >8</td><td align="center" valign="middle" >6011</td><td align="center" valign="middle" >0.301</td><td align="center" valign="middle" >0.896</td></tr><tr><td align="center" valign="middle" >9</td><td align="center" valign="middle" >5923</td><td align="center" valign="middle" >0.296</td><td align="center" valign="middle" >0.902</td></tr><tr><td align="center" valign="middle" >10</td><td align="center" valign="middle" >5889</td><td align="center" valign="middle" >0.294</td><td align="center" valign="middle" >0.892</td></tr></tbody></table></table-wrap><table-wrap id="table3" ><label><xref ref-type="table" rid="table3"><xref ref-type="table" rid="table">Table </xref>3</xref></label><caption><title> <xref ref-type="table" rid="table">Table </xref>indicating number of differentially expressed genes (DEG) and proportion of the total (POT) in real data sets using SAM. FDR for these results is lower than 0.0001</title></caption><table><tbody><thead><tr><th align="center" valign="middle" >Experiment</th><th align="center" valign="middle" >DEG</th><th align="center" valign="middle" >POT</th></tr></thead><tr><td align="center" valign="middle" >GSE67402</td><td align="center" valign="middle" >834</td><td align="center" valign="middle" >0.186</td></tr><tr><td align="center" valign="middle" >GSE76268</td><td align="center" valign="middle" >967</td><td align="center" valign="middle" >0.048</td></tr><tr><td align="center" valign="middle" >GSE72548</td><td align="center" valign="middle" >10,796</td><td align="center" valign="middle" >0.321</td></tr></tbody></table></table-wrap><table-wrap id="table4" ><label><xref ref-type="table" rid="table4"><xref ref-type="table" rid="table">Table </xref>4</xref></label><caption><title> <xref ref-type="table" rid="table">Table </xref>indicating number of differentially expressed genes (DEG), proportion of the total (POT) in simulations using SAM. FDR for these results is lower than 0.0001</title></caption><table><tbody><thead><tr><th align="center" valign="middle" >Simulation</th><th align="center" valign="middle" >DEG</th><th align="center" valign="middle" >POT</th></tr></thead><tr><td align="center" valign="middle" >1</td><td align="center" valign="middle" >5306</td><td align="center" valign="middle" >0.265</td></tr><tr><td align="center" valign="middle" >2</td><td align="center" valign="middle" >5392</td><td align="center" valign="middle" >0.269</td></tr><tr><td align="center" valign="middle" >3</td><td align="center" valign="middle" >5366</td><td align="center" valign="middle" >0.268</td></tr><tr><td align="center" valign="middle" >4</td><td align="center" valign="middle" >5311</td><td align="center" valign="middle" >0.266</td></tr><tr><td align="center" valign="middle" >5</td><td align="center" valign="middle" >5242</td><td align="center" valign="middle" >0.262</td></tr><tr><td align="center" valign="middle" >6</td><td align="center" valign="middle" >5378</td><td align="center" valign="middle" >0.2689</td></tr><tr><td align="center" valign="middle" >7</td><td align="center" valign="middle" >5302</td><td align="center" valign="middle" >0.265</td></tr><tr><td align="center" valign="middle" >8</td><td align="center" valign="middle" >5383</td><td align="center" valign="middle" >0.269</td></tr><tr><td align="center" valign="middle" >9</td><td align="center" valign="middle" >5342</td><td align="center" valign="middle" >0.267</td></tr><tr><td align="center" valign="middle" >10</td><td align="center" valign="middle" >5253</td><td align="center" valign="middle" >0.263</td></tr></tbody></table></table-wrap><p>developed for the analysis of them can be used and will conduct to almost similar results. Moreover, the analysis show that this statement is true: Results obtained using methods developed for count data were very similar to results obtained when count data was RPKM transformed and tools developed for continuous microarray data are used. Therefore, it is safe to conclude that RPKM transformation conducts to a similar normalization and that analysis tools developed for microarrays can be used. This also is encouraging regarding the combined analysis of RNA-seq and microarray data. Nevertheless, caution is still required, until an analytic characterization of RPKM transformation is done to confirm the here presented results.</p></sec><sec id="s2_3_3"><title>2.3.3. Tools for RNA-Seq Co-Expression Networks</title><p>Reconstructing gene or protein networks is a very important tool in deciphering molecular mechanisms. One of the most important data source for this reconstruction has been gene expression data because it reflects coordinated activity of different genes at the same time. Only few examples of gene reconstruction based on RNA-seq data exist. As for assess-ing differential expression, two separate pathways have to be taken, depending if counts or transformed data is used. Even less discussion on this subject than for detection of differential expression is found in literature. Additionally to the already mentioned challenges, difficulties with the gene profile similarity estimation appear, which should be a measure suitable for count data if counts are not transformed.</p><p>Some studies addressed the question of comparing gene co-expression network reconstruction with RNA-seq data, applying Pearson correlation to both types of data avoiding discussion on usefulness of this similarity measure. Iancu et al. [<xref ref-type="bibr" rid="scirp.69891-ref27">27</xref>] conducted a study comparing co-expression networks constructed with count data to networks constructed with microarray data using the same method for different data types. They concluded that the RNA-seq coexpression network displayed overlapping structure with the microarray network. Pearson correlations from RNA-seq data were higher and therefore, higher network connectivity, heterogeneity and centrality was observed in the RNA-seq network. A more recent study constructs co-expression networks using also Pearson correlation on a huge Arabidopsis thaliana data set [<xref ref-type="bibr" rid="scirp.69891-ref22">22</xref>] . The authors observe sensitivity to variance stabilizing transformations on RNA-seq data but overall similarities between RNAseq networks and microarray networks.</p><p>Gene co-expression construction is a complex procedure and still open questions exist when used on microarray data [<xref ref-type="bibr" rid="scirp.69891-ref28">28</xref>] . These difficulties need to be addressed for RNA-seq data as well, but the most important open question regarding RNA-seq networks, is: What similarity measure is more appropriate. Is it right to apply Pearson correlation? Or is it better to use similarity measures for count data, perhaps test and adapt association measures like has been done for other type of data [<xref ref-type="bibr" rid="scirp.69891-ref29">29</xref>] ? How do mutual information and other non-linear similarity measures behave? A research in this sense would also shed some light on gene classification and clustering, which is performed on similarity or distance matrices between genes and widely used for annotation purposes and functional prediction.</p></sec></sec></sec><sec id="s3"><title>3. Conclusion</title><p>Although, several studies have used and analyzed RNA-seq data with traditional or new developed statistical methods, open questions on data type and nature remain still unanswered and should receive more attention. Throughout this review we have highlighted some open questions and addressed one of them regarding the RPKM tranformation of count data. Adressing all statistical issues related to RNA-seq data analysis is especially important in order to achieve confident results for assessing biological knowledge extracted from gene expression data, which has proven so far to be highly informative.</p></sec><sec id="s4"><title>Ethical Statement</title><p>The authors declare that they do not have any conflict of interest. This article does not contain any studies with human participants or animals performed by any of the authors.</p></sec><sec id="s5"><title>Cite this paper</title><p>Liliana L&#243;pez-Kleine,Cristian Gonz&#225;lez-Prieto, (2016) Challenges Analyzing RNA-Seq Gene Expression Data. Open Journal of Statistics,06,628-636. doi: 10.4236/ojs.2016.64053</p></sec></body><back><ref-list><title>References</title><ref id="scirp.69891-ref1"><label>1</label><mixed-citation publication-type="other" xlink:type="simple">Wang, Z., Gerstein, M. and Snyder, M. (2009) RNA-Seq: A Revolutionary Tool for Transcriptomics. Nature Reviews Genetics, 10, 57-63. http://dx.doi.org/10.1038/nrg2484</mixed-citation></ref><ref id="scirp.69891-ref2"><label>2</label><mixed-citation publication-type="other" xlink:type="simple">Ozsolak, F. and Milos, P.M. (2011) RNA Sequencing: Advances, Challenges and Opportunities. Nature Reviews Genetics, 12, 87-98. http://dx.doi.org/10.1038/nrg2934</mixed-citation></ref><ref id="scirp.69891-ref3"><label>3</label><mixed-citation publication-type="book" xlink:type="simple">Datta, S. and Nettleton, D., Eds. (2014) Statistical Analysis of Next Generation Sequencing Data. Springer, New York.  
http://dx.doi.org/10.1007/978-3-319-07212-8</mixed-citation></ref><ref id="scirp.69891-ref4"><label>4</label><mixed-citation publication-type="other" xlink:type="simple">Auer, P.L. and Doerge, R.W. (2010) Statistical Design and Analysis of RNA Sequencing Data. Genetics, 185, 405-416.  
http://dx.doi.org/10.1534/genetics.110.114983</mixed-citation></ref><ref id="scirp.69891-ref5"><label>5</label><mixed-citation publication-type="other" xlink:type="simple">Mortazavi, A., Williams, B.A., McCue, K., Schaeffer, L. and Wold, B. (2008) Mapping and Quantifying Mammalian tran-Scriptomes by RNA-Seq. Nature Methods, 5, 621-628. http://dx.doi.org/10.1038/nmeth.1226</mixed-citation></ref><ref id="scirp.69891-ref6"><label>6</label><mixed-citation publication-type="other" xlink:type="simple">Kharchenko, P.V., Xi, R. and Park, P.J. (2011) Evidence for Dosage Compensation between X and Autosomes in Mammals. Nature Genetics, 43, 1167-1169. http://dx.doi.org/10.1038/ng.991</mixed-citation></ref><ref id="scirp.69891-ref7"><label>7</label><mixed-citation publication-type="other" xlink:type="simple">Wagner, G.P., Kin, K. and Lynch, V.J. (2012) Measurement of mRNA Abundance Using RNA-Seq Data: RPKM Measure is Inconsistent among Samples. Theory in Biosciences, 131, 281-285.  
http://dx.doi.org/10.1007/s12064-012-0162-3</mixed-citation></ref><ref id="scirp.69891-ref8"><label>8</label><mixed-citation publication-type="other" xlink:type="simple">Wang, L., Wang, S. and Li, W. (2012) RSeQC: Quality Control of RNA-Seq Experiments. Bioinformatics, 28, 2184-2185. http://dx.doi.org/10.1093/bioinformatics/bts356</mixed-citation></ref><ref id="scirp.69891-ref9"><label>9</label><mixed-citation publication-type="other" xlink:type="simple">Jiang, H. and Wong, W.H. (2009) Statistical Inferences for Isoform Expression in RNA-Seq. Bioinformatics, 25, 1026-1032. http://dx.doi.org/10.1093/bioinformatics/btp113</mixed-citation></ref><ref id="scirp.69891-ref10"><label>10</label><mixed-citation publication-type="other" xlink:type="simple">Trapnell, C., Pachter, L. and Salzberg, S.L. (2009) Tophat: Discovering Splice Junctions with RNA-Seq. Bioinformatics, 25, 1105-1111. http://dx.doi.org/10.1093/bioinformatics/btp120</mixed-citation></ref><ref id="scirp.69891-ref11"><label>11</label><mixed-citation publication-type="other" xlink:type="simple">Trapnell, C., Roberts, A., Goff, L., Pertea, G, Kim, D., Kelley, D.R., Pimentel, H., Salzberg, S.L., Rinn, J.L. and Pachter, L. (2012) Differential Gene and Transcript Expression Analysis of RNA-Seq Experiments with TopHat and Cufflinks. Nature Protocols, 7, 562-578. http://dx.doi.org/10.1038/nprot.2012.016</mixed-citation></ref><ref id="scirp.69891-ref12"><label>12</label><mixed-citation publication-type="other" xlink:type="simple">Bullard, J.H., Purdom, E., Hansen, K.D. and Dudoit, S. (2010) Evaluation of Statistical Methods for Normalization and Differential Expression in mRNA-Seq Experiments. BMC Bioinformatics, 11, 94.  
http://dx.doi.org/10.1186/1471-2105-11-94</mixed-citation></ref><ref id="scirp.69891-ref13"><label>13</label><mixed-citation publication-type="other" xlink:type="simple">Anders, S. and Huber, W. (2010) Differential Expression Analysis for Sequence Count Data. Genome Biology, 11, R106. http://dx.doi.org/10.1186/gb-2010-11-10-r106</mixed-citation></ref><ref id="scirp.69891-ref14"><label>14</label><mixed-citation publication-type="other" xlink:type="simple">Robinson, M.D. and Oshlack, A. (2010) A Scaling Normalization Method for Differential Expression Analysis of RNA-Seq Data. Genome Biology, 11, R25. http://dx.doi.org/10.1186/gb-2010-11-3-r25</mixed-citation></ref><ref id="scirp.69891-ref15"><label>15</label><mixed-citation publication-type="other" xlink:type="simple">Tovar, J.R., López-Kleine, L. and Ordo?ez, J.A. (2015) Identification of Global Gene Expression Shifts Using Microarray Data from Different Biological Conditions. Open Journal of Statistics, 5, 360-372.  
http://dx.doi.org/10.4236/ojs.2015.55038 </mixed-citation></ref><ref id="scirp.69891-ref16"><label>16</label><mixed-citation publication-type="other" xlink:type="simple">Robinson, M.D., McCarthy, D.J. and Smyth, G.K. (2010) EdgeR: A Bioconductor Package for Differential Expression Analysis of Digital Gene Expression Data. Bioinformatics, 26, 139-140.  
http://dx.doi.org/10.1093/bioinformatics/btp616</mixed-citation></ref><ref id="scirp.69891-ref17"><label>17</label><mixed-citation publication-type="other" xlink:type="simple">Hardcastle, T.J. and Kelly, K.A. (2010) BaySeq: Empirical Bayesian Methods for Identifying Differential Expression in Sequence Count Data. BMC Bioinformatics, 11, 422. http://dx.doi.org/10.1186/1471-2105-11-422</mixed-citation></ref><ref id="scirp.69891-ref18"><label>18</label><mixed-citation publication-type="other" xlink:type="simple">Kvam, V.M., Liu, P. and Si, Y. (2012) A Comparison of Statistical Methods for Detecting Differentially Expressed Genes from RNA-Seq Data. American Journal of Botany, 99, 248-256. http://dx.doi.org/10.3732/ajb.1100340</mixed-citation></ref><ref id="scirp.69891-ref19"><label>19</label><mixed-citation publication-type="other" xlink:type="simple">Guo, Y., Li, C.I., Ye, F. and Shyr, Y. (2013) Evaluation of Read Count Based RNAseq Analysis Methods. BMC Genomics, 14, S2. http://dx.doi.org/10.1186/1471-2164-14-S8-S2</mixed-citation></ref><ref id="scirp.69891-ref20"><label>20</label><mixed-citation publication-type="other" xlink:type="simple">Niemi, J., Mittman, E., Landau, W. and Nettleton, D. (2015) Empirical Bayes Analysis of RNA-Seq Data for Detection of Gene Expression Heterosis. Journal of Agricultural, Biological, and Environmental Statistics, 20, 614-628.  
http://dx.doi.org/10.1007/s13253-015-0230-5</mixed-citation></ref><ref id="scirp.69891-ref21"><label>21</label><mixed-citation publication-type="other" xlink:type="simple">Guffanti, A., Simchovitz, A. and Soreq, H. (2014) Emerging Bioinformatics Approaches for Analysis of NGS-Derived Coding and Non-Coding RNAs in Neurodegenerative Diseases. Frontiers in Cellular Neuroscience, 8, 89.  
http://dx.doi.org/10.3389/fncel.2014.00089</mixed-citation></ref><ref id="scirp.69891-ref22"><label>22</label><mixed-citation publication-type="other" xlink:type="simple">Giorgi, F.M., Del Fabbro, C. and Licausi, F. (2013) Comparative Study of RNA-Seq-and Microarray Coexpression Net-Works in Arabidopsis Thaliana. Bioinformatics, 29, 717-724. http://dx.doi.org/10.1093/bioinformatics/btt053</mixed-citation></ref><ref id="scirp.69891-ref23"><label>23</label><mixed-citation publication-type="other" xlink:type="simple">Houser, J.R., Barnhart, C., Boutz, D.R., et al. (2015) Controlled Measurement and Comparative Analysis of Cellular Components in E. coli Reveals Broad Regulatory Changes in Response to Glucose Starvation. PLoS Computational Biology, 11, e1004400. http://dx.doi.org/10.1371/journal.pcbi.1004400</mixed-citation></ref><ref id="scirp.69891-ref24"><label>24</label><mixed-citation publication-type="other" xlink:type="simple">Ackermann, A.M., Wang, Z., Schug, J., Naji, A. and Kaestner, K.H. (2016) Integration of ATAC-Seq and RNA-Seq Identi-Fieses Human Alpha Cell and Beta Cell Signature Genes. Molecular Metabolism, 5, 233-244.  
http://dx.doi.org/10.1016/j.molmet.2016.01.002</mixed-citation></ref><ref id="scirp.69891-ref25"><label>25</label><mixed-citation publication-type="other" xlink:type="simple">Shanks, C.M., Rice, J.H., Zubo, Y., Schaller, G.E., Hewezi, T. and Kieber, J.J. (2015) The Role of Cytokinin during Infection of Arabidopsis thaliana by the Cyst Nematode Heterodera schachtii. Molecular Plant-Microbe Interactions, 29, 57-68. http://dx.doi.org/10.1094/MPMI-07-15-0156-R</mixed-citation></ref><ref id="scirp.69891-ref26"><label>26</label><mixed-citation publication-type="other" xlink:type="simple">Tusher, V.G., Tibshirani, R. and Chu, G. (2001) Significance Analysis of Microarrays Applied to the Ionizing Radiation Response. Proceedings of the National Academy of Sciences of the United States of America, 98, 5116-5121.  
http://dx.doi.org/10.1073/pnas.091062498</mixed-citation></ref><ref id="scirp.69891-ref27"><label>27</label><mixed-citation publication-type="other" xlink:type="simple">Iancu, O.D., Kawane, S., Bottoml, D., Searles, R., Hitzemann, R. and McWeeney, S. (2012) Utilizing RNA-Seq Data for de Novo Coexpression Network Inference. Bioinformatics, 28, 1592-1597.  
http://dx.doi.org/10.1093/bioinformatics/bts245</mixed-citation></ref><ref id="scirp.69891-ref28"><label>28</label><mixed-citation publication-type="other" xlink:type="simple">López-Kleine, L., Leal, L. and López, C. (2013) Biostatistical Approaches for the Reconstruction of Gene Co-Expression Networks Based on Transcriptomic Data. Briefings in Functional Genomics, 12, 457-467.  
http://dx.doi.org/10.1093/bfgp/elt003</mixed-citation></ref><ref id="scirp.69891-ref29"><label>29</label><mixed-citation publication-type="other" xlink:type="simple">Riaz, M., Munir, S. and Asghar, Z. (2014) On the Performance Evaluation of Different Measures of Association. Revista Colombiana de Estadística, 37, 1-24. http://dx.doi.org/10.15446/rce.v37n1.44353</mixed-citation></ref></ref-list></back></article>