<?xml version="1.0" encoding="UTF-8"?><!DOCTYPE article  PUBLIC "-//NLM//DTD Journal Publishing DTD v3.0 20080202//EN" "http://dtd.nlm.nih.gov/publishing/3.0/journalpublishing3.dtd"><article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" dtd-version="3.0" xml:lang="en" article-type="research article"><front><journal-meta><journal-id journal-id-type="publisher-id">JBiSE</journal-id><journal-title-group><journal-title>Journal of Biomedical Science and Engineering</journal-title></journal-title-group><issn pub-type="epub">1937-6871</issn><publisher><publisher-name>Scientific Research Publishing</publisher-name></publisher></journal-meta><article-meta><article-id pub-id-type="doi">10.4236/jbise.2013.64059</article-id><article-id pub-id-type="publisher-id">JBiSE-30509</article-id><article-categories><subj-group subj-group-type="heading"><subject>Articles</subject></subj-group><subj-group subj-group-type="Discipline-v2"><subject>Biomedical&amp;Life Sciences</subject></subj-group></article-categories><title-group><article-title>
 
 
  Evaluation of RNA-Seq software in gene expression quantification
 
</article-title></title-group><contrib-group><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>an</surname><given-names>Ji</given-names></name><xref ref-type="aff" rid="aff1"><sup>1</sup></xref></contrib><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Ziliang</surname><given-names>Qian</given-names></name><xref ref-type="aff" rid="aff2"><sup>2</sup></xref></contrib><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Jia</surname><given-names>Wei</given-names></name><xref ref-type="aff" rid="aff1"><sup>1</sup></xref><xref ref-type="corresp" rid="cor1"><sup>*</sup></xref></contrib></contrib-group><aff id="aff2"><addr-line>Innovation Center China, AstraZeneca, Shanghai, China</addr-line></aff><aff id="aff1"><addr-line>R &amp;amp; D Information, AstraZeneca, Shanghai, China</addr-line></aff><author-notes><corresp id="cor1">* E-mail:<email>jenny.wei@astrazeneca.com(JW)</email>;</corresp></author-notes><pub-date pub-type="epub"><day>29</day><month>04</month><year>2013</year></pub-date><volume>06</volume><issue>04</issue><fpage>473</fpage><lpage>477</lpage><history><date date-type="received"><day>22</day>	<month>February</month>	<year>2013</year></date><date date-type="rev-recd"><day>27</day>	<month>March</month>	<year>2013</year>	</date><date date-type="accepted"><day>6</day>	<month>April</month>	<year>2013</year></date></history><permissions><copyright-statement>&#169; Copyright  2014 by authors and Scientific Research Publishing Inc. </copyright-statement><copyright-year>2014</copyright-year><license><license-p>This work is licensed under the Creative Commons Attribution International License (CC BY). http://creativecommons.org/licenses/by/4.0/</license-p></license></permissions><abstract><p>
 
 
   High-throughput RNA sequencing (RNA-Seq) promises a complete annotation and quantification of all genes and their isoforms across samples. Because sequencing reads from this new technology are shorter than transcripts from which they are derived, expression estimation with RNA-Seq requires increasingly complex computational methods. In recent years, a number of expression quantification methods have been published from both public and commercial sources. Here we presented an overview of these attempts on quantifying gene expression. We then defined a set of criteria and compared the performance of several programs based on these criteria, and we further provided advices on selecting suitable tools for different biological applications. 
 
</p></abstract><kwd-group><kwd>RNA-Seq; Next Generation Sequencing</kwd></kwd-group></article-meta></front><body><sec id="s1"><title>1. INTRODUCTION</title><p>Next-generation sequencing (NGS) platforms have been widely available recently [<xref ref-type="bibr" rid="scirp.30509-ref1">1</xref>]. A massively parallel sequencing technology termed RNA-Seq has made it possible to sequence cDNA derived from cellular RNA [<xref ref-type="bibr" rid="scirp.30509-ref2">2</xref>]. Compared to previous technologies for gene mapping with their alternative isoforms and expression detection across diverse cell types, RNA-Seq is more promising in building a complete transcriptome across cell types and states.</p><p>Recently, many studies have applied RNA-Seq to various biological and medical research. Quantification of alternative splicing in tissues [<xref ref-type="bibr" rid="scirp.30509-ref3">3</xref>], discovery of new fusion genes in cancer [<xref ref-type="bibr" rid="scirp.30509-ref4">4</xref>], and new transcript identification [<xref ref-type="bibr" rid="scirp.30509-ref5">5</xref>] have all benefited from this new technology. To fully enable RNA-Seq technology to solve biological problems, powerful computational tools are required. In the past two years, software applications for RNA-Seq analysis have been flooding the market from public domains as well as commercial organizations. How to identify and use the suitable tools for RNA-Seq analysis becomes critical.</p><p>Here we focus on the computational methods for gene expression quantification by RNA-Seq. Using Google Scholar citation, as shown in <xref ref-type="table" rid="table1">Table 1</xref>, we selected two popular analysis pipelines from public domains and two workflows from commercial products. We applied them to a human gastric cancer RNA-Seq dataset consisting of 40 million paired-end 100-base reads from Illumina Hiseq 2000 platform. We also compared RNA-Seq quantification with Affy quantification using Affymetrix human genome U133A2.0 array on the same human gastric cancer sample.</p></sec><sec id="s2"><title>2. RESULTS</title><sec id="s2_1"><title>2.1. Descriptions of Chosen RNA-Seq Quantification Tools</title><p>In general, current transcriptome assembly tools belong to either a reference-based strategy or de novo strategy or both [<xref ref-type="bibr" rid="scirp.30509-ref6">6</xref>]. When a reference genome is available, RNASeq reads are firstly mapped by a splice-aware aligner and an output alignment file is used as the input file of a transcriptome assembly tool. Two reference-based transcriptome assembly tools, Cufflinks [<xref ref-type="bibr" rid="scirp.30509-ref7">7</xref>] and Scripture [<xref ref-type="bibr" rid="scirp.30509-ref5">5</xref>], are selected according to their average citation numbers per month (CPM) calculated by the total number of citations retrieved from Google Scholar divided by the number of months since their publication date. Another tool Alexa-seq [<xref ref-type="bibr" rid="scirp.30509-ref8">8</xref>] is not chosen because of both the small CPM 1.6 and difficulties to install it on a Linux server.</p><p>Cufflinks can be launched in two modes using options -G/--GTF and -g/--GTF-guide. Both modes need a reference GFF annotation file from mainly three data sources Ensembl (www.ensembl.org), NCBI (www.ncbi.nlm.nih.gov) and UCSC (http://genome.ucsc.edu/). The first option -G/--GTF tells Cufflinks to use the supplied reference annotation to estimate isoform expression. The latter option -g/--GTFguide tells Cufflinks to use the supplied reference annotation to guide RABT assembly [<xref ref-type="bibr" rid="scirp.30509-ref9">9</xref>]. However, Scripture is a method for transcriptome reconstruction that relies solely on RNA-Seq reads and an assembled genome to build a transcriptome ab initio.</p><p>Array Studio is a suite of tools developed by OmicSoft (www.omicsoft.com) in which an RNA-Seq analysis workflow is provided. Expression quantification analysis of RNA-Seq can be performed in two ways by mapping to either genome or transcriptome.</p><p>CLC Genomics Workbench is a Desktop application for NGS analysis developed by CLCbio (www.clcbio.com).</p></sec><sec id="s2_2"><title>2.2. Summaries of Results of RNA-Seq Analysis</title><p>An in-house RNA-Seq dataset was used and six types of results were generated. For Cufflinks, two results from both -G/--GTF and -g/--GTF-guide modes which are denoted by Cuff.(-G) and Cuff.(-g). For Array Studio, by against both genome and transcriptome, two results were shown and denoted by OMIC(G) and OMIC(T). The result of CLC Genomics Workbench was CLC GW, and the last result is from Scripture. Summaries about the six results can be found in Tables 2 and 3. Note that CLC Genomics Workbench only gives gene information, so genes with only one transcript were counted.</p><p>There are some differences between Tables 2 and 3. In</p><p><xref ref-type="table" rid="table2">Table 2</xref>, numbers of total features found by tools are directly counted from their output files without any preprocessing. Transcripts are features with at least two exons. Most of results of OMIC(G) are transcripts. Results of OMIC(T), Cuff.(-G), Cuff.(-g) and Scripture are transcripts and exons. Results of CLC GW are genes, and only genes with only one transcript were counted. In <xref ref-type="table" rid="table3">Table 3</xref>, genes of non-zero expression values are counted. The cell values of OMIC(G) and “RNAseq.total” are numbers of features which can be annotated with known gene names by ArrayStudio. Cufflinks can output a file recording both gene names and their FPKM values.</p><p>In <xref ref-type="table" rid="table3">Table 3</xref>, Affy data were processed using Affymetrix Expression Console. Most of genes in Affy are included by RNA-Seq results. Cufflinks adopts a naming mechanism when -g/--GTF-guide option is turned on, so the number of common genes is very small when merging two datasets according gene names. Actually, we showed that 77% of results of Cuff(-g) can reflect (i.e., match or include) 100% of that of Cuff(-G).</p></sec><sec id="s2_3"><title>2.3. Comparisons between RNA-Seq and Affy in Terms of Expression Values</title><p>There are always three types of expression values used in the RNA-Seq analysis, RPKM/TPKM, TPM and na&#239;ve counts. Because Affy presents expression values on the gene level, only expression values of RNA-Seq on the gene level were shown in <xref ref-type="table" rid="table3">Table 3</xref>.</p><p>In <xref ref-type="fig" rid="fig1">Figure 1</xref>, six correlation scatter plots and their Pearson values were shown. The y-axis values of Figures 1(a) and (b) are calculated by Cufflinks with dif-</p><p><xref ref-type="table" rid="table1">Table 1</xref>. Open source tools selection criteria.</p><p><img src="6-9101646\f8196ede-bded-4da2-9068-be8ae4ec8d31.jpg" /></p><p><xref ref-type="table" rid="table2">Table 2</xref>. Numbers of total features and transcripts given by tools.</p><p><img src="6-9101646\8dab74cc-6f96-44b9-836d-7c8581677d55.jpg" /></p><p><xref ref-type="table" rid="table3">Table 3</xref>. Comparisons of numbers of genes of Affy and numbers of genes found by tools.</p><p><img src="6-9101646\6b92e798-2405-476a-9df6-b5e8aee6863e.jpg" /></p><p>ferent parameters. The y-axis value of figure C is calculated by CLC Genomic workbench. The y-axis values of Figures 1(d)-(f) are calculated by OmicSoft software with different calculation methods for gene expression values based on RNA-Seq reads. Because the biggest two correlation values with Affy are Cuff.(-g) and Cuff. (-G), Cufflinks has the best performance for the calculations of expression values from RNA-Seq reads. Except OMIC(G, na&#239;ve count) with Pearson value 0.68, Array studio (OMIC(G, TPM), OMIC(G, RPKM)) has a better performance than CLC Genomics Workbench (CLC GW).</p></sec><sec id="s2_4"><title>2.4. Comparisons between RNA-Seq and Public Annotations for Gene Structures</title><p>In this section, results of RNAseq analysis were assessed from the perspective of the linear structure of genomic</p><p>features on the level of transcripts. In this paper, USSC hg19 annotation file was used (ftp://igenome:G3nom3s4u@ftp.illumina.com/Homo_sapiens/UCSC/hg19/Homo_sapiens_UCSC_hg19.tar.gz).</p><p>In order to explore differences among results of chosen RNA-Seq analysis, four aspects were illustrated by following terms. First, “Match” means transcripts of both the result and public annotations are the same if all their exons are matched according to chromosome coordinates. Second, “Including” means exons of a transcript in the public annotations contains exons of a transcript in the results of RNA-Seq analysis. Third, “Included” means exons of a transcript in the public annotations are a subset of exons of a transcript in the result of RNA-Seq analysis. Fourth, “Overlap” means they share common exons.</p><p>As shown in <xref ref-type="fig" rid="fig2">Figure 2</xref>, the x-axis values of <xref ref-type="fig" rid="fig2">Figure 2</xref>(a)</p><p>and (b) are calculated by Omicsoft software with different parameters. The x-axis values of Figures 2(c) and (d) are calculated by cufflinks with different parameters. The x-axis values of Figures 2(e) and (f) are respectively calculated by CLC Genomic workbench and Scripture. From comparisons between Cuff.(-G) and OMIC(G) on the “Match” aspect, the bigger values of Cuff.(-G) may explain its higher correlation scores with expression values of Affy than OMIC(G). On the “Including” aspect, both Cuff(-G) and OMIC(T) have very small values; Scripture has the biggest values. This can be explained by Scripture using only RNA-Seq data without public annotations as reference. Both Cuff(-G) and OMIC(T) fully utilize reference annotations for transcriptome assembly.</p></sec></sec><sec id="s3"><title>3. CONCLUSION</title><p>In this article, we have shown the evaluation results of a set of public and commercial tools for gene expression quantification by RNA-Seq. Because of rapid improvements in RNA-Seq data generation, more efforts need to be done in the areas of transcriptome analysis, mutation detection, and fusion identification. New questions will continue to emerge and novel programs will evolve. The tool evaluation needs to keep up with the pace of these changes in order to apply RNA-Seq technologies to drug discovery and development.</p></sec><sec id="s4"><title>4. ACKNOWLEDGEMENTS</title><p>This work was supported in part by a project from AstraZeneca.</p></sec><sec id="s5"><title>REFERENCES</title></sec></body><back><ref-list><title>References</title><ref id="scirp.30509-ref1"><label>1</label><mixed-citation publication-type="other" xlink:type="simple">Metzker, M.L. (2010) Sequencing technologies—The next generation. Nature Reviews Genetics, 11, 31-46.  
doi:10.1038/nrg2626</mixed-citation></ref><ref id="scirp.30509-ref2"><label>2</label><mixed-citation publication-type="other" xlink:type="simple">Mortazavi, A., Williams, B.A., McCue, K., Schaeffer, L. and Wold, B. (2008) Mapping and quantifying mammalian transcriptomes by RNA-Seq. Nature Methods, 5, 621- 628. doi:10.1038/nmeth.1226</mixed-citation></ref><ref id="scirp.30509-ref3"><label>3</label><mixed-citation publication-type="other" xlink:type="simple">Wang, E.T., Sandberg, R., Luo, S., Khrebtukova, I., Zhang, L., Mayr, C., et al. (2008) Alternative isoform regulation in human tissue transcriptomes. Nature, 456, 470-476.  
doi:10.1038/nature07509</mixed-citation></ref><ref id="scirp.30509-ref4"><label>4</label><mixed-citation publication-type="other" xlink:type="simple">Maher, C.A., Kumar-Sinha, C., Cao, X., Kalyana-Sundaram, S., Han, B., Jing, X., et al. (2009) Transcriptome sequencing to detect gene fusions in cancer. Nature, 458, 97-101. doi:10.1038/nature07638</mixed-citation></ref><ref id="scirp.30509-ref5"><label>5</label><mixed-citation publication-type="other" xlink:type="simple">Guttman, M., Garber, M., Levin, J.Z., Donaghey, J., Robinson, J. and Adiconis, X. (2010) Ab initio reconstruction of cell type-specific transcriptomes in mouse reveals the conserved multi-exonic structure of lincRNAs. Nature Biotechnology, 28, 503-510. doi:10.1038/nbt.1633</mixed-citation></ref><ref id="scirp.30509-ref6"><label>6</label><mixed-citation publication-type="other" xlink:type="simple">Martin, J.A. and Wang, Z. (2011) Next-generation transcriptome assembly. Nature Reviews Genetics, 12, 671- 682. doi:10.1038/nrg3068</mixed-citation></ref><ref id="scirp.30509-ref7"><label>7</label><mixed-citation publication-type="other" xlink:type="simple">Trapnell, C., Williams, B.A., Pertea, G., Mortazavi, A., Kwan, G., van Baren, M.J., et al. (2010) Transcript assembly and quantification by RNA-Seq reveals unannotated transcripts and isoform switching during cell differentiation. Nature Biotechnology, 28, 511-515.  
doi:10.1038/nbt.1621</mixed-citation></ref><ref id="scirp.30509-ref8"><label>8</label><mixed-citation publication-type="other" xlink:type="simple">Griffith, M., Griffith, O.L., Mwenifumbo, J., Goya, R., Morrissy, A.S., Morin, R.D., et al. (2010) Alternative expression analysis by RNA sequencing. Nature Methods, 7, 843-847. doi:10.1038/nmeth.1503</mixed-citation></ref><ref id="scirp.30509-ref9"><label>9</label><mixed-citation publication-type="other" xlink:type="simple">Roberts, A., Pimentel, H., Trapnell, C. and Pachter, L. (2011) Identification of novel transcripts in annotated genomes using RNA-Seq. Bioinformatics, 27, 2325-2329.  
doi:10.1093/bioinformatics/btr355</mixed-citation></ref></ref-list></back></article>