<?xml version="1.0" encoding="UTF-8"?><!DOCTYPE article  PUBLIC "-//NLM//DTD Journal Publishing DTD v3.0 20080202//EN" "http://dtd.nlm.nih.gov/publishing/3.0/journalpublishing3.dtd"><article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" dtd-version="3.0" xml:lang="en" article-type="research article"><front><journal-meta><journal-id journal-id-type="publisher-id">JBiSE</journal-id><journal-title-group><journal-title>Journal of Biomedical Science and Engineering</journal-title></journal-title-group><issn pub-type="epub">1937-6871</issn><publisher><publisher-name>Scientific Research Publishing</publisher-name></publisher></journal-meta><article-meta><article-id pub-id-type="doi">10.4236/jbise.2013.65072</article-id><article-id pub-id-type="publisher-id">JBiSE-31900</article-id><article-categories><subj-group subj-group-type="heading"><subject>Articles</subject></subj-group><subj-group subj-group-type="Discipline-v2"><subject>Biomedical&amp;Life Sciences</subject></subj-group></article-categories><title-group><article-title>
 
 
  In silico tests on sequence motif significances for human tissue specific genes
 
</article-title></title-group><contrib-group><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>iujun</surname><given-names>Gong</given-names></name><xref ref-type="aff" rid="aff1"><sup>1</sup></xref><xref ref-type="corresp" rid="cor1"><sup>*</sup></xref></contrib><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Hualin</surname><given-names>Xu</given-names></name><xref ref-type="aff" rid="aff1"><sup>1</sup></xref><xref ref-type="corresp" rid="cor1"><sup>*</sup></xref></contrib></contrib-group><aff id="aff1"><addr-line>School of Computer Science and Technology, Tianjin University, Tianjin, China</addr-line></aff><author-notes><corresp id="cor1">* E-mail:<email>gongxj@tju.edu.cn(IG)</email>;<email>lin_snowing@yahoo.com.cn(HX)</email>;</corresp></author-notes><pub-date pub-type="epub"><day>20</day><month>05</month><year>2013</year></pub-date><volume>06</volume><issue>05</issue><fpage>572</fpage><lpage>578</lpage><history><date date-type="received"><day>13</day>	<month>February</month>	<year>2013</year></date><date date-type="rev-recd"><day>3</day>	<month>April</month>	<year>2013</year>	</date><date date-type="accepted"><day>9</day>	<month>May</month>	<year>2013</year></date></history><permissions><copyright-statement>&#169; Copyright  2014 by authors and Scientific Research Publishing Inc. </copyright-statement><copyright-year>2014</copyright-year><license><license-p>This work is licensed under the Creative Commons Attribution International License (CC BY). http://creativecommons.org/licenses/by/4.0/</license-p></license></permissions><abstract><p>
 
 
  Identification and analysis of tissue-specific (TS) genes 
  and their regulatory activities play an important role in understanding the mechanisms of the organism, disease diagnosis and drug design. Although so far we are not clear about the mechanisms totally, the sequence features of TS genes are becoming an important clue. In this paper we used an integrated pipeline to discover sequences motifs for the promoter regions of TS genes. To test the significances of those motifs in a specific tissue, we used hypotheses test approaches including Bayesian hypothesis, Binomial distribution and traditional z-test. We finally got 2784, 1204 and 703 motifs respectively out of 3244 motifs obtained in discovery phase using above three tests from 3954 TS genes across 83 human tissues. 52.7% of those motifs can be found in public databases available. 
 
</p></abstract><kwd-group><kwd>Tissue Specific Genes; Hypothesis Test; Tissue Rich Motif; Tissue Even Motif</kwd></kwd-group></article-meta></front><body><sec id="s1"><title>1. INTRODUCTION</title><p>Identification and analysis of tissue-specific (TS) genes and their regulatory activities play an important role in understanding mechanisms of the organism, disease diagnosis and drug design [<xref ref-type="bibr" rid="scirp.31900-ref1">1</xref>]. In last years, many research projects were performed to study expressions and regulatory mechanisms of TS genes including transcription factor and their binding sites, sequence features of promoter regions [<xref ref-type="bibr" rid="scirp.31900-ref2">2</xref>], alternative splicing [<xref ref-type="bibr" rid="scirp.31900-ref3">3</xref>] and Epigenetics features [<xref ref-type="bibr" rid="scirp.31900-ref4">4</xref>] of those genes.</p><p>Although until now we are not completely clear about the mechanisms of the gene tissue specificity, the sequence features of TS genes are becoming an important clue [<xref ref-type="bibr" rid="scirp.31900-ref2">2</xref>]. P. FitzGerald et al. calculated the statistics of Simple Sequence Repeats (SSR) and identified that the SSR could be an important factor to the tissue specificity [<xref ref-type="bibr" rid="scirp.31900-ref5">5</xref>]. F. Song et al. pointed that methylation changes during development are dynamic, involve demethylation and methylation, and may occur at late stages of embryonic development or even postnatally using mouse genome data [<xref ref-type="bibr" rid="scirp.31900-ref6">6</xref>]. C. Heber et al. showed that Nucleosome rotational setting is associated with transcriptional regulation in promoters of tissue-specific human genes [<xref ref-type="bibr" rid="scirp.31900-ref4">4</xref>].</p><p>With the completion of the whole human genome project, various algorithms have been developed for discovering patterns or motifs of huge volume genome sequences. Those typical algorithms include three phases: motif searching, redundant motif pruning and motif significance testing. The methods for motif discovery may be grouped into two categories [<xref ref-type="bibr" rid="scirp.31900-ref7">7</xref>]: enumerative methods and alignment-based methods. Enumerative methods typically involve exhaustive enumeration of words up to some maximum size in a dataset, and are thus best suited to consensus sequence motif models, like Consensus, PROJECTION, PDEM. Alignment methods take on a wide variety of forms, but often involve the development of a probabilistic model of the observed sequence data and optimization to find motifs common to all input sequences, such as MEME [<xref ref-type="bibr" rid="scirp.31900-ref8">8</xref>] program, the expectationmaximization (EM) algorithm and Gibbs sampling [<xref ref-type="bibr" rid="scirp.31900-ref9">9</xref>]. Each algorithm has its unique advantage on individual species or datasets. Tompa et al. [<xref ref-type="bibr" rid="scirp.31900-ref7">7</xref>] conducted a study that compares the performance of 13 different motif finders by using a variety of real and synthetic sequence sets covering a range of genomes. A common practice is to apply several such algorithms simultaneously to improve coverage at the cost of increased redundancy [<xref ref-type="bibr" rid="scirp.31900-ref10">10</xref>].</p><p>In this paper, we first applied an integrated motif searching approach to find motifs for TS genes. As we known, it is the first time to search sequence motifs for tissue specific genes. Then we merged the similar motifs using the method in literature [<xref ref-type="bibr" rid="scirp.31900-ref7">7</xref>]. To test the significances of those motifs in each tissue, we used three hypothesis test methods: Bayesian hypothesis, Binomial distribution and traditional z-test. We also distinguish two kinds of significant motifs: tissue rich motifs (TIM) and tissue even motifs (TEM). The former refer to motifs only showing significance in few tissues, and the later refer to motifs in most of the tissues. We finally got 2784, 1204 and 703 motifs respectively out of 3244 motifs obtained in discovery phase using above three tests from 3954 TS genes across 83 human tissues. 52.7% those motifs can be found available in databases public.</p><sec id="s1_1"><title>2.1. Date Preparing</title><p>Tissue specific genes were obtained mainly by querying the tissue specific gene expression database TiGER [<xref ref-type="bibr" rid="scirp.31900-ref11">11</xref>] against the tissue names. Some of them came from TisGED [<xref ref-type="bibr" rid="scirp.31900-ref12">12</xref>] database. All of the TS genes with PubMed IDs were used in the experiment. We finally got 3954 human tissue specific genes across 83 human tissues. The gene’s promoter sequences are downloaded from DBTSS [<xref ref-type="bibr" rid="scirp.31900-ref13">13</xref>] and EPD [<xref ref-type="bibr" rid="scirp.31900-ref14">14</xref>]. The promoter region with 1500 bp (−499 bp - 1000 bp around TSS) length is used for motif searching.</p></sec><sec id="s1_2"><title>2.2. Motif Searching</title><sec id="s1_2_1"><title>2.2.1. Motif Searching</title><p>In this phase, we integrated three motif searching programs: MEME, AlignACE and Gibbs Sampler. The length of candidate motifs is fixed to 6 - 12 bp, other parameters as the default setting. In this phase, we get 6794 motifs.</p></sec><sec id="s1_2_2"><title>2.2.2. PWM Representations of Motifs</title><p>Since different motif search programs have their own motif formats as outputs, we have to define a uniform format for motifs to compare their similarities in motif merging phase. A common used representation is the Position Specific Weight Matrix (PWM or PSWM) [<xref ref-type="bibr" rid="scirp.31900-ref15">15</xref>], which is a matrix of nucleotide frequencies in each position of the motif (i.e. the frequencies of the nucleotides A, C, G and T in each position). We transformed all the motifs to the PWM representation.</p></sec><sec id="s1_2_3"><title>2.2.3. Motif Merging</title><p>In motif merging phase, we used the method similar with in literature [<xref ref-type="bibr" rid="scirp.31900-ref16">16</xref>] to remove motif redundancies. Because this step isn’t the emphasis of this paper, we skip the details of the merging process. After motif merging, 3244 motifs were obtained.</p></sec></sec><sec id="s1_3"><title>2.3. Motif Tissue Significance Testing</title><p>To identify whether a motif is really related with tissue specificity or not, we statistically distinguish two kinds of motifs: tissue rich motifs (TRM) and tissue even motifs (TEM). The former refer to motifs only showing statics significance in less than 3 tissues, and the later refer to motifs in more than 70 tissues. We used hypothesis approaches to test the significance of motifs in each tissue. To do the hypothesis test, the distributions of motifs in a given sequence must be estimated. Therefore, a key step is to calculate the statistic of a motif in a given sequence.</p><p>For a given motif m with length w from tissue T<sub>0</sub>, in which the motif is discovered, our purpose is to judge whether its occurrence in tissue T<sub>1</sub> is significant or not. Therefore we have to take a measure on the motif occurrences. Based on the requirements of different hypothesis tests, we applied scoring schemas.</p><p>Definition 1: for a given motif m, its matching Score with a Promoter sequence segment x of the gene from tissue T<sub>1</sub> PMS1 is defined:</p><p><img src="8-72258\9e349c69-5244-49f5-b1b3-a35d9fd8c46c.jpg" /></p><p>where <img src="8-72258\5d3982a6-4d6f-4223-93a9-e5158dae67a9.jpg" /> is the score between m and x in position i, which can be calculated through the PWM of the motif.</p><p>Definition 2: for a given motif m, its matching Score with a Promoter Sequence S of the gene from tissue T<sub>1</sub> PSS1 is defined:</p><p><img src="8-72258\3e386732-17e6-423d-a36a-72e4faa6b353.jpg" /></p><p>where <img src="8-72258\37fe6335-a310-4ab9-b442-ec013aaa23a2.jpg" /> with PMS1 more than a predefined threshold is a segment of S by sliding a widow with length w, n is the number of<img src="8-72258\8fecb0af-a2e5-474a-8d08-957809cdb98b.jpg" />.</p><p>PSS1 is used in classical z-test and binomial test.</p><p>Definition 3: for a given motif m, its matching Score with a Promoter sequence segMent x of the gene from tissue T<sub>1</sub> PMS2 is defined [<xref ref-type="bibr" rid="scirp.31900-ref16">16</xref>]:</p><p><img src="8-72258\65f6a211-d688-4c88-b342-4d291d96c2ff.jpg" />where<img src="8-72258\039a3f2b-e01c-4d30-ba18-9828ed0b9121.jpg" />, <img src="8-72258\c584bef2-52a4-4088-b8be-e4730859c20d.jpg" />, <img src="8-72258\68a34f8a-0a3b-4e6d-b882-2754e7281b73.jpg" />.</p><p><img src="8-72258\045effb3-2213-4459-b1b6-9e5dd6dc3bff.jpg" />is the frequency of residue B at position i, which is from PWM; <img src="8-72258\64cd3f0f-4218-49e0-9125-053dc9d0f861.jpg" />is the smallest/largest frequency of the residue at position i and</p><p><img src="8-72258\96f86040-3c4f-4fdc-921f-62ec4ae06421.jpg" />describe the information content of residue B at position i.</p><p>Definition 4: for a given motif m, its matching Score with a Promoter Sequence S of the gene from tissue T<sub>1</sub> PSS2 is defined:</p><p><img src="8-72258\e509cf23-dca5-46ba-8851-30675657d367.jpg" /></p><p>where <img src="8-72258\482b5304-c86a-4385-b500-ebff75722dd5.jpg" /> with PMS2 more than a predefined threshold is a segment of S by sliding a widow with length w, n is the number of<img src="8-72258\143b8185-e862-4b70-bb3a-327713f5a3b3.jpg" />.</p><p>PSS2 is used in Bayesian hypothesis test.</p><sec id="s1_3_1"><title>2.3.1. Classical Z-Test</title><p>In the classical z-test, we estimated the mean and variance of the match score PSS1 in tissue T<sub>1</sub>, and then calculated the z-value:</p><p><img src="8-72258\dd1caf88-7559-4363-8edc-f4be4a29f019.jpg" />.</p><p>where <img src="8-72258\d3e92f1b-ef50-49be-829e-e7c3a9f64ca4.jpg" /> and <img src="8-72258\afe2f6d4-8537-4b21-8785-c739eb623fc6.jpg" /> are the mean and variance of the PSS1 in tissue T<sub>0</sub>.</p><p>In the experiment, we set the confidence degree 0.05.</p></sec><sec id="s1_3_2"><title>2.3.2. Bayes Hypothesis Test</title><p>Assumed that the PSS2 of a motif at tissue T<sub>0</sub> follows a Gaussian distribution<img src="8-72258\e5940235-e434-43fc-9d17-57e7345af481.jpg" />. To test that whether the motif is significant at tissue T<sub>1</sub>, we constructed two hypothesizes as the followings:</p><p><img src="8-72258\cfb60cd1-51e1-4a0e-9429-9d6358ce0eaf.jpg" /></p><p>where <img src="8-72258\0497682e-f142-40f4-9133-98c146d30592.jpg" /> is the mean of PSS2 in tissue T<sub>1</sub>.</p><p>Assumed that<img src="8-72258\5b8b6ce5-a806-4ec1-b939-7c63cd50eb31.jpg" />, where <img src="8-72258\e19ee851-91ae-44ee-9b1e-96066a60c922.jpg" /> is unknown and <img src="8-72258\2a0bfe4c-4844-4dbb-9829-76a68cd6850a.jpg" /> is known, <img src="8-72258\56e1fd5c-dc89-4626-80c8-76d318823977.jpg" />, where both <img src="8-72258\097ae62f-c5f7-4740-91c1-2a099c3de0c4.jpg" /> and <img src="8-72258\efb74d96-cd04-4afd-b103-16cdb2cbe6c3.jpg" /> are known. The post distribution of <img src="8-72258\1f3331ab-a206-49a8-a95f-8389c2396ecf.jpg" /> is followed <img src="8-72258\b7a83689-8d22-4a3a-97c2-e27358717888.jpg" /> according [<xref ref-type="bibr" rid="scirp.31900-ref16">16</xref>], where</p></sec><sec id="s1_3_3"><title>2.3.3. Binomial Distribution Test</title><p>In Binomial distribution test, instead of PSS1 value, we need the number of matches between the motif and the promoter sequence of a gene. A match between a motif and a sequence is defined if the PMS1 of the motif with a segment of the sequence is larger than a predefined value. We counted all the matches in tissue T<sub>0</sub> and T<sub>1</sub>, represented the numbers of matches by K<sub>0</sub> and K<sub>1</sub> respectively. The Binomial distribution test is to seek a value K-value holding:</p><p><img src="8-72258\36e37be1-a2df-4d43-a52b-4619fdc56bc9.jpg" /></p><p>where n<sub>0</sub> and n<sub>1</sub> are the numbers of promoter sequences in tissue T<sub>0</sub> and T<sub>1</sub> respectively and p is fixed to 0.5 in the experiment.</p></sec></sec><sec id="s1_4"><title>3.1. Data Sources</title><p>The gene expression datasets, such as GNF, SAGE, and EST, are very widely used as data sources for the identifications of TS genes. However, because of the noise in expression datasets and human involvement in defining thresholds, the reliability of the identifications is often not high. In this paper, we use the specific genes obtained mainly by querying the tissue specific gene expression database TiGER against the tissue names. Some of them came from TisGED database. All of the TS genes with PubMed IDs were used in the experiment. We obtained 3954 TS genes across 83 human tissues. Because of the limitation of page size, the gene lists for all the tissues are available on request to the authors.</p><p>The gene’s promoter sequences were downloaded from DBTSS and EPD. The promoter region with length 1500 bp (−499 bp - 1000 bp around TSS) is used for motif discovery.</p></sec><sec id="s1_5"><title>3.2. Motifs Discovered by Three Test Methods</title><p>After merging phase, we get total 3244 motifs. The number of motifs in each tissue is shown in <xref ref-type="table" rid="table1">Table 1</xref>.</p><p>With Bayes Hypothesis Test method, we get 1534 TRMs and 1270 TEMs. With Classic z-test method, 539 TRMs and 164 TEMs are obtained. With Binomial Distribution test method, the numbers of two kinds of motifs are 270 and 925 respectively. For the details, see in Figures 1 and 2.</p></sec><sec id="s1_6"><title>3.3. Overlap Motifs in Three Test Methods</title><p>In all the TRMs, 5 TRMs are covered by three methods, 150 TRMs are covered by two methods. In all the TEMs, 39 TEMs covered by three methods, 264 TEMs covered by two methods. For the details, see <xref ref-type="fig" rid="fig3">Figure 3</xref>.</p><p>We also compared the overlapped 5 TRMs and 39 TEMotif with JASPAR [<xref ref-type="bibr" rid="scirp.31900-ref17">17</xref>]. 4 TRMs (see <xref ref-type="table" rid="table2">Table 2</xref>) out of 5 TRMs are found in the JASPAR. For an example, [CCCCNCCCCC] is a motif which was discovered by previous researches in JASPAR ID MA0079.2_SP1, and [GGGGAATCCCC] with JASPAR ID MA0105.1_ NFKB1. 19 TEMs out of 39 TEMs are found in the JASPAR. For an example, the motif [NGNNGCRSCG] has JASPAR ID MA0123.1_abi4. For the details see <xref ref-type="table" rid="table3">Table 3</xref>.</p></sec></sec><sec id="s2"><title>4. CONCLUSIONS</title><p>Tissue specificity is the foundation for cells form specific</p><p>tissues and functional organs. Identification and analysis of tissue-specific genes and their regulatory activities play an important role in understanding mechanisms of the organism, disease diagnosis and drug design. And finding accurate and meaningful motif with tissue specificity still remains a big challenge.</p><p>In this paper we used an integrated pipeline to discover sequence motifs for the promoter regions of TS genes. To test the significances of those motifs in a specific tissue, we used hypotheses test approaches including Bayesian hypothesis, Binomial distribution and traditional</p><p>z-test. We finally got 2784, 1204 and 703 motifs respectively out of 3244 motifs obtained in discovery phase using above three tests from 3954 TS genes across 83 human tissues. 52.7% of those motifs can be found available in databases public.</p></sec><sec id="s3"><title>REFERENCES</title></sec><sec id="s4"><title>NOTES</title></sec></body><back><ref-list><title>References</title><ref id="scirp.31900-ref1"><label>1</label><mixed-citation publication-type="other" xlink:type="simple">Dezso, Z., et al. (2008) A comprehensive functional analysis of tissue specificity of human gene expression. BMC Biology, 6, 49.</mixed-citation></ref><ref id="scirp.31900-ref2"><label>2</label><mixed-citation publication-type="other" xlink:type="simple">Kuzmin, D., et al. (2010) Novel strong tissue specific promoter for gene expression in human germ cells. BMC Biotechnology, 10, 58. doi:10.1186/1472-6750-10-58</mixed-citation></ref><ref id="scirp.31900-ref3"><label>3</label><mixed-citation publication-type="other" xlink:type="simple">Grosso, A., Gomes, A. and Barbosa, N. (2008) Tissuespecific splicing factor gene expression signatures. Nucleic Acids, 36, 4823-4832. doi:10.1093/nar/gkn463</mixed-citation></ref><ref id="scirp.31900-ref4"><label>4</label><mixed-citation publication-type="other" xlink:type="simple">Hebert, C. (2010) Nucleosome rotational setting is associated with transcriptional regulation in promoters of tissue-specific human genes. Genome Biology, 11, R51.  
doi:10.1186/gb-2010-11-5-r51</mixed-citation></ref><ref id="scirp.31900-ref5"><label>5</label><mixed-citation publication-type="other" xlink:type="simple">Lawson, M.J. and Zhang, L. (2008) Housekeeping and tissue-specific genes differ in simple sequence repeats in the 5'-UTR region. Gene, 407, 54-62.  
doi:10.1016/j.gene.2007.09.017</mixed-citation></ref><ref id="scirp.31900-ref6"><label>6</label><mixed-citation publication-type="other" xlink:type="simple">Song, F., et al. (2009) Tissue specific differentially methylated regions (TDMR): Changes in DNA methylation during development. Genomics, 93, 130-139.  
doi:10.1016/j.ygeno.2008.09.003</mixed-citation></ref><ref id="scirp.31900-ref7"><label>7</label><mixed-citation publication-type="other" xlink:type="simple">Tompa, M., et al. (2005) Assessing computational tools for the discovery of transcription factor binding sites. Nature Biotechnology, 23, 137-144. doi:10.1038/nbt1053</mixed-citation></ref><ref id="scirp.31900-ref8"><label>8</label><mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Bailey</surname><given-names> T.L.</given-names></name>,<name name-style="western"><surname> et al. </surname><given-names>  </given-names></name>,<etal>et al</etal>. (<year>2009</year>)<article-title>MEME SUITE: Tools for motif discovery and searching</article-title><source> Nucleic Acids Research</source><volume> 37</volume>,<fpage> W202</fpage>-<lpage>W208</lpage>.<pub-id pub-id-type="doi"></pub-id></mixed-citation></ref><ref id="scirp.31900-ref9"><label>9</label><mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Neuwald</surname><given-names> F.</given-names></name>,<name name-style="western"><surname> Liu</surname><given-names> J.S. and Lawrence</given-names></name>,<name name-style="western"><surname> C.E. </surname><given-names>  </given-names></name>,<etal>et al</etal>. (<year>1995</year>)<article-title>Gibbs motif sampling detection of bacterial outer membrane protein repeats</article-title><source> Protein Science: A Publication of the Protein Society</source><volume> 4</volume>,<fpage> 1618</fpage>-<lpage>1632</lpage>.<pub-id pub-id-type="doi"></pub-id></mixed-citation></ref><ref id="scirp.31900-ref10"><label>10</label><mixed-citation publication-type="other" xlink:type="simple">Clements, M. (2007) Creating motifs with LocoMotif.  
Scanning.</mixed-citation></ref><ref id="scirp.31900-ref11"><label>11</label><mixed-citation publication-type="other" xlink:type="simple">Liu, X., Yu, X., Zack, D.J., Zhu, H. and Qian, J. (2008) TiGER: A database for tissue-specific gene expression and regulation. BMC Bioinformatics, 9, 271.  
doi:10.1186/1471-2105-9-271</mixed-citation></ref><ref id="scirp.31900-ref12"><label>12</label><mixed-citation publication-type="other" xlink:type="simple">Xiao, S.J., Zhang, C. and Zou, Q. (2010) TiSGeD: A database for tissue-specific genes. Bioinformatics, 26, 12731275. doi:10.1093/bioinformatics/btq109</mixed-citation></ref><ref id="scirp.31900-ref13"><label>13</label><mixed-citation publication-type="other" xlink:type="simple">Suzuki, Y., Yamashita, R., Nakai, K. and Sugano, S. (2002) DBTSS: DataBase of human transcriptional start sites and full-length cDNAs. Nucleic Acids Research, 30, 328-331. doi:10.1093/nar/30.1.328</mixed-citation></ref><ref id="scirp.31900-ref14"><label>14</label><mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Périer</surname><given-names> R.C.</given-names></name>,<name name-style="western"><surname> Praz</surname><given-names> V.</given-names></name>,<name name-style="western"><surname> Junier</surname><given-names> T.</given-names></name>,<name name-style="western"><surname> Bonnard</surname><given-names> C. and Bucher</given-names></name>,<name name-style="western"><surname> P. </surname><given-names>  </given-names></name>,<etal>et al</etal>. (<year>2000</year>)<article-title>The eukaryotic promoter database (EPD)</article-title><source> Nucleic Acids Research</source><volume> 28</volume>,<fpage> 302</fpage>-<lpage>303</lpage>.<pub-id pub-id-type="doi"></pub-id></mixed-citation></ref><ref id="scirp.31900-ref15"><label>15</label><mixed-citation publication-type="other" xlink:type="simple">Zare-Mirakabad, F., Ahrabian, H., Sadeghi, M., Hashemifar, S., Nowzari-Dalini, A. and Goliaei, B. (2009) Genetic algorithm for dyad pattern finding in DNA sequences. Genes &amp; Genetic Systems, 84, 81-93.  
doi:10.1266/ggs.84.81</mixed-citation></ref><ref id="scirp.31900-ref16"><label>16</label><mixed-citation publication-type="other" xlink:type="simple">Habib, N., Kaplan, T., Margalit, H. and Friedman, N. (2008) A novel bayesian DNA motif comparison method for clustering and retrieval. PLoS Computational Biology, 4. doi:10.1371/journal.pcbi.1000010</mixed-citation></ref><ref id="scirp.31900-ref17"><label>17</label><mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Sandelin</surname><given-names> A.</given-names></name>,<name name-style="western"><surname> Alkema</surname><given-names> W.</given-names></name>,<name name-style="western"><surname> Engstr&amp;#246;m</surname><given-names> P.</given-names></name>,<name name-style="western"><surname> Wasserman</surname><given-names> W.W. and Lenhard</given-names></name>,<name name-style="western"><surname> B. </surname><given-names>  </given-names></name>,<etal>et al</etal>. (<year>2004</year>)<article-title>JASPAR: An open-access database for eukaryotic transcription factor binding profiles</article-title><source> Nucleic Acids Research</source><volume> 32</volume>,<fpage> D91</fpage>-<lpage>D94</lpage>.<pub-id pub-id-type="doi"></pub-id></mixed-citation></ref></ref-list></back></article>