<?xml version="1.0" encoding="UTF-8"?><!DOCTYPE article  PUBLIC "-//NLM//DTD Journal Publishing DTD v3.0 20080202//EN" "http://dtd.nlm.nih.gov/publishing/3.0/journalpublishing3.dtd"><article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" dtd-version="3.0" xml:lang="en" article-type="research article"><front><journal-meta><journal-id journal-id-type="publisher-id">CMB</journal-id><journal-title-group><journal-title>Computational Molecular Bioscience</journal-title></journal-title-group><issn pub-type="epub">2165-3445</issn><publisher><publisher-name>Scientific Research Publishing</publisher-name></publisher></journal-meta><article-meta><article-id pub-id-type="doi">10.4236/cmb.2016.62003</article-id><article-id pub-id-type="publisher-id">CMB-67895</article-id><article-categories><subj-group subj-group-type="heading"><subject>Articles</subject></subj-group><subj-group subj-group-type="Discipline-v2"><subject>Biomedical&amp;Life Sciences</subject></subj-group></article-categories><title-group><article-title>
 
 
  Use of FFT in Protein Sequence Comparison under Their Binary Representations
 
</article-title></title-group><contrib-group><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Jayanta</surname><given-names>Pal</given-names></name><xref ref-type="aff" rid="aff1"><sup>1</sup></xref></contrib><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Soumen</surname><given-names>Ghosh</given-names></name><xref ref-type="aff" rid="aff2"><sup>2</sup></xref></contrib><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Bansibadan</surname><given-names>Maji</given-names></name><xref ref-type="aff" rid="aff3"><sup>3</sup></xref></contrib><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Dilip</surname><given-names>Kumar Bhattacharya</given-names></name><xref ref-type="aff" rid="aff4"><sup>4</sup></xref></contrib></contrib-group><aff id="aff4"><addr-line>Department of Pure Mathematics, Calcutta University, Kolkata, India</addr-line></aff><aff id="aff2"><addr-line>Department of Information Technology, Narula Institute of Technology, Kolkata, India</addr-line></aff><aff id="aff3"><addr-line>Department of Electronics &amp;amp; Communication Engineering, National Institute of Technology, Durgapur, India</addr-line></aff><aff id="aff1"><addr-line>Department of Computer Science &amp;amp; Engineering, Narula Institute of Technology, Kolkata, India</addr-line></aff><pub-date pub-type="epub"><day>17</day><month>06</month><year>2016</year></pub-date><volume>06</volume><issue>02</issue><fpage>33</fpage><lpage>40</lpage><history><date date-type="received"><day>14</day>	<month>February</month>	<year>2016</year></date><date date-type="rev-recd"><day>accepted</day>	<month>27</month>	<year>June</year>	</date><date date-type="accepted"><day>30</day>	<month>June</month>	<year>2016</year></date></history><permissions><copyright-statement>&#169; Copyright  2014 by authors and Scientific Research Publishing Inc. </copyright-statement><copyright-year>2014</copyright-year><license><license-p>This work is licensed under the Creative Commons Attribution International License (CC BY). http://creativecommons.org/licenses/by/4.0/</license-p></license></permissions><abstract><p>
 
 
  The paper considers Voss type representation of amino acids and uses FFT on the represented binary sequences to get the spectrum in the frequency domain. Based on the analysis of this spectrum by using the method of inter coefficient difference (ICD), it compares protein sequences of ND5 and ND6 category. Results obtained agree with the standard ones. The purpose of the paper is to extend the ICD method of comparison of DNA sequences to comparison of protein sequences. The topic of discussion is to develop a novel method of comparing protein sequences. The main achievements of the work are that the method applied is completely new of its kind, so far as protein sequence comparison is concerned and moreover the results of comparison agree with the previous results obtained by other methods for the same category of protein sequences.
 
</p></abstract><kwd-group><kwd>Voss Type Representation</kwd><kwd> Inter-Coefficient Difference (ICD) Method</kwd><kwd> Distance Matrix</kwd><kwd> Phylogenetic Tree</kwd><kwd> Fast Fourier Transform (FFT)</kwd><kwd> ND5 and ND6 Category of Protein</kwd></kwd-group></article-meta></front><body><sec id="s1"><title>1. Introduction</title><p>Among the numerous available amino acids only 20 are generally found in living beings and every protein sequence is expressed by these 20 amino acids. The representation of protein in terms of its amino acids is called its primary sequence. Based on this primary sequence representation, protein sequence comparison involves basically two types of methods: 1) Alignment Based Method and 2) Alignment Free Method. Protein sequence comparison was primarily done by different alignment based methods [<xref ref-type="bibr" rid="scirp.67895-ref1">1</xref>] - [<xref ref-type="bibr" rid="scirp.67895-ref3">3</xref>] . But especially due to execution time and comparatively difficult procedure, alignment free methods were preferred subsequently. So far as alignment free methods are concerned, a good literature up to 2003 is available in [<xref ref-type="bibr" rid="scirp.67895-ref4">4</xref>] . So we start with highlighting some of the most important contribution in protein sequence comparisons by alignment free methods from 2003 onwards [<xref ref-type="bibr" rid="scirp.67895-ref5">5</xref>] - [<xref ref-type="bibr" rid="scirp.67895-ref25">25</xref>] . Obviously in most cases, protein sequence comparison also follows similar approach as is considered in genome sequence analysis, because the role of four nucleotides is the same as the role of 20 amino acids in a protein sequence. In details, first of all, numerical representations of the protein sequences are obtained from the numerical values given to the individual amino acids, then graphical representation of the protein sequences is obtained; from these graphs descriptors are derived. These are finally used in comparing protein sequences. All the papers from [<xref ref-type="bibr" rid="scirp.67895-ref7">7</xref>] - [<xref ref-type="bibr" rid="scirp.67895-ref24">24</xref>] involve graphical representations. But another completely different approach is also followed in protein sequence comparison. These are based on classification of amino acids in different groups with different cardinality [<xref ref-type="bibr" rid="scirp.67895-ref8">8</xref>] [<xref ref-type="bibr" rid="scirp.67895-ref26">26</xref>] [<xref ref-type="bibr" rid="scirp.67895-ref27">27</xref>] . Again application of Discrete Fourier Transform in Bioinformatics is also well known. Discrete Fourier Transform (DFT) is nicely used in signal and image processing [<xref ref-type="bibr" rid="scirp.67895-ref28">28</xref>] - [<xref ref-type="bibr" rid="scirp.67895-ref34">34</xref>] . The main areas of its application in DNA research are found in gene prediction, hierarchical analysis and such others [<xref ref-type="bibr" rid="scirp.67895-ref35">35</xref>] - [<xref ref-type="bibr" rid="scirp.67895-ref40">40</xref>] . It is effectively used in identification of protein coding regions, because a DFT spectrum of a DNA sequence reflects the distribution and periodic pattern of the sequence [<xref ref-type="bibr" rid="scirp.67895-ref41">41</xref>] . Use of DFT on binary sequence is found in [<xref ref-type="bibr" rid="scirp.67895-ref42">42</xref>] , where the binary sequence is generated from genome sequences by Voss type of representation. Naturally to find similar use of DFT in protein sequence analysis, corresponding Voss type representation of amino acids is to be known priori. Fortunately Voss representation of DNA sequences involving 4 nucleotides has already been generalized to Voss type representation of 20 amino acids in protein sequences [<xref ref-type="bibr" rid="scirp.67895-ref43">43</xref>] . Such representation of amino acids has already been used in obtaining fuzzy representation of amino acids [<xref ref-type="bibr" rid="scirp.67895-ref43">43</xref>] . These are found to be effective in classification of amino acids in 6 different groups. Finally protein sequence classification has been obtained based on such classified groups of amino acids [<xref ref-type="bibr" rid="scirp.67895-ref8">8</xref>] [<xref ref-type="bibr" rid="scirp.67895-ref25">25</xref>] - [<xref ref-type="bibr" rid="scirp.67895-ref27">27</xref>] [<xref ref-type="bibr" rid="scirp.67895-ref44">44</xref>] . Thus Voss type representation of amino acids is an important contribution in protein sequence analysis. But use of FFT on the binary representations of protein sequences generated by such Voss type representation of amino acids has not yet been attempted in protein sequence comparison. This is the motivation of the paper to consider such binary sequences in comparing protein sequences.</p></sec><sec id="s2"><title>2. Methodology</title><sec id="s2_1"><title>2.1. Voss Type Binary Representation of Amino Acid</title><p>20 amino acids are taken in the following order:</p><p>Alanine (A), Cysteine (C), Aspartic acid (D), Glutamic acid (E), Phenylalanine (F), Glycine (G), Histidine (H), Isoleucine (I), Lysine (K), Leucine (L), Methionine (M), Asparagine (N), Proline (P), Glutamine (Q), Arginine (R), Serine (S), Tyrosine (T), Valine (V), Tryptophan (W) and Threonine (Y).</p><p>Each amino acid is represented by a 20 component vector of which one bit is 1 and others are 0. But the representation follows the order of amino acid taken. For example amino acid Alanine(a) is represented by 10000000000000000000. The same rule is applied for other amino acids also, so that the last amino acid Threonine (Y) is represented by 00000000000000000001.</p><p>From each protein sequence S we get 20 different representations corresponding to 20 different amino acids by putting in the protein sequence 1 for the particular amino acid considered and the rest all 0 for the remaining amino acids. Thus 20 different binary representations viz., U<sub>A</sub>, U<sub>C</sub>, U<sub>D</sub>, U<sub>E</sub>, U<sub>F</sub>, U<sub>G</sub>, U<sub>H</sub>, U<sub>I</sub>, U<sub>K</sub>, U<sub>L</sub>, U<sub>M</sub>, U<sub>N</sub>, U<sub>P</sub>, U<sub>Q</sub>, U<sub>R</sub>, U<sub>S</sub>, U<sub>T</sub>, U<sub>V</sub>, U<sub>W</sub> and U<sub>Y</sub> are obtained.</p></sec><sec id="s2_2"><title>2.2. ICD Method for Protein Sequence Analysis</title><p>The ICD method of DNA sequence and Protein sequence analysis basically remains the same as both deals with binary sequence only. So we describe ICD method as described in [<xref ref-type="bibr" rid="scirp.67895-ref43">43</xref>] . First of all FFT is applied on the binary represented protein sequences of length N say. In the Fourier spectrum the amplitudes are taken, which are N/2 distinct numbers. We normalize these N/2 components by their lengths. On these N/2 normalized components, we take absolute value of the inter coefficient difference (ICD) by calculating the differences of the succeeding terms from the preceding ones. Thus we get (N/2 − 1) distinct elements corresponding to each amino acid. Now 20 such (N/2) − 1 distinct components are concatenated to give a descriptor of length 20*((N/2) − 1)). From such descriptors distance matrix is formed by considering Euclidian Distance measures as follows.</p><p>If <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/2-2220054x6.png" xlink:type="simple"/></inline-formula> and <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/2-2220054x7.png" xlink:type="simple"/></inline-formula> are two sequences for two proteins X and Y, then the distance between X and Y is given by</p><disp-formula id="scirp.67895-formula1102"><graphic  xlink:href="http://html.scirp.org/file/2-2220054x8.png"  xlink:type="simple"/></disp-formula><p>This is the Euclidean distance between X and Y. The smaller is the distance; more similar are the protein sequences. On the basis of this formula the distances between pair of proteins are calculated and they are used to form the diagonal matrix. Due to similarity, only the lower half of the matrix is taken. Now using the UPGMA software on this matrix the Phylogenetic Tree for all the species is obtained. For comparison of protein sequences of different lengths the question of making all the lengths same does not arise normally in FFT. But if necessary, the length may be manually adjusted by putting additional zeros. For example, suppose two protein sequences are of lengths M and N. Then the descriptors for the first and second sequences are of lengths 20*((N/2) − 1) and 20*((M/2) − 1) respectively. As the descriptors are of unequal lengths, so comparison becomes infeasible. Hence if M = N − 2, say, then we first make the lengths of both the sequences equal to N, by putting two additional zeros to the second sequence. But there is no problem in doing so, as the Fourier transform of zeros gives zero spectrum.</p></sec></sec><sec id="s3"><title>3. Sequences for Comparison</title><p>We have used the NADH dehydrogenase subunit 5 (ND5) and subunit 6 (ND6) protein sequences of nine species for comparison as shown in <xref ref-type="table" rid="table1">Table 1</xref>.</p></sec><sec id="s4"><title>4. Results and Discussions</title><sec id="s4_1"><title>4.1. Results</title><p>Distance matrix obtained by applying our method for 9 protein sequences of ND5 and ND6 category have been presented in <xref ref-type="table" rid="table2">Table 2</xref> and <xref ref-type="table" rid="table3">Table 3</xref> respectively. Phylogenetic tree obtained from these data have been presented in <xref ref-type="fig" rid="fig1">Figure 1</xref> and <xref ref-type="fig" rid="fig2">Figure 2</xref> for ND5 and ND6 category respectively.</p></sec><sec id="s4_2"><title>4.2. Discussion</title><p> ICD method, which is dependent on Voss type representation of DNA sequences, is already known to be very much successful in comparing DNA sequences. Voss type representation for protein sequences is comparatively a newer concept. As Voss type representation for protein sequences has been applied recently in different areas and found to be very much successful there, so it is expected that this type of representation might be useful in protein sequence comparison also. This is why; in our paper ICD method based on Voss type representation for protein sequences has been developed and used for protein sequence comparison. No doubt that the present method is a new contribution to the literature of protein sequence comparison.</p><table-wrap id="table1" ><label><xref ref-type="table" rid="table1">Table 1</xref></label><caption><title> List of nine species with their versions and lengths</title></caption><table><tbody><thead><tr><th align="center" valign="middle"  rowspan="2"  >Sl. No.</th><th align="center" valign="middle"  rowspan="2"  >Species</th><th align="center" valign="middle"  colspan="2"  >ND5</th><th align="center" valign="middle"  colspan="2"  >ND6</th></tr></thead><tr><td align="center" valign="middle" >NCBI Reference</td><td align="center" valign="middle" >Length</td><td align="center" valign="middle" >NCBI Reference</td><td align="center" valign="middle" >Length</td></tr><tr><td align="center" valign="middle" >1</td><td align="center" valign="middle" >HUMAN</td><td align="center" valign="middle" >AP-000649.1</td><td align="center" valign="middle" >603</td><td align="center" valign="middle" >AP-000650.1</td><td align="center" valign="middle" >174</td></tr><tr><td align="center" valign="middle" >2</td><td align="center" valign="middle" >GORILLA</td><td align="center" valign="middle" >NP-008222.1</td><td align="center" valign="middle" >603</td><td align="center" valign="middle" >NP-008223.1</td><td align="center" valign="middle" >174</td></tr><tr><td align="center" valign="middle" >3</td><td align="center" valign="middle" >COMMON CHIMPANZEE</td><td align="center" valign="middle" >NP-008196.1</td><td align="center" valign="middle" >603</td><td align="center" valign="middle" >NP-008197.1</td><td align="center" valign="middle" >174</td></tr><tr><td align="center" valign="middle" >4</td><td align="center" valign="middle" >PYGMY CHIMPANZEE</td><td align="center" valign="middle" >NP-008209.1</td><td align="center" valign="middle" >603</td><td align="center" valign="middle" >NP-008210.1</td><td align="center" valign="middle" >174</td></tr><tr><td align="center" valign="middle" >5</td><td align="center" valign="middle" >FIN WHALE</td><td align="center" valign="middle" >NP-006899.1</td><td align="center" valign="middle" >606</td><td align="center" valign="middle" >NP-006900.1</td><td align="center" valign="middle" >175</td></tr><tr><td align="center" valign="middle" >6</td><td align="center" valign="middle" >BLUE WHALE</td><td align="center" valign="middle" >NP-007066.1</td><td align="center" valign="middle" >606</td><td align="center" valign="middle" >NP-007067.1</td><td align="center" valign="middle" >175</td></tr><tr><td align="center" valign="middle" >7</td><td align="center" valign="middle" >RAT</td><td align="center" valign="middle" >AP-004902.1</td><td align="center" valign="middle" >610</td><td align="center" valign="middle" >AP-004903.1</td><td align="center" valign="middle" >172</td></tr><tr><td align="center" valign="middle" >8</td><td align="center" valign="middle" >MOUSE</td><td align="center" valign="middle" >NP-904338.1</td><td align="center" valign="middle" >607</td><td align="center" valign="middle" >NP-904339.1</td><td align="center" valign="middle" >172</td></tr><tr><td align="center" valign="middle" >9</td><td align="center" valign="middle" >OPOSSUM</td><td align="center" valign="middle" >NP-007105.1</td><td align="center" valign="middle" >602</td><td align="center" valign="middle" >NP-007106.1</td><td align="center" valign="middle" >168</td></tr></tbody></table></table-wrap><fig id="fig1"  position="float"><label><xref ref-type="fig" rid="fig1">Figure 1</xref></label><caption><title> Phylogenetic tree obtained for 9 protein sequences of ND5 category</title></caption><graphic mimetype="image"   position="float"  xlink:type="simple"  xlink:href="http://html.scirp.org/file/2-2220054x9.png"/></fig><fig id="fig2"  position="float"><label><xref ref-type="fig" rid="fig2">Figure 2</xref></label><caption><title> Phylogenetic tree obtained for 9 protein sequences of ND6 category</title></caption><graphic mimetype="image"   position="float"  xlink:type="simple"  xlink:href="http://html.scirp.org/file/2-2220054x10.png"/></fig><table-wrap id="table2" ><label><xref ref-type="table" rid="table2">Table 2</xref></label><caption><title> Distance matrix (lower triangular) for 9 protein sequences of ND5 category</title></caption><table><tbody><thead><tr><th align="center" valign="middle" ></th><th align="center" valign="middle" >Human</th><th align="center" valign="middle" >Gorilla</th><th align="center" valign="middle" >P. Chim</th><th align="center" valign="middle" >C. Chim</th><th align="center" valign="middle" >Rat</th><th align="center" valign="middle" >Mouse</th><th align="center" valign="middle" >B_Whale</th><th align="center" valign="middle" >F_Whale</th><th align="center" valign="middle" >Opossum</th></tr></thead><tr><td align="center" valign="middle" >Human</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td></tr><tr><td align="center" valign="middle" >Gorilla</td><td align="center" valign="middle" >0.88011</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td></tr><tr><td align="center" valign="middle" >P. Chim</td><td align="center" valign="middle" >0.78392</td><td align="center" valign="middle" >0.83493</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td></tr><tr><td align="center" valign="middle" >C. Chim</td><td align="center" valign="middle" >0.80377</td><td align="center" valign="middle" >0.86392</td><td align="center" valign="middle" >0.70748</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td></tr><tr><td align="center" valign="middle" >Rat</td><td align="center" valign="middle" >1.33343</td><td align="center" valign="middle" >1.32515</td><td align="center" valign="middle" >1.32102</td><td align="center" valign="middle" >1.33353</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td></tr><tr><td align="center" valign="middle" >Mouse</td><td align="center" valign="middle" >1.31723</td><td align="center" valign="middle" >1.3021</td><td align="center" valign="middle" >1.29413</td><td align="center" valign="middle" >1.31185</td><td align="center" valign="middle" >1.16305</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td></tr><tr><td align="center" valign="middle" >B_Whale</td><td align="center" valign="middle" >1.28719</td><td align="center" valign="middle" >1.28518</td><td align="center" valign="middle" >1.27026</td><td align="center" valign="middle" >1.29161</td><td align="center" valign="middle" >1.32947</td><td align="center" valign="middle" >1.3423</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td></tr><tr><td align="center" valign="middle" >F_Whale</td><td align="center" valign="middle" >1.27388</td><td align="center" valign="middle" >1.2825</td><td align="center" valign="middle" >1.26707</td><td align="center" valign="middle" >1.29284</td><td align="center" valign="middle" >1.31818</td><td align="center" valign="middle" >1.32474</td><td align="center" valign="middle" >0.73081</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" ></td></tr><tr><td align="center" valign="middle" >Opossum</td><td align="center" valign="middle" >1.39053</td><td align="center" valign="middle" >1.38974</td><td align="center" valign="middle" >1.37859</td><td align="center" valign="middle" >1.3812</td><td align="center" valign="middle" >1.39501</td><td align="center" valign="middle" >1.3896</td><td align="center" valign="middle" >1.4018</td><td align="center" valign="middle" >1.40041</td><td align="center" valign="middle" >0</td></tr></tbody></table></table-wrap><p> Obviously ICD method, may be for DNA sequence comparison or Protein sequence comparison, is comparatively easier and straight forward to apply.</p><p> To compare our results with those obtained earlier by other methods on the same species, we first mention them as far as possible. The phylogenetic tree obtained in [<xref ref-type="bibr" rid="scirp.67895-ref25">25</xref>] for 9 species of ND5 category is given in <xref ref-type="fig" rid="fig3">Figure 3</xref>. Similarly the phylogenetic trees obtained in [<xref ref-type="bibr" rid="scirp.67895-ref26">26</xref>] for 9 species of ND5 category and ND6 category are given in <xref ref-type="fig" rid="fig4">Figure 4</xref> and <xref ref-type="fig" rid="fig5">Figure 5</xref> respectively and the phylogenetic trees obtained in [<xref ref-type="bibr" rid="scirp.67895-ref44">44</xref>] for 9 species of ND5 category and ND6 category are given in <xref ref-type="fig" rid="fig6">Figure 6</xref> and <xref ref-type="fig" rid="fig7">Figure 7</xref> respectively.</p><fig id="fig3"  position="float"><label><xref ref-type="fig" rid="fig3">Figure 3</xref></label><caption><title> Phylogenetic tree obtained in [<xref ref-type="bibr" rid="scirp.67895-ref25">25</xref>] for 9 species of ND5 category</title></caption><graphic mimetype="image"   position="float"  xlink:type="simple"  xlink:href="http://html.scirp.org/file/2-2220054x11.png"/></fig><fig id="fig4"  position="float"><label><xref ref-type="fig" rid="fig4">Figure 4</xref></label><caption><title> Phylogenetic tree obtained in [<xref ref-type="bibr" rid="scirp.67895-ref26">26</xref>] for 9 species of ND5 category</title></caption><graphic mimetype="image"   position="float"  xlink:type="simple"  xlink:href="http://html.scirp.org/file/2-2220054x12.png"/></fig><table-wrap id="table3" ><label><xref ref-type="table" rid="table3">Table 3</xref></label><caption><title> Distance matrix (lower triangular) for 9 protein sequences of ND6 category</title></caption><table><tbody><thead><tr><th align="center" valign="middle" ></th><th align="center" valign="middle" >Human</th><th align="center" valign="middle" >Gorilla</th><th align="center" valign="middle" >P. Chim</th><th align="center" valign="middle" >C. Chim</th><th align="center" valign="middle" >Rat</th><th align="center" valign="middle" >Mouse</th><th align="center" valign="middle" >B_Whale</th><th align="center" valign="middle" >F_Whale</th><th align="center" valign="middle" >Opossum</th></tr></thead><tr><td align="center" valign="middle" >Human</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td></tr><tr><td align="center" valign="middle" >Gorilla</td><td align="center" valign="middle" >0.61527</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td></tr><tr><td align="center" valign="middle" >P. Chim</td><td align="center" valign="middle" >0.66602</td><td align="center" valign="middle" >0.60809</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td></tr><tr><td align="center" valign="middle" >C. Chim</td><td align="center" valign="middle" >0.65876</td><td align="center" valign="middle" >0.60947</td><td align="center" valign="middle" >0.45713</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td></tr><tr><td align="center" valign="middle" >Rat</td><td align="center" valign="middle" >1.40056</td><td align="center" valign="middle" >1.3828</td><td align="center" valign="middle" >1.37694</td><td align="center" valign="middle" >1.38812</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td></tr><tr><td align="center" valign="middle" >Mouse</td><td align="center" valign="middle" >1.44023</td><td align="center" valign="middle" >1.43475</td><td align="center" valign="middle" >1.43451</td><td align="center" valign="middle" >1.44952</td><td align="center" valign="middle" >1.10948</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td></tr><tr><td align="center" valign="middle" >B_Whale</td><td align="center" valign="middle" >1.29625</td><td align="center" valign="middle" >1.28337</td><td align="center" valign="middle" >1.26906</td><td align="center" valign="middle" >1.27875</td><td align="center" valign="middle" >1.35032</td><td align="center" valign="middle" >1.37165</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td></tr><tr><td align="center" valign="middle" >F_Whale</td><td align="center" valign="middle" >1.28281</td><td align="center" valign="middle" >1.28369</td><td align="center" valign="middle" >1.265</td><td align="center" valign="middle" >1.28498</td><td align="center" valign="middle" >1.35476</td><td align="center" valign="middle" >1.36975</td><td align="center" valign="middle" >0.86639</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" ></td></tr><tr><td align="center" valign="middle" >Opossum</td><td align="center" valign="middle" >1.48984</td><td align="center" valign="middle" >1.46302</td><td align="center" valign="middle" >1.4753</td><td align="center" valign="middle" >1.47082</td><td align="center" valign="middle" >1.37682</td><td align="center" valign="middle" >1.43112</td><td align="center" valign="middle" >1.46181</td><td align="center" valign="middle" >1.44397</td><td align="center" valign="middle" >0</td></tr></tbody></table></table-wrap><fig id="fig5"  position="float"><label><xref ref-type="fig" rid="fig5">Figure 5</xref></label><caption><title> Phylogenetic tree obtained in [<xref ref-type="bibr" rid="scirp.67895-ref26">26</xref>] for 9 species of ND6 category</title></caption><graphic mimetype="image"   position="float"  xlink:type="simple"  xlink:href="http://html.scirp.org/file/2-2220054x13.png"/></fig><fig id="fig6"  position="float"><label><xref ref-type="fig" rid="fig6">Figure 6</xref></label><caption><title> Phylogenetic tree obtained in [<xref ref-type="bibr" rid="scirp.67895-ref44">44</xref>] for 9 species of ND5 category</title></caption><graphic mimetype="image"   position="float"  xlink:type="simple"  xlink:href="http://html.scirp.org/file/2-2220054x14.png"/></fig><fig id="fig7"  position="float"><label><xref ref-type="fig" rid="fig7">Figure 7</xref></label><caption><title> Phylogenetic tree obtained in [<xref ref-type="bibr" rid="scirp.67895-ref44">44</xref>] for 9 species of ND6 category</title></caption><graphic mimetype="image"   position="float"  xlink:type="simple"  xlink:href="http://html.scirp.org/file/2-2220054x15.png"/></fig><p>From the above phylogenetic trees obtained for ND5 and ND6 categories of protein, it is revealed that in both the cases the phylogenetic trees obtained by our method almost agree with the earlier phylogenetic trees obtained by other methods.</p></sec></sec><sec id="s5"><title>5. Conclusion</title><p>Our method is effective and easier to apply in protein sequence comparison.</p></sec><sec id="s6"><title>Cite this paper</title><p>Jayanta Pal,Soumen Ghosh,Bansibadan Maji,Dilip Kumar Bhattacharya, (2016) Use of FFT in Protein Sequence Comparison under Their Binary Representations. Computational Molecular Bioscience,06,33-40. doi: 10.4236/cmb.2016.62003</p></sec></body><back><ref-list><title>References</title><ref id="scirp.67895-ref1"><label>1</label><mixed-citation publication-type="other" xlink:type="simple">Phillips, A., Janies, D. and Wheeler, W. (2000) Multiple Sequence Alignment in Phylogenetic Analysis. Molecular Phylogenetics and Evolution, 16, 317-330. http://dx.doi.org/10.1006/mpev.2000.0785</mixed-citation></ref><ref id="scirp.67895-ref2"><label>2</label><mixed-citation publication-type="other" xlink:type="simple">Thompson, J.D., Higgins, D.G. and Gibson, T.J. (1994) CLUSTAL W: Improving the Sensitivity of Progressive Multiple Sequence Alignment through Sequence Weighting, Position-Specific Gap Penalties and Weight Matrix Choice. Nucleic Acids Research, 22, 4673-4680. http://dx.doi.org/10.1093/nar/22.22.4673</mixed-citation></ref><ref id="scirp.67895-ref3"><label>3</label><mixed-citation publication-type="other" xlink:type="simple">Katoh, K., Misawa, K., Kuma, K. and Miyata, T. (2002) MAFFT: A Novel Method for Rapid Multiple Sequence Alignment Based on Fast Fourier Transform. Nucleic Acids Research, 30, 3059-3066. http://dx.doi.org/10.1093/nar/gkf436</mixed-citation></ref><ref id="scirp.67895-ref4"><label>4</label><mixed-citation publication-type="other" xlink:type="simple">Vinga, S. and Almeida, J. (2003) Alignment-Free Sequence Comparison—A Review. Bioinformatics, 19, 513-523.http://dx.doi.org/10.1093/bioinformatics/btg005</mixed-citation></ref><ref id="scirp.67895-ref5"><label>5</label><mixed-citation publication-type="other" xlink:type="simple">Pinello, L., Lo Bosco, G. and Yuan, G.-C. (2013) Applications of Alignment-Free Methods in Epigenomics. Briefings in Bioinformatics, 15, 419-430. http://dx.doi.org/10.1093/bib/bbt078</mixed-citation></ref><ref id="scirp.67895-ref6"><label>6</label><mixed-citation publication-type="other" xlink:type="simple">Domazet-Loso, M. and Haubold, B. (2011) Alignment-Free Detection of Local Similarity among Viral and Bacterial Genomes. Bioinformatics, 27, 1466-1472. http://dx.doi.org/10.1093/bioinformatics/btr176</mixed-citation></ref><ref id="scirp.67895-ref7"><label>7</label><mixed-citation publication-type="other" xlink:type="simple">Ghosh, S., Pal, J., Maji, B. and Bhattacharya, D.K. (2016) Condensed Matrix Descriptor for Proteinb Sequence Comparison. International Journal of Analytical Mass Spectrometry and Chromatography, 4, 1-13.http://dx.doi.org/10.4236/ijamsc.2016.41001</mixed-citation></ref><ref id="scirp.67895-ref8"><label>8</label><mixed-citation publication-type="other" xlink:type="simple">Li, C., Xing, L.L. and Wang, X. (2008) 2-D Graphical Representation of Protein Sequences and Its Application to Coronavirus Phylogeny. BMB Reports, 41, 217-222. http://dx.doi.org/10.5483/BMBRep.2008.41.3.217</mixed-citation></ref><ref id="scirp.67895-ref9"><label>9</label><mixed-citation publication-type="other" xlink:type="simple">Randic, M., Mehulic, K., Vukicevic, D., Pisanski, T., Vikic-Topic, D. and Plavsic, D. (2009) Graphical Representation of Proteins as Four-Color Maps and Their Numerical Characterization. Journal of Molecular Graphics and Modelling, 27, 637-641. http://dx.doi.org/10.1016/j.jmgm.2008.10.004</mixed-citation></ref><ref id="scirp.67895-ref10"><label>10</label><mixed-citation publication-type="other" xlink:type="simple">Bai, F. and Wang, T. (2006) On Graphical and Numerical Representation of Protein Sequences. Journal of Biomolecular Structure and Dynamics, 23, 537-545. http://dx.doi.org/10.1080/07391102.2006.10507078</mixed-citation></ref><ref id="scirp.67895-ref11"><label>11</label><mixed-citation publication-type="other" xlink:type="simple">Randic, M. (2007) 2-D Graphical Representation of Proteins Based on Physico-Chemical Properties of Amino Acids. Chemical Physics Letters, 440, 291-295. http://dx.doi.org/10.1016/j.cplett.2007.04.037</mixed-citation></ref><ref id="scirp.67895-ref12"><label>12</label><mixed-citation publication-type="other" xlink:type="simple">Ghosh, A. and Nandy, A. (2011) Graphical Representation and Mathematical Characterization of Protein Sequences and Applications to Viral Proteins. Advances in Protein Chemistry and Structural Biology, 83, 1-42.http://dx.doi.org/10.1016/B978-0-12-381262-9.00001-X</mixed-citation></ref><ref id="scirp.67895-ref13"><label>13</label><mixed-citation publication-type="other" xlink:type="simple">Randic, M., Zupan, J. and Vikic-Topic, D. (2007) On Representation of Proteins by Star-Like Graphs. Journal of Molecular Graphics and Modelling, 26, 290-305. http://dx.doi.org/10.1016/j.jmgm.2006.12.006</mixed-citation></ref><ref id="scirp.67895-ref14"><label>14</label><mixed-citation publication-type="other" xlink:type="simple">Wen, J. and Zhang, Y. (2009) A 2D Graphical Representation of Protein Sequence and Its Numerical Characterization. Chemical Physics Letters, 476, 281-286. http://dx.doi.org/10.1016/j.cplett.2009.06.017</mixed-citation></ref><ref id="scirp.67895-ref15"><label>15</label><mixed-citation publication-type="other" xlink:type="simple">Liao, B., Sun, X. and Zeng, Q. (2010) A Novel Method for Similarity Analysis and Protein Sub-Cellular Localization Prediction. Bio-Informatics, 26, 2678-2683. http://dx.doi.org/10.1093/bioinformatics/btq521</mixed-citation></ref><ref id="scirp.67895-ref16"><label>16</label><mixed-citation publication-type="other" xlink:type="simple">Novic, M. and Randic, M. (2008) Representation of Proteins as Walks in 20-D Space. SAR and QSAR in Environmental Research, 19, 317-337. http://dx.doi.org/10.1080/10629360802085066</mixed-citation></ref><ref id="scirp.67895-ref17"><label>17</label><mixed-citation publication-type="other" xlink:type="simple">Yu, H.-J. and Huang, D.-S. (2012) Novel 20-D Descriptors of Protein Sequences and Its Applications in Similarity Analysis. Chemical Physics Letters, 531, 261-266. http://dx.doi.org/10.1016/j.cplett.2012.02.030</mixed-citation></ref><ref id="scirp.67895-ref18"><label>18</label><mixed-citation publication-type="other" xlink:type="simple">He, P.-A., Wei, J., Yao, Y. and Tie, Z. (2012) A Novel Graphical Representation of Proteins and Its Application. Physica A: Statistical Mechanics and Its Applications, 391, 93-99.</mixed-citation></ref><ref id="scirp.67895-ref19"><label>19</label><mixed-citation publication-type="other" xlink:type="simple">Randic, M., Novic, M. and Vracko, M. (2008) On Novel Representation of Proteins Based on Amino Acid Adjacency Matrix. SAR and QSAR in Environmental Research, 19, 339-349. http://dx.doi.org/10.1080/10629360802085082</mixed-citation></ref><ref id="scirp.67895-ref20"><label>20</label><mixed-citation publication-type="other" xlink:type="simple">Abo-Elkhier, M.M. (2012) Similarity/Dissimilarity Analysis of Protein Sequences Using the Spatial Median as a Descriptor. Journal of Biophysical Chemistry, 3, 142-148. http://dx.doi.org/10.4236/jbpc.2012.32016</mixed-citation></ref><ref id="scirp.67895-ref21"><label>21</label><mixed-citation publication-type="other" xlink:type="simple">Randic, M., Zupan, J. and Balaban, A.T. (2004) Unique Graphical Representation of Protein Sequences Based on Nucleotide Triplet Codons. Chemical Physics Letters, 397, 247-252. http://dx.doi.org/10.1016/j.cplett.2004.08.118</mixed-citation></ref><ref id="scirp.67895-ref22"><label>22</label><mixed-citation publication-type="other" xlink:type="simple">El-Lakkani, A. and El-Sherif, S. (2013) Similarity Analysis of Protein Sequences Based on 2D and 3D Amino Acid Adjacency Matrices. Chemical Physics Letters, 590, 192-195. http://dx.doi.org/10.1016/j.cplett.2013.10.032</mixed-citation></ref><ref id="scirp.67895-ref23"><label>23</label><mixed-citation publication-type="other" xlink:type="simple">Feng, Z.-P. and Zhang, C.-T. (2002) A Graphic Representation of Protein Sequence and Predicting the Sub-Cellular Locations of Prokaryotic Proteins. International Journal of Biochemistry and Cell Biology, 34, 298-307. http://dx.doi.org/10.1016/S1357-2725(01)00121-2</mixed-citation></ref><ref id="scirp.67895-ref24"><label>24</label><mixed-citation publication-type="other" xlink:type="simple">Yao, Y.H., Kong, F., Dai, Q. and He, P.-A. (2013) A Sequence-Segmented Method Applied to the Similarity Analysis of Long Protein Sequence. MATCH: Communications in Mathematical and in Computer Chemistry, 70, 431-450.</mixed-citation></ref><ref id="scirp.67895-ref25"><label>25</label><mixed-citation publication-type="other" xlink:type="simple">He, P.-A., Li, X.-F., Yang, J.-L. and Wang, J. (2011) A Novel Descriptor for Protein Similarity Analysis. MATCH: Communications in Mathematical and in Computer Chemistry, 65, 445-458.</mixed-citation></ref><ref id="scirp.67895-ref26"><label>26</label><mixed-citation publication-type="other" xlink:type="simple">Ghosh, S., Pal, J., Das, S. and Bhattacharya, D.K. (2015) Differentiation of Protein Sequence Comparison Based on Biological and Theoretical Classifications of Amino Acids in Six Groups. International Journal of Computer Science and Software Engineering, 5, 695-698.</mixed-citation></ref><ref id="scirp.67895-ref27"><label>27</label><mixed-citation publication-type="other" xlink:type="simple">Zhang, Y.S. and Yu, X.T. (2010) Analysis of Protein Sequence Similarity. IEEE, 1255-1258.</mixed-citation></ref><ref id="scirp.67895-ref28"><label>28</label><mixed-citation publication-type="other" xlink:type="simple">Wu, Y.-L., Agrawal, D. and El Abbadi, A. (2000) A Comparison of DFT and DWT Based Similarity Search in Time-Series Databases. Proceedings of the 9th International Conference on Information and Knowledge Management, McLean, 6-11 November 2000, 488-495.</mixed-citation></ref><ref id="scirp.67895-ref29"><label>29</label><mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Anastassiou</surname><given-names> D. </given-names></name>,<etal>et al</etal>. (<year>2000</year>)<article-title>Frequency-Domain Analysis of Bimolecular Sequences</article-title><source> Bioinformatics</source><volume> 16</volume>,<fpage> 1073</fpage>-<lpage>1081</lpage>.<pub-id pub-id-type="doi"></pub-id></mixed-citation></ref><ref id="scirp.67895-ref30"><label>30</label><mixed-citation publication-type="other" xlink:type="simple">Vaidyanathan, P. and Yoon, B.-J. (2004) The Role of Signal Processing Concepts in Genomics and Proteomics. Journal of the Franklin Institute, 341, 111-135. http://dx.doi.org/10.1016/j.jfranklin.2003.12.001</mixed-citation></ref><ref id="scirp.67895-ref31"><label>31</label><mixed-citation publication-type="other" xlink:type="simple">Brigham, E.O. and Morrow, R.E. (1967) The Fast Fourier Transform. IEEE Spectrum, 4, 63-70. http://dx.doi.org/10.1109/MSPEC.1967.5217220</mixed-citation></ref><ref id="scirp.67895-ref32"><label>32</label><mixed-citation publication-type="other" xlink:type="simple">Lyons, R.G. (2004) Understanding Digital Signal Processing. Pearson Education, Upper Saddle River.</mixed-citation></ref><ref id="scirp.67895-ref33"><label>33</label><mixed-citation publication-type="other" xlink:type="simple">Oppenheim, A.V. and Schafer, R.W. (2010) Discrete-Time Signal Processing. 3rd Edition, Prentice Hall, Upper Saddle River.</mixed-citation></ref><ref id="scirp.67895-ref34"><label>34</label><mixed-citation publication-type="other" xlink:type="simple">Akhtar, M., Epps, J. and Ambikairajah, E. (2008) Signal Processing in Sequence Analysis: Advances in Eukaryotic Gene Prediction. IEEE Journal of Selected Topics in Signal Processing, 2, 310-321. http://dx.doi.org/10.1109/JSTSP.2008.923854</mixed-citation></ref><ref id="scirp.67895-ref35"><label>35</label><mixed-citation publication-type="other" xlink:type="simple">Yin, C.C. and Yau, S.S.-T. (2007) Prediction of Protein Coding Regions by the 3-Base Periodicity Analysis of a DNA Sequence. Journal of Theoretical Biology, 247, 687-894. http://dx.doi.org/10.1016/j.jtbi.2007.03.038</mixed-citation></ref><ref id="scirp.67895-ref36"><label>36</label><mixed-citation publication-type="other" xlink:type="simple">Tiwari, S., Ramchandran, S., Bhattacharya, A., Bhattacharya, S. and Ramaswami, R. (1997) Prediction of Probable Genes by Fourier Analysis of Genome Sequences. Computer Applications in the Biosciences, 13, 263-270.</mixed-citation></ref><ref id="scirp.67895-ref37"><label>37</label><mixed-citation publication-type="other" xlink:type="simple">Afreixo, V., Bastos, C.A., Garcia, S.P. and Ferrieira, P.J. (2009) Genome Analysis with Inter-Nucleotide Distances. Bioinformatics, 25, 3064-3070.</mixed-citation></ref><ref id="scirp.67895-ref38"><label>38</label><mixed-citation publication-type="other" xlink:type="simple">Abu-Zahhad, M., Ahmed, S.M. and Abd-Elrahman, S.A. (2012) Genomic Analysis and Classification of Exon and Intron Sequences Using DNA Nu-merical Mapping Techniques. International Journal of Information Technology and Computer Science (IJITCS), 8, 22-36.</mixed-citation></ref><ref id="scirp.67895-ref39"><label>39</label><mixed-citation publication-type="other" xlink:type="simple">Sitansu, S.S. and Panda, G. (2010) A DSP Approach for Protein Coding Region Identification in DNA Sequence. International Journal of Signal and Image Processing, 1, 75-79.</mixed-citation></ref><ref id="scirp.67895-ref40"><label>40</label><mixed-citation publication-type="other" xlink:type="simple">Saberkari, H., Shamsi, M., Sedaaghi, M. and Golabi, F. (2012) Prediction of Protein Coding Regions in DNA Sequences Using Signal Processing Methods. IEEE Symposium on Industrial Electronics and Applications (ISIEA), Bandung, 23-26 September 2012, 355-360.</mixed-citation></ref><ref id="scirp.67895-ref41"><label>41</label><mixed-citation publication-type="other" xlink:type="simple">Hoang, T., Yin, C.C., Zheng, H., Yu, C.L. and He, R.L. (2015) A New Method to Cluster DNA Sequences Using Fourier Power Spectrum. Journal of Theoretical Biology, 372, 135-145. http://dx.doi.org/10.1016/j.jtbi.2015.02.026</mixed-citation></ref><ref id="scirp.67895-ref42"><label>42</label><mixed-citation publication-type="other" xlink:type="simple">King, B.R., Aburdene, M., Thompson, A. and Warres, Z. (2014) Application of Discrete Fourier Inter-Coefficient Difference for Assessing Genetic Sequence Similarity. EURASIP Journal on Bioinformatics and Systems Biology, 2014, 8.</mixed-citation></ref><ref id="scirp.67895-ref43"><label>43</label><mixed-citation publication-type="other" xlink:type="simple">Ghosh, S., Pal, J. and Bhattacharya, D.K. (2014) Classi-fication of Amino Acids of a Protein on the Basis of Fuzzy Set Theory. International Journal of Modern Sciences and Engineering Technology, 1, 30-35.</mixed-citation></ref><ref id="scirp.67895-ref44"><label>44</label><mixed-citation publication-type="other" xlink:type="simple">Jafarzadeh, N. and Iranmanesh, A. (2015) A New Measure for Pairwise Comparison of Protein Sequences. MATCH: Communications in Mathematical and in Computer Chemistry, 74, 563-574.</mixed-citation></ref></ref-list></back></article>