<?xml version="1.0" encoding="UTF-8"?><!DOCTYPE article  PUBLIC "-//NLM//DTD Journal Publishing DTD v3.0 20080202//EN" "http://dtd.nlm.nih.gov/publishing/3.0/journalpublishing3.dtd"><article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" dtd-version="3.0" xml:lang="en" article-type="research article"><front><journal-meta><journal-id journal-id-type="publisher-id">JBiSE</journal-id><journal-title-group><journal-title>Journal of Biomedical Science and Engineering</journal-title></journal-title-group><issn pub-type="epub">1937-6871</issn><publisher><publisher-name>Scientific Research Publishing</publisher-name></publisher></journal-meta><article-meta><article-id pub-id-type="doi">10.4236/jbise.2019.127027</article-id><article-id pub-id-type="publisher-id">JBiSE-93817</article-id><article-categories><subj-group subj-group-type="heading"><subject>Articles</subject></subj-group><subj-group subj-group-type="Discipline-v2"><subject>Biomedical&amp;Life Sciences</subject></subj-group></article-categories><title-group><article-title>
 
 
  Predicting pH Optimum for Activity of Beta-Glucosidases
 
</article-title></title-group><contrib-group><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Shaomin</surname><given-names>Yan</given-names></name><xref ref-type="aff" rid="aff1"><sup>1</sup></xref></contrib><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Guang</surname><given-names>Wu</given-names></name><xref ref-type="aff" rid="aff1"><sup>1</sup></xref><xref ref-type="corresp" rid="cor1"><sup>*</sup></xref></contrib></contrib-group><aff id="aff1"><addr-line>State Key Laboratory of Non-Food Biomass and Enzyme Technology, Guangxi Key Laboratory of Bio-Refinery, National Engineering Research Center for Non-Food Biorefinery, Guangxi Biomass Engineering Technology Research Center, Guangxi Academy of Sciences, Nanning, China</addr-line></aff><pub-date pub-type="epub"><day>08</day><month>07</month><year>2019</year></pub-date><volume>12</volume><issue>07</issue><fpage>354</fpage><lpage>367</lpage><history><date date-type="received"><day>10,</day>	<month>June</month>	<year>2019</year></date><date date-type="rev-recd"><day>20,</day>	<month>July</month>	<year>2019</year>	</date><date date-type="accepted"><day>23,</day>	<month>July</month>	<year>2019</year></date></history><permissions><copyright-statement>&#169; Copyright  2014 by authors and Scientific Research Publishing Inc. </copyright-statement><copyright-year>2014</copyright-year><license><license-p>This work is licensed under the Creative Commons Attribution International License (CC BY). http://creativecommons.org/licenses/by/4.0/</license-p></license></permissions><abstract><p>
 
 
  The working conditions for enzymatic reaction are elegant, but not many optimal conditions are documented in literatures. For newly mutated and newly found enzymes, the optimal working conditions can only be extrapolated from our previous experience. Therefore a question raised here is whether we can use the knowledge on enzyme structure to predict the optimal working conditions. Although working conditions for enzymes can be easily measured in experiments, the predictions of working conditions for enzymes are still important because they can minimize the experimental cost and time. In this study, we develop a 20-1 feedforward backpropagation neural network with information on amino acid sequence to predict the pH optimum for the activity of beta-glucosidase, because this enzyme has drawn much attention for its role in bio-fuel industries. Among 25 features of amino acids being screened, the results show that 11 features can be used as predictors in this model and the amino-acid distribution probability is the best in predicting the pH optimum for the activity of beta-glucosidases. Our study paves the way for predicting the optimal working conditions of enzymes based on the amino-acid features.
 
</p></abstract><kwd-group><kwd>Beta-Glucosidase</kwd><kwd> Enzyme</kwd><kwd> pH Optimum</kwd><kwd> Prediction</kwd></kwd-group></article-meta></front><body><sec id="s1"><title>1. Introduction</title><p>Enzymatic reactions require many elegant conditions, which are usually determined through experiments. Those elegant experimental conditions are valuable for any new experiments with new enzymes because they can save much time and money for experimenters. On the other hand, many elegant experimental conditions are not always available in literature, so the valuable experience could not be fully useful for fellow researchers.</p><p>Still, the modern protein designing produces numerous new enzymes, whose optimal working conditions are totally unknown. Although we can extrapolate our previous experience to new enzymes, they are generally empirical.</p><p>With fast development on computational chemistry and bioinformatics, it could be possible to use models to predict the optimal working conditions for enzymatic reactions with newly designed enzymes. This is plausible because currently a lot of information on primary, secondary, tertiary, and quaternary structures is readily available, and many studies have been done on account of structure-function relationship of proteins [1,2]. Actually the optimal working conditions are adjusted in order to be suitable for enzymatic function, more exactly for enzyme structure, therefore we could assume that there is a certain relationship between enzyme structure and working conditions in enzymatic reaction. This assumption would lay the foundation for predicting the working condition in enzymatic reaction using enzyme structure.</p><p>Another effort made by scientific community is to build a comprehensive database to include enzymes with their functional parameters in enzymatic reactions, for example, Km and pH. However, even such comprehensive database cannot include all the parameters for all enzymes simply because many enzymatic parameters are not documented in literature.</p><p>Although a measurement of working condition is not difficult during experiments, a measurement is different from a prediction, not only because they are different along the time course, i.e. a measurement is related to the past while a prediction is related to the future; but also because they are different in mechanism, i.e. a measurement is related to mechanism of enzymatic reaction while a prediction is related to enzyme structure-function relationship.</p><p>The β-glucosidase (EC 3.2.1 .21) plays an important role in biological processes because it cuts the β-bond linkage in glucose molecules [<xref ref-type="bibr" rid="scirp.93817-ref3">3</xref>], of which celluloses got much recent attention because of interests in its role in biofuels [<xref ref-type="bibr" rid="scirp.93817-ref4">4</xref>]. With such great interest, more efforts are made not only to search for new β-glucosidases but also to mutate current β-glucosidases, so we have more and more β-glucosidases with clear annotations of their primary structures but without their working conditions for enzymatic reactions, for example, pH optimum. This would provide a good case for developing models to predict the pH optimum for the activity of newly mutated and newly found β-glucosidases because the predictions of pH for protein stability have been the research focus for years [5 - 8]. Therefore, the prediction of pH optimum for enzyme reaction would advance our current knowledge from structure-function relationship to structure-environment relationship.</p><p>In this study, we attempted to use the knowledge about amino-acid features from β-glucosidase sequences to predict the pH optimum for the activity of β-glucosidases.</p></sec><sec id="s2"><title>2. MATERIALS AND METHODS</title><sec id="s2_1"><title>2.1. Data</title><p>The β-glucosidases (EC 3.2.1 .21) are found in the Comprehensive Enzyme Information System BRENDA [<xref ref-type="bibr" rid="scirp.93817-ref9">9</xref>]. In this databank, only 34 β-glucosidases were found under the category of pH optimum as functional parameter, of which two β-glucosidases are documented with their mutants [10,11]. Also, two pH values are documented in each of the β-glucosidase A9UIG0, B5TWK3, and P96316, respectively. In total, this databank provides 44 matched β-glucosidases with their pH values, while information on sequences of β-glucosidases was found in the UniProt [<xref ref-type="bibr" rid="scirp.93817-ref12">12</xref>].</p></sec><sec id="s2_2"><title>2.2. Predictors</title><p>For most enzymes, we generally have only their primary structure because the knowledge on secondary, tertiary, and quaternary structures would require considerable amount of experiments. Therefore, the prediction at this stage would focus on using knowledge of amino acids from enzyme sequences. We use several amino-acid properties listed in <xref ref-type="table" rid="table1">Table 1</xref> as predictors. The knowledge in <xref ref-type="table" rid="table1">Table 1</xref> is actually the values reflecting various aspects of amino acids [<xref ref-type="bibr" rid="scirp.93817-ref13">13</xref>], for example, the spatial properties listed from row 3 to row 6 [14,15]; the hydrophobic properties listed from row 7 to row 11 [16 - 18]; the electronic properties</p><p>listed from row 12 to row 18 [<xref ref-type="bibr" rid="scirp.93817-ref19">19</xref>], and the amino-acid based secondary structure predictions listed from row 17 to 25 row, which are depended on assigning a set of predicted values to a residue and then calculated by applying a simple algorithm [<xref ref-type="bibr" rid="scirp.93817-ref20">20</xref>].</p><p>A particular characteristic in <xref ref-type="table" rid="table1">Table 1</xref> is that those values are constant regardless amino-acid position in a protein, neighboring amino acids, protein length, etc. This is understandable because those properties would not be changed in these regards, for example, an amino acid’s physicochemical property would not be different no matter where this amino acid is located in a protein. As the amino-acid composition is different one from another in β-glucosidases, we weigh the values listed in <xref ref-type="table" rid="table1">Table 1</xref> by multiplying their amino-acid composition of each β-glucosidase.</p><p>Besides the classical knowledge listed in <xref ref-type="table" rid="table1">Table 1</xref>, there is also the amino-acid distribution probability that is based on the occupancy of subpopulations and partitions [<xref ref-type="bibr" rid="scirp.93817-ref21">21</xref>] and reflects the random aspect of amino-acid distribution along a protein (for review and textbook, see [22 - 26]). The difference is that the amino-acid distribution probability does not give each amino acid a constant value as shown in <xref ref-type="table" rid="table1">Table 1</xref>, but the value subject to the length of enzyme and position of each amino acid. <xref ref-type="table" rid="table2">Table 2</xref> shows such a difference.</p></sec><sec id="s2_3"><title>2.3. Predictive Model</title><p>As the predictors are directly related to 20 types of amino acids, so it is natural to consider a predictive model to couple 20 inputs of knowledge on amino acids with single output with documented pH optimum. As this predictive model advances a step from structure-function to structure-environment relationship, we choose a 20-1 feedforward backpropagation neural network [27,28] to account this hidden and implicit relationship after large workings on model selection in <xref ref-type="fig" rid="fig1">Figure 1</xref>.</p></sec><sec id="s2_4"><title>2.4. Validation of Predictions</title><p>The second column in <xref ref-type="table" rid="table3">Table 3</xref> lists all 44 β-glucosidases obtained from the databank, of which 30 were used to generate the model parameters, weights and biases in neural network as the training group, and 14 were used to validate the neural network with generated weights and biases as the validation group. This is a very traditional approach for validation in neural network.</p><p>The second approach for validation is the delete-1 observation jackknife, each time we use 43 β-glucosidases as the training group to generate model parameters, and then to validate the prediction in omitted β-glucosidase until all 44 β-glucosidases undergo the same procedure. It is said to be the most effective approach for validation [<xref ref-type="bibr" rid="scirp.93817-ref29">29</xref>] although it is labor-intensive and time-demanding.</p><p>The third approach for validation is the cross-validation, through which 44 β-glucosidases were split into 11 subsets containing 4 cases each or 4 subsets containing 11 cases each. Each time, ten or three subsets were used to generate the model parameters, and one subset was used for validation, such a procedure was conducted in turn until each subset has served for validation [<xref ref-type="bibr" rid="scirp.93817-ref29">29</xref>].</p></sec><sec id="s2_5"><title>2.5. Statistics</title><p>For each predictor, we generated 100 sets of model parameters in order that the predictions based on 100 sets of model parameters to have well normally distributed mean&#177;SD to compare with the documented pH optimum of activity for each β-glucosidase [<xref ref-type="bibr" rid="scirp.93817-ref30">30</xref>]. For the data with normal distribution, the Student’s t-test was used, and for the data with abnormal distribution, the non-parametric Mann-Whitney U-test was used. P &lt; 0.05 is considered statistically significant. For visual comparison, linear regression was also used to evaluate the predicted pH values with their documented ones.</p></sec></sec><sec id="s3"><title>3. RESULTS AND DISCUSSION</title><p>It is highly likely that the relationship between feature of amino acids and pH optimum of activity is at least one step beyond the structure-function relationship, which might imply that we need at least two layers in the neural network to account this hidden and implicit relationship (<xref ref-type="fig" rid="fig1">Figure 1</xref>). Technically, the development of predictive method includes (1) selection of predictors, and (2) selection of predictive models, while the general and efficient practice is to select predictors at first, and then to select predictive models.</p><p>Working with this network model, the next consideration is the training process, which once again guarantees a fair selection of predictors. Technically, both initialization of weights and biases and number of training epochs govern whether the neural network can converge. We used the random initialization</p><table-wrap id="table1" ><label><xref ref-type="table" rid="table2">Table 2</xref></label><caption><title> Inductive effect scale, amino-acid number and distribution probability in β-glucosidase A9UIG0 and Q4U4W7. The amino-acid distribution probability, is computed according to the equation, r!/(q<sub>0</sub>! &#215; q<sub>1</sub>! &#215; ... &#215; q<sub>n</sub>!) &#215; r!/(r<sub>1</sub>! &#215; r<sub>2</sub>! &#215; ... &#215; r<sub>n</sub>!) &#215; n<sup>−r</sup>, where! is the factorial function, r is the number of a type of amino acid, q is the number of partitions with the same number of amino acids and n is the number of partitions in the protein for a type of amino acid. The computation can be found in the web site (http://www.gxas.cn/dp.htm)</title></caption><table><tbody><thead><tr><th align="center" valign="middle"  rowspan="2"  >Amino Acid</th><th align="center" valign="middle"  colspan="2"  >Inductive effect scale</th><th align="center" valign="middle"  colspan="2"  >Amino-acid number</th><th align="center" valign="middle"  colspan="2"  >Distribution probability</th></tr></thead><tr><td align="center" valign="middle" >A9UIG0</td><td align="center" valign="middle" >Q4U4W7</td><td align="center" valign="middle" >A9UIG0</td><td align="center" valign="middle" >Q4U4W7</td><td align="center" valign="middle" >A9UIG0</td><td align="center" valign="middle" >Q4U4W7</td></tr><tr><td align="center" valign="middle" >A</td><td align="center" valign="middle" >0.05</td><td align="center" valign="middle" >0.05</td><td align="center" valign="middle" >82</td><td align="center" valign="middle" >74</td><td align="center" valign="middle" >0.0021</td><td align="center" valign="middle" >0.0015</td></tr><tr><td align="center" valign="middle" >R</td><td align="center" valign="middle" >−0.26</td><td align="center" valign="middle" >−0.26</td><td align="center" valign="middle" >29</td><td align="center" valign="middle" >36</td><td align="center" valign="middle" >0.0043</td><td align="center" valign="middle" >0.0012</td></tr><tr><td align="center" valign="middle" >N</td><td align="center" valign="middle" >−0.14</td><td align="center" valign="middle" >−0.14</td><td align="center" valign="middle" >56</td><td align="center" valign="middle" >46</td><td align="center" valign="middle" >0.0202</td><td align="center" valign="middle" >0.0174</td></tr><tr><td align="center" valign="middle" >D</td><td align="center" valign="middle" >0.51</td><td align="center" valign="middle" >0.51</td><td align="center" valign="middle" >50</td><td align="center" valign="middle" >50</td><td align="center" valign="middle" >0.0039</td><td align="center" valign="middle" >0.0180</td></tr><tr><td align="center" valign="middle" >C</td><td align="center" valign="middle" >−0.01</td><td align="center" valign="middle" >−0.01</td><td align="center" valign="middle" >8</td><td align="center" valign="middle" >8</td><td align="center" valign="middle" >0.1682</td><td align="center" valign="middle" >0.0841</td></tr><tr><td align="center" valign="middle" >E</td><td align="center" valign="middle" >0.68</td><td align="center" valign="middle" >0.68</td><td align="center" valign="middle" >35</td><td align="center" valign="middle" >36</td><td align="center" valign="middle" >0.0218</td><td align="center" valign="middle" >0.0224</td></tr><tr><td align="center" valign="middle" >Q</td><td align="center" valign="middle" >−0.10</td><td align="center" valign="middle" >−0.10</td><td align="center" valign="middle" >28</td><td align="center" valign="middle" >25</td><td align="center" valign="middle" >0.0642</td><td align="center" valign="middle" >0.0051</td></tr><tr><td align="center" valign="middle" >G</td><td align="center" valign="middle" >0.00</td><td align="center" valign="middle" >0.00</td><td align="center" valign="middle" >94</td><td align="center" valign="middle" >86</td><td align="center" valign="middle" >0.0006</td><td align="center" valign="middle" >0.0027</td></tr><tr><td align="center" valign="middle" >H</td><td align="center" valign="middle" >−0.01</td><td align="center" valign="middle" >−0.01</td><td align="center" valign="middle" >16</td><td align="center" valign="middle" >11</td><td align="center" valign="middle" >0.0715</td><td align="center" valign="middle" >0.0808</td></tr><tr><td align="center" valign="middle" >I</td><td align="center" valign="middle" >0.06</td><td align="center" valign="middle" >0.06</td><td align="center" valign="middle" >35</td><td align="center" valign="middle" >43</td><td align="center" valign="middle" >0.0194</td><td align="center" valign="middle" >0.0240</td></tr><tr><td align="center" valign="middle" >L</td><td align="center" valign="middle" >0.02</td><td align="center" valign="middle" >0.02</td><td align="center" valign="middle" >58</td><td align="center" valign="middle" >65</td><td align="center" valign="middle" >0.0002</td><td align="center" valign="middle" >0.0054</td></tr><tr><td align="center" valign="middle" >K</td><td align="center" valign="middle" >−0.16</td><td align="center" valign="middle" >−0.16</td><td align="center" valign="middle" >29</td><td align="center" valign="middle" >29</td><td align="center" valign="middle" >0.0317</td><td align="center" valign="middle" >0.0317</td></tr><tr><td align="center" valign="middle" >M</td><td align="center" valign="middle" >0.08</td><td align="center" valign="middle" >0.08</td><td align="center" valign="middle" >11</td><td align="center" valign="middle" >19</td><td align="center" valign="middle" >0.0404</td><td align="center" valign="middle" >0.0138</td></tr><tr><td align="center" valign="middle" >F</td><td align="center" valign="middle" >0.04</td><td align="center" valign="middle" >0.04</td><td align="center" valign="middle" >34</td><td align="center" valign="middle" >33</td><td align="center" valign="middle" >0.0285</td><td align="center" valign="middle" >0.0193</td></tr><tr><td align="center" valign="middle" >P</td><td align="center" valign="middle" >0.00</td><td align="center" valign="middle" >0.00</td><td align="center" valign="middle" >53</td><td align="center" valign="middle" >65</td><td align="center" valign="middle" >0.0058</td><td align="center" valign="middle" >0.0010</td></tr><tr><td align="center" valign="middle" >S</td><td align="center" valign="middle" >−0.03</td><td align="center" valign="middle" >−0.03</td><td align="center" valign="middle" >62</td><td align="center" valign="middle" >61</td><td align="center" valign="middle" >0.0029</td><td align="center" valign="middle" >0.0112</td></tr><tr><td align="center" valign="middle" >T</td><td align="center" valign="middle" >−0.05</td><td align="center" valign="middle" >−0.05</td><td align="center" valign="middle" >57</td><td align="center" valign="middle" >56</td><td align="center" valign="middle" >0.0023</td><td align="center" valign="middle" >0.0005</td></tr><tr><td align="center" valign="middle" >W</td><td align="center" valign="middle" >0.06</td><td align="center" valign="middle" >0.06</td><td align="center" valign="middle" >18</td><td align="center" valign="middle" >16</td><td align="center" valign="middle" >0.0023</td><td align="center" valign="middle" >0.1362</td></tr><tr><td align="center" valign="middle" >Y</td><td align="center" valign="middle" >0.05</td><td align="center" valign="middle" >0.05</td><td align="center" valign="middle" >41</td><td align="center" valign="middle" >33</td><td align="center" valign="middle" >0.0142</td><td align="center" valign="middle" >0.0174</td></tr><tr><td align="center" valign="middle" >V</td><td align="center" valign="middle" >0.01</td><td align="center" valign="middle" >0.01</td><td align="center" valign="middle" >70</td><td align="center" valign="middle" >74</td><td align="center" valign="middle" >0.0067</td><td align="center" valign="middle" >0.0008</td></tr></tbody></table></table-wrap><p>function to initialize weights and biases, and 250 training epochs for convergence. <xref ref-type="fig" rid="fig2">Figure 2</xref> displays the performance of convergence in the training group, where each line represents a training process with random initialization of weights and biases running 250 training epochs. Different predictor has different profiles of its convergence. As seen, the convergence of 11 predictors can be reached within 250 training epochs with any random initialization, which guarantees our training process. However, the convergence of other predictors is not possible after 100 epochs, as the amino-acid composition shown in the top-left panel of <xref ref-type="fig" rid="fig2">Figure 2</xref>.</p><p><xref ref-type="fig" rid="fig3">Figure 3</xref> shows that the percentage of correctly predicted pH optimum ranges from about 70% to 90% with respect to different features. Actually, <xref ref-type="fig" rid="fig3">Figure 3</xref> gives us a basic concept on which predictor has better effect on predicting pH optimum of activity. Accordingly, the amino-acid distribution probability is the better one than others. This is because the amino acid features except for amino-acid distribution probability are not subject to their positions in amino acid sequence and their neighboring amino acids, whereas the amino-acid distribution probability is sensitive to these conditions.</p><p>Proteins evolve to function in specific cellular environment; thus pH of activity is subject to evolutionary pressure [<xref ref-type="bibr" rid="scirp.93817-ref5">5</xref>]. If the pH level could change the conformation of β-glucosidase; then different pH levels could have slight difference in conformation of β-glucosidase. In this context, the explanation for the results in <xref ref-type="fig" rid="fig3">Figure 3</xref> would be that the amino-acid distribution probability as it reflects the randomness in enzyme would more accurately reflect the changes in the conformation of β-glucosidase due to different pH levels, while the other predictors due to the fact that they have constant values, for example, physicochemical property, would less accurately reflect the changes in conformation of β-glucosidase.</p><p><xref ref-type="table" rid="table3">Table 3</xref> shows the comparison between documented and predicted pH optimum for each β-glucosidase found in the database [<xref ref-type="bibr" rid="scirp.93817-ref9">9</xref>]. We should consider a predictor workable if there is no statistical difference between documented and predicted pH optimum, and the last row of <xref ref-type="table" rid="table3">Table 3</xref> shows the overall performance, where we can see that the amino-acid distribution probability gives better predictive results than other predictors, whose results are similar.</p><p>In <xref ref-type="fig" rid="fig4">Figure 4</xref>, we used the regression between documented and predicted pH optimum of activity to visualize the predictive performance using the amino-acid distribution probability as predictor in order to confirm our observation visually.</p><p>To furthermore validate the above findings, we used the delete-1 jackknife validation and 3-fold, 10-fold cross-validation to treat these predictors as shown in <xref ref-type="fig" rid="fig5">Figure 5</xref>, where we once again find that the</p><p>best predictor is the amino-acid distribution probability.</p><p>The predictors used in this study include some related to the amino-acid based secondary structure of β-glucosidase; however, these features do not render better predictions than others. This may open the possibility to use the amino-acid features to predict various working conditions for enzymes, even the possibility to use the information about primary structure to predict the changing environments. This is so because other studies have confirmed that a physicochemical metric of charge distribution correlates better with subcellular pH [<xref ref-type="bibr" rid="scirp.93817-ref6">6</xref>]. The amino-acid composition is one of two factors that influence the pH of maximal protein stability [<xref ref-type="bibr" rid="scirp.93817-ref7">7</xref>] and can empirically model the pH optimum of protein-protein binding [<xref ref-type="bibr" rid="scirp.93817-ref8">8</xref>].</p><p>If we pay our attention to validation, the technique detail would puzzle us, that is, the Jackknife validation is worse than traditional validation by comparing <xref ref-type="fig" rid="fig3">Figure 3</xref> with <xref ref-type="fig" rid="fig5">Figure 5</xref>. This is interesting because the Jackknife uses almost all the samples to generate model parameters, but produces a worse prediction. This is very counter-intuitive, because the current knowledge indicates that the larger the trained data, the larger the chance that predicted sample would be included, the better the prediction. Clearly, this technique should require many more studies to deal with.</p><p>Currently it is not very clear whether different pH optimums would suggest that β-glucosidases would have different structures. If so, our prediction still falls into the so-called structure-function relationship; if not, our prediction would suggest a more sophisticated mechanism between structure and enzymatic working condition, i.e. structure-environment relationship.</p><p>In conclusion, this study suggests that we can use the features of amino acids of β-glucosidases to predict their pH optimum of activity. Among 25 amino-acid features screened, 11 can be serve as predictors to estimate the pH optimum, including 6/7 of the electronic properties and 4/7 of the amino-acid based secondary structure predictions, and the amino-acid distribution probability reaches better prediction than other predictors. However, the amino-acid composition, electronic properties and secondary structure predictions themselves cannot work in the neural network model, and they must be weighted with the amino-acid composition. Thus, the amino-acid distribution probability reveals its advantage in the prediction, indicating the random mechanisms may underline in enzyme reaction. The model provides the possibility to use the amino-acid features to predict various working conditions for enzymes.</p></sec><sec id="s4"><title>FUND</title><p>This study was supported by National Natural Science Foundation of China (31560315), and Key Project of Guangxi Scientific Research and Technology Development Plan (AB17190534).</p></sec><sec id="s5"><title>CONFLICTS OF INTEREST</title><p>The authors declare no conflicts of interest regarding the publication of this paper.</p></sec></body><back><ref-list><title>References</title><ref id="scirp.93817-ref1"><label>1</label><mixed-citation publication-type="other" xlink:type="simple">Wang, T., Wu, M.B., Lin, J.P. and Yang, L.R. (2015) Quantitative Structure-Activity Relationship: Promising Advances in Drug Discovery Platforms. Expert Opinion on Drug Discovery, 10, 1283-1300.  
https://doi.org/10.1517/17460441.2015.1083006</mixed-citation></ref><ref id="scirp.93817-ref2"><label>2</label><mixed-citation publication-type="other" xlink:type="simple">Xue, L.C., Dobbs, D., Bonvin, A.M. and Honavar, V. (2015) Computational Prediction of Protein Interfaces: A Review of Data Driven Methods. FEBS Letters, 589, 3516-3526. https://doi.org/10.1016/j.febslet.2015.10.003</mixed-citation></ref><ref id="scirp.93817-ref3"><label>3</label><mixed-citation publication-type="other" xlink:type="simple">Jeng, W.Y., Wang, N.C., Lin, M.H., Lin, C.T., Liaw, Y.C., Chang, W.J., Liu, C.I., Liang, P.H. and Wang, A.H.J. (2011) Structural and Functional Analysis of Three Beta-Glucosidases from Bacterium Clostridium cellulovorans, fungus Trichoderma reesei and Termite Neotermes koshunensis. Journal of Structural Biology, 173, 46-56.  
https://doi.org/10.1016/j.jsb.2010.07.008</mixed-citation></ref><ref id="scirp.93817-ref4"><label>4</label><mixed-citation publication-type="other" xlink:type="simple">Sticklen, M. (2006) Plant Genetic Engineering to Improve Biomass Characteristics for Biofuels. Current Opinion in Biotechnology, 17, 315-319. https://doi.org/10.1016/j.copbio.2006.05.003</mixed-citation></ref><ref id="scirp.93817-ref5"><label>5</label><mixed-citation publication-type="other" xlink:type="simple">Talley, K. and Alexov, E. (2010) On the pH-Optimum of Activity and Stability of Proteins. Proteins: Structure, Function, and Bioinformatics, 78, 2699-706. https://doi.org/10.1002/prot.22786</mixed-citation></ref><ref id="scirp.93817-ref6"><label>6</label><mixed-citation publication-type="other" xlink:type="simple">Garcia-Moreno, B. (2009) Adaptations of Proteins to Cellular and Subcellular pH. Journal of Biology, 8, 98. 
https://doi.org/10.1186/jbiol199</mixed-citation></ref><ref id="scirp.93817-ref7"><label>7</label><mixed-citation publication-type="other" xlink:type="simple">Alexov, E. (2004) Numerical Calculations of the pH of Maximal Protein Stability. The Effect of the Sequence Composition and Three-Dimensional Structure. European Journal of Biochemistry, 271, 173-185. 
https://doi.org/10.1046/j.1432-1033.2003.03917.x</mixed-citation></ref><ref id="scirp.93817-ref8"><label>8</label><mixed-citation publication-type="other" xlink:type="simple">Mitra, R.C., Zhang, Z. and Alexov, E. (2011) In Silico Modeling of pH-Optimum of Protein-Protein Binding. Proteins: Structure, Function, and Bioinformatics, 79, 925-936. https://doi.org/10.1002/prot.22931</mixed-citation></ref><ref id="scirp.93817-ref9"><label>9</label><mixed-citation publication-type="other" xlink:type="simple">Placzek, S., Schomburg, I., Chang, A., Jeske, L., Ulbrich, M., Tillack, J. and Schomburg, D. (2017) BRENDA in 2017: New Perspectives and New Tools in BRENDA. Nucleic Acids Research, 45, D380-D388.  
https://doi.org/10.1093/nar/gkw952</mixed-citation></ref><ref id="scirp.93817-ref10"><label>10</label><mixed-citation publication-type="other" xlink:type="simple">Berrin, J.G., Czjzek, M., Kroon, P.A., McLauchlan, W.R., Puigserver, A., Williamson, G. and Juge, N. (2003) Substrate (Aglycone) Specificity of Human Cytosolic Beta-Glucosidase. Biochemical Journal, 373, 41-48. 
https://doi.org/10.1042/bj20021876</mixed-citation></ref><ref id="scirp.93817-ref11"><label>11</label><mixed-citation publication-type="other" xlink:type="simple">Tsukada, T., Igarashi, K., Fushinobu, S. and Samejima, M. (2008) Role of Subsite +1 Residues in pH Dependence and Catalytic Activity of the Glycoside Hydrolase Family 1 β-Glucosidase BGL1A from the Basidiomycete Phanerochaete chrysosporium. Biotechnology and Bioengineering, 99, 1295-1302. 
https://doi.org/10.1002/bit.21717</mixed-citation></ref><ref id="scirp.93817-ref12"><label>12</label><mixed-citation publication-type="other" xlink:type="simple">UniProt Consortium (2019) UniProt: A Worldwide Hub of Protein Knowledge. Nucleic Acids Research, 47, D506-D515. https://doi.org/10.1093/nar/gky1049</mixed-citation></ref><ref id="scirp.93817-ref13"><label>13</label><mixed-citation publication-type="other" xlink:type="simple">Burlingame, A.L. and Carr, S.A. (1996) Mass Spectrometry in the Biological Sciences. Humana Press, Totowa, NJ. https://doi.org/10.1007/978-1-4612-0229-5</mixed-citation></ref><ref id="scirp.93817-ref14"><label>14</label><mixed-citation publication-type="other" xlink:type="simple">Zamyatin, A.A. (1972) Protein Volume in Solution. Progress in Biophysics &amp; Molecular Biology, 24, 107-123. 
https://doi.org/10.1016/0079-6107(72)90005-3</mixed-citation></ref><ref id="scirp.93817-ref15"><label>15</label><mixed-citation publication-type="other" xlink:type="simple">Darby, N.J. and Creighton, T.E. (1993) Dissecting the Disulphide-Coupled Folding Pathway of Bovine Pancreatic Trypsin Inhibitor. Forming the First Disulphide Bonds in Analogues of the Reduced Protein. Journal of Molecular Biology, 232, 873-896. https://doi.org/10.1006/jmbi.1993.1437</mixed-citation></ref><ref id="scirp.93817-ref16"><label>16</label><mixed-citation publication-type="other" xlink:type="simple">Kyte, J. and Doolittle, R.F. (1982) A Simple Method for Displaying the Hydropathic Character of a Protein. Journal of Molecular Biology, 157, 105-132. https://doi.org/10.1016/0022-2836(82)90515-0</mixed-citation></ref><ref id="scirp.93817-ref17"><label>17</label><mixed-citation publication-type="other" xlink:type="simple">Trinquier, G., Sanejouand, Y.H. and Hausman, R.E. (1998) Which Effective Property of Amino Acids Is Best Preserved by the Genetic Code? Protein Engineering, Design and Selection, 11, 153-169.  
https://doi.org/10.1093/protein/11.3.153</mixed-citation></ref><ref id="scirp.93817-ref18"><label>18</label><mixed-citation publication-type="other" xlink:type="simple">Cooper, G.M. (2004) The Cell: A Molecular Approach. ASM Press, Washington DC, 51.</mixed-citation></ref><ref id="scirp.93817-ref19"><label>19</label><mixed-citation publication-type="other" xlink:type="simple">Dwyer, D.S. (2005) Electronic Properties of Amino Acid Side Chains: Quantum Mechanics Calculation of Substituent Effects. BMC Chemical Biology, 5, 2. https://doi.org/10.1186/1472-6769-5-2</mixed-citation></ref><ref id="scirp.93817-ref20"><label>20</label><mixed-citation publication-type="other" xlink:type="simple">Chou, P.Y. and Fasman, G.D. (1978) Prediction of Secondary Structure of Proteins from Amino Acid Sequence. Advances in Enzymology and Related Subjects of Biochemistry, 47, 45-148.  
https://doi.org/10.1002/9780470122921.ch2</mixed-citation></ref><ref id="scirp.93817-ref21"><label>21</label><mixed-citation publication-type="other" xlink:type="simple">Feller, W. (1968) An Introduction to Probability Theory and Its Applications. 3rd Edition, Wiley, New York.</mixed-citation></ref><ref id="scirp.93817-ref22"><label>22</label><mixed-citation publication-type="other" xlink:type="simple">Wu, G. and Yan, S.M. (2002) Randomness in the Primary Structure of Protein: Methods and Implications. Molecular Biology Today, 3, 55-69.</mixed-citation></ref><ref id="scirp.93817-ref23"><label>23</label><mixed-citation publication-type="other" xlink:type="simple">Wu, G. and Yan, S. (2006) Mutation Trend of Hemagglutinin of Influenza a Virus: A Review from Computational Mutation Viewpoint. Acta Pharmacologia Sinica, 27, 513-526.  
https://doi.org/10.1111/j.1745-7254.2006.00329.x</mixed-citation></ref><ref id="scirp.93817-ref24"><label>24</label><mixed-citation publication-type="other" xlink:type="simple">Wu, G. and Yan, S. (2006) Fate of Influenza a Virus Proteins. Protein &amp; Peptide Letters, 13, 377-384. 
https://doi.org/10.2174/092986606775974474</mixed-citation></ref><ref id="scirp.93817-ref25"><label>25</label><mixed-citation publication-type="other" xlink:type="simple">Yan, S. and Wu, G. (2010) Creation and Application of Computational Mutation. Journal of Guangxi Academy of Sciences, 17, 145-150.</mixed-citation></ref><ref id="scirp.93817-ref26"><label>26</label><mixed-citation publication-type="other" xlink:type="simple">Wu, G. and Yan, S. (2008) Lecture Notes on Computational Mutation. Nova Science Publishers, New York.</mixed-citation></ref><ref id="scirp.93817-ref27"><label>27</label><mixed-citation publication-type="other" xlink:type="simple">Demuth, H. and Beale, M. (2001) Neural Network Toolbox for Use with MatLab. User’s Guide, Version 4.</mixed-citation></ref><ref id="scirp.93817-ref28"><label>28</label><mixed-citation publication-type="other" xlink:type="simple">MathWorks Inc (1984-2001) MatLab—The Language of Technical Computing.</mixed-citation></ref><ref id="scirp.93817-ref29"><label>29</label><mixed-citation publication-type="other" xlink:type="simple">Chou, K.C. (2011) Some Remarks on Protein Attribute Prediction and Pseudo Amino Acid Composition. Journal of Theoretical Biology, 273, 236-247. https://doi.org/10.1016/j.jtbi.2010.12.024</mixed-citation></ref><ref id="scirp.93817-ref30"><label>30</label><mixed-citation publication-type="other" xlink:type="simple">Sokal, R.R. and Rohlf, F.J. (1995) Biometry: The Principles and Practices of Statistics in Biological Research. 3rd Edition, W. H. Freeman, New York, 203-218.</mixed-citation></ref></ref-list></back></article>