<?xml version="1.0" encoding="UTF-8"?><!DOCTYPE article  PUBLIC "-//NLM//DTD Journal Publishing DTD v3.0 20080202//EN" "http://dtd.nlm.nih.gov/publishing/3.0/journalpublishing3.dtd"><article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" dtd-version="3.0" xml:lang="en" article-type="research article"><front><journal-meta><journal-id journal-id-type="publisher-id">JDAIP</journal-id><journal-title-group><journal-title>Journal of Data Analysis and Information Processing</journal-title></journal-title-group><issn pub-type="epub">2327-7211</issn><publisher><publisher-name>Scientific Research Publishing</publisher-name></publisher></journal-meta><article-meta><article-id pub-id-type="doi">10.4236/jdaip.2024.122012</article-id><article-id pub-id-type="publisher-id">JDAIP-133191</article-id><article-categories><subj-group subj-group-type="heading"><subject>Articles</subject></subj-group><subj-group subj-group-type="Discipline-v2"><subject>Computer Science&amp;Communications</subject><subject> Physics&amp;Mathematics</subject></subj-group></article-categories><title-group><article-title>
 
 
  High Dimension Multivariate Data Analysis for Small Group Samples of Chemical Volatile Profiles of African Nightshade Species
 
</article-title></title-group><contrib-group><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Lorna</surname><given-names>Chepkemoi</given-names></name><xref ref-type="aff" rid="aff1"><sup>1</sup></xref></contrib><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Daisy</surname><given-names>Salifu</given-names></name><xref ref-type="aff" rid="aff1"><sup>1</sup></xref></contrib><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Lucy</surname><given-names>Kananu Murungi</given-names></name><xref ref-type="aff" rid="aff2"><sup>2</sup></xref></contrib><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Henri</surname><given-names>E. Z. Tonnang</given-names></name><xref ref-type="aff" rid="aff1"><sup>1</sup></xref></contrib></contrib-group><aff id="aff1"><addr-line>International Centre of Insect Physiology and Ecology (ICIPE), Nairobi, Kenya</addr-line></aff><aff id="aff2"><addr-line>Department of Horticulture and Food Security, Jomo Kenyatta University of Agriculture and Technology (JKUAT), Nairobi, Kenya</addr-line></aff><pub-date pub-type="epub"><day>12</day><month>04</month><year>2024</year></pub-date><volume>12</volume><issue>02</issue><fpage>210</fpage><lpage>231</lpage><history><date date-type="received"><day>6,</day>	<month>February</month>	<year>2024</year></date><date date-type="rev-recd"><day>14,</day>	<month>May</month>	<year>2024</year>	</date><date date-type="accepted"><day>17,</day>	<month>May</month>	<year>2024</year></date></history><permissions><copyright-statement>&#169; Copyright  2014 by authors and Scientific Research Publishing Inc. </copyright-statement><copyright-year>2014</copyright-year><license><license-p>This work is licensed under the Creative Commons Attribution International License (CC BY). http://creativecommons.org/licenses/by/4.0/</license-p></license></permissions><abstract><p>
 
 
  Quantitative headspace analysis of volatiles emitted by plants or any other living organisms in chemical ecology studies generates large multidimensional data that require extensive mining and refining to extract useful information. More often the number of variables and the quantified volatile compounds exceed the number of observations or samples and hence many traditional statistical analysis methods become inefficient. Here, we employed machine learning algorithm, random forest (RF) in combination with distance-based procedure, similarity percentage (SIMPER) as preprocessing steps to reduce the data dimensionality in the chemical profiles of volatiles from three African nightshade plant species before subjecting the data to non-metric multidimensional scaling (NMDS). In addition, non-parametric methods namely permutational multivariate analysis of variance (PERMANOVA) and analysis of similarities (ANOSIM) were applied to test hypothesis of differences among the African nightshade species based on the volatiles profiles and ascertain the patterns revealed by NMDS plots. Our results revealed that there were significant differences among the African nightshade species when the data&amp;#8217;s dimension was reduced using RF variable importance and SIMPER, as also supported by NMDS plots that showed &lt;i&gt;S. &lt;/i&gt;&lt;i&gt;scabr&lt;/i&gt;&lt;i&gt;um&lt;/i&gt; being separated from &lt;i&gt;S. &lt;/i&gt;&lt;i&gt;villosum&lt;/i&gt; and &lt;i&gt;S. &lt;/i&gt;&lt;i&gt;sarrachoides&lt;/i&gt; based on the reduced data variables. The novelty of our work is on the merits of using data reduction techniques to successfully reveal differences in groups which could have otherwise not been the case if the analysis were performed on the entire original data matrix characterized by small samples. The R code used in the analysis has been shared herein for interested researchers to customise it for their own data of similar nature.
 
</p></abstract><kwd-group><kwd>Random Forest</kwd><kwd> Similarity Percentage</kwd><kwd> PERMANOVA</kwd><kwd> ANOSIM</kwd><kwd> Non-Metric Multi-Dimensional Scaling</kwd></kwd-group></article-meta></front><body><sec id="s1"><title>1. Introduction</title><p>Quantification of plant volatiles or any volatiles emitted by living organisms gives rise to high multidimensional data. Such studies often are set to compare volatile profiles of groups of organisms. The fundamental objective in such studies is to identify compounds that discriminate between groups. In plants, such studies are useful to understand host plant-insect pest interaction, as particular plant species could be host or non-host of insect pest [<xref ref-type="bibr" rid="scirp.133191-ref1">1</xref>] . These volatiles can be emitted from flowers, leaves, fruits, roots or any other part of the plant into the atmosphere or soil, allowing the plant to interact with other organisms. There has been extensive investigation on the significance of volatiles in plant physiology and ecology and their roles in mutualistic interaction with other organisms [<xref ref-type="bibr" rid="scirp.133191-ref2">2</xref>] . For instance, pollinators are attracted by volatiles emitted from floral tissues [<xref ref-type="bibr" rid="scirp.133191-ref1">1</xref>] and conversely play a crucial role in host finding by insect pests in agro-ecosystems [<xref ref-type="bibr" rid="scirp.133191-ref3">3</xref>] [<xref ref-type="bibr" rid="scirp.133191-ref4">4</xref>] .</p><p>Other studies have demonstrated the beneficial effect of herbivore-induced plant volatile compounds (HIVOCs) as host location signals for parasitoids and herbivore predators [<xref ref-type="bibr" rid="scirp.133191-ref2">2</xref>] [<xref ref-type="bibr" rid="scirp.133191-ref5">5</xref>] [<xref ref-type="bibr" rid="scirp.133191-ref6">6</xref>] . This indirect chemical defense is most likely as significant as direct chemical and physical defenses in inhibiting herbivore damage [<xref ref-type="bibr" rid="scirp.133191-ref5">5</xref>] . Furthermore, volatile signals emitted by injured plants may transmit a signal to surrounding plants, stimulating defensive responses [<xref ref-type="bibr" rid="scirp.133191-ref7">7</xref>] . Plants may release volatiles in response to changes in light, temperature, or other abiotic stressors [<xref ref-type="bibr" rid="scirp.133191-ref2">2</xref>] . Research on plant volatiles has therefore provided insights in understanding variations in plant species with regard to evolutionary origins and ecological consequences in terms of plant-insect interactions and functional responses.</p><p>High dimensional multivariate data obtained from chemical volatile analysis has usually been analyzed using usual linear methods such as principal component analysis (PCA) [<xref ref-type="bibr" rid="scirp.133191-ref8">8</xref>] [<xref ref-type="bibr" rid="scirp.133191-ref9">9</xref>] , linear discriminant analysis (LDA) [<xref ref-type="bibr" rid="scirp.133191-ref10">10</xref>] [<xref ref-type="bibr" rid="scirp.133191-ref11">11</xref>] , multivariate analysis of variance and other methods. In some cases, PCA has been used as a preprocess to reduce dimensionality of data before applying LDA on the principal components [<xref ref-type="bibr" rid="scirp.133191-ref12">12</xref>] . Principal component analysis and linear discriminant analysis are famous feature extraction methods that are subject to small sample sizes. In fact, the effect of small sample sizes for high dimensional data has been discussed by several authors [<xref ref-type="bibr" rid="scirp.133191-ref13">13</xref>] [<xref ref-type="bibr" rid="scirp.133191-ref14">14</xref>] [<xref ref-type="bibr" rid="scirp.133191-ref15">15</xref>] . Generally, sample size n must be greater than the number of variables or features, p. Small sample sizes with PCA tend to provide eigenvectors coefficients (also known as factor loadings) and eigenvalues that are unprecise estimate of population values while large sample sizes provide precise estimates [<xref ref-type="bibr" rid="scirp.133191-ref13">13</xref>] [<xref ref-type="bibr" rid="scirp.133191-ref16">16</xref>] . The effect of small sample size is even stronger for linear discriminant methods such as LDA which loses performance when faced with small sample sizes in high dimensional variable space [<xref ref-type="bibr" rid="scirp.133191-ref17">17</xref>] . The effect of small sample size on classical statistical methods is further aggravated by the methods’ reliance on model assumptions that can hardly be verified when sample sizes are small. This favors use of statistical methods having their model assumptions relaxed for the analysis of such small samples. Consequently, this has led to increased application of non-metric ordination techniques, especially in revealing patterns and producing meaningful results that are easy to interpret in multivariate data [<xref ref-type="bibr" rid="scirp.133191-ref18">18</xref>] . Specifically, non-metric multidimensional scaling (NMDS) has gained immense application in ecological data due to its ability to handle non-linear data [<xref ref-type="bibr" rid="scirp.133191-ref19">19</xref>] [<xref ref-type="bibr" rid="scirp.133191-ref20">20</xref>] .</p><p>Data on chemical profile of volatiles is characterized by small samples n and large number of variables p, thus often p &gt; n. The crucial problem then is the presence of variables not significantly contributing to discrimination of samples but capable of contributing to random noise which potentially could obscure differences in groups. The use of variable selection methods to reduce dimensionality can lead to improvements in discrimination of samples and enhanced visualization. Therefore, in this study we use random forest (RF) technique and similarity percentage (SIMPER) as variable reduction techniques on chemical profiles of volatile compounds of three African nightshade plants prior to application of NMDS and hypothesis testing using permutational multivariate analysis of variance (PERMANOVA) and analysis of similarities (ANOSIM).</p><p>The random forest (RF) technique is an ensemble classifier that generates several decision trees from a sample obtained from the original dataset [<xref ref-type="bibr" rid="scirp.133191-ref21">21</xref>] . Each decision tree uses a different bootstrap sample in building the tree by randomly selecting with replacement a sample from the dataset. The bootstrap sample is then fed as input to base learners which are combined using a majority vote [<xref ref-type="bibr" rid="scirp.133191-ref22">22</xref>] . Since the decision trees are unrelated, the decision made as a majority vote is better than the decision made by each individual tree [<xref ref-type="bibr" rid="scirp.133191-ref23">23</xref>] . Machine learning (ML) models have been shown to be robust in handling small sample size compared to other well-established models such as linear discriminant analysis [<xref ref-type="bibr" rid="scirp.133191-ref24">24</xref>] . In particular, RF has shown superior performance over other ML models especially with high dimensional small-sample datasets (p &gt; n) [<xref ref-type="bibr" rid="scirp.133191-ref25">25</xref>] [<xref ref-type="bibr" rid="scirp.133191-ref26">26</xref>] [<xref ref-type="bibr" rid="scirp.133191-ref27">27</xref>] . Further, RF offers variable importance measures which are used to rank variables based on their predictive ability and we leverage on this aspect of RF to reduce the dimension of the data in this study prior to application of NMDS and hypothesis tests. Gini importance and permutation importance (mean decrease in accuracy) are the two commonly used variable importance measures [<xref ref-type="bibr" rid="scirp.133191-ref28">28</xref>] and are described in the methodology of this study.</p><p>Similarity percentage (SIMPER), proposed by Clarke [<xref ref-type="bibr" rid="scirp.133191-ref29">29</xref>] compares groups of sampling units pairwise and calculates each samples’ contribution to the average between-group Bray-Curtis dissimilarity [<xref ref-type="bibr" rid="scirp.133191-ref30">30</xref>] [<xref ref-type="bibr" rid="scirp.133191-ref31">31</xref>] . The existence of good discriminator variables in samples result in high quantitative presence yielding high average dissimilarity [<xref ref-type="bibr" rid="scirp.133191-ref30">30</xref>] . This allows identifying variables that significantly contribute to the dissimilarity between samples [<xref ref-type="bibr" rid="scirp.133191-ref32">32</xref>] .</p><p>The study aims to assess the similarity or dissimilarity of three African nightshade species (Solanum sarrachoides Sendtner, S. scabrum Miller and S. villosum Mill.) using NMDS with data preprocessed for variable reduction using RF and SIMPER. We demonstrate a step-by-step analysis using NMDS and show the merits of reducing the variables using RF and SIMPER as also confirmed by hypothesis tests using permutational multivariate analysis of variance (PERMANOVA) and analysis of similarities (ANOSIM). We demonstrate that using RF and SIMPER to reduce data dimension, enhances the visualization of the projection of the nightshade species on the volatile compounds and increases the power of hypothesis tests in PERMANOVA and ANOSIM. The R code used in the analysis is shared here for interested researchers to customise it for their own datasets of similar nature.</p></sec><sec id="s2"><title>2. Materials and Methods</title><sec id="s2_1"><title>2.1. Data</title><p>Our study used secondary data on amount of volatile organic compounds (VOCs) obtained from intact plants of three African nightshade species namely, S. sarrachoides, S. scabrum and S. villosum. The volatile chemical analyses were performed using gas chromatography-mass spectrometry (GC/MS) on three samples from each African nightshade species. A total of 58 volatile organic compounds were identified. A full description of the data and methodology is found in Murungi et al. [<xref ref-type="bibr" rid="scirp.133191-ref33">33</xref>] .</p></sec><sec id="s2_2"><title>2.2. Data Reduction Techniques</title><p>In this study, we take advantage of the fundamental outcome of RF to reduce the dimension of the 58 volatile organic compounds of the three African nightshade species prior to application of NMDS, PERMANOVA and ANOSIM.</p><p>Random forest technique generates several decision trees from a sample obtained from the original dataset. The parameters under consideration in the implementation of the RF algorithm are therefore, number of features for growing each tree (mtry) and number of trees to be generated (ntree). Here, ntree was fixed at default value 500 while mtry was evaluated by searching for the optimal mtry value using the tune function implemented in random Forest package in R. Approximately two-thirds of the samples (in-bag samples) are used to train the decision trees, with the remaining one-third (out-of-bag samples) used during an internal cross-validation procedure to estimate the performance of RF algorithm [<xref ref-type="bibr" rid="scirp.133191-ref21">21</xref>] [<xref ref-type="bibr" rid="scirp.133191-ref34">34</xref>] . There is no pruning of trees in RF as ensemble and bootstrapping schemes help it to overcome overfitting issues [<xref ref-type="bibr" rid="scirp.133191-ref25">25</xref>] . RF produces variable importance measures namely Gini importance and permutation importance (mean decrease in accuracy), which are used to rank variables based on their predictive ability. Gini importance has been disputed as having undesirable properties such as being biased in favor of variables with many categories while permutation importance has been proposed as a corrective measure to this biasness [<xref ref-type="bibr" rid="scirp.133191-ref35">35</xref>] . Therefore, variable importance was derived using permutation importance.</p><p>Additionally, SIMPER was used to identify volatile compounds that showed significant difference between the African nightshade species (α = 0.05) to augment the volatile compounds selected under the RF variable importance measure. SIMPER analysis works at the univariate level by computing the relative contribution of each variable to the overall average Bray-Curtis dissimilarities by pairwise comparison of groups.</p></sec><sec id="s2_3"><title>2.3. Non-Metric Multidimensional Scaling and How It Works</title><p>NMDS is a rank-based approach whose algorithm works by first randomly placing samples in an ordination space, with the desired number of dimensions defined a priori. The placement of samples is by an iterative process that attempts to find an ordination based on a stress function, in which ordinated sample distance closely match the order of sample dissimilarities in the original distance matrix [<xref ref-type="bibr" rid="scirp.133191-ref36">36</xref>] . This means that the original distance data is substituted with ranks. Samples are represented as points in a two or three-dimensional space such that the relative distances of all points are in the same rank order as the relative similarities of the samples [<xref ref-type="bibr" rid="scirp.133191-ref37">37</xref>] [<xref ref-type="bibr" rid="scirp.133191-ref38">38</xref>] . The mapping of samples using ranks preserves their ranked differences which enhances rescaling or rotation of axes for better visualization and interpretation [<xref ref-type="bibr" rid="scirp.133191-ref36">36</xref>] . Several iterations are implemented in the algorithm to obtain the lowest stress value possible, thus the stress function measures the goodness of fit of the distance adjustment in the reduced variable space configuration. Therefore, the lower the stress value, the better the data are represented in an ordination. The commonly utilized stress measure is Kruskal’s stress [<xref ref-type="bibr" rid="scirp.133191-ref39">39</xref>] defined as;</p><p>stress 1 2 = ∑ i j ( d u v − d ^ u v ) ∑ u v d u v 2 (1)</p><p>where d u v represents the actual distance between samples u and v in an ordination space;</p><p>d ^ u v represents the fitted distance between samples u and v.</p><p>Stress values are measured on a scale of 0 to 1 [<xref ref-type="bibr" rid="scirp.133191-ref39">39</xref>] , with a stress value of 0 indicating similarity between all rank order distances in the input data and final ordination. Stress values reduce with increasing NMDS dimensionality. Stress value less than 0.05 gives an excellent representation with no prospect of misinterpretation while stress values greater than 0.2 are likely to yield NMDS plots that are hard to interpret [<xref ref-type="bibr" rid="scirp.133191-ref29">29</xref>] . In this study, we take advantage of the fundamental outcome of RF to reduce the dimension of the 58 volatile organic compounds of the three African nightshade species prior to application of NMDS, PERMANOVA and ANOSIM.</p></sec><sec id="s2_4"><title>2.4. Distance Measures Used in Non-Metric Multidimensional Scaling</title><p>NMDS uses distance measures for ordination and some of these distances include Euclidean, Manhattan, Bray-Curtis, Kulczynski as described below.</p><sec id="s2_4_1"><title>2.4.1. Euclidean</title><p>Euclidean distance measures distance between two samples in multidimensional space and is calculated as the square root of the sum, over all the variables, of the square of the difference between values of a pair of samples [<xref ref-type="bibr" rid="scirp.133191-ref40">40</xref>] . It has no upper limit and is strongly affected by large number of zeros in the data, which often lead to high similarities between samples not sharing same variables. Moreover, it is a symmetrical index which treats double zeros in the same way as double presences resulting to shrinking of distance between two samples. To make the resulting Euclidean distances asymmetrical, the data is first transformed using either Chord, Hellinger or chi-square transformation [<xref ref-type="bibr" rid="scirp.133191-ref41">41</xref>] . The Euclidean distance measure is given as:</p><p>D = ∑ v = 1 k ( y 1 v − y 2 v ) 2 (2)</p><p>where D is the distance measure, k is the number of variables, and y 1 v and y 2 v are values of variable v in sample 1 and 2, respectively.</p></sec><sec id="s2_4_2"><title>2.4.2. Manhattan</title><p>Manhattan distance is obtained by computing the sum of the absolute differences between distances of a pair of samples [<xref ref-type="bibr" rid="scirp.133191-ref42">42</xref>] . It has the same properties as Euclidean distance and is majorly dominated by variables with large values. The distance measure is given as:</p><p>D = ∑ v = i k | ( y 1 v − y 2 v ) | (3)</p></sec><sec id="s2_4_3"><title>2.4.3. Bray-Curtis</title><p>Bray-Curtis is a modification of Manhattan distance measure where the sum of differences between samples across variables is standardized by the sum of variable values across samples, also summed across variables. Standardization was introduced to ensure that each variable is maximum-adjusted to equalize their contributions, and to relativize samples to reduce the effect of differing summed quantities. Bray-Curtis distance ranges between zero (completely similar variables) and one (completely dissimilar variables) [<xref ref-type="bibr" rid="scirp.133191-ref43">43</xref>] . The distance measure is given as:</p><p>D = ∑ v = 1 k | y 1 v − y 2 v | ∑ v = 1 k ( y 1 v + y 2 v ) (4)</p></sec><sec id="s2_4_4"><title>2.4.4. Kulczynski</title><p>The distance measure calculates dissimilarities between pairs of samples. It is calculated by summing variable minima and dividing this value by each sampling unit’s total. The distance between the two sampling units is one minus the average of these two values [<xref ref-type="bibr" rid="scirp.133191-ref44">44</xref>] . The distance measure is given as:</p><p>D = 1 − ( ∑ v = 1 k min ( y 1 v , y 2 v ) ∑ v = 1 k ( y 1 v ) + ∑ v = 1 k min ( y 1 v , y 2 v ) ∑ v = 1 k ( y 2 v ) ) 2 (5)</p></sec></sec><sec id="s2_5"><title>2.5. Test of Difference in Groups</title><p>To compare overall variation in volatile compounds composition between the African nightshade species, analysis of similarities (ANOSIM) and permutational multivariate analysis of variance (PERMANOVA) were used.</p><p>Permutational multivariate analysis of variance (PERMANOVA) is a semiparametric method which tests and estimates sizes of main effects or interaction terms while retaining important statistical properties of rank based non-parametric multivariate methods such as flexibility to base the analysis on a dissimilarity measure of choice and distribution-free inferences achieved by permutations, with no assumption of multivariate normality [<xref ref-type="bibr" rid="scirp.133191-ref45">45</xref>] . Pseudo F-ratio is used as a test statistic in PERMANOVA [<xref ref-type="bibr" rid="scirp.133191-ref44">44</xref>] and is given as:</p><p>F = SSB / ( β − 1 ) SSW / ( N − β ) (6)</p><p>where SSB is the sum of squared dissimilarities between groups; SSW is the sum of squared dissimilarities within groups; (β − 1) is the degrees of freedom associated with grouping variable; and (N − β) is the degrees of freedom associated with residuals.</p><p>The test statistic compares the total sum of squared ranked dissimilarities among samples in different groups to those belonging to the same group. The p-value is used to validate the significance of Pseudo F-ratio. On the other hand, ANOSIM is a hypothesis testing procedure that uses a dissimilarity measure to test for differences among groups. The null hypothesis being tested is that the average rank dissimilarities among samples within groups are the same as the average rank dissimilarities among samples from different groups. ANOSIM test statistic (R) is based on the rank differences between the average between-group ( r &#175; B ) and within-group ( r &#175; W ) given as:</p><p>R = r &#175; B − r &#175; W n ( n − 1 ) / 4 (7)</p><p>R is scaled within the range −1 to 1 with values greater than zero suggesting differences between groups, with more dissimilarity between groups than within groups. R values less than zero indicate more dissimilarities within groups than between groups, while R values of zero indicate that the dissimilarity within groups is the same as dissimilarity from different groups.</p><p>The workflow of our study in terms of methodology is summarized in <xref ref-type="fig" rid="fig1">Figure 1</xref>. The data under study is subjected to variable reduction techniques; thus, random forest and similarity percentage (SIMPER) prior to analysis by NMDS,</p><p>PERMANOVA and ANOSIM. The NMDS, PERMANOVA and ANOSIM are likewise performed on the entire dataset to compare out with that of reduced dimension.</p></sec><sec id="s2_6"><title>2.6. The Analysis</title><p>Random forest technique was used to reduce the 58 volatile organic compounds (VOCs) under study based on variable importance. The top 12 compounds were selected to be used in the discrimination of the three African nightshade species (<xref ref-type="fig" rid="fig2">Figure 2</xref>). On the other hand, SIMPER was also used to identify variables that showed significant difference between the African nightshade species at α = 0.05 level of significance and a total of 13 compounds were obtained. The outcomes of RF and SIMPER were combined to give a total of 16 “relevant” variables displayed in the Venn diagram (<xref ref-type="fig" rid="fig3">Figure 3</xref>). The 16 volatile compounds were then used in NMDS analysis. Bray Curtis distance which was determined as the suitable distance for these data was used to obtain pairwise similarity matrix, which determines the ecological distance between all pairs of nightshade species. Suitable k dimension for the NMDS plot was determined using scree plot, which is a plot of stress values versus number of dimensions. PERMANOVA and ANOSIM were performed to test for significant difference in volatile compound profiles of the African nightshade species. The NMDS, PERMANOVA and ANOSIM output on the reduced dataset (16 VOCs) was compared to NMDS, PERMANOVA and ANOSIM implemented on the entire dataset (58 volatile compounds).</p><p>All analyses were implemented in R version 4.1.3 [<xref ref-type="bibr" rid="scirp.133191-ref46">46</xref>] using the following packages; randomForest [<xref ref-type="bibr" rid="scirp.133191-ref47">47</xref>] and vip [<xref ref-type="bibr" rid="scirp.133191-ref48">48</xref>] for Random Forest variable importance analysis, vegan [<xref ref-type="bibr" rid="scirp.133191-ref49">49</xref>] for NMDS, PERMANOVA and ANOSIM; goeveg [<xref ref-type="bibr" rid="scirp.133191-ref50">50</xref>] , ggplot2 [<xref ref-type="bibr" rid="scirp.133191-ref51">51</xref>] and ggforce [<xref ref-type="bibr" rid="scirp.133191-ref52">52</xref>] for scree plot and NMDS plots. The R script for commands used in the study is available at https://github.com/icipe-official/non-metric-multidimensional-scaling/blob/main/r-code</p></sec></sec><sec id="s3"><title>3. Results</title><sec id="s3_1"><title>3.1. Similarities of VOCs in the Three Nightshade Species</title><p>The pairwise similarity matrix of VOCs in S. sarrachoides, S. scabrum and S. villosum indicated that all distances ranged between 0.210 to 0.902. This explains the variation in the VOCs emitted by the three African nightshade species (<xref ref-type="table" rid="table1"><xref ref-type="table" rid="table">Table </xref>1</xref>).</p></sec><sec id="s3_2"><title>3.2. Variables Selected</title><p>From the RF and SIMPER procedures, a total of 16 volatile organic compounds out of 58 were selected, nine of which were the same, while three were unique to RF and four unique to SIMPER according to our selection criterion. The 16 variables are displayed in <xref ref-type="fig" rid="fig3">Figure 3</xref>. RF variable importance results had c46 as the</p><table-wrap id="table1" ><label><xref ref-type="table" rid="table1"><xref ref-type="table" rid="table">Table </xref>1</xref></label><caption><title> Pairwise similarity matrix of volatile organic compounds concentration in three Af-rican nightshade species (S. sarrachoides, S. scabrum and S. villosum) based on Bray-Curtis distance</title></caption><table><tbody><thead><tr><th align="center" valign="middle" ></th><th align="center" valign="middle" ></th><th align="center" valign="middle"  colspan="9"  >African nightshade species</th></tr></thead><tr><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle"  colspan="3"  >S. sarrachoides</td><td align="center" valign="middle"  colspan="3"  >S. scabrum</td><td align="center" valign="middle"  colspan="3"  >S. villosum</td></tr><tr><td align="center" valign="middle" ></td><td align="center" valign="middle" >rep</td><td align="center" valign="middle" >1</td><td align="center" valign="middle" >2</td><td align="center" valign="middle" >3</td><td align="center" valign="middle" >1</td><td align="center" valign="middle" >2</td><td align="center" valign="middle" >3</td><td align="center" valign="middle" >1</td><td align="center" valign="middle" >2</td><td align="center" valign="middle" >3</td></tr><tr><td align="center" valign="middle"  rowspan="3"  >S. sarrachoides</td><td align="center" valign="middle" >1</td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td></tr><tr><td align="center" valign="middle" >2</td><td align="center" valign="middle" >0.8962</td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td></tr><tr><td align="center" valign="middle" >3</td><td align="center" valign="middle" >0.7798</td><td align="center" valign="middle" >0.5532</td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td></tr><tr><td align="center" valign="middle"  rowspan="3"  >S. scabrum</td><td align="center" valign="middle" >1</td><td align="center" valign="middle" >0.4873</td><td align="center" valign="middle" >0.8102</td><td align="center" valign="middle" >0.5505</td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td></tr><tr><td align="center" valign="middle" >2</td><td align="center" valign="middle" >0.8738</td><td align="center" valign="middle" >0.3453</td><td align="center" valign="middle" >0.3856</td><td align="center" valign="middle" >0.7370</td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td></tr><tr><td align="center" valign="middle" >3</td><td align="center" valign="middle" >0.8080</td><td align="center" valign="middle" >0.5722</td><td align="center" valign="middle" >0.2101</td><td align="center" valign="middle" >0.6152</td><td align="center" valign="middle" >0.4046</td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td></tr><tr><td align="center" valign="middle"  rowspan="3"  >S. villosum</td><td align="center" valign="middle" >1</td><td align="center" valign="middle" >0.2813</td><td align="center" valign="middle" >0.9020</td><td align="center" valign="middle" >0.7664</td><td align="center" valign="middle" >0.4852</td><td align="center" valign="middle" >0.8743</td><td align="center" valign="middle" >0.8134</td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td></tr><tr><td align="center" valign="middle" >2</td><td align="center" valign="middle" >0.7141</td><td align="center" valign="middle" >0.7760</td><td align="center" valign="middle" >0.6048</td><td align="center" valign="middle" >0.6932</td><td align="center" valign="middle" >0.6989</td><td align="center" valign="middle" >0.5879</td><td align="center" valign="middle" >0.7664</td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td></tr><tr><td align="center" valign="middle" >3</td><td align="center" valign="middle" >0.7003</td><td align="center" valign="middle" >0.7375</td><td align="center" valign="middle" >0.4968</td><td align="center" valign="middle" >0.5552</td><td align="center" valign="middle" >0.6518</td><td align="center" valign="middle" >0.5732</td><td align="center" valign="middle" >0.6781</td><td align="center" valign="middle" >0.4922</td><td align="center" valign="middle" ></td></tr></tbody></table></table-wrap><p>top-most important volatile compound (<xref ref-type="fig" rid="fig2">Figure 2</xref>; <xref ref-type="table" rid="table">Table </xref>S2). This was consistent with SIMPER results as c46 was significantly different between S. sarrachoides and S. scabrum, and S. villosum and S. scabrum, respectively (<xref ref-type="table" rid="table">Table </xref>S1; <xref ref-type="table" rid="table">Table </xref>S2). These “relevant” volatile compounds were used to perform NMDS, PERMANOVA and ANOSIM analysis.</p></sec><sec id="s3_3"><title>3.3. NMDS on Reduced Dataset and Full Dataset</title><p>Non-metric multi-dimensional scaling based on Bray Curtis distance was performed with dimension k = 3 as suggested by the scree plot (<xref ref-type="fig" rid="fig4">Figure 4</xref>) on the full dataset (58 VOCs). Given a scree plot, the value of the dimension of NMDS is at the elbow of the line plot which is the value beyond which additional dimensions do not substantially lower the stress value. Such value provides a suitable dimension for visualizing the NMDS plot.</p><p>We evaluated the NMDS ordination at different dimensions to obtain the stress value that optimizes the ordination fit based on Bray Curtis distance for both reduced dataset and full dataset. The stress values obtained at different dimensions with different number of input variables are presented in <xref ref-type="table" rid="table">Table </xref>2.</p><p><xref ref-type="table" rid="table">Table </xref>2 indicates that as the number of dimensions increase, the stress value reduces. Stress values are also higher for high-dimensional data in variable space as compared to low-dimensional data in variable space. The NMDS ordination algorithm could not converge when the “relevant” variables with dimension k = 3 were used, as the stress value was nearly zero. Consequently, the reduced dataset (16 VOCs) with lower stress value for dimension k = 2 was the ordination of choice as was also supported by the shephard plot that indicated goodness of fit with linear fit, r<sup>2</sup> = 0.993 and non-metric fit, R<sup>2</sup> = 0.998 (<xref ref-type="fig" rid="fig5">Figure 5</xref>(b)) which were both higher compared to the NMDS ordination using full dataset (<xref ref-type="fig" rid="fig5">Figure 5</xref>(a)).</p><table-wrap id="table2" ><label><xref ref-type="table" rid="table">Table </xref>2</label><caption><title> NMDS ordination dimension and the corresponding stress values based on Bray Curtis distance for the full dataset of 58 VOCs and the reduced dataset of 16 VOCs</title></caption><table><tbody><thead><tr><th align="center" valign="middle" >Dimension</th><th align="center" valign="middle" >Stress value (full dataset 58 VOCs)</th><th align="center" valign="middle" >Stress value (reduced dataset 16 VOCs)</th></tr></thead><tr><td align="center" valign="middle" >1</td><td align="center" valign="middle" >0.2453</td><td align="center" valign="middle" >0.1998</td></tr><tr><td align="center" valign="middle" >2</td><td align="center" valign="middle" >0.0806</td><td align="center" valign="middle" >0.0403</td></tr><tr><td align="center" valign="middle" >3</td><td align="center" valign="middle" >0.0320</td><td align="center" valign="middle" >No convergence</td></tr></tbody></table></table-wrap><p>Stress value &lt; 0.05-Excellent; Stress value &gt; 0.2-Poor.</p></sec><sec id="s3_4"><title>3.4. NMDS Ordination Plots</title><p>When all volatile organic compounds were used, the ordination ellipses showed heavy overlaps among the three African nightshade species (<xref ref-type="fig" rid="fig6">Figure 6</xref>). On the other hand, when only the selected “relevant” variables were put into consideration, it was possible to distinguish the three African nightshade species (<xref ref-type="fig" rid="fig7">Figure 7</xref>). NMDS plots the distances between points in the same rank order as distances (or similarities) in the original matrix. The closer two samples are on the plot the more similar those samples are in terms of the underlying data. Further, the samples enclosed within an ellipse or close to the ellipse belong to the same group. The NMDS plots showed that S. scabrum had dissimilar volatile compounds as compared to S. villosum and S. sarrachoides. Further, S. sarrachoides had dissimilar volatiles profiles as compared to S. scabrum and S. villosum (<xref ref-type="fig" rid="fig7">Figure 7</xref>).</p><p>The dissimilarities of VOCs in the three African nightshade species showed by NMDS ordination plots, was supported by hypothesis testing using PERMANOVA and ANOSIM. PERMANOVA results showed that there was no significant difference between the three African nightshade species (p-value = 0.554) when all the VOCs were considered while there was significant difference between the plants (p-value = 0.022) when only “relevant” VOCs were considered. This was further supported by ANOSIM test that showed similar conclusion (<xref ref-type="table" rid="table">Table </xref>3).</p></sec></sec><sec id="s4"><title>4. Discussion</title><p>The data in this study are characterized by high dimensionality and small sample size, which tends to reduce the statistical power of tests. The classical statistical methods do not appropriately regulate type 1 error rate when sample sizes are</p><table-wrap id="table3" ><label><xref ref-type="table" rid="table">Table </xref>3</label><caption><title> PERMANOVA (number of permutations = 999) and ANOSIM hypothesis tests on volatile organic compounds for the three African nightshade species</title></caption><table><tbody><thead><tr><th align="center" valign="middle" >Test</th><th align="center" valign="middle" ></th><th align="center" valign="middle" >58 Volatile organic compounds</th><th align="center" valign="middle" >16 Volatile organic compounds</th></tr></thead><tr><td align="center" valign="middle"  rowspan="2"  >PERMANOVA</td><td align="center" valign="middle" >Pseudo F-ratio</td><td align="center" valign="middle" >0.8176</td><td align="center" valign="middle" >2.7595</td></tr><tr><td align="center" valign="middle" >p-value</td><td align="center" valign="middle" >0.554</td><td align="center" valign="middle" >0.022</td></tr><tr><td align="center" valign="middle"  rowspan="2"  >ANOSIM</td><td align="center" valign="middle" >ANOSIM R statistic</td><td align="center" valign="middle" >−0.07</td><td align="center" valign="middle" >0.4733</td></tr><tr><td align="center" valign="middle" >p-value</td><td align="center" valign="middle" >0.589</td><td align="center" valign="middle" >0.018</td></tr></tbody></table></table-wrap><p>very small as they require moderate to large sample sizes for analysis. Such methods as multivariate analysis of variance (MANOVA) either behave liberally and over-reject the null hypothesis, or behave conservatively [<xref ref-type="bibr" rid="scirp.133191-ref53">53</xref>] . Chang et al. [<xref ref-type="bibr" rid="scirp.133191-ref54">54</xref>] highlighted the need of a large sample size for an accurate type 1 error control.</p><p>Non-metric multidimensional scaling has been used in small sample situation to reveal patterns in multivariate datasets visualized in a reduced dimension space. As such its application in chemical ecology characterized by high dimensional small sample size datasets has become of notable interest. For instance, Hufnagel [<xref ref-type="bibr" rid="scirp.133191-ref55">55</xref>] visualized differences in the amount of glycoalkaloid α-solanine among Solanum tuberosum L., S. chacoense Bitter, S. pinnatisectum Dunal and S. immite Dunal using NMDS. Suinyuy et al. [<xref ref-type="bibr" rid="scirp.133191-ref56">56</xref>] analyzed volatile composition of male and female of African cycad species using headspace technique and gas chromatography-mass spectrometry (GC-MS) where the species were clustered using NMDS according to shared chemical volatiles. NMDS efficiency in the chemical ecology field has been contributed by the technique’s properties such as being less sensitive to variation in species response curve [<xref ref-type="bibr" rid="scirp.133191-ref57">57</xref>] and its requirement of only two dimensions to visualize similarity patterns compared to other ordination techniques which requires a minimum of three dimensions [<xref ref-type="bibr" rid="scirp.133191-ref43">43</xref>] .</p><p>In our study, NMDS results revealed that when all the chemical volatiles were used in the analysis, the nightshade species were highly overlapping. This might have been caused by the “curse of dimensionality” problem where, in high dimensional space an exponential increase in the space volume is experienced as the data becomes relatively small [<xref ref-type="bibr" rid="scirp.133191-ref58">58</xref>] . This makes it hard to find patterns in data samples shown by the ellipses overlap. Further, since the data was projected to a two-dimensional space from 58 dimensions, the similarities revealed by the NMDS plot might have been contributed by it.</p><p>To identify patterns in the volatile compounds, RF and SIMPER were employed to reduce the data dimensions. RF performance was contributed by its efficiency in recognizing data patterns, no assumptions related to data properties, user-friendly parameters and ability to flexibly address the interactions between predictive variables [<xref ref-type="bibr" rid="scirp.133191-ref27">27</xref>] . On the other hand, SIMPER determined the contribution of individual compounds to the separation of the three African nightshades as reflected in the NMDS plots. The performance of the data reduction techniques used in this study is supported by Muthoni [<xref ref-type="bibr" rid="scirp.133191-ref59">59</xref>] , who used SIMPER and one way ANOSIM to compare the chemical profiles of the leaf volatiles of healthy and infected tomato plants and further visualized the clustering of the volatiles using NMDS.</p><p>Bray Curtis distance was considered as the best distance measure in obtaining the NMDS plots since the other distance measures; Euclidean, Manhattan and Kulcynski had poor ordination fit. This might have been contributed by some of the properties of the distance measures. For instance, Junker [<xref ref-type="bibr" rid="scirp.133191-ref60">60</xref>] highlighted that Euclidean distance often lead to high similarities between samples not sharing the same variables as it is affected by large number of zeros in the data [<xref ref-type="bibr" rid="scirp.133191-ref43">43</xref>] [<xref ref-type="bibr" rid="scirp.133191-ref61">61</xref>] . This contradicts with what Legendre and Legendre [<xref ref-type="bibr" rid="scirp.133191-ref43">43</xref>] stated on the properties of a good ecological distance in describing differences in species composition. Species sharing the same or most of the volatiles should have a small ecological distance than those not sharing any volatiles. Tomašev et al. [<xref ref-type="bibr" rid="scirp.133191-ref62">62</xref>] observed that Euclidean and Manhattan distances had similar trends in their results, which was in agreement with the similar NMDS plots obtained based on the two distance measures.</p><p>Bray-Curtis distance only takes the value zero for identical variables and ignores other variables having zeros [<xref ref-type="bibr" rid="scirp.133191-ref63">63</xref>] . Ecological distances based on Bray Curtis range from 0 to 1, with 0 indicating complete similarity and 1 indicating complete dissimilarity. Hence, interpreting the distances based on Bray Curtis is easier than the other distances which do not have an upper bound making it difficult to understand how similar two species are, as they are only understood in a relative way [<xref ref-type="bibr" rid="scirp.133191-ref41">41</xref>] . This is not to say that the distance does not have its limitations; Bray Curtis and related measure such as Kulcynski tend to under estimate true ecological distances when distances become large. Therefore, the distance measure is as useful so far as it produces reasonable ecological ordinations through the ranks used for NMDS [<xref ref-type="bibr" rid="scirp.133191-ref45">45</xref>] [<xref ref-type="bibr" rid="scirp.133191-ref64">64</xref>] .</p><p>Although NMDS performed well in revealing the patterns and reducing the dimension, it did not rigorously express the nature and degree of uncertainty concerning a priori hypotheses. Therefore, non-parametric methods that tested hypothesis concerning the three nightshades species were required to make probabilistic statements about the VOCs data [<xref ref-type="bibr" rid="scirp.133191-ref45">45</xref>] . The results obtained from the two non-parametric tests (PERMANOVA and ANOSIM) were in agreement. These two tests as discussed by Somerfield et al. [<xref ref-type="bibr" rid="scirp.133191-ref65">65</xref>] are complementary tests rather than alternative. Additionally, Rojas et al. [<xref ref-type="bibr" rid="scirp.133191-ref66">66</xref>] used NMDS to visualize the seed disperser functional types, and their relationships with fruit traits, the patterns observed were supported by PERMANOVA results.</p><p>In spite of the patterns among sampling units being visualized with an NMDS plot, use of rank orders to represent points in low dimensional space makes the solution obtained unstable and can even degenerate when applied to a small dataset [<xref ref-type="bibr" rid="scirp.133191-ref67">67</xref>] . In using NMDS, normality assumption is not required, however this necessitates use of intensive iterative algorithm since optimal solution may not be obtained from a single run. Therefore, multiple NMDS solutions with specified dimensionality is necessary to ensure a stable and optimal ordination configuration [<xref ref-type="bibr" rid="scirp.133191-ref67">67</xref>] . Further, multivariate visualization of samples by any ordination technique is not the end point of analysis but should be viewed as a framework in which patterns of individual subjects can be interpreted.</p></sec><sec id="s5"><title>5. Conclusion</title><p>Our results showed that there was a dissimilarity between African nightshade species of S. sarrachoides, S. scabrum and S. villosum when the data dimension was reduced to only 16 volatile compounds as compared to using all the 58 volatile compounds depicted in the NMDS plots and outputs of PERMANOVA and ANOSIM hypothesis tests. Our study shows the merit of reducing variables using RF and SIMPER to enhance visualization and in turn increase the power of PERMANOVA or ANOSIM in analysis of high dimensional small sample dataset as encountered in chemical ecology. Based on our results, we recommend use of RF, SIMPER or any other applicable data reduction technique when dealing with small samples in high dimensional data. Although the data reduction techniques used in our study performed well in discriminating the African nightshade species, the selected variables that were identified by RF and SIMPER might not necessarily be biologically important for chemical ecologists.</p></sec><sec id="s6"><title>Acknowledgements</title><p>The authors gratefully acknowledge the financial support by the following organizations and agencies: UK’s Foreign, Commonwealth &amp; Development Office (FCDO); the Swedish International Development Cooperation Agency (Sida); the Swiss Agency for Development and Cooperation (SDC); the Federal Democratic Republic of Ethiopia; and the Government of the Republic of Kenya. We are grateful to Juma Meltus for creating the flow diagram <xref ref-type="fig" rid="fig1">Figure 1</xref>. The views expressed herein do not reflect the official opinion of the donors.</p></sec><sec id="s7"><title>Authors Contributions</title><p>L.C.: Data analysis, manuscript writing and editing; D.S.: Conceptualization, manuscript review and editing. L.K.M.: Data author, manuscript review and editing, H.T.: Manuscript review and editing.</p></sec><sec id="s8"><title>Competing Interests</title><p>The authors declare no competing interests.</p></sec><sec id="s9"><title>Cite this paper</title><p>Chepkemoi, L., Salifu, D., Murungi, L.K. and Tonnang, H.E.Z. (2024) High Dimension Multivariate Data Analysis for Small Group Samples of Chemical Volatile Profiles of African Nightshade Species. Journal of Data Analysis and Information Processing, 12, 210-231. https://doi.org/10.4236/jdaip.2024.122012</p></sec><sec id="s10"><title>Supplementary Information</title><table-wrap id="table4" ><label><xref ref-type="table" rid="table">Table </xref>S1</label><caption><title> Average contribution of the three nightshade species based on pairwise comparison of the three African nightshade species, S. sarrachoides (sarr) and S. scabrum (sca), and S. villosum (vill) using similarity percentage (SIMPER)</title></caption><table><tbody><thead><tr><th align="center" valign="middle" >Volatiles</th><th align="center" valign="middle" >Average contribution</th><th align="center" valign="middle" >Standard deviation</th><th align="center" valign="middle" >P-value</th><th align="center" valign="middle" >Pairwise comparison</th></tr></thead><tr><td align="center" valign="middle" >c54</td><td align="center" valign="middle" >0.0104</td><td align="center" valign="middle" >0.0087</td><td align="center" valign="middle" >0.027</td><td align="center" valign="middle" >sarr vs villo</td></tr><tr><td align="center" valign="middle" >c3</td><td align="center" valign="middle" >0.0063</td><td align="center" valign="middle" >0.0054</td><td align="center" valign="middle" >0.019</td><td align="center" valign="middle" >sarr vs villo</td></tr><tr><td align="center" valign="middle" >c40</td><td align="center" valign="middle" >0.0040</td><td align="center" valign="middle" >0.0043</td><td align="center" valign="middle" >0.027</td><td align="center" valign="middle" >sarr vs villo</td></tr><tr><td align="center" valign="middle" >c46</td><td align="center" valign="middle" >0.0129</td><td align="center" valign="middle" >0.0063</td><td align="center" valign="middle" >0.012</td><td align="center" valign="middle" >sarr vs sca</td></tr><tr><td align="center" valign="middle" >c45</td><td align="center" valign="middle" >0.0085</td><td align="center" valign="middle" >0.0035</td><td align="center" valign="middle" >0.011</td><td align="center" valign="middle" >sarr vs sca</td></tr><tr><td align="center" valign="middle" >c31</td><td align="center" valign="middle" >0.0001</td><td align="center" valign="middle" >0.0000695</td><td align="center" valign="middle" >0.035</td><td align="center" valign="middle" >sarr vs sca</td></tr><tr><td align="center" valign="middle" >c12</td><td align="center" valign="middle" >0.0471</td><td align="center" valign="middle" >0.0272</td><td align="center" valign="middle" >0.018</td><td align="center" valign="middle" >sarr vs sca</td></tr><tr><td align="center" valign="middle" >c32</td><td align="center" valign="middle" >0.0153</td><td align="center" valign="middle" >0.0132</td><td align="center" valign="middle" >0.022</td><td align="center" valign="middle" >sarr vs sca</td></tr><tr><td align="center" valign="middle" >c461</td><td align="center" valign="middle" >0.0143</td><td align="center" valign="middle" >0.0052</td><td align="center" valign="middle" >0.006</td><td align="center" valign="middle" >villo vs sca</td></tr><tr><td align="center" valign="middle" >c39</td><td align="center" valign="middle" >0.0103</td><td align="center" valign="middle" >0.0093</td><td align="center" valign="middle" >0.041</td><td align="center" valign="middle" >villo vs sca</td></tr><tr><td align="center" valign="middle" >c451</td><td align="center" valign="middle" >0.0094</td><td align="center" valign="middle" >0.0027</td><td align="center" valign="middle" >0.006</td><td align="center" valign="middle" >villo vs sca</td></tr><tr><td align="center" valign="middle" >c33</td><td align="center" valign="middle" >0.0065</td><td align="center" valign="middle" >0.0053</td><td align="center" valign="middle" >0.022</td><td align="center" valign="middle" >villo vs sca</td></tr><tr><td align="center" valign="middle" >c6</td><td align="center" valign="middle" >0.0053</td><td align="center" valign="middle" >0.0038</td><td align="center" valign="middle" >0.038</td><td align="center" valign="middle" >villo vs sca</td></tr><tr><td align="center" valign="middle" >c38</td><td align="center" valign="middle" >0.0028</td><td align="center" valign="middle" >0.0021</td><td align="center" valign="middle" >0.047</td><td align="center" valign="middle" >villo vs sca</td></tr><tr><td align="center" valign="middle" >c10</td><td align="center" valign="middle" >0.0020</td><td align="center" valign="middle" >0.0016</td><td align="center" valign="middle" >0.034</td><td align="center" valign="middle" >villo vs sca</td></tr></tbody></table></table-wrap><table-wrap-group id="5"><label><xref ref-type="table" rid="table">Table </xref>S2</label><caption><title> All 58 volatile organic compounds with their actual name and code as used in this study</title></caption><table-wrap id="5_1"><table><tbody><thead><tr><th align="center" valign="middle" >Chemical volatile</th><th align="center" valign="middle" >volatile code</th></tr></thead><tr><td align="center" valign="middle" >hexanal</td><td align="center" valign="middle" >c1</td></tr><tr><td align="center" valign="middle" >2-hexenal</td><td align="center" valign="middle" >c2</td></tr><tr><td align="center" valign="middle" >(Z)-3-hexen-1-ol</td><td align="center" valign="middle" >c3</td></tr><tr><td align="center" valign="middle" >heptanal</td><td align="center" valign="middle" >c4</td></tr><tr><td align="center" valign="middle" >3-Methyl-2-butenal</td><td align="center" valign="middle" >c5</td></tr><tr><td align="center" valign="middle" >1R-a-Pinene</td><td align="center" valign="middle" >c6</td></tr><tr><td align="center" valign="middle" >Benzaldehyde</td><td align="center" valign="middle" >c7</td></tr><tr><td align="center" valign="middle" >b-Pinene</td><td align="center" valign="middle" >c8</td></tr><tr><td align="center" valign="middle" >6-Methyl-5-hepten-2-one</td><td align="center" valign="middle" >c9</td></tr><tr><td align="center" valign="middle" >Beta Myrcene</td><td align="center" valign="middle" >c10</td></tr><tr><td align="center" valign="middle" >Octanal</td><td align="center" valign="middle" >c11</td></tr><tr><td align="center" valign="middle" >Limonene</td><td align="center" valign="middle" >c12</td></tr><tr><td align="center" valign="middle" >Benzyl alcohol</td><td align="center" valign="middle" >c13</td></tr><tr><td align="center" valign="middle" >Dihydromyrcenol</td><td align="center" valign="middle" >c14</td></tr><tr><td align="center" valign="middle" >3,7-Dimethyl-1-octanol (Geraniol tetrahydride)</td><td align="center" valign="middle" >c15</td></tr><tr><td align="center" valign="middle" >Methyl benzoate</td><td align="center" valign="middle" >c16</td></tr></tbody></table></table-wrap><table-wrap id="5_2"><table><tbody><thead><tr><th align="center" valign="middle" >Linalool</th><th align="center" valign="middle" >c17</th></tr></thead><tr><td align="center" valign="middle" >Nonanal</td><td align="center" valign="middle" >c18</td></tr><tr><td align="center" valign="middle" >1,2,4,5-Tetramethylbenzene (Durol)</td><td align="center" valign="middle" >c19</td></tr><tr><td align="center" valign="middle" >Isophorone</td><td align="center" valign="middle" >c20</td></tr><tr><td align="center" valign="middle" >2-Ethylhexanoic acid</td><td align="center" valign="middle" >c21</td></tr><tr><td align="center" valign="middle" >Octanoic acid</td><td align="center" valign="middle" >c22</td></tr><tr><td align="center" valign="middle" >&#224;-Terpineol</td><td align="center" valign="middle" >c23</td></tr><tr><td align="center" valign="middle" >Decanal</td><td align="center" valign="middle" >c24</td></tr><tr><td align="center" valign="middle" >2,3,3-Trimethyl-2-(3-methylbutyl)-cyclohexanone</td><td align="center" valign="middle" >c25</td></tr><tr><td align="center" valign="middle" >Isothymol methyl ether (anisole)</td><td align="center" valign="middle" >c26</td></tr><tr><td align="center" valign="middle" >Nonanoic acid</td><td align="center" valign="middle" >c27</td></tr><tr><td align="center" valign="middle" >Isobornyl acetate</td><td align="center" valign="middle" >c28</td></tr><tr><td align="center" valign="middle" >Isobutyl butanoate</td><td align="center" valign="middle" >c29</td></tr><tr><td align="center" valign="middle" >Butyl butanoate</td><td align="center" valign="middle" >c30</td></tr><tr><td align="center" valign="middle" >Copaene</td><td align="center" valign="middle" >c31</td></tr><tr><td align="center" valign="middle" >B-Elemene</td><td align="center" valign="middle" >c32</td></tr><tr><td align="center" valign="middle" >7-Epi-Sesquithujene</td><td align="center" valign="middle" >c33</td></tr><tr><td align="center" valign="middle" >Longifolene</td><td align="center" valign="middle" >c34</td></tr><tr><td align="center" valign="middle" >a-Cedrene</td><td align="center" valign="middle" >c35</td></tr><tr><td align="center" valign="middle" >Caryophyllene</td><td align="center" valign="middle" >c36</td></tr><tr><td align="center" valign="middle" >B-Cedrene</td><td align="center" valign="middle" >c37</td></tr><tr><td align="center" valign="middle" >Gemacrene B</td><td align="center" valign="middle" >c38</td></tr><tr><td align="center" valign="middle" >Geranyl acetone</td><td align="center" valign="middle" >c39</td></tr><tr><td align="center" valign="middle" >Humulene</td><td align="center" valign="middle" >c40</td></tr><tr><td align="center" valign="middle" >2,5-Di-tert-butylbenzoquinone</td><td align="center" valign="middle" >c41</td></tr><tr><td align="center" valign="middle" >Epizonarene</td><td align="center" valign="middle" >c42</td></tr><tr><td align="center" valign="middle" >Butylated hydoxytoluene</td><td align="center" valign="middle" >c43</td></tr><tr><td align="center" valign="middle" >D-Cadinene (+)-</td><td align="center" valign="middle" >c44</td></tr><tr><td align="center" valign="middle" >Geranyl linalool</td><td align="center" valign="middle" >c45</td></tr><tr><td align="center" valign="middle" >4,8,12-Trimethyl-1,3 (E), 7 (E)-11-Tridecatetraene</td><td align="center" valign="middle" >c46</td></tr><tr><td align="center" valign="middle" >Caryophyllene oxide</td><td align="center" valign="middle" >c47</td></tr><tr><td align="center" valign="middle" >Cedrol</td><td align="center" valign="middle" >c48</td></tr><tr><td align="center" valign="middle" >Humulene epoxide II</td><td align="center" valign="middle" >c49</td></tr><tr><td align="center" valign="middle" >Methyl dihydrojasmonate</td><td align="center" valign="middle" >c50</td></tr><tr><td align="center" valign="middle" >7-Methyl-Z-tetradecen-1-ol acetate</td><td align="center" valign="middle" >c51</td></tr><tr><td align="center" valign="middle" >Ethylhexyl benzoate</td><td align="center" valign="middle" >c52</td></tr><tr><td align="center" valign="middle" >Isopropyl myristate</td><td align="center" valign="middle" >c53</td></tr><tr><td align="center" valign="middle" >Hexadrofarnesyl acetone</td><td align="center" valign="middle" >c54</td></tr><tr><td align="center" valign="middle" >Hexadecanoic acid</td><td align="center" valign="middle" >c55</td></tr><tr><td align="center" valign="middle" >Isopropyl palmitate</td><td align="center" valign="middle" >c56</td></tr><tr><td align="center" valign="middle" >10,18-Bisnorabieta-8,11-triene</td><td align="center" valign="middle" >c57</td></tr><tr><td align="center" valign="middle" >(Z)-9-Octadecanoic acid</td><td align="center" valign="middle" >c58</td></tr></tbody></table></table-wrap></table-wrap-group></sec></body><back><ref-list><title>References</title><ref id="scirp.133191-ref1"><label>1</label><mixed-citation publication-type="other" xlink:type="simple">Tholl, D., Boland, W., Hansel, A., Loreto, F., Rose, U.S.R. and Schnitzler, J.P. (2006) Practical Approaches to Plant Volatile Analysis. &lt;i&gt;The Plant Journal&lt;/i&gt;, 45, 540-560. &lt;br&gt;https://doi.org/10.1111/j.1365-313X.2005.02612.x</mixed-citation></ref><ref id="scirp.133191-ref2"><label>2</label><mixed-citation publication-type="other" xlink:type="simple">Chung, S.H., Scully, E.D., Peiffer M., Geib S.M., Rosa, C., Hoover K. and Felton G.W. (2017) Host Plant Species Determines Symbiotic Bacterial Community Mediating Suppression of Plant Defenses. &lt;i&gt;Scientific Reports&lt;/i&gt;, 7, Article No. 39690. &lt;br&gt;https://doi.org/10.1038/srep39690</mixed-citation></ref><ref id="scirp.133191-ref3"><label>3</label><mixed-citation publication-type="other" xlink:type="simple">Salerno, G., Rebora, M., Piersanti, S., Gorb, E. and Gorb, S. (2020) Mechanical Ecology of Fruit-Insect Interaction in the Adult Mediterranean Fruit Fly &lt;i&gt;Ceratitis capitata&lt;/i&gt; (Diptera: Tephritidae). &lt;i&gt;Zoology&lt;/i&gt;, 139, Article 125748. &lt;br&gt;https://doi.org/10.1016/j.zool.2020.125748</mixed-citation></ref><ref id="scirp.133191-ref4"><label>4</label><mixed-citation publication-type="other" xlink:type="simple">Zu, P.J., Garc&amp;#237;a-Garc&amp;#237;a, R., Schuman, M.C., Saavedra, S. and Meli&amp;#225;n, C.J. (2023) Plant-Insect Chemical Communication in Ecological Communities: An Information Theory Perspective. &lt;i&gt;Journal of Systematics and Evolution&lt;/i&gt;, 61, 445-453. &lt;br&gt;https://doi.org/10.1111/jse.12841</mixed-citation></ref><ref id="scirp.133191-ref5"><label>5</label><mixed-citation publication-type="other" xlink:type="simple">War, A.R., Paulraj, H.C.S., Gabriel, M., War, M.Y. and Ignacimuthu, S. (2011) Herbivore Induced Plant Volatiles: Their Role in Plant Defense for Pest Management. &lt;i&gt;Plant Signaling &amp; Behavior&lt;/i&gt;, 6, 1973-1978. &lt;br&gt;https://doi.org/10.4161/psb.6.12.18053</mixed-citation></ref><ref id="scirp.133191-ref6"><label>6</label><mixed-citation publication-type="other" xlink:type="simple">Dicke, M., Van Poecke, R.M.P. and De Boer, J.G. (2003) Inducible Indirect Defence of Plants: From Mechanisms to Ecological Functions. &lt;i&gt;Basic and Applied Ecology&lt;/i&gt;, 4, 27-42. &lt;br&gt;https://doi.org/10.1078/1439-1791-00131</mixed-citation></ref><ref id="scirp.133191-ref7"><label>7</label><mixed-citation publication-type="other" xlink:type="simple">Engelberth, J., Alborn, H.T., Schmelz, E.A. and Tumlinson J.H. (2004) Airborne Signals Prime Plants Against Insect Herbivore Attack. Proc. &lt;i&gt;Proceedings of the N&lt;/i&gt;&lt;i&gt;a&lt;/i&gt;&lt;i&gt;tional Academy of Sciences&lt;/i&gt;, 101, 1781-1785. &lt;br&gt;https://doi.org/10.1073/pnas.0308037100</mixed-citation></ref><ref id="scirp.133191-ref8"><label>8</label><mixed-citation publication-type="other" xlink:type="simple">Chin, S.T., Nazimah, S.A.H., Quek, S.Y., Man, Y.B.C., Rahman, R.A. and Hashim, D.M. (2007) Analysis of Volatile Compounds from Malaysian Durians (&lt;i&gt;Durio zib&lt;/i&gt;&lt;i&gt;e&lt;/i&gt;&lt;i&gt;thinus&lt;/i&gt;) Using Headspace SPME Coupled to Fast GC-MS. &lt;i&gt;Journal of Food Compos&lt;/i&gt;&lt;i&gt;i&lt;/i&gt;&lt;i&gt;tion and Analysis&lt;/i&gt;, 20, 31-44. &lt;br&gt;https://doi.org/10.1016/j.jfca.2006.04.011</mixed-citation></ref><ref id="scirp.133191-ref9"><label>9</label><mixed-citation publication-type="other" xlink:type="simple">Drioiche, A., &lt;i&gt;et al&lt;/i&gt;. (2022) Correlation between the Chemical Composition and the Antimicrobial Properties of Seven Samples of Essential Oils of Endemic Thymes in Morocco against Multi-Resistant Bacteria and Pathogenic Fungi. &lt;i&gt;Saudi Pharm&lt;/i&gt;&lt;i&gt;a&lt;/i&gt;&lt;i&gt;ceutical Journal&lt;/i&gt;, 30, 1200-1214. &lt;br&gt;https://doi.org/10.1016/j.jsps.2022.06.022</mixed-citation></ref><ref id="scirp.133191-ref10"><label>10</label><mixed-citation publication-type="other" xlink:type="simple">Paliy, O. and Shankar, V. (2016) Application of Multivariate Statistical Techniques in Microbial Ecology. &lt;i&gt;Molecular Ecology&lt;/i&gt;, 25, 1032-1057. &lt;br&gt;https://doi.org/10.1111/mec.13536</mixed-citation></ref><ref id="scirp.133191-ref11"><label>11</label><mixed-citation publication-type="other" xlink:type="simple">Verma, S.P., Uscanga-Junco, O.A. and D&amp;#237;az-Gonz&amp;#225;lez, L. (2021) A Statistically Coherent Robust Multidimensional Classification Scheme for Water. &lt;i&gt;Science of the Total Environment&lt;/i&gt;, 750, Article 141704. &lt;br&gt;https://doi.org/10.1016/j.scitotenv.2020.141704</mixed-citation></ref><ref id="scirp.133191-ref12"><label>12</label><mixed-citation publication-type="other" xlink:type="simple">Ricciardi, C., &lt;i&gt;et al&lt;/i&gt;. (2020) Linear Discriminant Analysis and Principal Component Analysis to Predict Coronary Artery Disease. &lt;i&gt;Health Informatics Journal&lt;/i&gt;, 26, 2181-2192. &lt;br&gt;https://doi.org/10.1177/1460458219899210</mixed-citation></ref><ref id="scirp.133191-ref13"><label>13</label><mixed-citation publication-type="other" xlink:type="simple">Osborne, J.W. and Costello, A.B. (2004) Sample Size and Subject to Item Ratio in Principal Components Analysis. &lt;i&gt;Practical Assessment&lt;/i&gt;,&lt;i&gt; Research&lt;/i&gt;,&lt;i&gt; and Evaluation&lt;/i&gt;, 9, Article 11.</mixed-citation></ref><ref id="scirp.133191-ref14"><label>14</label><mixed-citation publication-type="other" xlink:type="simple">Kocovsky, P.M., Adams, J.V. and Bronte, C.R. (2009) The Effect of Sample Size on the Stability of Principal Components Analysis of Truss-Based Fish Morphometrics. &lt;i&gt;Transactions of the American Fisheries Society&lt;/i&gt;, 138, 487-496. &lt;br&gt;https://doi.org/10.1577/T08-091.1</mixed-citation></ref><ref id="scirp.133191-ref15"><label>15</label><mixed-citation publication-type="other" xlink:type="simple">Bj&amp;#246;rklund, M. (2019) Be Careful with Your Principal Components. &lt;i&gt;Evolution&lt;/i&gt;, 73, 2151-2158. &lt;br&gt;https://doi.org/10.1111/evo.13835</mixed-citation></ref><ref id="scirp.133191-ref16"><label>16</label><mixed-citation publication-type="other" xlink:type="simple">Shaukat, S.S., Rao, T.A. and Khan, M.A. (2016) Impact of Sample Size on Principal Component Analysis Ordination of an Environmental Data Set: Effects on Eigenstructure. &lt;i&gt;Ekologia &lt;/i&gt;(&lt;i&gt;Bratislava&lt;/i&gt;), 35, 173-190. &lt;br&gt;https://doi.org/10.1515/eko-2016-0014</mixed-citation></ref><ref id="scirp.133191-ref17"><label>17</label><mixed-citation publication-type="other" xlink:type="simple">Sharma, A. and Paliwal, K.K. (2015) Linear Discriminant Analysis for the Small Sample Size Problem: An Overview. &lt;i&gt;International Journal of Machine Learning and Cybernetics&lt;/i&gt;, 6, 443-454. &lt;br&gt;https://doi.org/10.1007/s13042-013-0226-9</mixed-citation></ref><ref id="scirp.133191-ref18"><label>18</label><mixed-citation publication-type="other" xlink:type="simple">Austin, M.P. (2013) Inconsistencies between Theory and Methodology: A Recurrent Problem in Ordination Studies. &lt;i&gt;Journal of Vegetation Science&lt;/i&gt;, 24, 251-268. &lt;br&gt;https://doi.org/10.1111/j.1654-1103.2012.01467.x</mixed-citation></ref><ref id="scirp.133191-ref19"><label>19</label><mixed-citation publication-type="other" xlink:type="simple">Damgaard, C. (2006) Modelling Ecological Presence-Absence Data along an Environmental Gradient: Threshold Levels of the Environment. &lt;i&gt;Environmental and Ecological Statistics&lt;/i&gt;, 13, 229-236. &lt;br&gt;https://doi.org/10.1007/s10651-005-0004-2</mixed-citation></ref><ref id="scirp.133191-ref20"><label>20</label><mixed-citation publication-type="other" xlink:type="simple">Jollife, I.T. and Cadima, J. (2016) Principal Component Analysis: A Review and Recent Developments. &lt;i&gt;Philosophical Transactions of the Royal Society A&lt;/i&gt;:&lt;i&gt; Mathemat&lt;/i&gt;&lt;i&gt;i&lt;/i&gt;&lt;i&gt;cal&lt;/i&gt;,&lt;i&gt; Physical and Engineering Sciences&lt;/i&gt;, 374, Article 20150202. &lt;br&gt;https://doi.org/10.1098/rsta.2015.0202</mixed-citation></ref><ref id="scirp.133191-ref21"><label>21</label><mixed-citation publication-type="other" xlink:type="simple">Belgiu, M. and Dr&amp;#259;gu, L. (2016) Random Forest in Remote Sensing: A Review of Applications and Future Directions. &lt;i&gt;ISPRS Journal of Photogrammetry and Remote Sensing&lt;/i&gt;, 114, 24-31. &lt;br&gt;https://doi.org/10.1016/j.isprsjprs.2016.01.011</mixed-citation></ref><ref id="scirp.133191-ref22"><label>22</label><mixed-citation publication-type="other" xlink:type="simple">Oshiro, T.M., Perez, P.S. and Baranauskas, J.A. (2012) How Many Trees in a Random Forest? &lt;i&gt;Machine&lt;/i&gt; &lt;i&gt;Learning&lt;/i&gt; &lt;i&gt;and&lt;/i&gt; &lt;i&gt;Data&lt;/i&gt; &lt;i&gt;Mining&lt;/i&gt; &lt;i&gt;in&lt;/i&gt; &lt;i&gt;Pattern&lt;/i&gt; &lt;i&gt;Recognition&lt;/i&gt;: 8&lt;i&gt;th&lt;/i&gt; &lt;i&gt;Inte&lt;/i&gt;&lt;i&gt;r&lt;/i&gt;&lt;i&gt;national&lt;/i&gt; &lt;i&gt;Conference&lt;/i&gt;, Berlin, 13-20 July 2012, 154-168.&lt;br&gt;https://doi.org/10.1007/978-3-642-31537-4_13</mixed-citation></ref><ref id="scirp.133191-ref23"><label>23</label><mixed-citation publication-type="other" xlink:type="simple">Chang, V., Bailey, J., Xu, Q.A. and Sun, Z. (2023) Pima Indians Diabetes Mellitus Classification Based on Machine Learning (ML) Algorithms. &lt;i&gt;Neural Computing and Applications&lt;/i&gt;, 35, 16157-16173. &lt;br&gt;https://doi.org/10.1007/s00521-022-07049-z</mixed-citation></ref><ref id="scirp.133191-ref24"><label>24</label><mixed-citation publication-type="other" xlink:type="simple">Olden, J.D. and Jackson, D.A. (2001) Fish-Habitat Relationships in Lakes: Gaining Predictive and Explanatory Insight by Using Artificial Neural Networks. &lt;i&gt;Transa&lt;/i&gt;&lt;i&gt;c&lt;/i&gt;&lt;i&gt;tions of the American Fisheries Society&lt;/i&gt;, 130, 878-897. &lt;br&gt;https://doi.org/10.1577/1548-8659(2001)130&lt;0878:FHRILG&gt;2.0.CO;2</mixed-citation></ref><ref id="scirp.133191-ref25"><label>25</label><mixed-citation publication-type="book" xlink:type="simple">Qi, Y. (2012) Random Forest for Bioinformatics. In: Zhang, C. and Ma, Y., Eds., &lt;i&gt;Ensemble Machine Learning&lt;/i&gt;. &lt;i&gt;Methods and Applications&lt;/i&gt;, Springer, New York, 307-323. &lt;br&gt;https://doi.org/10.1007/978-1-4419-9326-7</mixed-citation></ref><ref id="scirp.133191-ref26"><label>26</label><mixed-citation publication-type="other" xlink:type="simple">Wang, H., Yang, F. and Luo, Z. (2016) An Experimental Study of the Intrinsic Stability of Random Forest Variable Importance Measures. &lt;i&gt;BMC&lt;/i&gt; &lt;i&gt;Bioinformatics&lt;/i&gt;, 17, Article No. 60. &lt;br&gt;https://bmcbioinformatics.biomedcentral.com/articles/10.1186/s12859-016-0900-5&lt;br&gt;https://doi.org/10.1186/s12859-016-0900-5</mixed-citation></ref><ref id="scirp.133191-ref27"><label>27</label><mixed-citation publication-type="other" xlink:type="simple">Luan, J., Zhang, C., Xu, B., Xue, Y. and Ren, Y. (2020) The Predictive Performances of Random Forest Models with Limited Sample Size and Different Species Traits. &lt;i&gt;Fisheries Research&lt;/i&gt;, 227, Article 105534. &lt;br&gt;https://doi.org/10.1016/j.fishres.2020.105534</mixed-citation></ref><ref id="scirp.133191-ref28"><label>28</label><mixed-citation publication-type="other" xlink:type="simple">Janitza, S., Celik, E. and Boulesteix, A.L. (2018) A Computationally Fast Variable Importance Test for Random Forests for High-Dimensional Data. &lt;i&gt;Advances in Data Analysis and Classification&lt;/i&gt;, 12, 885-915. &lt;br&gt;https://doi.org/10.1007/s11634-016-0276-4</mixed-citation></ref><ref id="scirp.133191-ref29"><label>29</label><mixed-citation publication-type="other" xlink:type="simple">Clarke, K.R. (1993) Non-Parametric Multivariate Analyses of Changes in Community Structure. &lt;i&gt;Australian Journal of Ecology&lt;/i&gt;, 18, 117-143.&lt;br&gt;https://doi.org/10.1111/j.1442-9993.1993.tb00438.x</mixed-citation></ref><ref id="scirp.133191-ref30"><label>30</label><mixed-citation publication-type="other" xlink:type="simple">Hattas, D., Hj&amp;#228;lt&amp;#233;n, J., Julkunen-Tiitto, R., Scogings, P.F. and Rooke, T. (2011) Differential Phenolic Profiles in Six African Savanna Woody Species in Relation to Antiherbivore Defense. &lt;i&gt;Phytochemistry&lt;/i&gt;, 72, 1796-1803. &lt;br&gt;https://doi.org/10.1016/j.phytochem.2011.05.007</mixed-citation></ref><ref id="scirp.133191-ref31"><label>31</label><mixed-citation publication-type="other" xlink:type="simple">Gibert C. and Escarguel, G. (2019) PER-SIMPER&amp;#8212;A New Tool for Inferring Community Assembly Processes from Taxon Occurrences. &lt;i&gt;Global Ecology and Bioge&lt;/i&gt;&lt;i&gt;o&lt;/i&gt;&lt;i&gt;grap&lt;/i&gt;&lt;i&gt;hy&lt;/i&gt;, 28, 374-385. &lt;br&gt;https://doi.org/10.1111/geb.12859</mixed-citation></ref><ref id="scirp.133191-ref32"><label>32</label><mixed-citation publication-type="other" xlink:type="simple">Torok, V.A., Ophel-Keller, K., Loo, M. and Hughes, R.J. (2008) Application of Methods for Identifying Broiler Chicken Gut Bacterial Species Linked with Increased Energy Metabolism. &lt;i&gt;Applied and Environmental Microbiology&lt;/i&gt;, 74, 783-791. &lt;br&gt;https://doi.org/10.1128/AEM.01384-07</mixed-citation></ref><ref id="scirp.133191-ref33"><label>33</label><mixed-citation publication-type="other" xlink:type="simple">Murungi, L.K., Kirwa, H., Salifu, D. and Torto, B. (2016) Opposing Roles of Foliar and Glandular Trichome Volatile Components in Cultivated Nightshade Interaction with a Specialist Herbivore. &lt;i&gt;PLOS&lt;/i&gt; &lt;i&gt;ONE&lt;/i&gt;, 11, e0160383.&lt;br&gt;https://doi.org/10.1371/journal.pone.0160383</mixed-citation></ref><ref id="scirp.133191-ref34"><label>34</label><mixed-citation publication-type="other" xlink:type="simple">Kulkarni, V.Y. and Sinha, P.K. (2012) Pruning of Random Forest Classifiers: A Survey and Future Directions. 2012 &lt;i&gt;International&lt;/i&gt; &lt;i&gt;Conference&lt;/i&gt; &lt;i&gt;on&lt;/i&gt; &lt;i&gt;Data&lt;/i&gt; &lt;i&gt;Science&lt;/i&gt; &lt;i&gt;and&lt;/i&gt; &lt;i&gt;Engineering&lt;/i&gt; (&lt;i&gt;ICDSE&lt;/i&gt;), Cochin, 18-20 July 2012, 64-68. &lt;br&gt;https://doi.org/10.1109/ICDSE.2012.6282329</mixed-citation></ref><ref id="scirp.133191-ref35"><label>35</label><mixed-citation publication-type="other" xlink:type="simple">Strobl, C., Boulesteix, A.L., Kneib, T., Augustin, T. and Zeileis, A. (2008) Conditional Variable Importance for Random Forests. &lt;i&gt;BMC&lt;/i&gt; &lt;i&gt;Bioinformatics&lt;/i&gt;, 9, Article No. 307. &lt;br&gt;https://doi.org/10.1186/1471-2105-9-307</mixed-citation></ref><ref id="scirp.133191-ref36"><label>36</label><mixed-citation publication-type="other" xlink:type="simple">Ramette, A. (2007) Multivariate Analyses in Microbial Ecology. &lt;i&gt;FEMS Microbiology Ecology&lt;/i&gt;, 62, 142-160. &lt;br&gt;https://doi.org/10.1111/j.1574-6941.2007.00375.x</mixed-citation></ref><ref id="scirp.133191-ref37"><label>37</label><mixed-citation publication-type="other" xlink:type="simple">Van Der Gucht, K., &lt;i&gt;et al&lt;/i&gt;. (2005) Characterization of Bacterial Communities in Four Freshwater Lakes Differing in Nutrient Load and Food Web Structure. &lt;i&gt;FEMS M&lt;/i&gt;&lt;i&gt;i&lt;/i&gt;&lt;i&gt;crobiology Ecology&lt;/i&gt;, 53, 205-220. &lt;br&gt;https://doi.org/10.1016/j.femsec.2004.12.006</mixed-citation></ref><ref id="scirp.133191-ref38"><label>38</label><mixed-citation publication-type="other" xlink:type="simple">Salido, J.A. and Clemente, J. (2012) Non-Metric Multidimensional Scaling for Biological Characterization of Reduced Yeast Cell Cycle. 2012 &lt;i&gt;International&lt;/i&gt; &lt;i&gt;Conference&lt;/i&gt; &lt;i&gt;on&lt;/i&gt; &lt;i&gt;Biological&lt;/i&gt; &lt;i&gt;and&lt;/i&gt; &lt;i&gt;Life&lt;/i&gt; &lt;i&gt;Sciences&lt;/i&gt;, Singapore, 23-24 July 2012, 104-108.</mixed-citation></ref><ref id="scirp.133191-ref39"><label>39</label><mixed-citation publication-type="other" xlink:type="simple">Dexter, E., Rollwagen-Bollens, G. and Bollens, S.M. (2018) The Trouble with Stress: a Flexible Method for the Evaluation of Nonmetric Multidimensional Scaling. &lt;i&gt;Li&lt;/i&gt;&lt;i&gt;m&lt;/i&gt;&lt;i&gt;nology and Oceanography&lt;/i&gt;:&lt;i&gt; Methods&lt;/i&gt;, 16, 434-443. &lt;br&gt;https://doi.org/10.1002/lom3.10257</mixed-citation></ref><ref id="scirp.133191-ref40"><label>40</label><mixed-citation publication-type="other" xlink:type="simple">San Segundo, E., Tsanas, A. and G&amp;#243;mez-Vilda, P. (2017) Euclidean Distances as Measures of Speaker Similarity Including Identical Twin Pairs: A Forensic Investigation Using Source and Filter Voice Characteristics. &lt;i&gt;Forensic Science Internatio&lt;/i&gt;&lt;i&gt;n&lt;/i&gt;&lt;i&gt;al&lt;/i&gt;, 270, 25-38. &lt;br&gt;https://doi.org/10.1016/j.forsciint.2016.11.020</mixed-citation></ref><ref id="scirp.133191-ref41"><label>41</label><mixed-citation publication-type="other" xlink:type="simple">Legendre, P. and Gallagher, E.D. (2001) Ecologically Meaningful Transformations for Ordination of Species Data. &lt;i&gt;Oecologia&lt;/i&gt;, 129, 271-280.&lt;br&gt;https://doi.org/10.1007/s004420100716</mixed-citation></ref><ref id="scirp.133191-ref42"><label>42</label><mixed-citation publication-type="other" xlink:type="simple">Gomathi, V.V. and Karthikeyan, S. (2014) An Efficient Clustering Segmentation Algorithm for Computer Tomography Image Segmentation. &lt;i&gt;Journal of Biomedical Engineering and Medical Imaging&lt;/i&gt;, 1, 1-11. &lt;br&gt;https://doi.org/10.14738/jbemi.13.267</mixed-citation></ref><ref id="scirp.133191-ref43"><label>43</label><mixed-citation publication-type="other" xlink:type="simple">Legendre, P. and Legendre, L. (2012) Numerical Ecology, Developments in Environmental Modelling. 3rd Edition, Elsevier, Amsterdam, 419.</mixed-citation></ref><ref id="scirp.133191-ref44"><label>44</label><mixed-citation publication-type="other" xlink:type="simple">Gagn&amp;#233;, S.A. and Fahrig, L. (2011) Do Birds and Beetles Show Similar Responses to Urbanization? &lt;i&gt;Ecological Applications&lt;/i&gt;, 21, 2297-2312. &lt;br&gt;https://doi.org/10.1890/09-1905.1</mixed-citation></ref><ref id="scirp.133191-ref45"><label>45</label><mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Anderson</surname><given-names> M.J. </given-names></name>,<etal>et al</etal>. (<year>2001</year>)<article-title>A New Method for Non-Parametric Multivariate Analysis of Variance</article-title><source> &lt;i&gt;Austral Ecology&lt;/i&gt;</source><volume> 26</volume>,<fpage> 32</fpage>-<lpage>46</lpage>.<pub-id pub-id-type="doi"></pub-id></mixed-citation></ref><ref id="scirp.133191-ref46"><label>46</label><mixed-citation publication-type="other" xlink:type="simple">R Core Team (2022) R: A Language and Environment for Statistical Computing. R Foundation for Statistical Computing, Vienna, Austria. &lt;br&gt;https://www.r-project.org/</mixed-citation></ref><ref id="scirp.133191-ref47"><label>47</label><mixed-citation publication-type="other" xlink:type="simple">Liaw A. and Wiener, M. (2002) Classification and Regression by randomForest. &lt;i&gt;R&lt;/i&gt; &lt;i&gt;News&lt;/i&gt;, 2, 18-22. &lt;br&gt;https://cran.r-project.org/doc/Rnews/</mixed-citation></ref><ref id="scirp.133191-ref48"><label>48</label><mixed-citation publication-type="other" xlink:type="simple">Greenwell, B.M. and Boehmke, B.C. (2020) Variable Importance Plots&amp;#8212;An Introduction to the Vip Package. &lt;i&gt;The R Journal&lt;/i&gt;, 12, 343-366. &lt;br&gt;https://doi.org/10.32614/RJ-2020-013</mixed-citation></ref><ref id="scirp.133191-ref49"><label>49</label><mixed-citation publication-type="other" xlink:type="simple">Oksanen, J., &lt;i&gt;et al&lt;/i&gt;. (2022) Vegan: Community Ecology Package. &lt;br&gt;https://cran.r-project.org/package=vegan</mixed-citation></ref><ref id="scirp.133191-ref50"><label>50</label><mixed-citation publication-type="other" xlink:type="simple">Von Lampe, F. and Schellenberg, J. (2023) Goeveg: Functions for Community Data and Ordinations. &lt;br&gt;https://cran.r-project.org/package=goeveg</mixed-citation></ref><ref id="scirp.133191-ref51"><label>51</label><mixed-citation publication-type="other" xlink:type="simple">Wickham, H. (2016) Ggplot2: Elegant Graphics for Data Analysis. Springer-Verlag, New York. &lt;br&gt;https://ggplot2.tidyverse.org&lt;br&gt;https://doi.org/10.1007/978-3-319-24277-4</mixed-citation></ref><ref id="scirp.133191-ref52"><label>52</label><mixed-citation publication-type="other" xlink:type="simple">Pedersen, T.L. (2022) Ggforce: Accelerating &amp;#8216;Ggplot2&amp;#8217;. &lt;br&gt;https://cran.r-project.org/package=ggforce</mixed-citation></ref><ref id="scirp.133191-ref53"><label>53</label><mixed-citation publication-type="other" xlink:type="simple">Konietschke, F., Schwab, K. and Pauly, M. (2020) Small Sample Sizes&amp;#8239;: A Big Data Problem in High-Dimensional Data Analysis. &lt;i&gt;Statistical Methods in Medical R&lt;/i&gt;&lt;i&gt;e&lt;/i&gt;&lt;i&gt;search&lt;/i&gt;, 30, 687-701. &lt;br&gt;https://doi.org/10.1177/0962280220970228</mixed-citation></ref><ref id="scirp.133191-ref54"><label>54</label><mixed-citation publication-type="other" xlink:type="simple">Chang, J., Zheng, C., Zhou, W. and Zhou, W. (2017) Simulation-Based Hypothesis Testing of High Dimensional Means under Covariance Heterogeneity. &lt;i&gt;Biometrics&lt;/i&gt;, 73, 1300-1310. &lt;br&gt;https://doi.org/10.1111/biom.12695</mixed-citation></ref><ref id="scirp.133191-ref55"><label>55</label><mixed-citation publication-type="other" xlink:type="simple">Hufnagel, M.J. (2015) Chemical Ecology of Wild &lt;i&gt;Solanum spp&lt;/i&gt; and Their Interaction with the Colorado Potato Beetle. Master&amp;#8217;s Thesis, Michigan State University, East Lansing.</mixed-citation></ref><ref id="scirp.133191-ref56"><label>56</label><mixed-citation publication-type="other" xlink:type="simple">Suinyuy, T.N., Donaldson, J.S. and Johnson, S.D. (2012) Variation in the Chemical Composition of Cone Volatiles within the African Cycad Genus &lt;i&gt;Encephalartos&lt;/i&gt;. &lt;i&gt;Phytochemistry&lt;/i&gt;, 85, 82-91. &lt;br&gt;https://doi.org/10.1016/j.phytochem.2012.09.016</mixed-citation></ref><ref id="scirp.133191-ref57"><label>57</label><mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Ruokolainen L. and Salo</surname><given-names> K. </given-names></name>,<etal>et al</etal>. (<year>2006</year>)<article-title>Differences in Performance of Four Ordination Methods on a Complex Vegetation Dataset</article-title><source> &lt;i&gt;Annales Botanici&lt;/i&gt;&lt;i&gt; &lt;/i&gt;&lt;i&gt;Fennici&lt;/i&gt;</source><volume> 43</volume>,<fpage> 269</fpage>-<lpage>275</lpage>.<pub-id pub-id-type="doi"></pub-id></mixed-citation></ref><ref id="scirp.133191-ref58"><label>58</label><mixed-citation publication-type="other" xlink:type="simple">Wang, J., Liu, X. and Shen, H. (2019) High-Dimensional Data Analysis with Subspace Comparison Using Matrix Visualization. &lt;i&gt;Information Visualization&lt;/i&gt;, 18, 94-109. &lt;br&gt;https://doi.org/10.1177/1473871617733996</mixed-citation></ref><ref id="scirp.133191-ref59"><label>59</label><mixed-citation publication-type="other" xlink:type="simple">Muthoni, K.R. (2023) Identification and Mechanisms of Allelochemicals Regulating Root-Knot Nematode Parasitism. Ph.D. Thesis, Kenyatta University, Kahawa.</mixed-citation></ref><ref id="scirp.133191-ref60"><label>60</label><mixed-citation publication-type="other" xlink:type="simple">Junker, R.R. (2018) A Biosynthetically Informed Distance Measure to Compare Secondary Metabolite Profiles. &lt;i&gt;Chemoecology&lt;/i&gt;, 28, 29-37.&lt;br&gt;https://doi.org/10.1007/s00049-017-0250-4</mixed-citation></ref><ref id="scirp.133191-ref61"><label>61</label><mixed-citation publication-type="other" xlink:type="simple">Roberts, D.W. (2017) Distance, Dissimilarity, and Mean-Variance Ratios in Ordination. &lt;i&gt;Methods in Ecology and Evolution&lt;/i&gt;, 8, 1398-1407. &lt;br&gt;https://doi.org/10.1111/2041-210X.12739</mixed-citation></ref><ref id="scirp.133191-ref62"><label>62</label><mixed-citation publication-type="other" xlink:type="simple">Toma&amp;#353;ev, N., Radovanovi&amp;#263;, M., Mladeni&amp;#263;, D. and Ivanovi&amp;#263;, M. (2014) The Role of Hubness in Clustering High-Dimensional Data. &lt;i&gt;IEEE Transactions on Knowledge and Data Engineering&lt;/i&gt;, 26, 739-751. &lt;br&gt;https://doi.org/10.1109/TKDE.2013.25</mixed-citation></ref><ref id="scirp.133191-ref63"><label>63</label><mixed-citation publication-type="other" xlink:type="simple">Ricotta, C. and Podani, J. (2017) On Some Properties of the Bray-Curtis Dissimilarity and Their Ecological Meaning. &lt;i&gt;Ecological Complexity&lt;/i&gt;, 31, 201-205. &lt;br&gt;https://doi.org/10.1016/j.ecocom.2017.07.003</mixed-citation></ref><ref id="scirp.133191-ref64"><label>64</label><mixed-citation publication-type="other" xlink:type="simple">Faith, D.P., Minchin, P.R. and Belbin, L. (1987) Compositional Dissimilarity as a Robust Measure of Ecological Distance. &lt;i&gt;Vegetatio&lt;/i&gt;, 69, 57-68. &lt;br&gt;https://doi.org/10.1007/BF00038687</mixed-citation></ref><ref id="scirp.133191-ref65"><label>65</label><mixed-citation publication-type="other" xlink:type="simple">Somerfield, P.J., Clarke, K.R. and Gorley, R.N. (2021) Analysis of Similarities (ANOSIM) for 2-Way Layouts Using a Generalised ANOSIM Statistic, with Comparative Notes on Permutational Multivariate Analysis of Variance (PERMANOVA). &lt;i&gt;Austral Ecology&lt;/i&gt;, 46, 911-926. &lt;br&gt;https://doi.org/10.1111/aec.13059</mixed-citation></ref><ref id="scirp.133191-ref66"><label>66</label><mixed-citation publication-type="other" xlink:type="simple">Rojas, T.N., Zampini, I.C., Isla, M.I. and Blendiger, P.G. (2022) Fleshy Fruit Traits and Seed Dispersers: Which Traits Define Syndromes? &lt;i&gt;Annals of Botany&lt;/i&gt;, 129, 831-838. &lt;br&gt;https://doi.org/10.1093/aob/mcab150</mixed-citation></ref><ref id="scirp.133191-ref67"><label>67</label><mixed-citation publication-type="other" xlink:type="simple">Kenkel, N.C. (2006) On Selecting an Appropriate Multivariate Analysis. &lt;i&gt;Canadian Journal of Plant Science&lt;/i&gt;, 86, 663-676. &lt;br&gt;https://doi.org/10.4141/P05-164</mixed-citation></ref></ref-list></back></article>