<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.4 20241031//EN" "JATS-journalpublishing1-4.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article" dtd-version="1.4" xml:lang="en">
  <front>
    <journal-meta>
      <journal-id journal-id-type="publisher-id">gep</journal-id>
      <journal-title-group>
        <journal-title>Journal of Geoscience and Environment Protection</journal-title>
      </journal-title-group>
      <issn pub-type="epub">2327-4344</issn>
      <issn pub-type="ppub">2327-4336</issn>
      <publisher>
        <publisher-name>Scientific Research Publishing</publisher-name>
      </publisher>
    </journal-meta>
    <article-meta>
      <article-id pub-id-type="doi">10.4236/gep.2026.145009</article-id>
      <article-id pub-id-type="publisher-id">gep-151554</article-id>
      <article-categories>
        <subj-group>
          <subject>Article</subject>
        </subj-group>
        <subj-group>
          <subject>Earth</subject>
          <subject>Environmental Sciences</subject>
        </subj-group>
      </article-categories>
      <title-group>
        <article-title>Detection of Soil Chemical Profiles Using Supervised Learning: A Comparison of Ensemble Algorithms</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author" corresp="yes">
          <name name-style="western">
            <surname>Sarr</surname>
            <given-names>Ahmed Babacar</given-names>
          </name>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <name name-style="western">
            <surname>Ndiaye</surname>
            <given-names>Mapathé</given-names>
          </name>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <name name-style="western">
            <surname>Sarr</surname>
            <given-names>Sabou</given-names>
          </name>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <name name-style="western">
            <surname>Cisse</surname>
            <given-names>Abdoulaye</given-names>
          </name>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <name name-style="western">
            <surname>Camara</surname>
            <given-names>Ndiouga</given-names>
          </name>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
      </contrib-group>
      <aff id="aff1"><label>1</label> Engineering Science Department, University Iba Der THIAM, Thiès, Sénégal </aff>
      <aff id="aff2"><label>2</label> Laboratoire de Mécanique et Modélisation, Thiès, Sénégal </aff>
      <author-notes>
        <fn fn-type="conflict" id="fn-conflict">
          <p>The authors declare no conflicts of interest regarding the publication of this paper.</p>
        </fn>
      </author-notes>
      <pub-date pub-type="epub">
        <day>09</day>
        <month>05</month>
        <year>2026</year>
      </pub-date>
      <pub-date pub-type="collection">
        <month>05</month>
        <year>2026</year>
      </pub-date>
      <volume>14</volume>
      <issue>05</issue>
      <fpage>123</fpage>
      <lpage>136</lpage>
      <history>
        <date date-type="received">
          <day>06</day>
          <month>04</month>
          <year>2026</year>
        </date>
        <date date-type="accepted">
          <day>25</day>
          <month>05</month>
          <year>2026</year>
        </date>
        <date date-type="published">
          <day>28</day>
          <month>05</month>
          <year>2026</year>
        </date>
      </history>
      <permissions>
        <copyright-statement>© 2026 by the authors and Scientific Research Publishing Inc.</copyright-statement>
        <copyright-year>2026</copyright-year>
        <license license-type="open-access">
          <license-p> This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license ( <ext-link ext-link-type="uri" xlink:href="https://creativecommons.org/licenses/by/4.0/">https://creativecommons.org/licenses/by/4.0/</ext-link> ). </license-p>
        </license>
      </permissions>
      <self-uri content-type="doi" xlink:href="https://doi.org/10.4236/gep.2026.145009">https://doi.org/10.4236/gep.2026.145009</self-uri>
      <abstract>
        <p>Geophysical prospecting comprises a set of methods used to measure variations in a physical field or in the Earth’s chemical potential. It plays a crucial role in subsurface exploration to characterize these heterogeneities. This enables the location of buried structures, lithologies or geological features to be determined. The aim of this study is to evaluate the effectiveness of machine learning algorithms in soil classification based on chemical properties. A series of samples was collected from various regions of Senegal and examined in the laboratory to determine their chemical compositions. Data pre-processing, ranging from cleaning missing values and handling inconsistencies to normalization and checking statistical consistency, was carried out on the data prior to the application of machine learning algorithms to generate classification reports. The results obtained following data processing show that ensemble methods, particularly random forests, are effective classifiers, with the X-gradient classifier and the bagging classifier achieving the highest classification accuracy of over 98%. These results demonstrate the relevance of using machine learning algorithms as tools for soil classification, complementing geotechnical studies. However, it should be noted that this study was conducted on samples whose chemical composition is known, and which were collected in an environment conducive to producing a specific chemical content depending on the soil type and region.</p>
      </abstract>
      <kwd-group kwd-group-type="author-generated" xml:lang="en">
        <kwd>Soil Classification</kwd>
        <kwd>Geochemistry</kwd>
        <kwd>Cluster</kwd>
        <kwd>Geochemistry Factors</kwd>
        <kwd>Supervised Learning Techniques</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec1">
      <title>1. Introduction</title>
      <p>There is growing interest in the research and development of machine learning and data mining techniques designed to facilitate geophysical studies and decision-making in geoscience, such as the classification of soils based on their geophysical parameters ([<xref ref-type="bibr" rid="B10">10</xref>]). Generally speaking, statistical learning methods are applied to field data in order to identify broad trends at both large-scale and local levels, thereby assessing the behavior of specific subsets. These results can be used in future scenarios to establish a classification system or to strengthen the methodological approach. These trained models can be used to assist engineers in guiding their decision-making ([<xref ref-type="bibr" rid="B14">14</xref>]; [<xref ref-type="bibr" rid="B8">8</xref>]). Furthermore, these models can uncover previously unknown correlations between variables and the output, thereby improving knowledge and understanding of the soil. These correlations may lead to better interpretations or strategies for the utilization of the subsoil. Given that predictive models calculate forecasts based on information relating to a soil sample, but whose constituent minerals are common to that soil, they constitute promising tools for the purpose of soil classification. The potential of predictive models lies in their ability to generalize from training data. Despite their weakness in applying the same rules, they can process larger and more complex data whilst picking up on subtleties. This is because, in certain situations, they can prove counterintuitive ([<xref ref-type="bibr" rid="B12">12</xref>]). Predictive models rely on training data and depend on the quantity and quality of that data ([<xref ref-type="bibr" rid="B6">6</xref>]). A model extracts the existing signal from the data whilst ignoring noise. The measured data contain certain imperfections, some of which are related to the hardware and others to the system. Pre-processing these measurements involves cleaning and transforming the raw data into a coherent and usable signal. This improves the ability of predictive models to extract the actual signal from the measurements. Various methods exist depending on the specific requirements for improving the data ([<xref ref-type="bibr" rid="B17">17</xref>]), such as imputation, outlier handling or scaling to prevent class dominance. As predictive models are based on digital data, choices must be made regarding encoding methods, dimensionality, or feature engineering ([<xref ref-type="bibr" rid="B21">21</xref>]). These techniques aim to improve the performance of predictive models.</p>
    </sec>
    <sec id="sec2">
      <title>2. Material</title>
      <sec id="sec2dot1">
        <title>2.1. Data</title>
        <p>The data comes mainly from Senegal, The Gambia and Mali. These neighboring countries share a common geological heritage. This closely related geological heritage is linked to the geological continuity of a single basin. <bold>Table 1</bold> shows the distribution of acquisitions by country. These acquisitions are not made uniformly, but rather at sites carefully selected for their mineralogical potential.</p>
        <p>The data collected comes from different countries but pertains to a single continuous geological zone. This zone consists of clay, sand, and granite, as shown in <bold>Table 2</bold>, which presents the quantities collected by soil type.</p>
        <p><bold>Table 1</bold><bold>.</bold> Sample sizes by country.</p>
        <table-wrap id="tbl1">
          <label>Table 1</label>
          <table>
            <tbody>
              <tr>
                <td>
                </td>
                <td>
                  <bold>Sample</bold>
                </td>
                <td>
                  <bold>Pourcentage</bold>
                </td>
                <td>
                  <bold>%</bold>
                </td>
              </tr>
              <tr>
                <td>Sénégal</td>
                <td>406</td>
                <td>87.12446352</td>
                <td>%</td>
              </tr>
              <tr>
                <td>Mali</td>
                <td>8</td>
                <td>1.716738197</td>
                <td>%</td>
              </tr>
              <tr>
                <td>Gambie</td>
                <td>52</td>
                <td>11.15879828</td>
                <td>%</td>
              </tr>
              <tr>
                <td>Total</td>
                <td>466</td>
                <td>100</td>
                <td>%</td>
              </tr>
            </tbody>
          </table>
        </table-wrap>
        <p><bold>Table 2</bold><bold>.</bold> Sample sizes by soil type.</p>
        <table-wrap id="tbl2">
          <label>Table 2</label>
          <table>
            <tbody>
              <tr>
                <td>
                </td>
                <td>
                  <bold>Sample</bold>
                </td>
                <td>
                  <bold>Pourcentage</bold>
                </td>
                <td>
                  <bold>%</bold>
                </td>
              </tr>
              <tr>
                <td>Argile Noire</td>
                <td>130</td>
                <td>27.89699571</td>
                <td>%</td>
              </tr>
              <tr>
                <td>Granite</td>
                <td>123</td>
                <td>26.39484979</td>
                <td>%</td>
              </tr>
              <tr>
                <td>Argile Blanche</td>
                <td>72</td>
                <td>15.45064378</td>
                <td>%</td>
              </tr>
              <tr>
                <td>Argile Rouge</td>
                <td>68</td>
                <td>14.59227468</td>
                <td>%</td>
              </tr>
              <tr>
                <td>Sable</td>
                <td>73</td>
                <td>15.66523605</td>
                <td>%</td>
              </tr>
              <tr>
                <td>Total</td>
                <td>466</td>
                <td>100</td>
                <td>%</td>
              </tr>
            </tbody>
          </table>
        </table-wrap>
        <p>The samples collected reflect the diversity of soils, colors and, above all, the types of materials found in these different regions. This variation is explained by the fact that a large proportion of the data collected, particularly in the south, comes from lacustrine environments. These environments are characterized by a meandering river network, but also by significant plant decomposition of living organisms, making these soils very rich in organic matter. These data allow us to observe the ranges of variation in composition according to soil type and their mineral content. The diversity of the soils highlights their mineralogical variability depending on the conditions of formation. The data include, amongst other things: the content of silica (SiO<sub>2</sub>), alumina (Al<sub>2</sub>O<sub>3</sub>), titanium monoxide (TiO<sub>2</sub>), quicklime (CaO), sodium oxide (Na<sub>2</sub>O), hematite (Fe<sub>2</sub>O<sub>3</sub>), magnesia (MgO) and potassium oxide (K<sub>2</sub>O). Using these data, we calculated: the silica saturation index (IndSatSilice), soil alkalinity (indAlcalin), the proportion of iron relative to other oxides (IndFer), the basic potential (SomOxyBasik) and the total composition (Compotot). </p>
        <p>In our study, the response variables correspond to the five soil types defined based on their mineralogical characteristics. The indices used provide additional information that facilitates a better understanding of the soils and their behavior.</p>
        <p>This classification is based on soil type characteristics, without taking the region of origin into account. The Precambrian basement is an outcrop of hard rock located in eastern Senegal. This sedimentary basin provides a wealth of information on the soil science of Senegal. It stretches from Mauritania to Guinea-Bissau.</p>
      </sec>
      <sec id="sec2dot2">
        <title>2.2. Data Collection and Analysis</title>
        <p>As part of this study, data were collected from several targeted regions in Senegal, The Gambia and Mali. Sampling sites were selected along riverbanks, in areas particularly representative of alluvial deposits and local hydrogeological dynamics. The collection operations were carried out using a pick-up truck. The tools used included a GPS, a spade, a pickaxe, an auger and bags for packaging the collected materials. Sampling depths varied depending on the material but generally ranged from 30 cm to 1.5 m. Sampling was carried out according to spatial variability and soil friability. As part of this study, scalar robust normalization was used to control outliers, standardize the scales of the variables and improve the convergence of the algorithms. Exploratory statistical analyses enable the data to be visualized in order to understand the influences and dominance of the variables. These collected samples are predominantly rich in clay, granite and sand. These are the major soil types present in the area. <bold>Table 3</bold> shows the overall statistics for the data collected from the sites. It is on the basis of this data that the scaling will be applied.</p>
        <p><bold>Table 3</bold><bold>.</bold> Mean and variance of the measured chemical parameters.</p>
        <table-wrap id="tbl3">
          <label>Table 3</label>
          <table>
            <tbody>
              <tr>
                <td>Statistique</td>
                <td>
                  <bold>SiO</bold>
                  <bold>
                    <sub>2</sub>
                  </bold>
                </td>
                <td>
                  <bold>Al</bold>
                  <bold>
                    <sub>2</sub>
                  </bold>
                  <bold>O</bold>
                  <bold>
                    <sub>3</sub>
                  </bold>
                </td>
                <td>
                  <bold>TiO</bold>
                  <bold>
                    <sub>2</sub>
                  </bold>
                </td>
                <td>
                  <bold>CaO</bold>
                </td>
                <td>
                  <bold>MgO</bold>
                </td>
                <td>
                  <bold>Fe</bold>
                  <bold>
                    <sub>2</sub>
                  </bold>
                  <bold>O</bold>
                  <bold>
                    <sub>3</sub>
                  </bold>
                </td>
                <td>
                  <bold>K</bold>
                  <bold>
                    <sub>2</sub>
                  </bold>
                  <bold>O</bold>
                </td>
                <td>
                  <bold>Na</bold>
                  <bold>
                    <sub>2</sub>
                  </bold>
                  <bold>O</bold>
                </td>
              </tr>
              <tr>
                <td>
                  <bold>mean</bold>
                </td>
                <td>64.74</td>
                <td>18.86</td>
                <td>1.05</td>
                <td>0.55</td>
                <td>0.43</td>
                <td>2.24</td>
                <td>0.84</td>
                <td>0.86</td>
              </tr>
              <tr>
                <td>
                  <bold>std</bold>
                </td>
                <td>12.13</td>
                <td>6.91</td>
                <td>0.62</td>
                <td>2.45</td>
                <td>1.92</td>
                <td>3.26</td>
                <td>1.24</td>
                <td>0.74</td>
              </tr>
            </tbody>
          </table>
        </table-wrap>
        <p>The compressive strength of samples taken from different soils ranges from very dense/rigid soil to very soft rock, with typical compressive strength values ranging from 0.95 to 5 MPa. Soil analysis plays a crucial role in nutrient management. Samples taken are subjected to laboratory analysis to determine their chemical composition. Chemical soil analysis serves to quantify the nutrients that should be available to plants and for human needs. Soil analysis involving the characterization of its constituents and the composition of its inorganic phases is referred to as chemical analysis. A chemical analysis involves the identification of several parameters such as: pH, redox potential, organic matter content, total nitrogen, calcium carbonate content, available phosphorus, boron, metals (Cu, Zn, Mn and Fe), exchangeable potassium, calcium, magnesium, sodium, and anion content (<inline-formula><mml:math display="inline"><mml:mrow><mml:msubsup><mml:mrow><mml:mtext> NO </mml:mtext></mml:mrow><mml:mn> 3 </mml:mn><mml:mo> − </mml:mo></mml:msubsup></mml:mrow></mml:math></inline-formula> , <inline-formula><mml:math display="inline"><mml:mrow><mml:msubsup><mml:mrow><mml:mtext> SO </mml:mtext></mml:mrow><mml:mn> 4 </mml:mn><mml:mrow><mml:mn> 2 </mml:mn><mml:mo> − </mml:mo></mml:mrow></mml:msubsup></mml:mrow></mml:math></inline-formula> , <inline-formula><mml:math display="inline"><mml:mrow><mml:msubsup><mml:mrow><mml:mtext> PO </mml:mtext></mml:mrow><mml:mn> 4 </mml:mn><mml:mrow><mml:mn> 3 </mml:mn><mml:mo> − </mml:mo></mml:mrow></mml:msubsup></mml:mrow></mml:math></inline-formula> , Cl<sup>−</sup>) ([<xref ref-type="bibr" rid="B16">16</xref>]). The above parameters are generally measured by soil analysis laboratories. They determine the soil category based on its mineralogical composition. The content of chemical elements is determined by inductively coupled plasma atomic emission spectroscopy (ICP-AES), with soil samples first being treated with acids. <bold>Table 1</bold> shows the overall statistics for the data collected from the sites. It is on the basis of this data that the scaling will be applied.</p>
      </sec>
    </sec>
    <sec id="sec3">
      <title>3. Methodology</title>
      <p>Classification can be defined as a mapping of features to an object. A classification function <italic>y</italic> = f(<italic>X</italic>) is responsible for classifying data points into hyperplanes. The inputs are represented by a vector <italic>X</italic> and the output by a class label <italic>y</italic>. Using instances of <italic>X</italic> and <italic>y</italic>, supervised learning attempts to train a classification model. The aim of the function is to map features to outputs. Due to the size of the raw data and the requirements of the study, it is split into batches for training, evaluation and validation. </p>
      <p>The supervised classification task using machine learning is divided into four general stages: </p>
      <p>Pre-processing, which involves exploring the data,Training, which familiarizes the model with the classification features, Evaluation, which measures the model’s performance, Validation, which assesses the model’s ability to generalize.</p>
      <p>For this study, the data were first subjected to various preprocessing steps and then divided as follows: 70% for the training set, 20% for the test set, and 10% for the validation set.</p>
      <p>The strength of machine learning models lies in these hyperparameters. An optimal search helps to identify a more effective model. An unbiased evaluation quantifies the model’s ability to classify samples not used during training. In other words, it assesses how well the model generalizes. Performance metrics for classifiers or regression models are used to assess predictive ability. </p>
      <sec id="sec3dot1">
        <title>3.1. Random Forest Classifier</title>
        <p>Random forests are based on the principle known as the wisdom of crowds. They combine multiple decision trees built on samples obtained via bootstrapping and subsets of features ([<xref ref-type="bibr" rid="B3">3</xref>]). The idea is to produce robust predictions by aggregating the results of the subsets known as trees. For classification, this involves a majority vote, and for regression, an average. This approach is robust to outliers in various fields such as geophysics, geology and geochemistry. Visual geological mapping is slow and biased; random forests enable the production of more reliable maps useful for targeting areas of interest ([<xref ref-type="bibr" rid="B4">4</xref>]). Choosing the number of trees to ensure a balance between interpretability and algorithmic complexity remains one of the major challenges. Splitting the data into trees aims to reduce node impurity by using measures such as the Gini index or entropy. The optimal split is the one that maximizes the information gain.</p>
      </sec>
      <sec id="sec3dot2">
        <title>3.2. Logistic Regression</title>
        <p>Logistic regression is a primarily binary classification algorithm. It transforms a linear combination of variables into a probability using the sigmoid function. This probability is used to determine whether an example belongs to a given class. The characteristic parameters of a given sample are used to predict categorical outcomes ([<xref ref-type="bibr" rid="B5">5</xref>]). For multi-class problems, it is extended using the one-versus-all strategy. This strategy generates a model for each class using the softmax function ([<xref ref-type="bibr" rid="B2">2</xref>]). These extensions allow a probability vector, which sums to 1, to be assigned to several classes simultaneously whilst maintaining the model’s interpretability. However, performance depends on the choice of strategy and the quality of the data, particularly for imbalanced classes.</p>
      </sec>
      <sec id="sec3dot3">
        <title>3.3. Gaussian Naïve Bayes</title>
        <p>The Naive Bayes algorithm is a simple and effective classifier, well-suited to high-dimensional data such as that encountered in soil classification, particularly for food crops ([<xref ref-type="bibr" rid="B7">7</xref>]). Its Gaussian version (GNB) assumes that the features are continuous and follow a normal distribution. This simplifies calculations but limits its applicability to real-world data, which do not adhere to this assumption of a normal distribution. To overcome this limitation, a variant known as Stable Naive Bayes (SNB) replaces the Gaussian distribution with stable distributions, capable of modelling asymmetry and outliers in real-world data. These stable distributions generalize the normal distribution whilst retaining its property of stability under addition. This makes them suitable for real-world or non-symmetric data. This version is defined by four main parameters: α, which controls the tails or outliers; β, the asymmetry of the data; γ, their scale; and δ, the location. This offers increased flexibility for feature analysis ([<xref ref-type="bibr" rid="B19">19</xref>]). This extension of Naïve Bayes makes it more robust when dealing with complex distributions. Thus, SNB retains Bayesian simplicity whilst improving accuracy on real-world data.</p>
      </sec>
      <sec id="sec3dot4">
        <title>3.4. Adaboost Classifier</title>
        <p>This is an ensemble algorithm that combines weak classifiers to produce a more powerful predictive model. It operates through successive iterations, with each model being trained on weighted data. It demonstrates greater accuracy in lithology prediction ([<xref ref-type="bibr" rid="B15">15</xref>]). The subsequent iteration focuses on the misclassified samples from the previous one. The key hyperparameters of this estimator include the choice of base classifier, the number of iterations, and the learning rate, which controls the influence of each model. It is primarily used for binary or multi-class classification tasks. Its strength lies in its simplicity and efficiency, although it is sensitive to outliers in the measured data.</p>
      </sec>
      <sec id="sec3dot5">
        <title>3.5. Gradient Boosting Classifier</title>
        <p>Gradient Boosting is an ensemble method based on the principle of sequentially building weak models, known as bootstrapping. This process aims to correct the prediction errors of previous training runs. These models have improved the ability to predict the risk of landslides ([<xref ref-type="bibr" rid="B18">18</xref>]). There is a version for regression, but in the case of classification, the algorithm used is the Gradient Boosting Classifier. Thus, compared to other algorithms, Gradient Boosting has demonstrated good performance based on wave velocities, density and natural gamma to target mineral deposits ([<xref ref-type="bibr" rid="B1">1</xref>]). Where each new decision tree is trained to improve the separation between classes by following the gradient of the loss function. Both models apply the same logic, but regression tasks aim to progressively reduce the residual prediction error. The final model, in both classification and regression cases, is a weighted combination of several weak trees. This allows complex and non-linear relationships in real-world data to be captured. The key parameters are the loss function, the number of trees and the learning rate. These parameters must be carefully tuned to avoid overfitting. GBC and GBR offer great flexibility and remarkable performance.</p>
      </sec>
      <sec id="sec3dot6">
        <title>3.6. X Gradient Boosting Classifier</title>
        <p>XGB models are optimized variants of Gradient Boosting, designed to be faster and more efficient in terms of computational cost and algorithmic complexity on large datasets. As with the ensemble methods mentioned, they rely on the sequential construction of weak decision trees. Each new tree corrects the errors of the previous ones by following the gradient of the loss function. XGBoost introduces regularization mechanisms (L1 and L2) to reduce the risk of overfitting whilst improving the model’s generalization. In classification, it assigns probabilities to classes. Based on class probabilities, it makes a final decision, whereas for regression, the decision is based on an optimization of numerical predictions. For estimating above-ground biomass, XGBR has demonstrated its potential in terms of accuracy compared to other methods ([<xref ref-type="bibr" rid="B13">13</xref>]). Key parameters include those for gradient descent as well as metaheuristic optimization. This study demonstrates the ability of XGBC to predict improvements in academic achievement through time management ([<xref ref-type="bibr" rid="B9">9</xref>]). These results demonstrate that ensemble models are renowned for their performance, flexibility and ability to handle complex and large-scale data. In summary, XGBC and XGBR offer a combination of power, speed and robustness, which explains their widespread adoption in research, particularly in geoscience.</p>
      </sec>
      <sec id="sec3dot7">
        <title>3.7. Voting and Bagging Classifier</title>
        <p>The Voting Classifier is a technique based on the principle of the wisdom of crowds as applied to supervised models. It combines algorithms to produce predictions that are more robust than those obtained by a single model trained on the entire dataset. The decision is then made by majority vote (hard voting) or by combining probabilities (soft voting). Key parameters include the type of estimators and the weights assigned to them. Weighted voting is based on differential evolution, which improves classification accuracy and offers strong generalization ability as well as broad applicability ([<xref ref-type="bibr" rid="B20">20</xref>]). Bagging is based on the principle of stabilizing models by combining multiple predictors trained on different batches of samples drawn from the same dataset. Each subsample is obtained through bootstrap sampling, which introduces diversity and robustness into the constructed models. Bagging is a technique designed simply to reduce the prediction error of machine learning algorithms by reducing the variance of unstable prediction methods ([<xref ref-type="bibr" rid="B11">11</xref>]).</p>
      </sec>
    </sec>
    <sec id="sec4">
      <title>4. Result and Discussion</title>
      <p>In geochemistry, chemical elements act as indicators that provide clues about the physical, chemical or biological properties that led to the formation of the rock. The ratios between certain elements can indicate their relative abundance within a rock. <bold>Table 4</bold> shows the contributions of the scaled data for these markers in the soil.</p>
      <p><bold>Table 4</bold><bold>.</bold> Mean and variability of components by soil type.</p>
      <table-wrap id="tbl4">
        <label>Table 4</label>
        <table>
          <tbody>
            <tr>
              <td rowspan="2">
                <bold>Soil_Type</bold>
              </td>
              <td colspan="2">
                <bold>SiO</bold>
                <bold>
                  <sub>2</sub>
                </bold>
              </td>
              <td colspan="2">
                <bold>AL</bold>
                <bold>
                  <sub>2</sub>
                </bold>
                <bold>O</bold>
                <bold>
                  <sub>3</sub>
                </bold>
              </td>
              <td colspan="2">
                <bold>TiO</bold>
                <bold>
                  <sub>2</sub>
                </bold>
              </td>
              <td colspan="2">
                <bold>CaO</bold>
              </td>
              <td colspan="2">
                <bold>MgO</bold>
              </td>
              <td colspan="2">
                <bold>Fe</bold>
                <bold>
                  <sub>2</sub>
                </bold>
                <bold>O</bold>
                <bold>
                  <sub>3</sub>
                </bold>
              </td>
              <td colspan="2">
                <bold>K</bold>
                <bold>
                  <sub>2</sub>
                </bold>
                <bold>O</bold>
              </td>
              <td colspan="2">
                <bold>Na</bold>
                <bold>
                  <sub>2</sub>
                </bold>
                <bold>O</bold>
              </td>
            </tr>
            <tr>
              <td>
                <bold>mean</bold>
              </td>
              <td>
                <bold>std</bold>
              </td>
              <td>
                <bold>mean</bold>
              </td>
              <td>
                <bold>std</bold>
              </td>
              <td>
                <bold>mean</bold>
              </td>
              <td>
                <bold>std</bold>
              </td>
              <td>
                <bold>mean</bold>
              </td>
              <td>
                <bold>std</bold>
              </td>
              <td>
                <bold>mean</bold>
              </td>
              <td>
                <bold>std</bold>
              </td>
              <td>
                <bold>mean</bold>
              </td>
              <td>
                <bold>std</bold>
              </td>
              <td>
                <bold>mean</bold>
              </td>
              <td>
                <bold>std</bold>
              </td>
              <td>
                <bold>mean</bold>
              </td>
              <td>
                <bold>std</bold>
              </td>
            </tr>
            <tr>
              <td>
                <bold>Argile Blanche</bold>
              </td>
              <td>73.70</td>
              <td>15.44</td>
              <td>13.49</td>
              <td>9.41</td>
              <td>1.37</td>
              <td>0.58</td>
              <td>0.27</td>
              <td>0.17</td>
              <td>0.08</td>
              <td>0.07</td>
              <td>2.36</td>
              <td>5.01</td>
              <td>0.12</td>
              <td>0.05</td>
              <td>0.47</td>
              <td>0.12</td>
            </tr>
            <tr>
              <td>
                <bold>Argile Noire</bold>
              </td>
              <td>57.75</td>
              <td>6.55</td>
              <td>23.25</td>
              <td>3.79</td>
              <td>1.29</td>
              <td>0.31</td>
              <td>0.29</td>
              <td>0.09</td>
              <td>0.26</td>
              <td>0.16</td>
              <td>1.82</td>
              <td>0.60</td>
              <td>0.32</td>
              <td>0.08</td>
              <td>0.59</td>
              <td>0.09</td>
            </tr>
            <tr>
              <td>
                <bold>Argile Rouge</bold>
              </td>
              <td>68.32</td>
              <td>16.78</td>
              <td>11.87</td>
              <td>6.13</td>
              <td>1.44</td>
              <td>1.27</td>
              <td>1.53</td>
              <td>4.98</td>
              <td>0.30</td>
              <td>0.25</td>
              <td>11.03</td>
              <td>8.47</td>
              <td>0.24</td>
              <td>0.09</td>
              <td>0.73</td>
              <td>0.20</td>
            </tr>
            <tr>
              <td>
                <bold>Granite</bold>
              </td>
              <td>72.33</td>
              <td>9.06</td>
              <td>14.61</td>
              <td>3.08</td>
              <td>0.34</td>
              <td>0.36</td>
              <td>1.02</td>
              <td>4.33</td>
              <td>0.97</td>
              <td>3.68</td>
              <td>1.70</td>
              <td>1.75</td>
              <td>2.43</td>
              <td>1.55</td>
              <td>1.65</td>
              <td>1.09</td>
            </tr>
            <tr>
              <td>
                <bold>Sable</bold>
              </td>
              <td>73.32</td>
              <td>9.17</td>
              <td>13.81</td>
              <td>8.37</td>
              <td>0.64</td>
              <td>0.16</td>
              <td>1.78</td>
              <td>0.85</td>
              <td>0.45</td>
              <td>0.23</td>
              <td>2.42</td>
              <td>0.55</td>
              <td>0.33</td>
              <td>0.07</td>
              <td>0.60</td>
              <td>0.20</td>
            </tr>
          </tbody>
        </table>
      </table-wrap>
      <p>From this data, we will focus on three reports: </p>
      <p>1) Silica saturation index = SiO<sub>2</sub>/(SiO<sub>2</sub> + Al<sub>2</sub>O<sub>3</sub> + Fe<sub>2</sub>O<sub>3</sub>), </p>
      <p>2) Alkalinity index = Na<sub>2</sub>O + K<sub>2</sub>O/CaO + MgO,</p>
      <p>3) Iron index = Fe<sub>2</sub>O<sub>3</sub>/(SiO<sub>2</sub> + Al<sub>2</sub>O<sub>3</sub> + Fe<sub>2</sub>O<sub>3</sub>),</p>
      <p>4) Basic oxide index = CaO + MgO + Na<sub>2</sub>O + K<sub>2</sub>O.</p>
      <p>These indices characterize the cation saturation state of soil colloids. They have a significant influence on variations in the number of adsorbed cations. This makes it possible to compare the basic saturation index and the cation saturation index across a wide range of values. These indices enable the chemical state of a soil to be characterized, which determines the soil’s saturation. The indices provided indicate a methodological approach to quantifying soil saturation. They serve as indicators of soil moisture tension depending on their concentration. The multivariate datasets contain numerous variables that characterize the richness of the geochemical data. This richness enables us to explain complex soil behaviors that cannot be understood from a single observation. Multivariate methods allow us to simultaneously study variations in composition observed across multiple sites. It is highly likely that the observed variations are linked to a smaller number of underlying causes. By exploiting the correlations present in the data, we have been able to identify more parsimonious underlying models. The histograms on the diagonal show the distribution of each oxide according to soil type. Alumina (Al<sub>2</sub>O<sub>3</sub>) is concentrated in clays, reflecting the presence of aluminosilicates. </p>
      <p>Iron (Fe<sub>2</sub>O<sub>3</sub>) and titanium (TiO<sub>2</sub>) oxides are more pronounced in red clays, as shown in <xref ref-type="fig" rid="fig1">Figure 1</xref>. From the scatter plots, we can identify the major oxides, which are effective markers for soil differentiation. Silica distinguishes sands from clays. Meanwhile, the content of alumina and/or iron oxide indicates variations in the color of the clay or its formation conditions. Alumina, titanium oxide and iron oxide are strongly correlated. Meanwhile, sodium oxide and potassium oxide are also strongly correlated.</p>
      <fig id="fig1">
        <label>Figure 1</label>
        <graphic xlink:href="https://html.scirp.org/file/2173772-rId17.jpeg?20260528113301" />
      </fig>
      <p><bold>Figure 1</bold><bold>.</bold> Multivariate analysis of geochemical signals.</p>
      <sec id="sec4dot1">
        <title>4.1. Factor Analysis</title>
        <p>Before applying the machine learning model, the features must be rescaled through normalization. This involves rescaling the features in the dataset so that they are centered around the median using interquartile rescaling. Robust scaling algorithms enable features to be scaled in a way that is resilient to outliers in the measured data. This method is very similar to the MinMax scaling method, but its distinctive feature is that it uses interquartile ranges rather than the minimum and maximum values used in MinMax scaling. This is a scaling algorithm that reduces the influence of outliers by using the median and the extremes of the data based on the quantile range. For this scaling, each value of x is calculated relative to the first quartile <italic>Q</italic><sub>1</sub> and then normalized relative to the interquartile range, as shown in the following equation: </p>
        <disp-formula id="FD1">
          <mml:math display="inline">
            <mml:mrow>
              <mml:mi>x</mml:mi>
              <mml:mo>=</mml:mo>
              <mml:mfrac>
                <mml:mrow>
                  <mml:msub>
                    <mml:mi>x</mml:mi>
                    <mml:mi>i</mml:mi>
                  </mml:msub>
                  <mml:mo>−</mml:mo>
                  <mml:msub>
                    <mml:mi>Q</mml:mi>
                    <mml:mn>1</mml:mn>
                  </mml:msub>
                </mml:mrow>
                <mml:mrow>
                  <mml:msub>
                    <mml:mi>Q</mml:mi>
                    <mml:mn>1</mml:mn>
                  </mml:msub>
                  <mml:mo>−</mml:mo>
                  <mml:msub>
                    <mml:mi>Q</mml:mi>
                    <mml:mn>3</mml:mn>
                  </mml:msub>
                </mml:mrow>
              </mml:mfrac>
            </mml:mrow>
          </mml:math>
        </disp-formula>
        <p>An examination of the factors to interpret the underlying properties they represent shows, as illustrated in <xref ref-type="fig" rid="fig2">Figure 2</xref>, the following contributions by groups of variables:</p>
        <fig id="fig2">
          <label>Figure 2</label>
          <graphic xlink:href="https://html.scirp.org/file/2173772-rId20.jpeg?20260528113301" />
        </fig>
        <p><bold>Figure 2</bold><bold>.</bold> Factors interpretation.</p>
        <p>The variable components Fe<sub>2</sub>O<sub>3</sub>, MgO, CaO and IndFer reflect a dimension linked to coloring oxides and ferric indices. Soils with high levels of these components are often associated with the clays typical of central Senegal.The Al<sub>2</sub>O<sub>3</sub>, K<sub>2</sub>O and Na<sub>2</sub>O components indicate basic soils. These soils are the opposite of those dominated by compounds with a high iron index, reflecting a chemical contrast between ferric environments and the basic environments of the south.</p>
        <p>These factors indicate ferric, basic, alkaline or siliceous soil types, distinguishing between soils rich in iron oxides and those that are more basic, magnesium-rich or rich in sodium and potassium oxides. In practice, this reveals an important geochemical dimension for soil classification. In this study, exploratory factor analysis is used to reduce the number of sample characteristics and identify the underlying factors.</p>
      </sec>
      <sec id="sec4dot2">
        <title>4.2. Soil Clustering</title>
        <p>Factor analysis of the data reduced the initial set of nine chemical characteristics to a set of five. These characteristics, grouped into adjacent sub-groups, enable four new characteristics to be derived for core samples based on their shared geochemical traits. These groups represent similar geochemical facies. Cluster analysis is a suitable approach for assigning a common facies label to similar samples. This clustering makes it easy to distinguish between samples. This batch analysis is a form of classification that falls broadly under the category of unsupervised machine learning. It is an approach used to infer a structure from the data measured within the dataset itself. <xref ref-type="fig" rid="fig3">Figure 3</xref> illustrates the extent to which the silica and alumina content in soils in Senegal significantly influences their characterization. The data are reduced to principal components using the K-means approach, as shown in <xref ref-type="fig" rid="fig4">Figure 4</xref>. Each sample is thus assigned to a cluster representing geochemical facies. Examination of the signatures reveals characteristics specific to each cluster; for example, the iron index in the case of red clay soils. All these various transformations were applied to the raw data before it was divided into training, test, and validation sets.</p>
        <fig id="fig3">
          <label>Figure 3</label>
          <graphic xlink:href="https://html.scirp.org/file/2173772-rId21.jpeg?20260528113301" />
        </fig>
        <p><bold>Figure 3</bold><bold>.</bold> Retention threshold of the eigenvalue distribution.</p>
        <fig id="fig4">
          <label>Figure 4</label>
          <graphic xlink:href="https://html.scirp.org/file/2173772-rId22.jpeg?20260528113301" />
        </fig>
        <p><bold>Figure 4</bold><bold>.</bold> Comparison of k-means++ and random initialization.</p>
      </sec>
      <sec id="sec4dot3">
        <title>4.3. Comparison of Classification Model Performance</title>
        <p>In the context of the experiment, classification accuracy depends on the estimators’ hyperparameters and their tuning. <bold>Table 5</bold> compares the classification results, summarizing the performance of our models in terms of precision, recall, F1-score and accuracy. We observe that the random forest achieves the best overall classification performance with a precision of 99% and an accuracy of the same value. It is followed by X gradient boosting and the bagging classifier, which are very close behind (98% overall precision each). Logistic regression and the gradient boosting classifier follow closely in terms of overall precision. The Adaboost estimator provides the lowest overall accuracy in our classification.</p>
        <p><bold>Table 5</bold><bold>.</bold> Comparative analysis of supervised methods.</p>
        <table-wrap id="tbl5">
          <label>Table 5</label>
          <table>
            <tbody>
              <tr>
                <td>
                </td>
                <td>Setting</td>
                <td>P</td>
                <td>R</td>
                <td>F1</td>
                <td>Acc</td>
              </tr>
              <tr>
                <td>Regression Logistique (rl)</td>
                <td>C = 0.9</td>
                <td>97</td>
                <td>98</td>
                <td>98</td>
                <td>97</td>
              </tr>
              <tr>
                <td>Random Forest (rf)</td>
                <td>n_estimators = 200</td>
                <td>99</td>
                <td>99</td>
                <td>99</td>
                <td>99</td>
              </tr>
              <tr>
                <td>Bagging Classifier (bgc)</td>
                <td>n_estimators = 200</td>
                <td>97</td>
                <td>98</td>
                <td>98</td>
                <td>98</td>
              </tr>
              <tr>
                <td>Adaboost Classifer (abc)</td>
                <td>Algorithme = SAMME</td>
                <td>64</td>
                <td>76</td>
                <td>78</td>
                <td>76</td>
              </tr>
              <tr>
                <td>Gradient Boosting Classifier (gbc)</td>
                <td>Log_loss</td>
                <td>85</td>
                <td>92</td>
                <td>94</td>
                <td>97</td>
              </tr>
              <tr>
                <td>X gradient Boosting Classifier (xgbc)</td>
                <td>StritifiedKfold</td>
                <td>91</td>
                <td>97</td>
                <td>94</td>
                <td>98</td>
              </tr>
              <tr>
                <td>Voting Classifier (vc)</td>
                <td>rl, rf, gnb,</td>
                <td>96</td>
                <td>95</td>
                <td>95</td>
                <td>95</td>
              </tr>
            </tbody>
          </table>
        </table-wrap>
        <p>k = kernel; Acc = Accuracy; P = Precision; R = Recall; F1 = F1-score.</p>
      </sec>
    </sec>
    <sec id="sec5">
      <title>5. Conclusion</title>
      <p>Over the past decade, geoscientific modelling approaches have attracted increasing interest, gradually shifting from statistical models and “box” models to Earth system models. However, statistical models often rely on a limited number of environmental factors to describe physical or chemical processes. Due to dynamic processes linked to climate and human activities, these parameters are non-linear as they vary considerably from one spatio-temporal domain to another. In this study, we introduced an ensemble algorithm framework to address the classification problem; this helps to facilitate soil recognition criteria by individual models and enhances the ability of specific environmental variables used in these models to better characterize the processes. The results lead to the following conclusions:</p>
      <p>1) Ensemble learning methods have clearly demonstrated their potential for soil classification. Compared with traditional ensemble approaches, random forests achieved higher values for overall accuracy, precision, recall and F1-score in the estimation of soil classification parameters.</p>
      <p>2) Multimodal datasets contain a high degree of similarity, which leads the model to make classification errors. Reducing this similarity through an in-depth exploration of soil properties would improve classification accuracy. </p>
      <p>3) Soil is defined by these mechanical, physical or chemical properties. They help to improve the models, but rely on the effectiveness of machine learning classifiers. These classifiers require significant human intervention, such as the tuning of hyperparameters.</p>
      <p>4) Models based purely on deterministic statistical data may appear to be incompatible with the known physical laws governing the Earth. This leads to the conclusion that the performance of such models does not accurately reflect the rheological reality of soil properties. The determination of physical stress by parameter would be an excellent tool for validating estimators in geoscientific classification. </p>
    </sec>
  </body>
  <back>
    <ref-list>
      <title>References</title>
      <ref id="B1">
        <label>1.</label>
        <citation-alternatives>
          <mixed-citation publication-type="confproc">Atita, O., Durrheim, R., &amp; Saffou, E. (2022). Evaluation of Machine Learning Algorithms for the Classification of Lithology Using Geophysical Logs. In <italic>NSG2022 4th Conference on Geophysics for Mineral Exploration and Mining</italic> (pp. 1-5). European Association of Geoscientists &amp; Engineers. https://doi.org/10.3997/2214-4609.202220121 <pub-id pub-id-type="doi">10.3997/2214-4609.202220121</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.3997/2214-4609.202220121">https://doi.org/10.3997/2214-4609.202220121</ext-link></mixed-citation>
          <element-citation publication-type="confproc">
            <person-group person-group-type="author">
              <string-name>Atita, O.</string-name>
              <string-name>Durrheim, R.</string-name>
              <string-name>Saffou, E.</string-name>
            </person-group>
            <year>2022</year>
            <pub-id pub-id-type="doi">10.3997/2214-4609.202220121</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B2">
        <label>2.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Bewick, V., Cheek, L., &amp; Ball, J. (2005). <italic>Statistics</italic><italic>Review 14: Logistic Regression</italic>. <italic>Critical Care</italic><italic>,</italic><italic>9,</italic>Article Number 112. https://link.springer.com/article/10.1186/cc3045 <pub-id pub-id-type="doi">10.1186/cc3045</pub-id><pub-id pub-id-type="pmid">15693993</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1186/cc3045">https://doi.org/10.1186/cc3045</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Bewick, V.</string-name>
              <string-name>Cheek, L.</string-name>
              <string-name>Ball, J.</string-name>
            </person-group>
            <year>2005</year>
            <elocation-id>Number</elocation-id>
            <pub-id pub-id-type="doi">10.1186/cc3045</pub-id>
            <pub-id pub-id-type="pmid">15693993</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B3">
        <label>3.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Bose, S., &amp; Bose, S. (2025). Random Forests: The Wisdom of Crowds in Action. <italic>Journal</italic><italic>of</italic><italic>Emerging</italic><italic>Trends</italic><italic>in</italic><italic>Computer</italic><italic>Science</italic><italic>and</italic><italic>Applications,</italic><italic>1,</italic> 67-91. https://doi.org/10.65525/jetcsa.v1i1.5 <pub-id pub-id-type="doi">10.65525/jetcsa.v1i1.5</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.65525/jetcsa.v1i1.5">https://doi.org/10.65525/jetcsa.v1i1.5</ext-link></mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Bose, S.</string-name>
              <string-name>Bose, S.</string-name>
            </person-group>
            <year>2025</year>
            <pub-id pub-id-type="doi">10.65525/jetcsa.v1i1.5</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B4">
        <label>4.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Darijani, M., Farquharson, C. G., &amp; Perrouty, S. (2022). A Random Forest Approach to Predict Geology from Geophysics in the Pontiac Subprovince, Canada. <italic>Canadian</italic><italic>Journal</italic><italic>of</italic><italic>Earth</italic><italic>Sciences,</italic><italic>59,</italic> 489-503. https://doi.org/10.1139/cjes-2021-0089 <pub-id pub-id-type="doi">10.1139/cjes-2021-0089</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1139/cjes-2021-0089">https://doi.org/10.1139/cjes-2021-0089</ext-link></mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Darijani, M.</string-name>
              <string-name>Farquharson, C.</string-name>
              <string-name>Perrouty, S.</string-name>
              <string-name>Subprovince, C</string-name>
            </person-group>
            <year>2022</year>
            <pub-id pub-id-type="doi">10.1139/cjes-2021-0089</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B5">
        <label>5.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Giasson, E., Clarke, R. T., Inda Junior, A. V., Merten, G. H., &amp; Tornquist, C. G. (2006). Digital Soil Mapping Using Multiple Logistic Regression on Terrain Parameters in Southern Brazil. <italic>Scientia</italic><italic>Agricola,</italic><italic>63,</italic> 262-268. https://doi.org/10.1590/s0103-90162006000300008 <pub-id pub-id-type="doi">10.1590/s0103-90162006000300008</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1590/s0103-90162006000300008">https://doi.org/10.1590/s0103-90162006000300008</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Giasson, E.</string-name>
              <string-name>Clarke, R.</string-name>
              <string-name>Junior, A.</string-name>
              <string-name>Merten, G.</string-name>
              <string-name>Tornquist, C.</string-name>
            </person-group>
            <year>2006</year>
            <pub-id pub-id-type="doi">10.1590/s0103-90162006000300008</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B6">
        <label>6.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Johnson, T., &amp; Dasu, T. (2003). <italic>Data Quality and Data Cleaning: An Overview</italic>. https://doi.org/10.1145/872874.872875 <pub-id pub-id-type="doi">10.1145/872874.872875</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1145/872874.872875">https://doi.org/10.1145/872874.872875</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Johnson, T.</string-name>
              <string-name>Dasu, T.</string-name>
            </person-group>
            <year>2003</year>
            <pub-id pub-id-type="doi">10.1145/872874.872875</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B7">
        <label>7.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Kom, N. S., &amp; Kom, M. (2024). <italic>Classification System for Soil Types Suitable for Food Crops using Naïve Bayes Method</italic>. <italic>Sistemasi</italic><italic>:</italic><italic>Jurnal</italic><italic>Sistem</italic><italic>Informasi</italic><italic>, 13,</italic>1102-1113. https://doi.org/10.32520/stmsi.v13i3.3956 https://sistemasi.ftik.unisi.ac.id/index.php/stmsi/article/view/24 <pub-id pub-id-type="doi">10.32520/stmsi.v13i3.3956</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.32520/stmsi.v13i3.3956">https://doi.org/10.32520/stmsi.v13i3.3956</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Kom, N.</string-name>
              <string-name>Kom, M.</string-name>
            </person-group>
            <year>2024</year>
            <pub-id pub-id-type="doi">10.32520/stmsi.v13i3.3956</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B8">
        <label>8.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Kumar, S., Machireddy, J. S., Sankaran, T., &amp; Sholapurapu, P. K. (2025). Integration of Machine Learning and Data Science for Optimized Decision-Making in Computer Applications and Engineering. <italic>Journal of Information Systems Engineering and Management</italic><italic>, 10,</italic> 748-759. https://jisem-journal.com/index.php/journal/article/view/8990</mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Kumar, S.</string-name>
              <string-name>Machireddy, J.</string-name>
              <string-name>Sankaran, T.</string-name>
              <string-name>Sholapurapu, P.</string-name>
            </person-group>
            <year>2025</year>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B9">
        <label>9.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Li, S. (2024). Exploring the Impact of Time Management Skills on Academic Achievement with an XGBC Model and Metaheuristic Algorithm. <italic>International Journal of Advanced Computer Science and Applications, 15,</italic> 107-119. https://doi.org/10.14569/ijacsa.2024.0150710 <pub-id pub-id-type="doi">10.14569/ijacsa.2024.0150710</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.14569/ijacsa.2024.0150710">https://doi.org/10.14569/ijacsa.2024.0150710</ext-link></mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Li, S.</string-name>
            </person-group>
            <year>2024</year>
            <pub-id pub-id-type="doi">10.14569/ijacsa.2024.0150710</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B10">
        <label>10.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Lim, C. S., Mohamad, E. T., Motahari, M. R., Armaghani, D. J., &amp; Saad, R. (2020). Machine Learning Classifiers for Modeling Soil Characteristics by Geophysics Investigations: A Comparative Study. <italic>Applied</italic><italic>Sciences,</italic><italic>10,</italic> Article 5734. https://doi.org/10.3390/app10175734 <pub-id pub-id-type="doi">10.3390/app10175734</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.3390/app10175734">https://doi.org/10.3390/app10175734</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Lim, C.</string-name>
              <string-name>Mohamad, E.</string-name>
              <string-name>Motahari, M.</string-name>
              <string-name>Armaghani, D.</string-name>
              <string-name>Saad, R.</string-name>
            </person-group>
            <year>2020</year>
            <elocation-id>5734</elocation-id>
            <pub-id pub-id-type="doi">10.3390/app10175734</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B11">
        <label>11.</label>
        <citation-alternatives>
          <mixed-citation publication-type="confproc">Liu, H., Jia, J., &amp; Gong, N. Z. (2020). On the Intrinsic Differential Privacy of Bagging. <italic>Proceedings of the Thirtieth International Joint Conference on Artificial Intelligence Main Track</italic>, Yokohama, 7-15 January 2021, 2730-2736. https://doi.org/10.24963/ijcai.2021/376 <pub-id pub-id-type="doi">10.24963/ijcai.2021/376</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.24963/ijcai.2021/376">https://doi.org/10.24963/ijcai.2021/376</ext-link></mixed-citation>
          <element-citation publication-type="confproc">
            <person-group person-group-type="author">
              <string-name>Liu, H.</string-name>
              <string-name>Jia, J.</string-name>
              <string-name>Gong, N.</string-name>
              <string-name>Track, Y</string-name>
            </person-group>
            <year>2020</year>
            <pub-id pub-id-type="doi">10.24963/ijcai.2021/376</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B12">
        <label>12.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Naser, M. Z. (2026). When to Use Machine Learning? And Which Problems Stand to Benefit? <italic>Urban</italic><italic>Lifeline,</italic><italic>4,</italic> Article No. 3. https://doi.org/10.1007/s44285-025-00054-3 <pub-id pub-id-type="doi">10.1007/s44285-025-00054-3</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1007/s44285-025-00054-3">https://doi.org/10.1007/s44285-025-00054-3</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Naser, M.</string-name>
            </person-group>
            <year>2026</year>
            <elocation-id>No</elocation-id>
            <pub-id pub-id-type="doi">10.1007/s44285-025-00054-3</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B13">
        <label>13.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Pham, T. D., Le, N. N., Ha, N. T., Nguyen, L. V., Xia, J., Yokoya, N. et al. (2020). Estimating Mangrove Above-Ground Biomass Using Extreme Gradient Boosting Decision Trees Algorithm with Fused Sentinel-2 and ALOS-2 PALSAR-2 Data in Can Gio Biosphere Reserve, Vietnam. <italic>Remote Sensing, 12,</italic> Article 777. https://doi.org/10.3390/rs12050777 <pub-id pub-id-type="doi">10.3390/rs12050777</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.3390/rs12050777">https://doi.org/10.3390/rs12050777</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Pham, T.</string-name>
              <string-name>Le, N.</string-name>
              <string-name>Ha, N.</string-name>
              <string-name>Nguyen, L.</string-name>
              <string-name>Xia, J.</string-name>
              <string-name>Yokoya, N.</string-name>
              <string-name>Reserve, V</string-name>
            </person-group>
            <year>2020</year>
            <elocation-id>777</elocation-id>
            <pub-id pub-id-type="doi">10.3390/rs12050777</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B14">
        <label>14.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Reich, Y., Medina, M. A., Shieh, T., &amp; Jacobs, T. L. (1996). Modeling and Debugging Engineering Decision Procedures with Machine Learning. <italic>Journal</italic><italic>of</italic><italic>Computing</italic><italic>in</italic><italic>Civil</italic><italic>Engineering,</italic><italic>10,</italic> 157-166. https://doi.org/10.1061/(asce)0887-3801(1996)10:2(157) <pub-id pub-id-type="doi">10.1061/(asce)0887-3801(1996)10:2(157)</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1061/(asce)0887-3801(1996)10:2(157)">https://doi.org/10.1061/(asce)0887-3801(1996)10:2(157)</ext-link></mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Reich, Y.</string-name>
              <string-name>Medina, M.</string-name>
              <string-name>Shieh, T.</string-name>
              <string-name>Jacobs, T.</string-name>
            </person-group>
            <year>1996</year>
            <volume>3801</volume>
            <issue>1996</issue>
            <fpage>2</fpage>
            <pub-id pub-id-type="doi">10.1061/(asce)0887-3801(1996)10:2(157)</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B15">
        <label>15.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Sun, Y., Pang, S., &amp; Zhang, Y. (2024). Application of Adaboost-Transformer Algorithm for Lithology Identification Based on Well Logging Data. <italic>IEEE</italic><italic>Geoscience</italic><italic>and</italic><italic>Remote</italic><italic>Sensing Letter</italic><italic>s, 21,</italic> 1-5. https://doi.org/10.1109/lgrs.2024.3372513 <pub-id pub-id-type="doi">10.1109/lgrs.2024.3372513</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1109/lgrs.2024.3372513">https://doi.org/10.1109/lgrs.2024.3372513</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Sun, Y.</string-name>
              <string-name>Pang, S.</string-name>
              <string-name>Zhang, Y.</string-name>
            </person-group>
            <year>2024</year>
            <pub-id pub-id-type="doi">10.1109/lgrs.2024.3372513</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B16">
        <label>16.</label>
        <citation-alternatives>
          <mixed-citation publication-type="web">Tripathi, D., Alonso-Pérez, M. O., &amp; Tiwari, D. K. (2018). Rapid Diagnosis of Soil Nutrients Using Microscopic Techniques. <italic>Microscopy and Microanalysis, 24,</italic> 680-681. https://academic.oup.com/mam/article/24/S1/680/6945654</mixed-citation>
          <element-citation publication-type="web">
            <person-group person-group-type="author">
              <string-name>Tripathi, D.</string-name>
              <string-name>Tiwari, D.</string-name>
            </person-group>
            <year>2018</year>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B17">
        <label>17.</label>
        <citation-alternatives>
          <mixed-citation publication-type="confproc">Varga, M., &amp; Kurko, K. (2010). Some Methods of Data Improvement in EIS. In <italic>2010 32nd International Conference on Information Technology Interfaces (ITI)</italic> .</mixed-citation>
          <element-citation publication-type="confproc">
            <person-group person-group-type="author">
              <string-name>Varga, M.</string-name>
              <string-name>Kurko, K.</string-name>
            </person-group>
            <year>2010</year>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B18">
        <label>18.</label>
        <citation-alternatives>
          <mixed-citation publication-type="confproc">Yasin, K. C., Niasari, S. W., Nukman, M., Sudarmaji,, Irnaka, T. M., Hartantyo, E. et al. (2025). Integration of Geoscientific Data Mining and GBM Hyperparameter Optimization to Improve the Accuracy of Landslide Hazard Prediction. In <italic>2025 5th International Conference on Artificial Intelligence, Big Data and Algorithms (CAIBDA)</italic>(pp. 367-370). IEEE. https://doi.org/10.1109/caibda65784.2025.11183569 <pub-id pub-id-type="doi">10.1109/caibda65784.2025.11183569</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1109/caibda65784.2025.11183569">https://doi.org/10.1109/caibda65784.2025.11183569</ext-link></mixed-citation>
          <element-citation publication-type="confproc">
            <person-group person-group-type="author">
              <string-name>Yasin, K.</string-name>
              <string-name>Niasari, S.</string-name>
              <string-name>Nukman, M.</string-name>
              <string-name>Irnaka, T.</string-name>
              <string-name>Hartantyo, E.</string-name>
              <string-name>Intelligence, B</string-name>
            </person-group>
            <year>2025</year>
            <pub-id pub-id-type="doi">10.1109/caibda65784.2025.11183569</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B19">
        <label>19.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Zeng, X., &amp; Pinsky, E. (2025). Stable Distribution Naive Bayes Achieves Higher Accuracy than Traditional Naive Bayes Classification. <italic>International Journal on Cybernetics &amp; Informatics, 14,</italic> 71-86. https://doi.org/10.5121/ijci.2025.140205 <pub-id pub-id-type="doi">10.5121/ijci.2025.140205</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.5121/ijci.2025.140205">https://doi.org/10.5121/ijci.2025.140205</ext-link></mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Zeng, X.</string-name>
              <string-name>Pinsky, E.</string-name>
            </person-group>
            <year>2025</year>
            <pub-id pub-id-type="doi">10.5121/ijci.2025.140205</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B20">
        <label>20.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Zhang, Y., Zhang, H., Cai, J., &amp; Yang, B. (2014). A Weighted Voting Classifier Based on Differential Evolution. <italic>Abstract and Applied Analysis</italic><italic>,</italic><italic>2014</italic><italic>,</italic>376950. https://onlinelibrary.wiley.com/doi/10.1155/2014/376950 <pub-id pub-id-type="doi">10.1155/2014/376950</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1155/2014/376950">https://doi.org/10.1155/2014/376950</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Zhang, Y.</string-name>
              <string-name>Zhang, H.</string-name>
              <string-name>Cai, J.</string-name>
              <string-name>Yang, B.</string-name>
            </person-group>
            <year>2014</year>
            <pub-id pub-id-type="doi">10.1155/2014/376950</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B21">
        <label>21.</label>
        <citation-alternatives>
          <mixed-citation publication-type="confproc">Zhao, Z., Shi, D., Huo, H., &amp; Fang, T. (2018). Feature Encoding Methods Evaluation Based on Multiple Kernel Learning. In <italic>Proceedings of the 2018 10th International Conference on Machine Learning and Computing (pp. 209-213)</italic>. ACM. https://doi.org/10.1145/3195106.3195152 <pub-id pub-id-type="doi">10.1145/3195106.3195152</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1145/3195106.3195152">https://doi.org/10.1145/3195106.3195152</ext-link></mixed-citation>
          <element-citation publication-type="confproc">
            <person-group person-group-type="author">
              <string-name>Zhao, Z.</string-name>
              <string-name>Shi, D.</string-name>
              <string-name>Huo, H.</string-name>
              <string-name>Fang, T.</string-name>
            </person-group>
            <year>2018</year>
            <pub-id pub-id-type="doi">10.1145/3195106.3195152</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
    </ref-list>
  </back>
</article>