<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.4 20241031//EN" "JATS-journalpublishing1-4.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article" dtd-version="1.4" xml:lang="en">
  <front>
    <journal-meta>
      <journal-id journal-id-type="publisher-id">ojg</journal-id>
      <journal-title-group>
        <journal-title>Open Journal of Geology</journal-title>
      </journal-title-group>
      <issn pub-type="epub">2161-7589</issn>
      <issn pub-type="ppub">2161-7570</issn>
      <publisher>
        <publisher-name>Scientific Research Publishing</publisher-name>
      </publisher>
    </journal-meta>
    <article-meta>
      <article-id pub-id-type="doi">10.4236/ojg.2026.169029</article-id>
      <article-id pub-id-type="publisher-id">ojg-154038</article-id>
      <article-categories>
        <subj-group>
          <subject>Article</subject>
        </subj-group>
        <subj-group>
          <subject>Earth</subject>
          <subject>Environmental Sciences</subject>
        </subj-group>
      </article-categories>
      <title-group>
        <article-title>Explaining the Performance Hierarchy of Linear Regression, Support Vector Regression, and Random Forest on Physics-Generated PVT Data: A Real-Gas-Law and Gauss-Markov Perspective</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author" corresp="yes">
          <contrib-id contrib-id-type="orcid">0009-0003-7013-7168</contrib-id>
          <name name-style="western">
            <surname>Hie</surname>
            <given-names>Gnessoa Rene</given-names>
          </name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <name name-style="western">
            <surname>Ajienka</surname>
            <given-names>Joseph Atubokiki</given-names>
          </name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <name name-style="western">
            <surname>Amieibibama</surname>
            <given-names>Joseph</given-names>
          </name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
      </contrib-group>
      <aff id="aff1"><label>1</label> Department of Petroleum and Gas Engineering, University of Port Harcourt, Port Harcourt, Nigeria </aff>
      <author-notes>
        <fn fn-type="conflict" id="fn-conflict">
          <p>The authors declare no conflicts of interest regarding the publication of this paper.</p>
        </fn>
      </author-notes>
      <pub-date pub-type="epub">
        <day>17</day>
        <month>09</month>
        <year>2026</year>
      </pub-date>
      <pub-date pub-type="collection">
        <month>09</month>
        <year>2026</year>
      </pub-date>
      <volume>16</volume>
      <issue>09</issue>
      <fpage>577</fpage>
      <lpage>590</lpage>
      <history>
        <date date-type="received">
          <day>16</day>
          <month>08</month>
          <year>2026</year>
        </date>
        <date date-type="accepted">
          <day>18</day>
          <month>09</month>
          <year>2026</year>
        </date>
        <date date-type="published">
          <day>21</day>
          <month>09</month>
          <year>2026</year>
        </date>
      </history>
      <permissions>
        <copyright-statement>© 2026 by the authors and Scientific Research Publishing Inc.</copyright-statement>
        <copyright-year>2026</copyright-year>
        <license license-type="open-access">
          <license-p> This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license ( <ext-link ext-link-type="uri" xlink:href="https://creativecommons.org/licenses/by/4.0/">https://creativecommons.org/licenses/by/4.0/</ext-link> ). </license-p>
        </license>
      </permissions>
      <self-uri content-type="doi" xlink:href="https://doi.org/10.4236/ojg.2026.169029">https://doi.org/10.4236/ojg.2026.169029</self-uri>
      <abstract>
        <p>Machine learning (ML) regression models are increasingly proposed as fast surrogates for deterministic Pressure-Volume-Temperature (PVT) simulation in reservoir and production engineering, yet the choice among competing algorithms is rarely justified on physical or statistical grounds. This study trains, evaluates, and compares three regression algorithms, Linear Regression (LR), Support Vector Regression (SVR), and Random Forest Regression (RF), on a PVT dataset generated from an industry-standard Integrated Production Modelling (IPM) software platform (PROSPER) using the Glaso correlation for a representative mature onshore Niger Delta oil well. Gas density was predicted from pressure and associated fluid properties across a depletion range of 1000 - 3500 psig. Following a structured preprocessing pipeline (exploratory data analysis, Isolation Forest anomaly removal, Min-Max normalization, K-means regime clustering, and time-series smoothing), the three algorithms were trained on an 80:20 split and evaluated using Root Mean Squared Error (RMSE) and the coefficient of determination (R<sup>2</sup>). Linear Regression achieved the highest accuracy (R<sup>2</sup> = 0.9954, RMSE = 0.2240 lb/ft<sup>3</sup>), followed by SVR (R<sup>2</sup> = 0.9408, RMSE = 0.8042 lb/ft<sup>3</sup>) and Random Forest (R<sup>2</sup> = 0.8389, RMSE = 1.4326 lb/ft<sup>3</sup>). Physically consistent extrapolation beyond the training range (correctly predicting monotonically increasing gas density at 4000 and 5000 psig) confirmed that the Linear Regression model had captured the governing thermodynamic relationship rather than an artefact of the training data. We show that this performance hierarchy is not an empirical coincidence but a direct, predictable consequence of two established results: the real gas law, which imposes a near-linear pressure-density relationship (Pearson r = 1.00) at approximately constant reservoir temperature, and the Gauss-Markov theorem, which guarantees that Ordinary Least Squares is the best linear unbiased estimator when the true relationship is linear. A bias-variance decomposition further explains why SVR’s kernel mismatch and Random Forest’s piecewise-constant approximation both add prediction variance without a compensating reduction in bias on this near-linear, physics-generated data. These findings provide petroleum engineers and applied ML practitioners with a principled, physics-grounded criterion for algorithm selection when training data are generated from simulator-based PVT correlations, rather than defaulting to the most complex available model.</p>
      </abstract>
      <kwd-group kwd-group-type="author-generated" xml:lang="en">
        <kwd>Machine Learning</kwd>
        <kwd>PVT Modelling</kwd>
        <kwd>Linear Regression</kwd>
        <kwd>Support Vector Regression</kwd>
        <kwd>Random Forest</kwd>
        <kwd>Gauss-Markov Theorem</kwd>
        <kwd>Real Gas Law</kwd>
        <kwd>Physics-Informed Machine Learning</kwd>
        <kwd>Reservoir Engineering</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec1">
      <title>1. Introduction</title>
      <p>Machine learning (ML) has been increasingly adopted in petroleum engineering as a computationally efficient alternative to physics-based simulation for tasks including production forecasting, well performance optimization, and fluid property prediction [<xref ref-type="bibr" rid="B1">1</xref>][<xref ref-type="bibr" rid="B2">2</xref>]. A comprehensive review of 166 published studies by Tariq <italic>et</italic><italic>al.</italic> [<xref ref-type="bibr" rid="B2">2</xref>] found that ensemble and deep-learning architectures dominate the recent literature but also identified the absence of physical consistency as the most frequently reported limitation of published ML models: fewer than one in five reviewed studies explicitly addressed the physical plausibility of model predictions. A model that fits historical data well but violates a governing physical law, for example, predicting that gas density decreases as pressure increases, cannot be trusted for engineering decision-making regardless of its statistical goodness of fit [<xref ref-type="bibr" rid="B3">3</xref>].</p>
      <p>The physics-informed machine learning (PIML) paradigm has emerged as a principled response to this risk, integrating physical knowledge into the ML workflow either by embedding governing equations directly into the training objective, as in Physics-Informed Neural Networks [<xref ref-type="bibr" rid="B4">4</xref>], or by incorporating physical priors into model architecture and training data [<xref ref-type="bibr" rid="B3">3</xref>][<xref ref-type="bibr" rid="B5">5</xref>]. A comparatively underexploited route to physical consistency is to generate the ML training data itself from a validated, physics-based simulator, so that the training distribution is thermodynamically and fluid-mechanically consistent by construction, and any model trained on it inherits this consistency [<xref ref-type="bibr" rid="B6">6</xref>]. This approach requires no modification to standard ML algorithms and is directly compatible with existing commercial production modelling software, making it readily transferable to industry practice.</p>
      <p>This raises a question that has received little systematic attention in the petroleum engineering ML literature: when training data are generated from a physics-based correlation rather than sampled directly from noisy field measurements, does the statistical structure imposed by the underlying physics predictably favor one class of ML algorithm over another, and can that preference be explained rather than merely observed? Existing comparative studies typically report which algorithm performed best on a given dataset [<xref ref-type="bibr" rid="B7">7</xref>] without connecting the result to the governing physics of the data-generating process or to the statistical theory of the estimators being compared.</p>
      <p>This study addresses that question directly. Using a PVT dataset generated from an industry-standard Integrated Production Modelling (IPM) platform (PROSPER) with the Glaso [<xref ref-type="bibr" rid="B8">8</xref>] correlation for a representative mature onshore Niger Delta well, we train, evaluate, and compare three regression algorithms representing three distinct algorithmic paradigms, Linear Regression (parametric linear modelling), Support Vector Regression (kernel-based regularized regression), and Random Forest Regression (ensemble decision-tree aggregation), on the task of predicting gas density from pressure and associated fluid properties. We then provide a physical and statistical explanation, grounded in the real gas law and the Gauss-Markov theorem, for the resulting performance hierarchy. The specific objective of this study is to train, evaluate, and compare the three algorithms using RMSE and R<sup>2</sup> as performance metrics, and to explain the physical and statistical basis of the observed performance hierarchy.</p>
      <p>The remainder of this paper is organized as follows. Section 2 describes the physics-based data generation procedure, the preprocessing pipeline, and the model training and evaluation methodology. Section 3 presents the comparative performance results. Section 4 discusses the physical and statistical explanation for the observed hierarchy and its practical implications for algorithm selection in petroleum engineering ML applications. Section 5 concludes.</p>
    </sec>
    <sec id="sec2">
      <title>2. Materials and Methods</title>
      <sec id="sec2dot1">
        <title>2.1. Study Context and Physics-Based Data Generation</title>
        <p>The dataset used in this study was generated from a production system model of a representative mature onshore oil well in the Niger Delta petroleum province of Nigeria, constructed in PROSPER, an industry-standard Integrated Production Modelling software platform. Pressure-Volume-Temperature (PVT) fluid properties were computed using the Glaso [<xref ref-type="bibr" rid="B8">8</xref>] correlation, a generalized empirical PVT correlation widely applied for black-oil systems. Eight PVT parameters, pressure, Gas-Oil Ratio (GOR), oil density, oil viscosity, oil formation volume factor, oil compressibility, gas viscosity, and gas density, were extracted at 20 discrete pressure steps spanning the reservoir’s depletion range, from the initial reservoir pressure of 3500 psig down to a late-decline pressure of 1000 psig, at an approximately constant reservoir temperature. Gas density, the fluid property governing the hydrostatic contribution of gas to wellbore mixture density in gas lift and vertical lift performance calculations, was selected as the prediction target. The reservoir’s bubble point pressure was approximately 1325 psig, so the dataset spans both the undersaturated (single-phase) and saturated (two-phase, solution-gas-liberating) flow regimes.</p>
      </sec>
      <sec id="sec2dot2">
        <title>2.2. Data Preprocessing Pipeline</title>
        <p>A five-step preprocessing pipeline was applied prior to model training. 1) Exploratory Data Analysis characterized the distributional properties (mean, standard deviation, skewness, kurtosis, range) and inter-variable correlation structure of all eight PVT variables. 2) Isolation Forest anomaly detection [<xref ref-type="bibr" rid="B9">9</xref>], an ensemble-based unsupervised method that isolates anomalous observations through recursive random partitioning of the feature space, identified and removed two observations that were statistically inconsistent with the thermodynamic trend of the surrounding data, likely reflecting transcription artefacts from manual entry of PROSPER outputs into the Python environment. Anomaly detection is applied as the first preprocessing step, before normalization and clustering, because both Min-Max scaling and K-means clustering are sensitive to extreme values, and removing anomalous records first prevents outliers from distorting the scaling range or cluster centroids. The five steps are applied in strict sequence for this same reason: Isolation Forest first; normalization second (because K-means uses Euclidean distance, which requires scaled features); K-means third; and smoothing last, so regime labels reflect instantaneous thermodynamic state before smoothing blurs the undersaturated-saturated boundary. Within the training partition, Isolation Forest was applied after the train-test split, so the four test observations were never exposed to or affected by anomaly-removal decisions. 3) Min-Max normalization [<xref ref-type="bibr" rid="B10">10</xref>] rescaled all input features to the range [0, 1] using <inline-formula><mml:math><mml:mrow><mml:msup><mml:mi> x </mml:mi><mml:mo> ′ </mml:mo></mml:msup><mml:mo> = </mml:mo><mml:mrow><mml:mrow><mml:mrow><mml:mo> ( </mml:mo><mml:mrow><mml:mi> x </mml:mi><mml:mo> − </mml:mo><mml:msub><mml:mi> x </mml:mi><mml:mrow><mml:mi> min </mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo> ) </mml:mo></mml:mrow></mml:mrow><mml:mo> / </mml:mo><mml:mrow><mml:mrow><mml:mo> ( </mml:mo><mml:mrow><mml:msub><mml:mi> x </mml:mi><mml:mrow><mml:mi> max </mml:mi></mml:mrow></mml:msub><mml:mo> − </mml:mo><mml:msub><mml:mi> x </mml:mi><mml:mrow><mml:mi> min </mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo> ) </mml:mo></mml:mrow></mml:mrow></mml:mrow></mml:mrow></mml:math></inline-formula> , applied after the train-test split and parameterized using training-set statistics only, to prevent test-set information leakage; this step is essential for SVR, whose objective function is sensitive to the relative magnitude of input features. 4) K-means clustering [<xref ref-type="bibr" rid="B11">11</xref>], specified with k = 2, partitioned the dataset into the undersaturated and saturated production regimes, providing an additional regime-indicator feature. 5) A simple moving-average smoothing step reduced high-frequency variation not attributable to the underlying pressure-dependent thermodynamic trend.</p>
      </sec>
      <sec id="sec2dot3">
        <title>2.3. Machine Learning Model Selection and Training</title>
        <p>Three regression algorithms, selected to represent three distinct algorithmic paradigms, were trained on a stratified 80:20 train-test split of the full 20-observation dataset (16 training observations, 4 held-out test observations), stratified to ensure representation of both undersaturated and saturated pressure conditions in both partitions; a simple random split would risk placing all four test observations within a single pressure regime and understate generalisation performance. The train-test split was performed before K-means clustering and moving-average smoothing, which were applied to the training partition only, ensuring that no test-set information influenced these preprocessing parameters. Isolation Forest anomaly detection was subsequently applied within the 16-observation training partition, identifying and removing two anomalous records, so that the models were fitted on 14 effective training observations; the 4-observation test set was held entirely clean of anomaly-removal decisions. For each of the three models, the predictor set consisted of pressure as the primary input variable, with the regime-indicator feature (a binary label produced by K-means Step 4 encoding the undersaturated/saturated distinction) included as an additional predictor. For the multiple-predictor Linear Regression extension, pressure, GOR, and gas viscosity were used jointly. All predictor variables are available at deployment from standard PVT monitoring data without requiring gas density information.</p>
        <p>Linear Regression was applied as the baseline parametric model, using Ordinary Least Squares (OLS) to fit gas density as a function of pressure. Under the Gauss-Markov theorem [<xref ref-type="bibr" rid="B12">12</xref>], OLS is the Best Linear Unbiased Estimator (BLUE) when the true relationship is linear and the classical regression assumptions hold. A multiple-predictor extension incorporating pressure, GOR, and gas viscosity jointly was also evaluated.</p>
        <p>Support Vector Regression (SVR) was applied following Min-Max normalization, using a Radial Basis Function (RBF) kernel with scikit-learn’s default hyperparameters (C = 1.0, <italic>ε</italic> = 0.1, <italic>γ</italic> = ‘scale’). SVR was selected for its theoretical capacity to model nonlinear relationships through kernel transformation while controlling complexity via the epsilon-insensitive loss and the regularization parameter C [<xref ref-type="bibr" rid="B13">13</xref>][<xref ref-type="bibr" rid="B14">14</xref>].</p>
        <p>Random Forest Regression was applied with 100 trees, using bootstrap sampling and random feature selection at each split node, following Breiman [<xref ref-type="bibr" rid="B15">15</xref>]. Random Forest was selected for its established empirical robustness on complex, noisy engineering datasets and its capacity to generate feature importance rankings.</p>
        <p>All computations were implemented in Python 3 using the scikit-learn library [<xref ref-type="bibr" rid="B10">10</xref>] within a Jupyter Notebook environment.</p>
      </sec>
      <sec id="sec2dot4">
        <title>2.4. Performance Evaluation</title>
        <p>Model performance was evaluated on the held-out test set using two complementary metrics: the Root Mean Squared Error (RMSE), which quantifies average prediction error in the target variable’s original units (lb/ft<sup>3</sup>) and penalizes large errors disproportionately, and the coefficient of determination (R<sup>2</sup>), a scale-independent measure of the proportion of variance explained. The joint use of both metrics follows the recommendation of Chai and Draxler [<xref ref-type="bibr" rid="B16">16</xref>]. Physical consistency of the best-performing model was additionally assessed through an extrapolation test, evaluating predictions at pressures (4000 and 5000 psig) outside the training data range, on the principle that a model which has captured the true governing relationship, rather than overfit to the training sample, should extrapolate in a physically correct (monotonically increasing) direction [<xref ref-type="bibr" rid="B17">17</xref>].</p>
      </sec>
    </sec>
    <sec id="sec3">
      <title>3. Results</title>
      <sec id="sec3dot1">
        <title>3.1. Dataset Characteristics</title>
        <p>Gas density in the preprocessed dataset ranged from 3.95 lb/ft<sup>3</sup> at 1000 psig to 14.07 lb/ft<sup>3</sup> at 3500 psig, a factor of approximately 3.6 across the pressure depletion range, consistent with the real gas law’s prediction that gas density scales approximately with pressure at constant temperature. As we can see in <xref ref-type="fig" rid="fig1">Figure 1</xref>, Pearson correlation analysis revealed a near-perfect linear relationship between gas density and pressure (r = 1.00), a very high correlation between gas density and gas viscosity (r = 0.99, reflecting their shared dependence on molecular packing density via Chapman-Enskog kinetic theory), and a moderate correlation between gas density and GOR (r = 0.58, arising because the dataset spans both undersaturated and saturated regimes). <xref ref-type="fig" rid="fig2">Figure 2</xref> shows that Pairwise scatterplots across all eight PVT variables confirmed a predominantly linear inter-variable structure, with only mild, bubble-point-localized nonlinearity observed in oil viscosity and oil compressibility.</p>
        <fig id="fig1">
          <label>Figure 1</label>
          <graphic xlink:href="https://html.scirp.org/file/1211970-rId17.jpeg?20260921024833" />
        </fig>
        <p><bold>Figure 1</bold><bold>.</bold> Heatmap of Pearson correlations between PVT dataset variables.</p>
      </sec>
      <sec id="sec3dot2">
        <title>3.2. Comparative Model Performance</title>
        <p><bold>Table 1</bold> presents the consolidated test-set performance of the three algorithms. Linear Regression achieved the highest accuracy (R<sup>2</sup> = 0.9954, RMSE = 0.2240 lb/ft<sup>3</sup>), followed by SVR (R<sup>2</sup> = 0.9408, RMSE = 0.8042 lb/ft<sup>3</sup>) and Random Forest (R<sup>2</sup> = 0.8389, RMSE = 1.4326 lb/ft<sup>3</sup>). The Random Forest RMSE was approximately</p>
        <fig id="fig2">
          <label>Figure 2</label>
          <graphic xlink:href="https://html.scirp.org/file/1211970-rId18.jpeg?20260921024833" />
        </fig>
        <p><bold>Figure 2</bold><bold>.</bold> Pairplot of PVT dataset variable relationships.</p>
        <p>6.4 times higher, and the SVR RMSE approximately 3.6 times higher, than the Linear Regression RMSE. <xref ref-type="fig" rid="fig3">Figures 3-5</xref> show five predicted-versus-actual pairs; the additional point reflects a supplementary prediction at the boundary pressure included for visual completeness and was not used in computing the tabulated test-set metrics. The multiple-predictor Linear Regression model (pressure, GOR, and gas viscosity jointly) achieved a marginally improved R<sup>2</sup> of 0.9975, confirming that pressure alone is already the dominant predictor of gas density.</p>
        <p><bold>Table 1.</bold> Comparative test-set performance of the three regression algorithms.</p>
        <table-wrap id="tbl1">
          <label>Table 1</label>
          <table>
            <tbody>
              <tr>
                <td>
                  <bold>Model</bold>
                </td>
                <td>
                  <bold>RMSE (lb/ft</bold>
                  <bold>
                    <sup>3</sup>
                  </bold>
                  <bold>)</bold>
                </td>
                <td>
                  <bold>R</bold>
                  <bold>
                    <sup>2</sup>
                  </bold>
                </td>
                <td>
                  <bold>Interpretation</bold>
                </td>
              </tr>
              <tr>
                <td>Linear Regression</td>
                <td>0.2240</td>
                <td>0.9954</td>
                <td>Highest accuracy; best fit for a near-linear, physics-generated dataset</td>
              </tr>
              <tr>
                <td>Support Vector Regression</td>
                <td>0.8042</td>
                <td>0.9408</td>
                <td>High accuracy; RBF kernel introduces mild curvature bias at the range extremes</td>
              </tr>
              <tr>
                <td>Random Forest Regression</td>
                <td>1.4326</td>
                <td>0.8389</td>
                <td>Lowest accuracy; piecewise-constant partitioning penalized on near-linear data</td>
              </tr>
            </tbody>
          </table>
        </table-wrap>
      </sec>
      <sec id="sec3dot3">
        <title>3.3. Error Structure and Physical Consistency</title>
        <p>Examination of the predicted-versus-actual plots revealed distinct error structures for each algorithm. SVR exhibited systematic positive bias at low gas density (low pressure) and systematic negative bias at high gas density (high pressure), indicating that the RBF kernel introduced curvature where the true relationship is effectively linear (<xref ref-type="fig" rid="fig4">Figure 4</xref>). Random Forest exhibited the widest scatter of the three models (<xref ref-type="fig" rid="fig5">Figure 5</xref>), with the largest deviations concentrated at the extremes of the pressure range, consistent with its piecewise-constant approximation strategy: each decision tree partitions the input space into rectangular regions and predicts a constant value within each region, so approximation error is largest where the underlying gradient is steep relative to the local region width, that is, at the pressure extremes.</p>
        <p>The physical consistency of the Linear Regression model was directly tested by extrapolating beyond the training pressure range. At 4,000 psig, the model predicted a gas density of 16.95 lb/ft<sup>3</sup>; at 5,000 psig, 24.10 lb/ft<sup>3</sup>. Both predictions are physically correct, monotonically increasing with pressure, as required by the real gas law, despite lying outside the range of conditions represented in the training data. An overfitted model would typically fail to generalize correctly beyond its training range [<xref ref-type="bibr" rid="B17">17</xref>]; the correct extrapolation behavior therefore indicates that the Linear Regression model captured the governing thermodynamic relationship rather than a spurious pattern specific to the training sample.</p>
        <fig id="fig3">
          <label>Figure 3</label>
          <graphic xlink:href="https://html.scirp.org/file/1211970-rId19.jpeg?20260921024833" />
        </fig>
        <p><bold>Figure 3</bold><bold>.</bold> Predicted versus actual gas density—linear regression.</p>
        <fig id="fig4">
          <label>Figure 4</label>
          <graphic xlink:href="https://html.scirp.org/file/1211970-rId20.jpeg?20260921024833" />
        </fig>
        <p><bold>Figure 4</bold><bold>.</bold> Predicted versus actual gas density—SVR.</p>
        <fig id="fig5">
          <label>Figure 5</label>
          <graphic xlink:href="https://html.scirp.org/file/1211970-rId21.jpeg?20260921024833" />
        </fig>
        <p><bold>Figure 5</bold><bold>.</bold> Predicted versus actual gas density—random forest.</p>
      </sec>
    </sec>
    <sec id="sec4">
      <title>4. Discussion</title>
      <sec id="sec4dot1">
        <title>4.1. Physical and Statistical Explanation of the Performance Hierarchy</title>
        <p>The observed performance hierarchy, Linear Regression &gt; SVR &gt; Random Forest, comparing the OLS Linear Regression implementation with the scikit-learn default SVR (RBF kernel, C = 1.0, <italic>ε</italic> = 0.1, <italic>γ</italic> = ‘scale’) and the default Random Forest (100 trees, bootstrap sampling) implementations, on this single held-out test partition, is not an arbitrary empirical outcome but is fully explicable from the physics of the data-generating process and the statistical theory governing the three estimators. The explanation proceeds in three steps.</p>
        <p>First, the real gas law prescribes the underlying physical relationship: <inline-formula><mml:math><mml:mrow><mml:msub><mml:mi> ρ </mml:mi><mml:mrow><mml:mi> g </mml:mi><mml:mi> a </mml:mi><mml:mi> s </mml:mi></mml:mrow></mml:msub><mml:mo> = </mml:mo><mml:mrow><mml:mrow><mml:mtext> PM </mml:mtext></mml:mrow><mml:mo> / </mml:mo><mml:mrow><mml:mtext> ZRT </mml:mtext></mml:mrow></mml:mrow></mml:mrow></mml:math></inline-formula> , where P is pressure, M is the gas molecular weight, Z is the gas compressibility factor, R is the universal gas constant, and T is the (approximately constant) reservoir temperature. Over the pressure range studied (1000 - 3500 psig), the Z-factor for a typical Niger Delta gas composition varies only modestly, from approximately 0.82 to 0.95 [<xref ref-type="bibr" rid="B18">18</xref>], so gas density varies approximately linearly with pressure. This near-linearity is directly confirmed by the correlation coefficient r = 1.00 observed between gas density and pressure in the dataset.</p>
        <p>Second, for a dataset in which the true relationship between the target variable and the primary predictor is approximately linear, the linear hypothesis class is the physically correct functional form, and the Gauss-Markov theorem guarantees that the OLS estimator is the Best Linear Unbiased Estimator, the minimum-variance unbiased estimator within that hypothesis class [<xref ref-type="bibr" rid="B12">12</xref>]. Any departure from the linear functional form, whether the nonlinear kernel transformation applied by SVR or the piecewise-constant partitioning applied by Random Forest, adds model complexity, and therefore variance, without a compensating reduction in bias, because the residual bias of Linear Regression on this dataset is already negligible.</p>
        <p>Third, the bias-variance decomposition of prediction error [<xref ref-type="bibr" rid="B17">17</xref>][<xref ref-type="bibr" rid="B19">19</xref>] explains the magnitude of the observed hierarchy. Linear Regression achieves near-zero bias, because the linear model correctly represents the true relationship, combined with minimum variance, guaranteed by the Gauss-Markov theorem, yielding the lowest expected prediction error. Random Forest introduces a piecewise-constant approximation error that is only partly offset by bootstrap aggregation, producing the largest deviation from optimal performance among the three models (R<sup>2</sup> = 0.8389). SVR introduces variance through the mismatch between its RBF kernel’s nonlinear hypothesis class and the effectively linear target function, together with the equal weighting of all within-margin predictions imposed by the epsilon-insensitive loss, without any corresponding reduction in bias, resulting in intermediate performance (R<sup>2</sup> = 0.9408).</p>
      </sec>
      <sec id="sec4dot2">
        <title>4.2. Practical Implications for Algorithm Selection</title>
        <p>These findings carry a specific and generalizable implication for ML practitioners in petroleum engineering: the optimal algorithm for a given prediction task is determined, in part, by the statistical structure that the underlying physics imposes on the training data, and this structure depends on how the training data were generated. For datasets generated from simulator-based PVT correlations, such as PROSPER’s Glaso-correlation workflow, which produce approximately linear inter-variable relationships as a direct consequence of the real gas law and related thermodynamic constraints, Linear Regression is both the theoretically justified and empirically best-performing choice, and its use avoids the unnecessary variance cost of more complex algorithms. This is a case for principled, physics-informed model selection rather than a default preference for algorithmic complexity.</p>
        <p>The converse also holds. Field-measured PVT and production datasets incorporate additional noise sources, sensor error, sampling and operational variability, and potentially genuine nonlinear contributions from geological heterogeneity and multiphase interaction effects that are absent from a simulator-generated dataset. Under these conditions, the linearity assumption that favors Linear Regression in this study is weaker, and the added flexibility of Random Forest or SVR may become an asset rather than a source of unnecessary variance [<xref ref-type="bibr" rid="B2">2</xref>][<xref ref-type="bibr" rid="B7">7</xref>]. The performance hierarchy identified here should therefore be interpreted as specific to physics-generated training data rather than as a general recommendation against ensemble or kernel-based methods in petroleum engineering ML applications.</p>
        <p>More broadly, this study demonstrates a practical and readily transferable route to physically consistent machine learning: generating ML training data from an industry-standard, physics-based simulator rather than embedding physical constraints directly into the model architecture or loss function, as in Physics-Informed Neural Networks [<xref ref-type="bibr" rid="B4">4</xref>]. This route requires no modification to standard ML algorithms or software, is compatible with existing commercial IPM platforms, and, as shown by the correct extrapolation behavior of the trained Linear Regression model, produces models whose predictions remain physically consistent even outside the training data range.</p>
      </sec>
      <sec id="sec4dot3">
        <title>4.3. Limitations</title>
        <p>Several limitations qualify these findings. First, the dataset was generated from a single-well production system model using the Glaso [<xref ref-type="bibr" rid="B8">8</xref>] PVT correlation, which was originally calibrated on North Sea crude oils and may not perfectly represent Niger Delta fluid compositions; any systematic bias in the correlation propagates into the ML training data and, in turn, into the model predictions. Second, the training dataset, drawn from 20 simulated pressure steps, is small by conventional ML standards, reflecting the discrete nature of simulator-generated PVT data rather than continuously sampled field data; the resulting train-test split (16/4 observations) limits the statistical power of the test-set evaluation, and the reported metrics should be interpreted as indicative rather than as precise population estimates. Furthermore, the performance hierarchy reported here is based on a single stratified 80:20 train-test split. The dataset’s origin as physics-simulation output at discrete pressure steps, rather than a randomly sampled field time-series, constrains the resampling options available: standard k-fold cross-validation applied to an ordered pressure series would break temporal ordering and produce splits in which training observations are pressure-adjacent to test observations, inflating apparent generalisation. Time-series cross-validation requires minimum fold sizes that this small dataset cannot support. The stratified split used here was the methodologically appropriate choice given these constraints, ensuring both pressure regimes were represented in training and testing, but the four-observation test set limits the statistical precision of the reported metrics. Future work with larger simulator-generated datasets, spanning a wider pressure range at finer discretization, should validate the hierarchy using repeated resampling. Third, because the models were trained to reproduce the Glaso correlation’s predictions rather than direct field measurements of gas density, the reported accuracy reflects fidelity to the correlation, not independently validated accuracy against measured field gas density. Fourth, the SVR and Random Forest models were evaluated at scikit-learn default hyperparameter values rather than values tuned within the training data; the conclusions therefore compare these specific default implementations rather than the globally optimised algorithm classes. Hyperparameter tuning within the training partition may improve SVR and Random Forest accuracy and should be evaluated in future work. Importantly, however, the underperformance of SVR relative to Linear Regression on this dataset cannot be attributed to SVR as an algorithm class in general, but specifically to the mismatch between a flexible nonlinear learner and a near-linear data-generating process: the Gauss-Markov theorem and the bias-variance decomposition both predict that any expansion of the hypothesis class beyond the linear functional form, whether through kernel choice or hyperparameter tuning, adds variance without reducing the near-zero bias that OLS already achieves on this data structure. This is a structural prediction, not a contingent one, and it reinforces rather than undermines the algorithm-selection argument. Fifth, a minor numerical inconsistency was identified between the reported Random Forest R<sup>2</sup> (0.8389) and the value implied by the RMSE (1.4326 lb/ft<sup>3</sup>) relative to the test-set variance implied by the Linear Regression and SVR metrics; the reported values are those produced directly by the scikit-learn evaluation, and any rounding or precision difference should be interpreted in that context. Extension of this comparison to field-measured, multi-well datasets is an important direction for future work.</p>
      </sec>
    </sec>
    <sec id="sec5">
      <title>5. Conclusion</title>
      <p>This study trained, evaluated, and compared Linear Regression, Support Vector Regression, and Random Forest Regression on a physics-generated PVT dataset for gas density prediction, derived from an industry-standard Integrated Production Modelling platform using the Glaso [<xref ref-type="bibr" rid="B8">8</xref>] correlation for a mature Niger Delta oil well. Linear Regression achieved the highest predictive accuracy (R<sup>2</sup> = 0.9954, RMSE = 0.2240 lb/ft<sup>3</sup>), outperforming SVR (R<sup>2</sup> = 0.9408) and Random Forest (R<sup>2</sup> = 0.8389), and produced physically consistent extrapolations beyond the training pressure range. This performance hierarchy was shown to be a direct and predictable consequence of the real gas law, which imposes a near-linear pressure-density relationship at approximately constant reservoir temperature, combined with the Gauss-Markov theorem, which guarantees the optimality of Ordinary Least Squares under linear data conditions. The bias-variance decomposition further clarified why the added flexibility of SVR’s kernel transformation and Random Forest’s ensemble partitioning each introduced variance without a compensating reduction in bias on this near-linear dataset. These findings offer petroleum engineers and applied ML practitioners a principled, physics-grounded basis for machine learning algorithm selection, specifically comparing OLS Linear Regression with SVR and Random Forest at their scikit-learn default configurations, when training data are generated from simulator-based PVT correlations, and they illustrate a practical, readily transferable route to physically consistent machine learning that requires no modification to standard algorithms or existing commercial modelling software.</p>
    </sec>
    <sec id="sec6">
      <title>Acknowledgements</title>
      <p>The author gratefully acknowledges the guidance and supervision of Prof. Joseph A. Ajienka and Dr. Joseph Amieibibama, Department of Petroleum and Gas Engineering, University of Port Harcourt.</p>
    </sec>
    <sec id="sec7">
      <title>Author Contributions</title>
      <p>Conceptualization, Gnessoa Rene Hie and Joseph A. Ajienka; methodology, Gnessoa Rene Hie; software, Gnessoa Rene Hie; validation, Gnessoa Rene Hie, Joseph A. Ajienka, and Joseph Amieibibama; formal analysis, Gnessoa Rene Hie; investigation, Gnessoa Rene Hie; resources, Gnessoa Rene Hie; data curation, Gnessoa Rene Hie; writing—original draft preparation, Gnessoa Rene Hie; writing—review and editing, Gnessoa Rene Hie; visualization, Gnessoa Rene Hie; supervision, Gnessoa Rene Hie; project administration, Gnessoa Rene Hie. All authors have read and agreed to the published version of the manuscript.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <title>References</title>
      <ref id="B1">
        <label>1.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">LeCun, Y., Bengio, Y. and Hinton, G. (2015) Deep Learning. <italic>Nature</italic>, 521, 436-444. https://doi.org/10.1038/nature14539 <pub-id pub-id-type="doi">10.1038/nature14539</pub-id><pub-id pub-id-type="pmid">26017442</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1038/nature14539">https://doi.org/10.1038/nature14539</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>LeCun, Y.</string-name>
              <string-name>Bengio, Y.</string-name>
              <string-name>Hinton, G.</string-name>
            </person-group>
            <year>2015</year>
            <article-title>Deep Learning</article-title>
            <source>Nature</source>
            <volume>521</volume>
            <pub-id pub-id-type="doi">10.1038/nature14539</pub-id>
            <pub-id pub-id-type="pmid">26017442</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B2">
        <label>2.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Tariq, Z., Aljawad, M.S., Hasan, A., Murtaza, M., Mohammed, E., El-Husseiny, A., <italic>et al</italic>. (2021) A Systematic Review of Data Science and Machine Learning Applications to the Oil and Gas Industry. <italic>Journal</italic><italic>of</italic><italic>Petroleum</italic><italic>Exploration</italic><italic>and</italic><italic>Production</italic><italic>Technology</italic>, 11, 4339-4374. https://doi.org/10.1007/s13202-021-01302-2 <pub-id pub-id-type="doi">10.1007/s13202-021-01302-2</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1007/s13202-021-01302-2">https://doi.org/10.1007/s13202-021-01302-2</ext-link></mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Tariq, Z.</string-name>
              <string-name>Aljawad, M.S.</string-name>
              <string-name>Hasan, A.</string-name>
              <string-name>Murtaza, M.</string-name>
              <string-name>Mohammed, E.</string-name>
              <string-name>El-Husseiny, A.</string-name>
            </person-group>
            <year>2021</year>
            <article-title>A Systematic Review of Data Science and Machine Learning Applications to the Oil and Gas Industry</article-title>
            <source>Journal of Petroleum Exploration and Production Technology</source>
            <volume>11</volume>
            <pub-id pub-id-type="doi">10.1007/s13202-021-01302-2</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B3">
        <label>3.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Karpatne, A., Atluri, G., Faghmous, J.H., Steinbach, M., Banerjee, A., Ganguly, A., <italic>et al</italic>. (2017) Theory-Guided Data Science: A New Paradigm for Scientific Discovery from Data. <italic>IEEE</italic><italic>Transactions</italic><italic>on</italic><italic>Knowledge</italic><italic>and</italic><italic>Data</italic><italic>Engineering</italic>, 29, 2318-2331. https://doi.org/10.1109/tkde.2017.2720168 <pub-id pub-id-type="doi">10.1109/tkde.2017.2720168</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1109/tkde.2017.2720168">https://doi.org/10.1109/tkde.2017.2720168</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Karpatne, A.</string-name>
              <string-name>Atluri, G.</string-name>
              <string-name>Faghmous, J.H.</string-name>
              <string-name>Steinbach, M.</string-name>
              <string-name>Banerjee, A.</string-name>
              <string-name>Ganguly, A.</string-name>
            </person-group>
            <year>2017</year>
            <article-title>Theory-Guided Data Science: A New Paradigm for Scientific Discovery from Data</article-title>
            <source>IEEE Transactions on Knowledge and Data Engineering</source>
            <volume>29</volume>
            <pub-id pub-id-type="doi">10.1109/tkde.2017.2720168</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B4">
        <label>4.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Raissi, M., Perdikaris, P. and Karniadakis, G.E. (2019) Physics-Informed Neural Networks: A Deep Learning Framework for Solving Forward and Inverse Problems Involving Nonlinear Partial Differential Equations. <italic>Journal</italic><italic>of</italic><italic>Computational</italic><italic>Physics</italic>, 378, 686-707. https://doi.org/10.1016/j.jcp.2018.10.045 <pub-id pub-id-type="doi">10.1016/j.jcp.2018.10.045</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1016/j.jcp.2018.10.045">https://doi.org/10.1016/j.jcp.2018.10.045</ext-link></mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Raissi, M.</string-name>
              <string-name>Perdikaris, P.</string-name>
              <string-name>Karniadakis, G.E.</string-name>
            </person-group>
            <year>2019</year>
            <article-title>Physics-Informed Neural Networks: A Deep Learning Framework for Solving Forward and Inverse Problems Involving Nonlinear Partial Differential Equations</article-title>
            <source>Journal of Computational Physics</source>
            <volume>378</volume>
            <pub-id pub-id-type="doi">10.1016/j.jcp.2018.10.045</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B5">
        <label>5.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Karniadakis, G.E., Kevrekidis, I.G., Lu, L., Perdikaris, P., Wang, S. and Yang, L. (2021) Physics-Informed Machine Learning. <italic>Nature</italic><italic>Reviews</italic><italic>Physics</italic>, 3, 422-440. https://doi.org/10.1038/s42254-021-00314-5 <pub-id pub-id-type="doi">10.1038/s42254-021-00314-5</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1038/s42254-021-00314-5">https://doi.org/10.1038/s42254-021-00314-5</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Karniadakis, G.E.</string-name>
              <string-name>Kevrekidis, I.G.</string-name>
              <string-name>Lu, L.</string-name>
              <string-name>Perdikaris, P.</string-name>
              <string-name>Wang, S.</string-name>
              <string-name>Yang, L.</string-name>
            </person-group>
            <year>2021</year>
            <article-title>Physics-Informed Machine Learning</article-title>
            <source>Nature Reviews Physics</source>
            <volume>3</volume>
            <pub-id pub-id-type="doi">10.1038/s42254-021-00314-5</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B6">
        <label>6.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Nwachukwu, A., Jeong, H., Pyrcz, M. and Lake, L.W. (2018) Fast Evaluation of Well Placements in Heterogeneous Reservoir Models Using Machine Learning. <italic>Journal</italic><italic>of</italic><italic>Petroleum</italic><italic>Science</italic><italic>and</italic><italic>Engineering</italic>, 163, 463-475. https://doi.org/10.1016/j.petrol.2018.01.019 <pub-id pub-id-type="doi">10.1016/j.petrol.2018.01.019</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1016/j.petrol.2018.01.019">https://doi.org/10.1016/j.petrol.2018.01.019</ext-link></mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Nwachukwu, A.</string-name>
              <string-name>Jeong, H.</string-name>
              <string-name>Pyrcz, M.</string-name>
              <string-name>Lake, L.W.</string-name>
            </person-group>
            <year>2018</year>
            <article-title>Fast Evaluation of Well Placements in Heterogeneous Reservoir Models Using Machine Learning</article-title>
            <source>Journal of Petroleum Science and Engineering</source>
            <volume>163</volume>
            <pub-id pub-id-type="doi">10.1016/j.petrol.2018.01.019</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B7">
        <label>7.</label>
        <citation-alternatives>
          <mixed-citation publication-type="confproc">Temizel, C., Canbaz, C.H., Palabiyik, Y. and Ozcan, G. (2014) Machine Learning Applications in Unconventional Resources: An Analysis of Production Data. <italic>SPE Annual Technical Conference and Exhibition</italic>, Amsterdam, 27-29 October 2014.</mixed-citation>
          <element-citation publication-type="confproc">
            <person-group person-group-type="author">
              <string-name>Temizel, C.</string-name>
              <string-name>Canbaz, C.H.</string-name>
              <string-name>Palabiyik, Y.</string-name>
              <string-name>Ozcan, G.</string-name>
              <string-name>Exhibition, A</string-name>
            </person-group>
            <year>2014</year>
            <article-title>Machine Learning Applications in Unconventional Resources: An Analysis of Production Data</article-title>
            <source>SPE Annual Technical Conference and Exhibition</source>
            <volume>27</volume>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B8">
        <label>8.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Glaso, O. (1980) Generalized Pressure-Volume-Temperature Correlations. <italic>Journal</italic><italic>of</italic><italic>Petroleum</italic><italic>Technology</italic>, 32, 785-795. https://doi.org/10.2118/8016-pa <pub-id pub-id-type="doi">10.2118/8016-pa</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.2118/8016-pa">https://doi.org/10.2118/8016-pa</ext-link></mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Glaso, O.</string-name>
            </person-group>
            <year>1980</year>
            <article-title>Generalized Pressure-Volume-Temperature Correlations</article-title>
            <source>Journal of Petroleum Technology</source>
            <volume>32</volume>
            <pub-id pub-id-type="doi">10.2118/8016-pa</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B9">
        <label>9.</label>
        <citation-alternatives>
          <mixed-citation publication-type="confproc">Liu, F.T., Ting, K.M. and Zhou, Z. (2008) Isolation Forest. 2008 <italic>Eighth IEEE International Conference on Data Mining</italic>, Pisa, 15-19 December 2008, 413-422. https://doi.org/10.1109/icdm.2008.17 <pub-id pub-id-type="doi">10.1109/icdm.2008.17</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1109/icdm.2008.17">https://doi.org/10.1109/icdm.2008.17</ext-link></mixed-citation>
          <element-citation publication-type="confproc">
            <person-group person-group-type="author">
              <string-name>Liu, F.T.</string-name>
              <string-name>Ting, K.M.</string-name>
              <string-name>Zhou, Z.</string-name>
              <string-name>Mining, P</string-name>
            </person-group>
            <year>2008</year>
            <article-title>Isolation Forest</article-title>
            <source>2008 Eighth IEEE International Conference on Data Mining</source>
            <volume>15</volume>
            <pub-id pub-id-type="doi">10.1109/icdm.2008.17</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B10">
        <label>10.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Pedregosa, F., Varoquaux, G., Gramfort, A., Michel, V., Thirion, B., Grisel, O., et al. (2011) Scikit-Learn: Machine Learning in Python. <italic>Journal of Machine Learning Research</italic>, 12, 2825-2830.</mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Pedregosa, F.</string-name>
              <string-name>Varoquaux, G.</string-name>
              <string-name>Gramfort, A.</string-name>
              <string-name>Michel, V.</string-name>
              <string-name>Thirion, B.</string-name>
              <string-name>Grisel, O.</string-name>
            </person-group>
            <year>2011</year>
            <article-title>Scikit-Learn: Machine Learning in Python</article-title>
            <source>Journal of Machine Learning Research</source>
            <volume>12</volume>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B11">
        <label>11.</label>
        <citation-alternatives>
          <mixed-citation publication-type="confproc">MacQueen, J. (1967) Some Methods for Classification and Analysis of Multivariate Observations. <italic>Proceedings of the</italic>5 <italic>th Berkeley Symposium on Mathematical Statistics and Probability</italic>, Berkeley, 21 June-18 July 1965 and 27 December 1965-7 January 1966, 281-297.</mixed-citation>
          <element-citation publication-type="confproc">
            <person-group person-group-type="author">
              <string-name>MacQueen, J.</string-name>
              <string-name>Probability, B</string-name>
            </person-group>
            <year>1967</year>
            <article-title>Some Methods for Classification and Analysis of Multivariate Observations</article-title>
            <source>Proceedings of the 5th Berkeley Symposium on Mathematical Statistics and Probability</source>
            <volume>21</volume>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B12">
        <label>12.</label>
        <citation-alternatives>
          <mixed-citation publication-type="book">Greene, W.H. (2018) Econometric Analysis. 8th Edition, Pearson.</mixed-citation>
          <element-citation publication-type="book">
            <person-group person-group-type="author">
              <string-name>Greene, W.H.</string-name>
              <string-name>Edition, P</string-name>
            </person-group>
            <year>2018</year>
            <article-title>Econometric Analysis</article-title>
            <source>8th Edition</source>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B13">
        <label>13.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Vapnik, V.N. (1995) The Nature of Statistical Learning Theory. Springer. https://doi.org/10.1007/978-1-4757-2440-0 <pub-id pub-id-type="doi">10.1007/978-1-4757-2440-0</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1007/978-1-4757-2440-0">https://doi.org/10.1007/978-1-4757-2440-0</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Vapnik, V.N.</string-name>
            </person-group>
            <year>1995</year>
            <article-title>The Nature of Statistical Learning Theory</article-title>
            <pub-id pub-id-type="doi">10.1007/978-1-4757-2440-0</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B14">
        <label>14.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Smola, A.J. and Schölkopf, B. (2004) A Tutorial on Support Vector Regression. <italic>Statistics</italic><italic>and</italic><italic>Computing</italic>, 14, 199-222. https://doi.org/10.1023/b:stco.0000035301.49549.88 <pub-id pub-id-type="doi">10.1023/b:stco.0000035301.49549.88</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1023/b:stco.0000035301.49549.88">https://doi.org/10.1023/b:stco.0000035301.49549.88</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Smola, A.J.</string-name>
            </person-group>
            <year>2004</year>
            <article-title>A Tutorial on Support Vector Regression</article-title>
            <source>Statistics and Computing</source>
            <volume>14</volume>
            <pub-id pub-id-type="doi">10.1023/b:stco.0000035301.49549.88</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B15">
        <label>15.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Breiman, L. (2001) Random Forests. <italic>Machine</italic><italic>Learning</italic>, 45, 5-32. https://doi.org/10.1023/a:1010933404324 <pub-id pub-id-type="doi">10.1023/a:1010933404324</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1023/a:1010933404324">https://doi.org/10.1023/a:1010933404324</ext-link></mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Breiman, L.</string-name>
            </person-group>
            <year>2001</year>
            <article-title>Random Forests</article-title>
            <source>Machine Learning</source>
            <volume>45</volume>
            <fpage>101093</fpage>
            <pub-id pub-id-type="doi">10.1023/a:1010933404324</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B16">
        <label>16.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Chai, T. and Draxler, R.R. (2014) Root Mean Square Error (RMSE) or Mean Absolute Error (MAE)? Arguments against Avoiding RMSE in the Literature. <italic>Geoscientific</italic><italic>Model</italic><italic>Development</italic>, 7, 1247-1250. https://doi.org/10.5194/gmd-7-1247-2014 <pub-id pub-id-type="doi">10.5194/gmd-7-1247-2014</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.5194/gmd-7-1247-2014">https://doi.org/10.5194/gmd-7-1247-2014</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Chai, T.</string-name>
              <string-name>Draxler, R.R.</string-name>
            </person-group>
            <year>2014</year>
            <article-title>Root Mean Square Error (RMSE) or Mean Absolute Error (MAE)? Arguments against Avoiding RMSE in the Literature</article-title>
            <source>Geoscientific Model Development</source>
            <volume>7</volume>
            <pub-id pub-id-type="doi">10.5194/gmd-7-1247-2014</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B17">
        <label>17.</label>
        <citation-alternatives>
          <mixed-citation publication-type="book">Hastie, T., Tibshirani, R. and Friedman, J. (2009) The Elements of Statistical Learning: Data Mining, Inference, and Prediction. 2nd Edition, Springer. https://doi.org/10.1007/978-0-387-84858-7 <pub-id pub-id-type="doi">10.1007/978-0-387-84858-7</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1007/978-0-387-84858-7">https://doi.org/10.1007/978-0-387-84858-7</ext-link></mixed-citation>
          <element-citation publication-type="book">
            <person-group person-group-type="author">
              <string-name>Hastie, T.</string-name>
              <string-name>Tibshirani, R.</string-name>
              <string-name>Friedman, J.</string-name>
              <string-name>Mining, I</string-name>
              <string-name>Edition, S</string-name>
            </person-group>
            <year>2009</year>
            <article-title>The Elements of Statistical Learning: Data Mining, Inference, and Prediction</article-title>
            <source>2nd Edition</source>
            <pub-id pub-id-type="doi">10.1007/978-0-387-84858-7</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B18">
        <label>18.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Standing, M.B. and Katz, D.L. (1942) Density of Natural Gases. <italic>Transactions</italic><italic>of</italic><italic>the</italic><italic>AIME</italic>, 146, 140-149. https://doi.org/10.2118/942140-g <pub-id pub-id-type="doi">10.2118/942140-g</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.2118/942140-g">https://doi.org/10.2118/942140-g</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Standing, M.B.</string-name>
              <string-name>Katz, D.L.</string-name>
            </person-group>
            <year>1942</year>
            <article-title>Density of Natural Gases</article-title>
            <source>Transactions of the AIME</source>
            <volume>146</volume>
            <pub-id pub-id-type="doi">10.2118/942140-g</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B19">
        <label>19.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Geman, S., Bienenstock, E. and Doursat, R. (1992) Neural Networks and the Bias/Variance Dilemma. <italic>Neural</italic><italic>Computation</italic>, 4, 1-58. https://doi.org/10.1162/neco.1992.4.1.1 <pub-id pub-id-type="doi">10.1162/neco.1992.4.1.1</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1162/neco.1992.4.1.1">https://doi.org/10.1162/neco.1992.4.1.1</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Geman, S.</string-name>
              <string-name>Bienenstock, E.</string-name>
              <string-name>Doursat, R.</string-name>
            </person-group>
            <year>1992</year>
            <article-title>Neural Networks and the Bias/Variance Dilemma</article-title>
            <source>Neural Computation</source>
            <volume>4</volume>
            <pub-id pub-id-type="doi">10.1162/neco.1992.4.1.1</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
    </ref-list>
  </back>
</article>