<?xml version="1.0" encoding="UTF-8"?><!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v3.0 20080202//EN" "http://dtd.nlm.nih.gov/publishing/3.0/journalpublishing3.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" dtd-version="3.0" xml:lang="en" article-type="research article">
 <front>
  <journal-meta>
   <journal-id journal-id-type="publisher-id">
    acs
   </journal-id>
   <journal-title-group>
    <journal-title>
     Atmospheric and Climate Sciences
    </journal-title>
   </journal-title-group>
   <issn pub-type="epub">
    2160-0414
   </issn>
   <issn publication-format="print">
    2160-0422
   </issn>
   <publisher>
    <publisher-name>
     Scientific Research Publishing
    </publisher-name>
   </publisher>
  </journal-meta>
  <article-meta>
   <article-id pub-id-type="doi">
    10.4236/acs.2024.144026
   </article-id>
   <article-id pub-id-type="publisher-id">
    acs-136564
   </article-id>
   <article-categories>
    <subj-group subj-group-type="heading">
     <subject>
      Articles
     </subject>
    </subj-group>
    <subj-group subj-group-type="Discipline-v2">
     <subject>
      Earth 
     </subject>
     <subject>
       Environmental Sciences
     </subject>
    </subj-group>
   </article-categories>
   <title-group>
    Optimizing the LSTM Deep Learning Model for Arctic Sea Ice Melting Prediction
   </title-group>
   <contrib-group>
    <contrib contrib-type="author" xlink:type="simple">
     <name name-style="western">
      <surname>
       Victoria Pegkou
      </surname>
      <given-names>
       Christofi
      </given-names>
     </name>
    </contrib>
    <contrib contrib-type="author" xlink:type="simple">
     <name name-style="western">
      <surname>
       Xiaodi
      </surname>
      <given-names>
       Wang
      </given-names>
     </name>
    </contrib>
   </contrib-group> 
   <aff id="affnull">
    <addr-line>
     aDepartment of Mathematics, Western Connecticut State University, Danbury, USA
    </addr-line> 
   </aff> 
   <pub-date pub-type="epub">
    <day>
     30
    </day> 
    <month>
     08
    </month>
    <year>
     2024
    </year>
   </pub-date> 
   <volume>
    14
   </volume> 
   <issue>
    04
   </issue>
   <fpage>
    429
   </fpage>
   <lpage>
    449
   </lpage>
   <history>
    <date date-type="received">
     <day>
      7,
     </day>
     <month>
      April
     </month>
     <year>
      2024
     </year>
    </date>
    <date date-type="published">
     <day>
      11,
     </day>
     <month>
      April
     </month>
     <year>
      2024
     </year> 
    </date> 
    <date date-type="accepted">
     <day>
      11,
     </day>
     <month>
      October
     </month>
     <year>
      2024
     </year> 
    </date>
   </history>
   <permissions>
    <copyright-statement>
     © Copyright 2014 by authors and Scientific Research Publishing Inc. 
    </copyright-statement>
    <copyright-year>
     2014
    </copyright-year>
    <license>
     <license-p>
      This work is licensed under the Creative Commons Attribution International License (CC BY). http://creativecommons.org/licenses/by/4.0/
     </license-p>
    </license>
   </permissions>
   <abstract>
    The National Oceanic and Atmospheric Administration reports a 95% decline in the oldest Arctic ice over the last 33 years [1], while the National Aeronautics and Space Administration states that summer Arctic Sea Ice Extent (SIE) is shrinking by 12.2% per decade since 1979 due to warmer temperatures [2]. Given the rapidly changing Arctic conditions, accurate prediction models are crucial. Deep learning models developed for Arctic forecasts primarily focus on exploring convolutional neural networks (CNN) and convolutional Long Short-Term Memory (LSTM) networks, while the exploration of the power of LSTM networks is limited. In this research, we focus on enhancing the performance of an LSTM network for predicting monthly Arctic SIE. We leverage five climate and atmospheric variables, validated for their correlation with SIE in prior studies [3]. We utilize the Spearman’s rank correlation and ExtraTrees regressor to enhance our understanding of the importance of the five variables in predicting SIE. We further enhance our predictor variables with seasonal information, lagged time steps, and a linear regression simulated SIE that accounts for the influence of past SIE on current SIE. Statistical methods guide our selection of data scalers and best evaluation metrics for our model. By experimenting with hyperparameter optimization and advanced deep learning training techniques, such as batch sizes, number of neurons, early stopping, and model checkpoint, our model achieved a Mean Absolute Error (MAE) of 0.191 and R
    <sup>2</sup> of 0.996, underscoring its ability to account for nearly all the variance in our data and holds great promise for the prediction of SIE.
   </abstract>
   <kwd-group> 
    <kwd>
     Arctic
    </kwd> 
    <kwd>
      Sea Ice Extent
    </kwd> 
    <kwd>
      Deep Learning
    </kwd> 
    <kwd>
      Long Short-Term Memory Networks
    </kwd> 
    <kwd>
      Climate Change
    </kwd>
   </kwd-group>
  </article-meta>
 </front>
 <body>
  <sec id="s1">
   <title>1. Introduction</title>
   <p>The rapid decline of Arctic Sea ice has become increasingly relevant due to its profound and far-reaching effects on the environment, climate, ecosystems, and human communities <xref ref-type="bibr" rid="scirp.136564-2">
     [2]
    </xref>. The most common measure of sea ice is extent, defined as the area covered with an ice concentration of at least 15%. Extent increases and decreases with the seasons. As shown in <xref ref-type="fig" rid="fig1">
     Figure 1
    </xref>, the median Arctic extent, as assessed over the period 1981-2010, is greatest in mid-March. The minimum median extent falls in mid-September <xref ref-type="bibr" rid="scirp.136564-4">
     [4]
    </xref>.</p>
   <fig id="fig1" position="float">
    <label>Figure 1</label>
    <caption>
     <title>Figure 1. Daily Arctic Sea Ice Extent in millions of km<sup>2</sup> <xref ref-type="bibr" rid="scirp.136564-5">
       [5]
      </xref>.</title>
    </caption>
    <graphic mimetype="image" position="float" xlink:type="simple" xlink:href="https://html.scirp.org/file/4701247-rId13.jpeg?20241014013913" />
   </fig>
   <p>Since 2011, the decline of sea ice has persisted, and near-surface permafrost has continued to warm, signaling a transformation of the Arctic into a warmer, more humid, and less predictable environment <xref ref-type="bibr" rid="scirp.136564-6">
     [6]
    </xref>. The stability of marine ecosystems and the very survival of marine life are put at risk as sea ice continues to melt. Moreover, the alterations in the Arctic region will amplify disruptions in ocean currents, affecting global weather patterns, the viability of fishing industries, and the vulnerability of coastal cities to increased flooding risks. The consequences of melting sea ice underscore the urgent need for the development of precise predictive models that can capture the dynamic changes occurring in the Arctic. These models are essential for not only understanding the impacts of decreased sea ice on the Arctic but also for gaining invaluable insights into the broader issue of climate change.</p>
   <p>Various approaches have been made to develop prediction models for the Arctic, however they focus on exploring convolutional neural networks and Convolutional Long Short-Term Memory networks (ConvLSTM), while the exploration of Long Short-Term Memory (LSTM) networks is limited per our literature review. It’s important to understand that the melting of Arctic Sea ice is a multifaceted phenomenon, one that is influenced by a variety of climate and atmospheric variables. To address this complexity and enhance prediction accuracy, our study employs feature selection to encompass a wide array of variables, including sea surface temperature, surface pressure, total precipitation, latent heat, and sensible heat. Additionally, many predictive models underestimate the speed of melting in the Arctic and limit the range of techniques used to enhance their LSTM models.</p>
   <p>According to the Sea Ice Outlook by Sea Ice Prediction network (SIPN), before 2020, there were more submissions with statistical approaches and dynamical models in comparison to those utilizing machine learning <xref ref-type="bibr" rid="scirp.136564-7">
     [7]
    </xref>. In a pioneering study, Chi and Kim <xref ref-type="bibr" rid="scirp.136564-8">
     [8]
    </xref> utilized machine learning and developed a 1-month forecast LSTM model to predict sea ice concentration (SIC). Their model predicted monthly SIC with less than 9% average monthly prediction errors. Achieving accurate predictions during the melting season (June to October) was more challenging compared to the freezing season (November to May). This heightened difficulty was attributed to the impact of sea ice thinning <xref ref-type="bibr" rid="scirp.136564-8">
     [8]
    </xref>.</p>
   <p>This study branches off our previous research <xref ref-type="bibr" rid="scirp.136564-9">
     [9]
    </xref>, in which we focused on exploring the combination of image and numerical data when constructing predictive models for the Arctic Sea Ice Extent (SIE). In this study, we aim to further explore the potential of LSTM models using exclusively numerical data. This approach allows us to employ a variety of techniques to optimize the accuracy of our sequential LSTM model and make informed comparisons on the most effective methods for SIE prediction.</p>
   <p>This research not only enhances our comprehension of sea ice dynamics, but also equips us with the essential knowledge pertaining to LSTM models and their application in preserving and protecting the Earth’s delicate environmental balance.</p>
  </sec><sec id="s2">
   <title>2. Model</title>
   <sec id="s2_1">
    <title>2.1. Long Short-Term Memory Networks (LSTM)</title>
    <p>LSTM network is a special type of Recurrent Neural Network (RNN) which was developed to overcome the limitations of traditional RNNs, particularly their inability to capture and maintain long term dependencies <xref ref-type="bibr" rid="scirp.136564-10">
      [10]
     </xref>. Consequently, LSTMs are exceptionally well-suited for tasks involving the classification and prediction of time-series data.</p>
    <p>A typical LSTM structure is composed of distinct memory blocks known as cells. These cells act as transporters of essential information through the network’s chain and thus effectively serve as the network’s memory. This architecture enables the transfer of information from earlier time steps to subsequent ones, generating the capacity for long-term memory retention. An LSTM cell consists of the cell state (C), the hidden state (h) and three gates: the forget gate, the input gate, and the output gate. The cell state is the long-term memory of the LSTM whereas the hidden state is the short-term representation or output of the current time step. The gates regulate the flow of information by determining what to forget, update, and output, enabling the network to effectively capture and utilize long-term dependencies in sequential data.</p>
    <p>Activation functions play a crucial role in controlling the flow of information through LSTM networks and learning complex patterns in sequential data. They allow LSTM to selectively update and utilize information in the cell state based on the input, previous states, and the network’s learned parameters. For the gate and cell state operations, LSTM networks traditionally employ the sigmoid and hyperbolic tangent (tanH) activation functions due to specific gate-like characteristics and suitability of these activations for handling sequential data.</p>
    <p>The sigmoid function confines data to a range of 0 to 1; when the input is large and positive, the output approaches 1 else the output approaches 0. The tanH function maintains values between −1 and 1 to regulate data flow; when the input is a large positive number, the tanH function approaches 1, whereas it approaches −1 for large negative numbers.</p>
    <p>Referring to <xref ref-type="fig" rid="fig2">
      Figure 2
     </xref>, the forget gate decides which information to keep or forget from the cell state. To make this decision, it multiplies the current input (x<sub>t</sub>) and the previous hidden state (h<sub>t</sub><sub>−</sub><sub>1</sub>) by weight matrices and adds a bias. The sigmoid function is then applied and produces a vector with values from 0 to 1, corresponding to each number present in the cell state. An output of 0 indicates forgetting and an output of 1 indicates keeping that piece of information. The vector output (f<sub>t</sub>) produced by the sigmoid function is then multiplied<sup>1</sup> to the previous cell state (C<sub>t</sub><sub>−</sub><sub>1</sub>) so that knowledge no longer needed is forgotten.</p>
    <fig id="fig2" position="float">
     <label>Figure 2</label>
     <caption>
      <title>Figure 2. Basic Structure of an LSTM unit.</title>
     </caption>
     <graphic mimetype="image" position="float" xlink:type="simple" xlink:href="https://html.scirp.org/file/4701247-rId16.jpeg?20241014013914" />
    </fig>
    <p>
     <xref ref-type="bibr" rid="scirp.136564-"></xref>Then, the input gate determines the addition of new information in the current cell state. It begins by subjecting the current input (x<sub>t</sub>) and the previous hidden state (h<sub>t</sub><sub>-1</sub>) to a sigmoid function, producing values between 0 and 1 and thus creating a regulatory filter (i<sub>t</sub>). This process is similar to the one of the forget gate. Moreover, the current input (x<sub>t</sub>) and the previous hidden state (h<sub>t</sub><sub>−</sub><sub>1</sub>) also pass through a tanH function, producing a new cell state 
     <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mover accent="true"> 
         <mi>
           C 
         </mi> 
         <mo>
           ˜ 
         </mo> 
        </mover> 
        <mi>
          t 
        </mi> 
       </msub> 
      </mrow> 
     </math>. Finally, 
     <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mover accent="true"> 
         <mi>
           C 
         </mi> 
         <mo>
           ˜ 
         </mo> 
        </mover> 
        <mi>
          t 
        </mi> 
       </msub> 
      </mrow> 
     </math>. and the value of the regulatory filter (i<sub>t</sub>) are multiplied, and this information is added to the current cell state.</p>
    <p>The final gate, known as the output gate, utilizes relevant information from the forget and input gates to decide whether to update and store new state data. The previous cell state (C<sub>t</sub><sub>−</sub><sub>1</sub>) goes through the tanH function and the result is multiplied by the output of the sigmoid function (o<sub>t</sub>). The resulting vector is the new hidden state (h<sub>t</sub>) that is output by the LSTM for the current time step.</p>
    <p>The final gate, known as the output gate, utilizes relevant information from the forget and input gates to decide whether to update and store new state data. The previous cell state (C<sub>t</sub><sub>−</sub><sub>1</sub>) goes through the tanH function and the result is multiplied by the output of the sigmoid function (o<sub>t</sub>). The resulting vector is the new hidden state (h<sub>t</sub>) that is output by the LSTM for the current time step.</p>
    <p>For our research we implemented our LSTM model with TensorFlow, a flexible and high performing end-to-end machine learning platform <xref ref-type="bibr" rid="scirp.136564-11">
      [11]
     </xref>.</p>
   </sec>
   <sec id="s2_2">
    <title>2.2. Research Dataset</title>
    <p>The SIE data used for our model was obtained from the National Snow and Ice Data Centre (NSIDC) and consisted of monthly numerical data from 1989 to 2022 <xref ref-type="bibr" rid="scirp.136564-12">
      [12]
     </xref>. NSIDC monitors and provides data on various aspects of Earth’s cryosphere, including sea ice, glaciers, and snow cover. They use a variety of methods, such as satellite observations and ground-based measurements, to track and analyze changes and other characteristics of sea ice. The SIE data was derived from satellite sensors that measured the presence of ice cover over large areas. From there, the satellite measurements were combined and processed to calculate the monthly extent of sea ice coverage in the Arctic available in numerical form.</p>
    <p>To enrich our dataset with variables influencing changes in sea ice, we used ERA5 data from Copernicus <xref ref-type="bibr" rid="scirp.136564-13">
      [13]
     </xref>. ERA5 is a reanalysis dataset that offers consistent historical weather variable estimates. ERA5 integrates data from observation satellites, ground-based weather stations, and ocean buoys to create a comprehensive dataset. ERA5 data is processed by organizing it into a grid that covers the entire Earth. This grid structure process makes the data easily accessible and usable for various applications, as it provides a structured and consistent format for the data, simplifying data analysis and interpretation.</p>
    <p>Specifically, from ERA5 we extracted the following predictor variables: sea surface temperature (SST), total precipitation (TP), surface pressure (SP), surface latent heat flux (SLHF), and surface sensible heat flux (SSHF). The selection of these variables was based on analysis from Chen <xref ref-type="bibr" rid="scirp.136564-3">
      [3]
     </xref> that confirmed correlation with SIE. For example, SST directly affects sea ice dynamics, as higher temperatures can promote further ice melt. TP, in the form of snowfall, can slow down the rate of melting through insulating ice from atmospheric heat. However, precipitation in the form of rainfall can accelerate ice melt. SP can modulate wind patterns, thus indirectly affecting sea ice. Finally, SLHF and SSHF impact the energy balance of the ice, influencing freezing and melting rate.</p>
   </sec>
   <sec id="s2_3">
    <title>2.3. Data Pre-Processing</title>
    <p>To enhance our data integrity, we applied multiple data preprocessing techniques, including normalization, dimensionality adjustments, and dataset merging. Notably, we addressed spaces in target labels and handled missing values in the data. Specifically, we handled such values with polynomial interpolation, which is regarded as a good method for interpolation of monthly and yearly climate elements <xref ref-type="bibr" rid="scirp.136564-14">
      [14]
     </xref>. This meticulous handling is crucial to ensure the robustness and accuracy of our dataset for our machine learning model.</p>
    <p>After verifying our SIE data, we shifted our focus onto the ERA5 climate variable data. ERA5 data is stored in specialized formats like NetCDF (Network Common Data Form) or GRIB (Gridded Binary). As TensorFlow doesn’t provide built-in functionality to read ERA5 climate variable data, we extracted the data from the climate dataset. After loading the ERA5 data, we first analyzed the data structure and dimensions. Originally, our dataset contained the following dimensions: longitude, latitude, EXPVER, and time. EXPVER is a dimension representing the version or ensemble member of the experiment in climate or weather modeling. Our dataset included 2 EXPVER values, indicating that there were two ensemble members or experiment versions in the subset of the NetCDF file. In addition to EXPVER, each climate variable in the NetCDF file had its own 2D array of values representing the data for each combination of longitude and latitude. Our first step was to split the dataset based on its EXPVER. We then conducted spatial aggregation to lower data dimensions based on latitude and longitude. This effectively transformed our dataset, enabling conversion to CSV format for merging with our numerical SIE data. SIE was loaded as our target variable, while ERA5 data served as our control variables.</p>
   </sec>
   <sec id="s2_4">
    <title>2.4. Data Analysis</title>
    <p>In the initial stages of our data analysis, we recognized the potential impact of extreme values, or outliers, on the performance and robustness of our LSTM model.</p>
    <p>The distribution of data plays a significant role in outlier detection. The Shapiro-Wilk test results showed that all our variables except for SP did not follow a normal distribution. Thus, for outlier detection we decided to use the Interquartile Range (IQR) as it doesn’t depend on a specific distribution and uses percentiles.</p>
    <p>To determine the presence of outliers, we performed the following calculations: first, we solved for the first quartile value (Q1) and the third quartile (Q3). Subsequently, the IQR was derived by subtracting Q1 from Q3. We utilized the following formulas to find our outliers: Q1 − 1.5 (IQR) for the lower bounds and Q3 + 1.5 (IQR) for upper bounds. Our findings are displayed in the box plots of <xref ref-type="fig" rid="fig3">
      Figure 3
     </xref> below: all our variables, except SIE and TP, contained outliers.</p>
    <p>Our next goal was to check for the relation between our target variable, SIE, and climate variables. We utilized Spearman’s rank correlation coefficient as this method is robust against the presence of outliers and doesn’t require normal distributional shape when calculating relationships between variables <xref ref-type="bibr" rid="scirp.136564-15">
      [15]
     </xref>. A Spearman coefficient is abbreviated as ρ and is calculated with the ranks of the values of each variable, instead of the actual values <xref ref-type="bibr" rid="scirp.136564-16">
      [16]
     </xref>. The value of the Spearman coefficient ranges between –1 and +1. A value of –1 indicates strong negative correlation, meaning that as one variable increases, the other tends to decrease. A value 0 means no association, and a value of +1 indicates strong positive correlation, meaning that as values of one variable increase, the values of the other variable also increase.</p>
    <fig-group id="fig3" position="float">
     <fig id="fig3" position="float">
      <label>Figure 3</label>
      <caption>
       <title>Figure 3. Box plots for the SIE and predictor variables.</title>
      </caption>
      <graphic mimetype="image" position="float" xlink:type="simple" xlink:href="https://html.scirp.org/file/4701247-rId21.jpeg?20241014013914" />
     </fig>
     <fig id="fig3" position="float">
      <label>Figure 3</label>
      <caption>
       <title>Figure 3. Box plots for the SIE and predictor variables.</title>
      </caption>
      <graphic mimetype="image" position="float" xlink:type="simple" xlink:href="https://html.scirp.org/file/4701247-rId22.jpeg?20241014013914" />
     </fig>
     <fig id="fig3" position="float">
      <label>Figure 3</label>
      <caption>
       <title>Figure 3. Box plots for the SIE and predictor variables.</title>
      </caption>
      <graphic mimetype="image" position="float" xlink:type="simple" xlink:href="https://html.scirp.org/file/4701247-rId23.jpeg?20241014013915" />
     </fig>
    </fig-group>
    <fig id="fig4" position="float">
     <label>Figure 4</label>
     <caption>
      <title>Figure 4. Heat map for SIE and predictor variables.</title>
     </caption>
     <graphic mimetype="image" position="float" xlink:type="simple" xlink:href="https://html.scirp.org/file/4701247-rId24.jpeg?20241014013914" />
    </fig>
    <p>As shown in <xref ref-type="fig" rid="fig4">
      Figure 4
     </xref>, the analysis indicates a strong negative correlation between SST and TP and SIE (correlation coefficients of −0.95 and −0.86 respectively), indicating that as SST and TP increase, SIE decreases. There is a moderate positive monotonic relationship between SLHF and SIE (correlation coefficient of 0.58), and a weak positive monotonic relationship for SP and SIE (correlation coefficient of 0.25). For SSHF and SIE, the Spearman’s rank correlation coefficient is −0.04, suggesting a very weak or negligible monotonic relationship between the variables.</p>
    <p>Next, we utilized the ExtraTrees regressor, a meta-estimator that employs randomized decision trees on dataset subsets to analyze feature contributions for target variable predictions. This method assesses both linear and nonlinear relationships and enhances predictive accuracy as it averages results and helps to mitigate overfitting risks. The output is a numerical value of each variable’s importance, where all values sum to 1. The results were as follows:</p>
    <p>Our SST variable retained its position as the most influential feature for predicting SIE, despite changes in the importance of other variables. For this analysis, our SP variable resulted as the least impactful for our prediction model, followed closely by SSHF.</p>
    <p>Combining the results from the Spearman’s rank correlation coefficients with ones from Extra Trees Regressor feature importance can provide valuable insights into the relationships between variables and confirm their importance in predicting SIE (see <xref ref-type="table" rid="table1">
      Table 1
     </xref>):</p>
    <table-wrap id="table1">
     <label>
      <xref ref-type="table" rid="table1">
       Table 1
      </xref></label>
     <caption>
      <title>
       <xref ref-type="bibr" rid="scirp.136564-"></xref>Table 1. Spearman’s correlation and extra trees regressor results.</title>
     </caption>
     <table class="MsoTableGrid custom-table" border="0" cellspacing="0" cellpadding="0"> 
      <tr> 
       <td class="custom-bottom-td acenter" width="20.26%"><p style="text-align:center"></p></td> 
       <td class="custom-bottom-td acenter" width="44.33%"><p style="text-align:center">Spearman’s Correlation</p></td> 
       <td class="custom-bottom-td acenter" width="35.41%"><p style="text-align:center">Extra Trees Regressor</p></td> 
      </tr> 
      <tr> 
       <td class="custom-top-td acenter" width="20.26%"><p style="text-align:center">Sea Surface Temperature (SST)</p></td> 
       <td class="custom-top-td aleft" width="44.33%"><p style="text-align:left">−0.95: Strong negative correlation. This suggests that as SST increases, SIE decreases.</p></td> 
       <td class="custom-top-td aleft" width="35.41%"><p style="text-align:left">0.600: High importance, the most influential feature in our dataset.</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="20.26%"><p style="text-align:center">Surface Latent Heat Flux (SLHF)</p></td> 
       <td class="aleft" width="44.33%"><p style="text-align:left">0.58: Moderate positive correlation, indicating that higher SLHF is associated with increased SIE.</p></td> 
       <td class="aleft" width="35.41%"><p style="text-align:left">0.075: Moderate importance contributing to the model’s predictive power.</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="20.26%"><p style="text-align:center">Surface Temperature (SP)</p></td> 
       <td class="aleft" width="44.33%"><p style="text-align:left">0.25 Positive correlation but weaker than SST and SLHF.</p></td> 
       <td class="aleft" width="35.41%"><p style="text-align:left">0.012: Very low importance but still contributes to the model</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="20.26%"><p style="text-align:center">Surface Sensible Heat Flux (SSHF)</p></td> 
       <td class="aleft" width="44.33%"><p style="text-align:left">−0.04: Very weak negative correlation, indicating a negligible relationship.</p></td> 
       <td class="aleft" width="35.41%"><p style="text-align:left">0.025: Low importance. It contributes to the model but is not a major influencer.</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="20.26%"><p style="text-align:center">Total Precipitation (TP)</p></td> 
       <td class="aleft" width="44.33%"><p style="text-align:left">−0.86: Strong negative correlation, indicating that higher precipitation is associated with reduced SIE.</p></td> 
       <td class="aleft" width="35.41%"><p style="text-align:left">0.288: Significant importance, the second most influential feature in our dataset.</p></td> 
      </tr> 
     </table>
    </table-wrap>
    <p>As a summary:</p>
   </sec>
  </sec><sec id="s3">
   <title>3. Building the LSTM Model</title>
   <p>In our initial study <xref ref-type="bibr" rid="scirp.136564-9">
     [9]
    </xref> we developed a LSTM model tailored for handling numerical climate data. The model demonstrated satisfactory performance with R<sup>2</sup> score of 95.48, MAE of 0.58, MSE of 0.49 and RMSE of 0.70, showcasing its efficacy in SIE prediction. Experimentation showed that the best performance was achieved with 78 LSTM units and 64 batch size.</p>
   <p>In our current research, we systematically investigate several other crucial dimensions of our model, in addition to LSTM units, batch sizes, and adjustable learning rate to evaluate the impact on its predictive capabilities.</p>
   <sec id="s3_1">
    <title>3.1. Model Optimization</title>
    <p>Our LSTM model’s optimization journey encompasses the key areas described below.</p>
    <p>We explore the impact of varying historical context by creating data sequences with 6 and 12 past months. Chen <xref ref-type="bibr" rid="scirp.136564-3">
      [3]
     </xref> showed that when using multiple linear regression (MLR), support vector regression (SVR), and random forest regression (RFR) techniques, a 12-month sea ice data resulted in a better forecast for SIE than using 6-month data. As shown further below, our analysis provides valuable insights into the significance of historical data length for accurate predictions in LSTM models.</p>
    <p>The number of LSTM units determines the complexity of the model. More units allow the model to learn more complex patterns in the data but may also lead to overfitting if the model becomes too complex for the given dataset. We also experiment with batch size, an important hyper-parameter in deep learning that represents the number of samples that work through the network on each iteration before updating the internal model parameters. The batch size has a direct impact on the accuracy and computational efficiency of the training process. Large batch sizes can provide a smoother gradient signal and lead to faster training times. However, it has been observed that using larger batches lead to the decline of a model’s ability to generalize <xref ref-type="bibr" rid="scirp.136564-17">
      [17]
     </xref>. On the other hand, it has been demonstrated that smaller batches allow a significantly smaller memory footprint and introduce noise in the optimization, leading in improved generalization performance <xref ref-type="bibr" rid="scirp.136564-18">
      [18]
     </xref>. This yields a more stable and reliable training.</p>
    <p>As presented earlier, our data doesn’t follow a Normal (Gaussian) distribution and contains outliers. Standardization is suitable for data that has a normal distribution, whereas Normalization (MinMax) is preferred when data does not follow a normal distribution but is sensitive to outliers. Robust scaling is used when data has many outliers <xref ref-type="bibr" rid="scirp.136564-19">
      [19]
     </xref>-<xref ref-type="bibr" rid="scirp.136564-21">
      [21]
     </xref>. For our data scaling, we experiment with all three scalers.</p>
    <p>To capture the temporal patterns inherent in climate data, we test the inclusion of the month variable into our data series. We aim to refine the model’s understanding of seasonal variations, a critical factor in climate prediction. We also experiment with the inclusion of past SIE information to check how this could help enhance our model’s ability to capture temporal dependencies in the data.</p>
    <p>In addition to the adjustable learning rate, we experiment with model checkpoint and early stopping techniques. Model checkpoint ensures the preservation of the best-performing model during training, while early stopping prevents overfitting by stopping training when validation loss starts increasing, leading to a more generalizable and robust LSTM model. Overfitting occurs when a model learns to fit the training data very closely, but it performs poorly on unseen or test data.</p>
    <p>Regularization techniques are also used to prevent overfitting and improve the generalization of a model. For example, the Elastic Net regularization combines the advantages of both L1 and L2 regularization. The L1 encourages the model to focus on a subset of the most relevant features, effectively ignoring irrelevant or redundant ones, thus reducing the complexity of the model. On the other hand, the L2 regularization prevents any one feature from dominating the model’s predictions and improves the model’s stability <xref ref-type="bibr" rid="scirp.136564-22">
      [22]
     </xref> <xref ref-type="bibr" rid="scirp.136564-23">
      [23]
     </xref>.</p>
    <p>As mentioned earlier, LSTM networks traditionally employ the sigmoid function for the gate activation and tanH function for the cell state operations, due to specific gate-like characteristics and suitability of these activations for handling sequential data. However, both sigmoid and tanH functions present the vanishing gradient problem. The Rectified Linear Unit (ReLU) activation was introduced to solve this problem. It is linear for all values greater than or equal to zero, which enables it to learn faster. However, any input value that is below or equal to zero causes the ReLU to produce an output of zero, leading to zero gradients during backpropagation, a problem known as the dying ReLU problem. The Parametric ReLU (PReLU) improves ReLU by adding a slope for negative values, thus allowing the algorithm to learn from negative data regions.</p>
    <p>Experimenting with activation functions won’t be a focus in this study. Nevertheless, due to having both positive and negative values in our dataset, we checked our best performing model with the Parametric ReLU (PReLU) for the LSTM layer.</p>
    <p>The primary goal is to build a model that predicts a continuous numerical value, that is of the SIE. Since the focus is on producing an output that directly represents the SIE value, we will only use the default linear function for the output layer.</p>
   </sec>
   <sec id="s3_2">
    <title>3.2. Evaluation Metrics</title>
    <p>In our study, we conduct a comprehensive evaluation of our model’s performance using the following standard deterministic accuracy metrics: Mean Absolute Error (MAE), R-squared value (R<sup>2</sup>), Mean Square Error (MSE), and Root Mean Square Error (RMSE).</p>
    <p>MAE measures the average absolute differences between predicted values and actual values in a dataset. R<sup>2</sup> represents the proportion of the variance in the dependent variable that is predictable from the independent variables in a regression model. MSE measures the average of the squared differences between predicted values and actual values in a regression problem. Finally, RMSE provides a more interpretable measure by taking the square root of the MSE.</p>
    <p>As our data analysis demonstrates, our data contains outliers. Since the squaring of errors enforces a higher importance on outliers, MSE and RMSE may be less robust for our model compared to MAE <xref ref-type="bibr" rid="scirp.136564-24">
      [24]
     </xref>.</p>
   </sec>
   <sec id="s3_3">
    <title>3.3. Data Preparation</title>
    <p>Data splitting into training, validation, and test sets is essential for assessing and fine-tuning machine learning models, preventing overfitting, choosing the best model, and ensuring they can make accurate predictions on new, unseen data. It’s a fundamental practice for robust and reliable model development. Since our data is a time series, we used a temporal split: allocated the earlier portion of the data for training (80%), a middle portion for validation (10%), and the most recent portion for testing (10%).</p>
    <p>We organize our data in sequences, with each sequence containing a series of past data points of the SIE and predicting variables, and the corresponding target representing the next data point for the SIE. The inclusion of the SIE from the previous month was important to account for temporal dependencies in the data.</p>
    <p>For example, in the case of sequences considering three lagged time steps, the structure of the inputs and outputs would be as follows:</p>
    <table class="MsoTableGrid custom-table" border="0" cellspacing="0" cellpadding="0"> 
     <tr> 
      <td rowspan="3" class="aleft"><p style="text-align:left">X</p></td> 
      <td class="aleft"><p style="text-align:left">[[SIE(t-3), var1(t-3), var2(t-3) …]</p></td> 
     </tr> 
     <tr> 
      <td class="aleft"><p style="text-align:left">[SIE(t-2), var1(t-2), var2(t-2) …]</p></td> 
     </tr> 
     <tr> 
      <td class="aleft"><p style="text-align:left">[SIE(t-1), var1(t-1), var2(t-1) …]]</p></td> 
     </tr> 
     <tr> 
      <td class="aleft"><p style="text-align:left">Y</p></td> 
      <td class="aleft"><p style="text-align:left">SIE(t)</p></td> 
     </tr> 
    </table>
   </sec>
  </sec><sec id="s4">
   <title>4. LSTM Model</title>
   <p>In this section we present the steps we take for exploring performance improvements to our supervised regression learning problem: predicting the SIE at the current month (t) using the SIE and environmental variables values from prior time steps. We will only highlight the configurations that yield the best results. Our baseline, stateful LSTM model was built with two dense layers and utilized the default activation methods. The first is the LSTM layer, using sigmoid and tanH activations, followed by an output layer with a single unit that uses the linear activation function.</p>
   <p>model = Sequential()</p>
   <p>model.add(LSTM(lstm_units, input_shape=(n_past_months, n_features)))</p>
   <p>model.add(Dense(1)) # Output layer for regression</p>
   <p>model.compile(loss='mae', optimizer='adam')</p>
   <p>We also included the ReduceLROnPlateau scheduler, a learning rate adjustment technique that operates by monitoring a specific metric, in our case the validation loss. If the validation loss does not improve after the number of epochs defined by the “patience” variable, the scheduler reduces the learning rate to 10% of its current value (factor parameter). The learning rate will never go below min_lr.</p>
   <p>ReduceLROnPlateau(monitor='val_loss', factor=0.1, patience=20, verbose=1, min_lr=1e-6)</p>
   <p>The first improvement we wanted to make was to train our model with an extended historical context of 6 and 12 months (i.e., 6 and 12 lagged time steps). We also experimented with different scalers to optimize our model’s performance by tailoring the data preprocessing to our specific dataset. The Robust Scaler scales the data in such a way that makes it more robust to outliers by using the median and interquartile range (IQR) rather than the mean and standard deviation. Standard and MinMax scalers are both not robust to outliers.</p>
   <p>
    <xref ref-type="table" rid="table2">
     Table 2
    </xref> below shows the best performance results for 6 and 12 lagged time steps for all scalers we tested. We obtained best MAE for 12 lagged time steps, indicating better predictive performance. This observation is consistent with the findings of Chen et all for their support vector, random forest, and multiple linear regression models. Out of the two scalers that are not robust to outliers, Standard scaler had superior results. Moreover, considering all performance indicators, scaling with Robust scaler provided better results than with MinMax scaler. Thus, we decided to check performance with both Standard and Robust scalers for the next tests.</p>
   <table-wrap id="table2">
    <label>
     <xref ref-type="table" rid="table2">
      Table 2
     </xref></label>
    <caption>
     <title>
      <xref ref-type="bibr" rid="scirp.136564-"></xref>Table 2. Best performance results with learning rate adjustment technique.</title>
    </caption>
    <table class="MsoTableGrid custom-table" border="0" cellspacing="0" cellpadding="0"> 
     <tr> 
      <td class="custom-bottom-td acenter" width="15.86%"><p style="text-align:center">Data Scaler</p></td> 
      <td class="custom-bottom-td acenter" width="15.87%"><p style="text-align:center">Months</p></td> 
      <td class="custom-bottom-td acenter" width="13.49%"><p style="text-align:center">LSTM Units</p></td> 
      <td class="custom-bottom-td acenter" width="15.41%"><p style="text-align:center">Batch Size</p></td> 
      <td class="custom-bottom-td acenter" width="10.55%"><p style="text-align:center">MAE</p></td> 
      <td class="custom-bottom-td acenter" width="8.65%"><p style="text-align:center">R<sup>2</sup></p></td> 
      <td class="custom-bottom-td acenter" width="9.29%"><p style="text-align:center">MSE</p></td> 
      <td class="custom-bottom-td acenter" width="10.88%"><p style="text-align:center">RMSE</p></td> 
     </tr> 
     <tr> 
      <td class="custom-top-td acenter" width="15.86%"><p style="text-align:center">Standard</p></td> 
      <td class="custom-top-td acenter" width="15.87%"><p style="text-align:center">6</p></td> 
      <td class="custom-top-td acenter" width="13.49%"><p style="text-align:center">64</p></td> 
      <td class="custom-top-td acenter" width="15.41%"><p style="text-align:center">32</p></td> 
      <td class="custom-top-td acenter" width="10.55%"><p style="text-align:center">0.485</p></td> 
      <td class="custom-top-td acenter" width="8.65%"><p style="text-align:center">0.98</p></td> 
      <td class="custom-top-td acenter" width="9.29%"><p style="text-align:center">0.235</p></td> 
      <td class="custom-top-td acenter" width="10.88%"><p style="text-align:center">0.484</p></td> 
     </tr> 
     <tr> 
      <td class="acenter" width="15.86%"><p style="text-align:center">Standard</p></td> 
      <td class="acenter" width="15.87%"><p style="text-align:center">12</p></td> 
      <td class="acenter" width="13.49%"><p style="text-align:center">96</p></td> 
      <td class="acenter" width="15.41%"><p style="text-align:center">64</p></td> 
      <td class="acenter" width="10.55%"><p style="text-align:center">0.334</p></td> 
      <td class="acenter" width="8.65%"><p style="text-align:center">0.99</p></td> 
      <td class="acenter" width="9.29%"><p style="text-align:center">0.181</p></td> 
      <td class="acenter" width="10.88%"><p style="text-align:center">0.425</p></td> 
     </tr> 
     <tr> 
      <td class="acenter" width="15.86%"><p style="text-align:center">MinMax</p></td> 
      <td class="acenter" width="15.87%"><p style="text-align:center">6</p></td> 
      <td class="acenter" width="13.49%"><p style="text-align:center">78</p></td> 
      <td class="acenter" width="15.41%"><p style="text-align:center">64</p></td> 
      <td class="acenter" width="10.55%"><p style="text-align:center">0.743</p></td> 
      <td class="acenter" width="8.65%"><p style="text-align:center">0.95</p></td> 
      <td class="acenter" width="9.29%"><p style="text-align:center">0.575</p></td> 
      <td class="acenter" width="10.88%"><p style="text-align:center">0.758</p></td> 
     </tr> 
     <tr> 
      <td class="acenter" width="15.86%"><p style="text-align:center">MinMax</p></td> 
      <td class="acenter" width="15.87%"><p style="text-align:center">12</p></td> 
      <td class="acenter" width="13.49%"><p style="text-align:center">78</p></td> 
      <td class="acenter" width="15.41%"><p style="text-align:center">32</p></td> 
      <td class="acenter" width="10.55%"><p style="text-align:center">0.467</p></td> 
      <td class="acenter" width="8.65%"><p style="text-align:center">0.98</p></td> 
      <td class="acenter" width="9.29%"><p style="text-align:center">0.273</p></td> 
      <td class="acenter" width="10.88%"><p style="text-align:center">0.522</p></td> 
     </tr> 
     <tr> 
      <td class="acenter" width="15.86%"><p style="text-align:center">Robust</p></td> 
      <td class="acenter" width="15.87%"><p style="text-align:center">6</p></td> 
      <td class="acenter" width="13.49%"><p style="text-align:center">64</p></td> 
      <td class="acenter" width="15.41%"><p style="text-align:center">32</p></td> 
      <td class="acenter" width="10.55%"><p style="text-align:center">0.571</p></td> 
      <td class="acenter" width="8.65%"><p style="text-align:center">0.97</p></td> 
      <td class="acenter" width="9.29%"><p style="text-align:center">0.330</p></td> 
      <td class="acenter" width="10.88%"><p style="text-align:center">0.574</p></td> 
     </tr> 
     <tr> 
      <td class="acenter" width="15.86%"><p style="text-align:center">Robust</p></td> 
      <td class="acenter" width="15.87%"><p style="text-align:center">12</p></td> 
      <td class="acenter" width="13.49%"><p style="text-align:center">64</p></td> 
      <td class="acenter" width="15.41%"><p style="text-align:center">64</p></td> 
      <td class="acenter" width="10.55%"><p style="text-align:center">0.480</p></td> 
      <td class="acenter" width="8.65%"><p style="text-align:center">0.98</p></td> 
      <td class="acenter" width="9.29%"><p style="text-align:center">0.232</p></td> 
      <td class="acenter" width="10.88%"><p style="text-align:center">0.482</p></td> 
     </tr> 
    </table>
   </table-wrap>
   <p>LSTM networks are well-suited for sequence prediction tasks where the order and timing of events matter. To enhance the model’s performance and provide additional context regarding how data changes over time, we decided to conduct further experiments with the incorporation of seasonal information. Seasonal information can offer valuable insights into recurring patterns within the data.</p>
   <p>To achieve this, we included month information using cyclical encoding. Cyclical encoding is a methodology used to represent cyclical patterns, such as those related to time of day, months, or seasons. This technique ensures that the values “wrap around” in a circular fashion, preserving the inherent cyclical nature of the data. By encoding the months in this manner, we aimed to provide our LSTM model with a deeper understanding of how seasonal changes influence the data, ultimately testing its impact on prediction accuracy. Cyclical encoding can be done using the following transformations <xref ref-type="bibr" rid="scirp.136564-25">
     [25]
    </xref>.</p>
   <p>
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
      <msub> 
       <mi>
         x 
       </mi> 
       <mrow> 
        <mi>
          sin 
        </mi> 
       </mrow> 
      </msub> 
      <mo>
        = 
      </mo> 
      <mi>
        sin 
      </mi> 
      <mrow> 
       <mo>
         ( 
       </mo> 
       <mrow> 
        <mrow> 
         <mrow> 
          <mn>
            2 
          </mn> 
          <mi>
            π 
          </mi> 
          <mo>
            * 
          </mo> 
          <mi>
            x 
          </mi> 
         </mrow> 
         <mo>
           / 
         </mo> 
         <mrow> 
          <mi>
            max 
          </mi> 
          <mrow> 
           <mo>
             ( 
           </mo> 
           <mi>
             x 
           </mi> 
           <mo>
             ) 
           </mo> 
          </mrow> 
         </mrow> 
        </mrow> 
       </mrow> 
       <mo>
         ) 
       </mo> 
      </mrow> 
     </mrow> 
    </math> (1)</p>
   <p>
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
      <msub> 
       <mi>
         x 
       </mi> 
       <mrow> 
        <mi>
          cos 
        </mi> 
       </mrow> 
      </msub> 
      <mo>
        = 
      </mo> 
      <mi>
        cos 
      </mi> 
      <mrow> 
       <mo>
         ( 
       </mo> 
       <mrow> 
        <mrow> 
         <mrow> 
          <mn>
            2 
          </mn> 
          <mi>
            π 
          </mi> 
          <mo>
            * 
          </mo> 
          <mi>
            x 
          </mi> 
         </mrow> 
         <mo>
           / 
         </mo> 
         <mrow> 
          <mi>
            max 
          </mi> 
          <mrow> 
           <mo>
             ( 
           </mo> 
           <mi>
             x 
           </mi> 
           <mo>
             ) 
           </mo> 
          </mrow> 
         </mrow> 
        </mrow> 
       </mrow> 
       <mo>
         ) 
       </mo> 
      </mrow> 
     </mrow> 
    </math> (2)</p>
   <p>For our monthly data, let 
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
      <mi>
        M 
      </mi> 
      <mo>
        ≡ 
      </mo> 
      <mtext>
        month 
      </mtext> 
     </mrow> 
    </math>, where 
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
      <mi>
        M 
      </mi> 
      <mo>
        = 
      </mo> 
      <mn>
        1 
      </mn> 
      <mo>
        , 
      </mo> 
      <mo>
        ⋯ 
      </mo> 
      <mo>
        , 
      </mo> 
      <mn>
        12 
      </mn> 
     </mrow> 
    </math>, and</p>
   <p>
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
      <msub> 
       <mi>
         M 
       </mi> 
       <mrow> 
        <mi>
          sin 
        </mi> 
       </mrow> 
      </msub> 
      <mo>
        = 
      </mo> 
      <mi>
        sin 
      </mi> 
      <mrow> 
       <mo>
         ( 
       </mo> 
       <mrow> 
        <mrow> 
         <mrow> 
          <mi>
            π 
          </mi> 
          <mo>
            * 
          </mo> 
          <mi>
            M 
          </mi> 
         </mrow> 
         <mo>
           / 
         </mo> 
         <mn>
           6 
         </mn> 
        </mrow> 
       </mrow> 
       <mo>
         ) 
       </mo> 
      </mrow> 
     </mrow> 
    </math> (3)</p>
   <p>
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
      <msub> 
       <mi>
         M 
       </mi> 
       <mrow> 
        <mi>
          cos 
        </mi> 
       </mrow> 
      </msub> 
      <mo>
        = 
      </mo> 
      <mi>
        cos 
      </mi> 
      <mrow> 
       <mo>
         ( 
       </mo> 
       <mrow> 
        <mrow> 
         <mrow> 
          <mi>
            π 
          </mi> 
          <mo>
            * 
          </mo> 
          <mi>
            M 
          </mi> 
         </mrow> 
         <mo>
           / 
         </mo> 
         <mn>
           6 
         </mn> 
        </mrow> 
       </mrow> 
       <mo>
         ) 
       </mo> 
      </mrow> 
     </mrow> 
    </math> (4)</p>
   <p>For a 12-month historical context, our model performed best with 96 LSTM units and 64 batch size (see <xref ref-type="fig" rid="fig5">
     Figure 5
    </xref> and <xref ref-type="fig" rid="fig6">
     Figure 6
    </xref>). Specifically, we obtained the following results:</p>
   <p>Standard Scaler: MAE: 0.399 || R<sup>2</sup>: 0.993 || MSE: 0.086 || RMSE: 0.294</p>
   <p>Robust Scaler: MAE: 0.230 || R<sup>2</sup>: 0.993 || MSE: 0.088 || RMSE: 0.296</p>
   <p>We also experimented with the inclusion of yearly information. However, this did not improve our results and therefore we excluded this from our dataset.</p>
   <fig id="fig5" position="float">
    <label>Figure 5</label>
    <caption>
     <title>Figure 5. Inclusion of Month in Dataset, performance for 12-month historical context—Standard Scaler.</title>
    </caption>
    <graphic mimetype="image" position="float" xlink:type="simple" xlink:href="https://html.scirp.org/file/4701247-rId37.jpeg?20241014013916" />
   </fig>
   <p>The next step was to utilize the early stopping technique. Early stopping monitors the model’s performance on a validation dataset during training and stops training when the performance starts to degrade, preventing overfitting. If the validation loss stops improving or begins to worsen, considering the “patience”, training is halted. Therefore, it helps in preventing the model from learning noise in the training data, saving time and resources, and often results in better generalization to unseen data.</p>
   <fig id="fig6" position="float">
    <label>Figure 6</label>
    <caption>
     <title>Figure 6. Inclusion of Month in Dataset, performance for 12-month historical context—Robust Scaler.</title>
    </caption>
    <graphic mimetype="image" position="float" xlink:type="simple" xlink:href="https://html.scirp.org/file/4701247-rId38.jpeg?20241014013916" />
   </fig>
   <p>from tensorflow.keras.callbacks import EarlyStopping</p>
   <p>early_stopping = EarlyStopping(monitor='val_loss', patience=10, restore_best_weights=True)</p>
   <p>history = model.fit(train_X, train_y, epochs=250, batch_size=64,</p>
   <p>validation_data=(validation_X, validation_y),</p>
   <p>callbacks=[early_stopping, reduce_lr], verbose=0,</p>
   <p>shuffle=False)</p>
   <p>We tested three different values for patience: 10, 15, and 20 (see <xref ref-type="table" rid="table3">
     Table 3
    </xref>). For Standard scaler we obtained better performance results for patience equal to 20. For Robust scaler, the performance results for 15 and 20 patience were very close. There was a slight improvement on MAE with patience 15, which indicated that our model was able to achieve better performance in terms of absolute prediction accuracy on the validation data.</p>
   <p>Thus, for our subsequent testing we decided to use patience 20 for Standard scaler and patience 15 for Robust scaler.</p>
   <table-wrap id="table3">
    <label>
     <xref ref-type="table" rid="table3">
      Table 3
     </xref></label>
    <caption>
     <title>
      <xref ref-type="bibr" rid="scirp.136564-"></xref>Table 3. Inclusion of early stopping—Performance results for patience 10, 15, 20.</title>
    </caption>
    <table class="MsoTableGrid custom-table" border="0" cellspacing="0" cellpadding="0"> 
     <tr> 
      <td class="custom-bottom-td acenter" width="18.07%"><p style="text-align:center">Data Scaler</p></td> 
      <td class="custom-bottom-td acenter" width="32.79%"><p style="text-align:center">Early Stopping Patience</p></td> 
      <td class="custom-bottom-td acenter" width="10.69%"><p style="text-align:center">MAE</p></td> 
      <td class="custom-bottom-td acenter" width="14.95%"><p style="text-align:center">R<sup>2</sup></p></td> 
      <td class="custom-bottom-td acenter" width="10.69%"><p style="text-align:center">MSE</p></td> 
      <td class="custom-bottom-td acenter" width="12.81%"><p style="text-align:center">RMSE</p></td> 
     </tr> 
     <tr> 
      <td class="custom-top-td acenter" width="18.07%"><p style="text-align:center">Standard</p></td> 
      <td class="custom-top-td acenter" width="32.79%"><p style="text-align:center">10</p></td> 
      <td class="custom-top-td acenter" width="10.69%"><p style="text-align:center">0.272</p></td> 
      <td class="custom-top-td acenter" width="14.95%"><p style="text-align:center">0.991</p></td> 
      <td class="custom-top-td acenter" width="10.69%"><p style="text-align:center">0.109</p></td> 
      <td class="custom-top-td acenter" width="12.81%"><p style="text-align:center">0.33</p></td> 
     </tr> 
     <tr> 
      <td class="acenter" width="18.07%"><p style="text-align:center">Standard</p></td> 
      <td class="acenter" width="32.79%"><p style="text-align:center">15</p></td> 
      <td class="acenter" width="10.69%"><p style="text-align:center">0.305</p></td> 
      <td class="acenter" width="14.95%"><p style="text-align:center">0.989</p></td> 
      <td class="acenter" width="10.69%"><p style="text-align:center">0.133</p></td> 
      <td class="acenter" width="12.81%"><p style="text-align:center">0.364</p></td> 
     </tr> 
     <tr> 
      <td class="acenter" width="18.07%"><p style="text-align:center">Standard</p></td> 
      <td class="acenter" width="32.79%"><p style="text-align:center">20</p></td> 
      <td class="acenter" width="10.69%"><p style="text-align:center">0.260</p></td> 
      <td class="acenter" width="14.95%"><p style="text-align:center">0.992</p></td> 
      <td class="acenter" width="10.69%"><p style="text-align:center">0.095</p></td> 
      <td class="acenter" width="12.81%"><p style="text-align:center">0.309</p></td> 
     </tr> 
     <tr> 
      <td class="acenter" width="18.07%"><p style="text-align:center">Robust</p></td> 
      <td class="acenter" width="32.79%"><p style="text-align:center">10</p></td> 
      <td class="acenter" width="10.69%"><p style="text-align:center">0.396</p></td> 
      <td class="acenter" width="14.95%"><p style="text-align:center">0.979</p></td> 
      <td class="acenter" width="10.69%"><p style="text-align:center">0.267</p></td> 
      <td class="acenter" width="12.81%"><p style="text-align:center">0.517</p></td> 
     </tr> 
     <tr> 
      <td class="acenter" width="18.07%"><p style="text-align:center">Robust</p></td> 
      <td class="acenter" width="32.79%"><p style="text-align:center">15</p></td> 
      <td class="acenter" width="10.69%"><p style="text-align:center">0.236</p></td> 
      <td class="acenter" width="14.95%"><p style="text-align:center">0.994</p></td> 
      <td class="acenter" width="10.69%"><p style="text-align:center">0.076</p></td> 
      <td class="acenter" width="12.81%"><p style="text-align:center">0.275</p></td> 
     </tr> 
     <tr> 
      <td class="acenter" width="18.07%"><p style="text-align:center">Robust</p></td> 
      <td class="acenter" width="32.79%"><p style="text-align:center">20</p></td> 
      <td class="acenter" width="10.69%"><p style="text-align:center">0.242</p></td> 
      <td class="acenter" width="14.95%"><p style="text-align:center">0.994</p></td> 
      <td class="acenter" width="10.69%"><p style="text-align:center">0.074</p></td> 
      <td class="acenter" width="12.81%"><p style="text-align:center">0.272</p></td> 
     </tr> 
    </table>
   </table-wrap>
   <p>The next step was to further explore our model’s performance with the inclusion of the model checkpoint technique. Model checkpoint saves the model’s weights and architecture periodically during training, typically after each epoch or after a certain number of training steps. This feature enables the ability to resume training from a specific point if the training process is interrupted or to select the best performing model for later use. Therefore, this provides access to the model’s best state during training, even if the training process is interrupted. Additionally, it allows for the evaluation of the model’s performance on different datasets or at different points in time.</p>
   <p>Setting the monitor equal to “val_loss” allowed the callback to save the weights when it observed an improvement in validation loss. The callback only saved the model weights if the validation loss had improved compared to the previous best value. We then loaded the best performing weights to make the predictions.</p>
   <p>The order in which we utilized early_stopping, reduce_lr, and checkpoint callbacks is important. First, we wanted to monitor the model’s performance with early_stopping. If training continues without early stopping, our model could consider reducing the learning rate with reduce_lr, when needed. Finally, we placed checkpoint last to save the model’s weights after the other callbacks run.</p>
   <p>checkpoint=ModelCheckpoint(filepath='best_model_weights.h5', monitor='val_loss', save_best_only=True)</p>
   <p>history = model.fit(train_X, train_y, epochs=250, batch_size=64,</p>
   <p>validation_data=(validation_X, validation_y),</p>
   <p>callbacks=[early_stopping, reduce_lr, checkpoint], verbose=0, shuffle=False)</p>
   <p># Load the best model weights</p>
   <p>model.load_weights('best_model_weights.h5')</p>
   <p>The combination of model checkpoint and early stopping enabled us to balance two important but somewhat conflicting objectives: preventing overfitting and ensuring that we save the best model. Therefore, we expected that it would not only ensure the stability and reliability of our model but would also optimize the use of computational resources, as unnecessary iterations would be prevented.</p>
   <p>The performance results for standard scaler deteriorated when we combined early stopping and model checkpoint. Both the MAE and MSE increased by approximately 8.46% and 25.26% respectively. In contrast, the results with Robust scaler improved; MAE and MSE decreased by approximately 10% and 6.7%. The difference in model performance with different scalers alongside model checkpoint and early stopping is something worth investigating more.</p>
   <p>Standard Scaler: MAE: 0.282 || R<sup>2</sup>: 0.99 || MSE: 0.119 || RMSE: 0.345</p>
   <p>Robust Scaler: MAE: 0.212 || R<sup>2</sup>: 0.994 || MSE: 0.069 || RMSE: 0.263</p>
   <p>Overall, with Robust scaler, combined with early stopping and model checkpoint, the model’s performance results prove that the predictions are very close to the actual values (MAE: 0.212, MSE: 0.069) and the model captures almost all the underlying patterns in the data (R<sup>2</sup>: 0.994). The low RMSE (0.263) also indicates a tight fit of our model to the data. <xref ref-type="fig" rid="fig7">
     Figure 7
    </xref> below shows the train and test loss and the actual versus predicted values SIE graph. <xref ref-type="table" rid="table4">
     Table 4
    </xref> also shows the actual (test dataset) versus predicted SIE values.</p>
   <fig id="fig7" position="float">
    <label>Figure 7</label>
    <caption>
     <title>Figure 7. Robust scaler: Inclusion of Month in Dataset—Performance for 12 month window, Early Stopping and Model checkpoint.</title>
    </caption>
    <graphic mimetype="image" position="float" xlink:type="simple" xlink:href="https://html.scirp.org/file/4701247-rId39.jpeg?20241014013916" />
   </fig>
   <table-wrap id="table4">
    <label>
     <xref ref-type="table" rid="table4">
      Table 4
     </xref></label>
    <caption>
     <title>
      <xref ref-type="bibr" rid="scirp.136564-"></xref>Table 4. Actual versus Predicted values for Robust scaler.</title>
    </caption>
    <table class="MsoTableGrid custom-table" border="0" cellspacing="0" cellpadding="0"> 
     <tr> 
      <td class="custom-bottom-td acenter" width="24.98%"><p style="text-align:center">Actual</p></td> 
      <td class="custom-bottom-td acenter" width="25.01%"><p style="text-align:center">Predicted</p></td> 
      <td class="custom-bottom-td acenter" width="25.01%"><p style="text-align:center">Actual</p></td> 
      <td class="custom-bottom-td acenter" width="25.01%"><p style="text-align:center">Predicted</p></td> 
     </tr> 
     <tr> 
      <td class="custom-top-td acenter" width="24.98%"><p style="text-align:center">7.29</p></td> 
      <td class="custom-top-td acenter" width="25.01%"><p style="text-align:center">7.98</p></td> 
      <td class="custom-top-td acenter" width="25.01%"><p style="text-align:center">6.82</p></td> 
      <td class="custom-top-td acenter" width="25.01%"><p style="text-align:center">6.88</p></td> 
     </tr> 
     <tr> 
      <td class="acenter" width="24.98%"><p style="text-align:center">5.07</p></td> 
      <td class="acenter" width="25.01%"><p style="text-align:center">5.34</p></td> 
      <td class="acenter" width="25.01%"><p style="text-align:center">9.83</p></td> 
      <td class="acenter" width="25.01%"><p style="text-align:center">9.79</p></td> 
     </tr> 
     <tr> 
      <td class="acenter" width="24.98%"><p style="text-align:center">4</p></td> 
      <td class="acenter" width="25.01%"><p style="text-align:center">4.26</p></td> 
      <td class="acenter" width="25.01%"><p style="text-align:center">12.15</p></td> 
      <td class="acenter" width="25.01%"><p style="text-align:center">12.19</p></td> 
     </tr> 
     <tr> 
      <td class="acenter" width="24.98%"><p style="text-align:center">5.33</p></td> 
      <td class="acenter" width="25.01%"><p style="text-align:center">5.91</p></td> 
      <td class="acenter" width="25.01%"><p style="text-align:center">13.87</p></td> 
      <td class="acenter" width="25.01%"><p style="text-align:center">13.58</p></td> 
     </tr> 
     <tr> 
      <td class="acenter" width="24.98%"><p style="text-align:center">8.99</p></td> 
      <td class="acenter" width="25.01%"><p style="text-align:center">9.04</p></td> 
      <td class="acenter" width="25.01%"><p style="text-align:center">14.61</p></td> 
      <td class="acenter" width="25.01%"><p style="text-align:center">14.45</p></td> 
     </tr> 
     <tr> 
      <td class="acenter" width="24.98%"><p style="text-align:center">11.73</p></td> 
      <td class="acenter" width="25.01%"><p style="text-align:center">11.57</p></td> 
      <td class="acenter" width="25.01%"><p style="text-align:center">14.59</p></td> 
      <td class="acenter" width="25.01%"><p style="text-align:center">14.84</p></td> 
     </tr> 
     <tr> 
      <td class="acenter" width="24.98%"><p style="text-align:center">13.5</p></td> 
      <td class="acenter" width="25.01%"><p style="text-align:center">13.14</p></td> 
      <td class="acenter" width="25.01%"><p style="text-align:center">13.99</p></td> 
      <td class="acenter" width="25.01%"><p style="text-align:center">13.82</p></td> 
     </tr> 
     <tr> 
      <td class="acenter" width="24.98%"><p style="text-align:center">14.39</p></td> 
      <td class="acenter" width="25.01%"><p style="text-align:center">14.21</p></td> 
      <td class="acenter" width="25.01%"><p style="text-align:center">12.88</p></td> 
      <td class="acenter" width="25.01%"><p style="text-align:center">12.68</p></td> 
     </tr> 
     <tr> 
      <td class="acenter" width="24.98%"><p style="text-align:center">14.66</p></td> 
      <td class="acenter" width="25.01%"><p style="text-align:center">14.46</p></td> 
      <td class="acenter" width="25.01%"><p style="text-align:center">10.88</p></td> 
      <td class="acenter" width="25.01%"><p style="text-align:center">11.03</p></td> 
     </tr> 
     <tr> 
      <td class="acenter" width="24.98%"><p style="text-align:center">13.79</p></td> 
      <td class="acenter" width="25.01%"><p style="text-align:center">13.97</p></td> 
      <td class="acenter" width="25.01%"><p style="text-align:center">8.29</p></td> 
      <td class="acenter" width="25.01%"><p style="text-align:center">8.19</p></td> 
     </tr> 
     <tr> 
      <td class="acenter" width="24.98%"><p style="text-align:center">12.68</p></td> 
      <td class="acenter" width="25.01%"><p style="text-align:center">12.35</p></td> 
      <td class="acenter" width="25.01%"><p style="text-align:center">5.95</p></td> 
      <td class="acenter" width="25.01%"><p style="text-align:center">5.53</p></td> 
     </tr> 
     <tr> 
      <td class="acenter" width="24.98%"><p style="text-align:center">10.77</p></td> 
      <td class="acenter" width="25.01%"><p style="text-align:center">10.35</p></td> 
      <td class="acenter" width="25.01%"><p style="text-align:center">4.9</p></td> 
      <td class="acenter" width="25.01%"><p style="text-align:center">4.85</p></td> 
     </tr> 
     <tr> 
      <td class="acenter" width="24.98%"><p style="text-align:center">7.65</p></td> 
      <td class="acenter" width="25.01%"><p style="text-align:center">7.91</p></td> 
      <td class="acenter" width="25.01%"><p style="text-align:center">6.66</p></td> 
      <td class="acenter" width="25.01%"><p style="text-align:center">6.70</p></td> 
     </tr> 
     <tr> 
      <td class="acenter" width="24.98%"><p style="text-align:center">5.71</p></td> 
      <td class="acenter" width="25.01%"><p style="text-align:center">5.83</p></td> 
      <td class="acenter" width="25.01%"><p style="text-align:center">9.73</p></td> 
      <td class="acenter" width="25.01%"><p style="text-align:center">9.64</p></td> 
     </tr> 
     <tr> 
      <td class="acenter" width="24.98%"><p style="text-align:center">4.95</p></td> 
      <td class="acenter" width="25.01%"><p style="text-align:center">5.10</p></td> 
      <td class="acenter" width="25.01%"><p style="text-align:center">11.89</p></td> 
      <td class="acenter" width="25.01%"><p style="text-align:center">12.00</p></td> 
     </tr> 
     <tr> 
      <td class="acenter" width="24.98%"><p style="text-align:center">7.29</p></td> 
      <td class="acenter" width="25.01%"><p style="text-align:center">7.98</p></td> 
      <td class="acenter" width="25.01%"><p style="text-align:center">6.82</p></td> 
      <td class="acenter" width="25.01%"><p style="text-align:center">6.88</p></td> 
     </tr> 
    </table>
   </table-wrap>
   <p>Since the robust scaler is designed to handle outliers and due to the existence of outliers in our data, we believe that the use of Robust scaler resulted in a more stable performance. Model checkpoint saves the model weights when the validation metric improves. Based on our observations, we believe that Robust Scaler leads to more stable training, thus making easier for early stopping to identify the optimal stopping point and for model checkpoint to capture more robust model snapshots.</p>
   <p>Similar to past research results <xref ref-type="bibr" rid="scirp.136564-8">
     [8]
    </xref>, the model with Robust scaler also achieved more accurate predictions during freezing season compared to the ones during melting season, with July having the least accurate prediction. <xref ref-type="table" rid="table5">
     Table 5
    </xref> below includes the performance results for each season and <xref ref-type="table" rid="table6">
     Table 6
    </xref> includes the performance results per month.</p>
   <table-wrap id="table5">
    <label>
     <xref ref-type="table" rid="table5">
      Table 5
     </xref></label>
    <caption>
     <title>
      <xref ref-type="bibr" rid="scirp.136564-"></xref>Table 5. Evaluation Metrics for Melting and Freezing season (96 LSTM units, 64 batch size, Robust scaler, adjustable learning rate, early stopping with patience 15 and model checkpoint).</title>
    </caption>
    <table class="MsoTableGrid custom-table" border="0" cellspacing="0" cellpadding="0"> 
     <tr> 
      <td class="custom-bottom-td acenter" width="49.98%" colspan="2"><p style="text-align:center">Melting Season</p><p style="text-align:center">(June to October)</p></td> 
      <td class="custom-bottom-td acenter" width="50.02%" colspan="2"><p style="text-align:center">Freezing Season</p><p style="text-align:center">(November to May)</p></td> 
     </tr> 
     <tr> 
      <td class="custom-top-td acenter" width="24.98%"><p style="text-align:center">MAE</p></td> 
      <td class="custom-top-td acenter" width="25.01%"><p style="text-align:center">0.254</p></td> 
      <td class="custom-top-td acenter" width="25.01%"><p style="text-align:center">MAE</p></td> 
      <td class="custom-top-td acenter" width="25.01%"><p style="text-align:center">0.174</p></td> 
     </tr> 
     <tr> 
      <td class="acenter" width="24.98%"><p style="text-align:center">R<sup>2</sup></p></td> 
      <td class="acenter" width="25.01%"><p style="text-align:center">0.975</p></td> 
      <td class="acenter" width="25.01%"><p style="text-align:center">R<sup>2</sup></p></td> 
      <td class="acenter" width="25.01%"><p style="text-align:center">0.988</p></td> 
     </tr> 
     <tr> 
      <td class="acenter" width="24.98%"><p style="text-align:center">MSE</p></td> 
      <td class="acenter" width="25.01%"><p style="text-align:center">0.103</p></td> 
      <td class="acenter" width="25.01%"><p style="text-align:center">MSE</p></td> 
      <td class="acenter" width="25.01%"><p style="text-align:center">0.039</p></td> 
     </tr> 
     <tr> 
      <td class="acenter" width="24.98%"><p style="text-align:center">RMSE</p></td> 
      <td class="acenter" width="25.01%"><p style="text-align:center">0.321</p></td> 
      <td class="acenter" width="25.01%"><p style="text-align:center">RMSE</p></td> 
      <td class="acenter" width="25.01%"><p style="text-align:center">0.199</p></td> 
     </tr> 
    </table>
   </table-wrap>
   <table-wrap id="table6">
    <label>
     <xref ref-type="table" rid="table6">
      Table 6
     </xref></label>
    <caption>
     <title>
      <xref ref-type="bibr" rid="scirp.136564-"></xref>Table 6. Evaluation Metrics per month (96 LSTM units, 64 batch size, Robust scaler, adjustable learning rate, early stopping with patience 15 and model checkpoint).</title>
    </caption>
    <table class="MsoTableGrid custom-table" border="0" cellspacing="0" cellpadding="0"> 
     <tr> 
      <td class="custom-bottom-td acenter" width="33.28%"><p style="text-align:center">1989-2002</p></td> 
      <td class="custom-bottom-td acenter" width="33.36%"><p style="text-align:center">MAE</p></td> 
      <td class="custom-bottom-td acenter" width="33.36%"><p style="text-align:center">MSE</p></td> 
     </tr> 
     <tr> 
      <td class="custom-top-td acenter" width="33.28%"><p style="text-align:center">Jul</p></td> 
      <td class="custom-top-td acenter" width="33.36%"><p style="text-align:center">0.349</p></td> 
      <td class="custom-top-td acenter" width="33.36%"><p style="text-align:center">0.185</p></td> 
     </tr> 
     <tr> 
      <td class="acenter" width="33.28%"><p style="text-align:center">Jan</p></td> 
      <td class="acenter" width="33.36%"><p style="text-align:center">0.323</p></td> 
      <td class="acenter" width="33.36%"><p style="text-align:center">0.115</p></td> 
     </tr> 
     <tr> 
      <td class="acenter" width="33.28%"><p style="text-align:center">Jun</p></td> 
      <td class="acenter" width="33.36%"><p style="text-align:center">0.281</p></td> 
      <td class="acenter" width="33.36%"><p style="text-align:center">0.106</p></td> 
     </tr> 
     <tr> 
      <td class="acenter" width="33.28%"><p style="text-align:center">Aug</p></td> 
      <td class="acenter" width="33.36%"><p style="text-align:center">0.266</p></td> 
      <td class="acenter" width="33.36%"><p style="text-align:center">0.097</p></td> 
     </tr> 
     <tr> 
      <td class="acenter" width="33.28%"><p style="text-align:center">May</p></td> 
      <td class="acenter" width="33.36%"><p style="text-align:center">0.265</p></td> 
      <td class="acenter" width="33.36%"><p style="text-align:center">0.086</p></td> 
     </tr> 
     <tr> 
      <td class="acenter" width="33.28%"><p style="text-align:center">Oct</p></td> 
      <td class="acenter" width="33.36%"><p style="text-align:center">0.228</p></td> 
      <td class="acenter" width="33.36%"><p style="text-align:center">0.074</p></td> 
     </tr> 
     <tr> 
      <td class="acenter" width="33.28%"><p style="text-align:center">Mar</p></td> 
      <td class="acenter" width="33.36%"><p style="text-align:center">0.224</p></td> 
      <td class="acenter" width="33.36%"><p style="text-align:center">0.051</p></td> 
     </tr> 
     <tr> 
      <td class="acenter" width="33.28%"><p style="text-align:center">Apr</p></td> 
      <td class="acenter" width="33.36%"><p style="text-align:center">0.175</p></td> 
      <td class="acenter" width="33.36%"><p style="text-align:center">0.031</p></td> 
     </tr> 
     <tr> 
      <td class="acenter" width="33.28%"><p style="text-align:center">Feb</p></td> 
      <td class="acenter" width="33.36%"><p style="text-align:center">0.168</p></td> 
      <td class="acenter" width="33.36%"><p style="text-align:center">0.030</p></td> 
     </tr> 
     <tr> 
      <td class="acenter" width="33.28%"><p style="text-align:center">Sept</p></td> 
      <td class="acenter" width="33.36%"><p style="text-align:center">0.153</p></td> 
      <td class="acenter" width="33.36%"><p style="text-align:center">0.028</p></td> 
     </tr> 
     <tr> 
      <td class="acenter" width="33.28%"><p style="text-align:center">Dec</p></td> 
      <td class="acenter" width="33.36%"><p style="text-align:center">0.102</p></td> 
      <td class="acenter" width="33.36%"><p style="text-align:center">0.013</p></td> 
     </tr> 
     <tr> 
      <td class="acenter" width="33.28%"><p style="text-align:center">Nov</p></td> 
      <td class="acenter" width="33.36%"><p style="text-align:center">0.057</p></td> 
      <td class="acenter" width="33.36%"><p style="text-align:center">0.004</p></td> 
     </tr> 
    </table>
   </table-wrap>
   <p>As a next step, we wanted to test the impact of adding additional past SIE as a feature in our model’s input sequence. Our hypothesis was that inclusion of more past SIE values could enhance capturing temporal dependencies in the data. If there were patterns or trends in the historical SIE that influence future values, the model could learn from these patterns. Also, if there was a sudden increase or decrease in SIE, it might affect the subsequent month’s extent, and the model could learn to recognize such patterns.</p>
   <p>Specifically, we first tried including the SIE of the corresponding month in the previous year, which did not improve the model’s performance. So, instead, we incorporated a simple linear regression simulated SIE of the previous year.</p>
   <p>We build a linear regression model to predict SIE based on past SIE. We then extracted the slope coefficient which we multiplied to the SIE of the corresponding month of the previous year. This new feature captures the influence of past SIE on current SIE, as learned by the model.</p>
   <p>
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
      <msub> 
       <mrow> 
        <mtext>
          SIE 
        </mtext> 
       </mrow> 
       <mrow> 
        <mtext>
          SIM 
        </mtext> 
       </mrow> 
      </msub> 
      <mrow> 
       <mo>
         ( 
       </mo> 
       <mi>
         t 
       </mi> 
       <mo>
         ) 
       </mo> 
      </mrow> 
      <mo>
        = 
      </mo> 
      <mtext>
        LinRegSlopeCoeff 
      </mtext> 
      <mo>
        * 
      </mo> 
      <mtext>
        SIE 
      </mtext> 
      <mrow> 
       <mo>
         ( 
       </mo> 
       <mrow> 
        <mi>
          t 
        </mi> 
        <mo>
          − 
        </mo> 
        <mn>
          12 
        </mn> 
       </mrow> 
       <mo>
         ) 
       </mo> 
      </mrow> 
     </mrow> 
    </math> (5)</p>
   <p>where 
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
      <mi>
        t 
      </mi> 
      <mo>
        = 
      </mo> 
      <mi>
        M 
      </mi> 
      <mo> 
      </mo> 
      <mi>
        o 
      </mi> 
      <mi>
        f 
      </mi> 
      <mo> 
      </mo> 
      <mi>
        Y 
      </mi> 
     </mrow> 
    </math></p>
   <p>
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
      <msub> 
       <mrow> 
        <mtext>
          SIE 
        </mtext> 
       </mrow> 
       <mrow> 
        <mtext>
          SIM 
        </mtext> 
       </mrow> 
      </msub> 
      <mrow> 
       <mo>
         ( 
       </mo> 
       <mi>
         t 
       </mi> 
       <mo>
         ) 
       </mo> 
      </mrow> 
      <mo>
        = 
      </mo> 
      <mtext>
        Slope 
      </mtext> 
      <mo>
        * 
      </mo> 
      <mtext>
        SIE 
      </mtext> 
      <mrow> 
       <mo>
         ( 
       </mo> 
       <mrow> 
        <mi>
          t 
        </mi> 
        <mo>
          − 
        </mo> 
        <mn>
          12 
        </mn> 
       </mrow> 
       <mo>
         ) 
       </mo> 
      </mrow> 
     </mrow> 
    </math> (6)</p>
   <p>where 
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
      <mi>
        t 
      </mi> 
      <mo>
        = 
      </mo> 
      <mi>
        M 
      </mi> 
      <mo> 
      </mo> 
      <mi>
        o 
      </mi> 
      <mi>
        f 
      </mi> 
      <mo> 
      </mo> 
      <mi>
        Y 
      </mi> 
     </mrow> 
    </math>.</p>
   <p>As shown in <xref ref-type="fig" rid="fig8">
     Figure 8
    </xref> and <xref ref-type="table" rid="table7">
     Table 7
    </xref>, MAE improved by 9.91% and MSE by 20.3%. Moreover, the predictions for melting season improved and outperformed the predictions for freezing season, overcoming the limitations of the previous model configuration and those in past research results that we are aware of.</p>
   <p>Robust Scaler: MAE: 0.191 || R<sup>2</sup>: 0.996 || MSE: 0.055 || RMSE: 0.234</p>
   <fig id="fig8" position="float">
    <label>Figure 8</label>
    <caption>
     <title>Figure 8. Robust scaler: Inclusion of SIE in Dataset—Performance for 12 month window, early stopping and model checkpoint.</title>
    </caption>
    <graphic mimetype="image" position="float" xlink:type="simple" xlink:href="https://html.scirp.org/file/4701247-rId47.jpeg?20241014013916" />
   </fig>
   <table-wrap id="table7">
    <label>
     <xref ref-type="table" rid="table7">
      Table 7
     </xref></label>
    <caption>
     <title>
      <xref ref-type="bibr" rid="scirp.136564-"></xref>Table 7. Evaluation Metrics for Melting and Freezing season with inclusion of past year SIE (96 LSTM units, 64 batch size, Robust scaler, adjustable learning rate, early stopping with patience 15 and model checkpoint).</title>
    </caption>
    <table class="MsoTableGrid custom-table" border="0" cellspacing="0" cellpadding="0"> 
     <tr> 
      <td class="custom-bottom-td acenter" width="49.98%" colspan="2"><p style="text-align:center">Melting Season</p><p style="text-align:center">(June to October)</p></td> 
      <td class="custom-bottom-td acenter" width="50.02%" colspan="2"><p style="text-align:center">Freezing Season</p><p style="text-align:center">(November to May)</p></td> 
     </tr> 
     <tr> 
      <td class="custom-top-td acenter" width="24.98%"><p style="text-align:center">MAE</p></td> 
      <td class="custom-top-td acenter" width="25.01%"><p style="text-align:center">0.170</p></td> 
      <td class="custom-top-td acenter" width="25.01%"><p style="text-align:center">MAE</p></td> 
      <td class="custom-top-td acenter" width="25.01%"><p style="text-align:center">0.207</p></td> 
     </tr> 
     <tr> 
      <td class="acenter" width="24.98%"><p style="text-align:center">R<sup>2</sup></p></td> 
      <td class="acenter" width="25.01%"><p style="text-align:center">0.989</p></td> 
      <td class="acenter" width="25.01%"><p style="text-align:center">R<sup>2</sup></p></td> 
      <td class="acenter" width="25.01%"><p style="text-align:center">0.981</p></td> 
     </tr> 
     <tr> 
      <td class="acenter" width="24.98%"><p style="text-align:center">MSE</p></td> 
      <td class="acenter" width="25.01%"><p style="text-align:center">0.048</p></td> 
      <td class="acenter" width="25.01%"><p style="text-align:center">MSE</p></td> 
      <td class="acenter" width="25.01%"><p style="text-align:center">0.060</p></td> 
     </tr> 
     <tr> 
      <td class="acenter" width="24.98%"><p style="text-align:center">RMSE</p></td> 
      <td class="acenter" width="25.01%"><p style="text-align:center">0.220</p></td> 
      <td class="acenter" width="25.01%"><p style="text-align:center">RMSE</p></td> 
      <td class="acenter" width="25.01%"><p style="text-align:center">0.245</p></td> 
     </tr> 
    </table>
   </table-wrap>
   <p>We also tested the impact on performance when including the L1 and L2 regularization. L1 regularization (Lasso) adds a penalty term based on the absolute values of the model’s coefficients and can help with feature selection by driving some coefficients to exactly zero. L2 regularization (Ridge) adds a penalty based on the squared values of the coefficients, which helps prevent overfitting and can improve generalization. We iteratively tested different combinations of L1 and L2 values and identified that the best performing values for our model were 0.001 and 0.001 respectively. The inclusion of L1 and L2 did not have any positive impact on our model’s performance. Although on average MSE remained the same, MAE slightly deteriorated by 3%. Since L1 did not lead to an improvement in the model performance, it may suggest that most of our features are contributing to our model and thus there aren’t many irrelevant features to eliminate. The fact that L2 did not help improve performance may suggest that early stopping is effectively controlling our model’s complexity leading to no overfitting.</p>
   <p>Lastly, we tested PReLU activation function both at the LSTM layer and as a separate layer. Both had a significant negative impact on the performance. In the first case, MAE increased by approximately 120% and in the second it increased by 50%.</p>
  </sec><sec id="s5">
   <title>5. Conclusions</title>
   <p>In this research, we optimized an LSTM model for the prediction of Arctic Sea ice extent on a monthly scale by carefully choosing several key hyperparameters and then properly tuning them through feature re-engineering and the inclusion of advanced training techniques. The development of accurate Arctic SIE predictive models is crucial due to the rapid decline of sea ice in recent years and provides essential insights for understanding and mitigating the impact of climate change on polar regions.</p>
   <p>In the initial stages of our data analysis, we focused on identifying outliers present in our data and exploring the relationship between SIE and our five climate variables: sea surface temperature, surface pressure, total precipitation, latent heat, and sensible heat. Our experiments revealed the presence of outliers in the data, influencing our choice of evaluation metrics and data scaling techniques. Feature importance analysis also confirmed that all our features significantly contributed to SIE, a finding further supported by the results of L1 Regularization.</p>
   <p>Our original model configuration included 96 LSTM units, 64 batch size, one-month time lag, and an adjustable learning rate. Our model delivered predictions with a MAE of 0.480, R<sup>2</sup> of 0.98 and MSE of 0.232.</p>
   <p>We trained our model with an extended historical context of 6 and 12 months (i.e., 6 and 12 lagged time steps). We experimented with three different scalers: Standard, MinMax, and Robust Scaler. The model’s performance with Standard and Robust Scaler was best for the 12 lagged time steps, with Standard outperforming Robust. We incorporated month information with cyclical encoding to introduce temporal and seasonal information. After these enhancements, the model’s performance with Robust Scaler was superior in terms of MAE. The configuration with Robust scaling continued to outperform after the incorporation of advanced training techniques like early stopping and model checkpoint, resulting in more accurate predictions for the freezing season (November to May) compared to the melting season (June to October).</p>
   <p>However, our results changed when we included one more feature: specifically, a linear regression simulated SIE that accounted for the influence of past SIE on current SIE. This increased the accuracy of predictions for the melting season compared to the freezing season and overcame limitations in not only model configuration but also those in past research. Further, our observations following L2 Regularization suggested to us that our early stopping was effectively controlling our model’s complexity and prevented overfitting.</p>
   <p>Overall, following meticulous hyperparameter tuning, feature re-engineering and inclusion of the advanced training techniques of early stopping and model checkpoint, MAE improved by 60% (0.191) and MSE by 76% (0.055). R<sup>2</sup> also reached an impressive value of 0.996, underscoring the model’s remarkable ability to explain and account for nearly all the variance in the data.</p>
  </sec><sec id="s6">
   <title>NOTES</title>
   <p><sup>1</sup>Element-wise multiplication: this type of multiplication takes in two vectors (or more generally two matrices) of the same dimensions and returns a vector (or matrix) of the multiplied corresponding elements.</p>
  </sec>
 </body><back>
  <ref-list>
   <title>References</title>
   <ref id="scirp.136564-ref1">
    <label>1</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     The National Oceanic and Atmospheric Administration (2018) Arctic Report Card: Update for 2018. &gt;https://arctic.noaa.gov/report-card/report-card-2018/executive-summary-5/ 
    </mixed-citation>
   </ref>
   <ref id="scirp.136564-ref2">
    <label>2</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     The National Aeronautics and Space Administration (2022) Arctic Sea Ice Minimum Extent Vital Signs—Climate Change: Vital Signs of the Planet. NASA Climate Change. &gt;https://climate.nasa.gov/vital-signs/arctic-sea-ice/ 
    </mixed-citation>
   </ref>
   <ref id="scirp.136564-ref3">
    <label>3</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Chen, S., Li, K., Fu, H., Wu, Y.C. and Huang, Y. (2023) Sea Ice Extent Prediction with Machine Learning Methods and Subregional Analysis in the Arctic. Atmosphere, 14, Article 1023. 
    </mixed-citation>
   </ref>
   <ref id="scirp.136564-ref4">
    <label>4</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Serreze, M.C. and Meier, W.N. (2019) The Arctic’s Sea Ice Cover: Trends, Variability, Predictability, and Comparisons to the Antarctic. Annals of the New York Academy of Sciences, 1436, 36-53. 
    </mixed-citation>
   </ref>
   <ref id="scirp.136564-ref5">
    <label>5</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     NASA Earth Observatory (2021) World of Change: Arctic Sea Ice. Earth Observatory. &gt;https://earthobservatory.nasa.gov/world-of-change/sea-ice-arctic 
    </mixed-citation>
   </ref>
   <ref id="scirp.136564-ref6">
    <label>6</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Arctic Monitoring and Assessment Programme (2017) Snow, Water, Ice and Permafrost in the Arctic (SWIPA). Assessment Summary for Policy-Makers. &gt;https://swipa.amap.no 
    </mixed-citation>
   </ref>
   <ref id="scirp.136564-ref7">
    <label>7</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Mu, B., Luo, X., Yuan, S. and Liang, X. (2023) IceTFT v1.0.0: Interpretable Long-Term Prediction of Arctic Sea Ice Extent with Deep Learning. Geoscientific Model Development, 16, 4677-4697.
    </mixed-citation>
   </ref>
   <ref id="scirp.136564-ref8">
    <label>8</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Chi, J. and Kim, H. (2017) Prediction of Arctic Sea Ice Concentration Using a Fully Data Driven Deep Neural Network. Remote Sensing, 9, Article 1305. 
    </mixed-citation>
   </ref>
   <ref id="scirp.136564-ref9">
    <label>9</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Pegkou Christofi, V., Li, A. and Lin, A. (2024) Wavelet Based Multiscale Deep Learning Algorithms for Arctic Sea Ice Melting Prediction. Proceedings of the 2024 6th International Symposium on Signal Processing Systems, Xi’an, 22-24 March 2024, 29-39. &gt;https://doi.org/10.1145/3665053.3665054 
    </mixed-citation>
   </ref>
   <ref id="scirp.136564-ref10">
    <label>10</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Sherstinsky, A. (2020) Fundamentals of Recurrent Neural Network (RNN) and Long Short-Term Memory (LSTM) Network. Physica D: Nonlinear Phenomena, 404, Article 132306. &gt;https://doi.org/10.1016/j.physd.2019.132306 
    </mixed-citation>
   </ref>
   <ref id="scirp.136564-ref11">
    <label>11</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     TensorFlow. &gt;https://www.tensorflow.org/ 
    </mixed-citation>
   </ref>
   <ref id="scirp.136564-ref12">
    <label>12</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Fetterer, F., Knowles, K., Meier, W.N., Savoie, M. and Windnagel, A.K. (2017) Sea Ice Index, Version 3 [Data Set]. Boulder, Colorado USA. National Snow and Ice Data Center. 
    </mixed-citation>
   </ref>
   <ref id="scirp.136564-ref13">
    <label>13</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Copernicus Climate Change Service (2023) Reanalysis ERA5 Single Levels Monthly Means. &gt;https://cds.climate.copernicus.eu/datasets/reanalysis-era5-single-levels-monthly-means?tab=download 
    </mixed-citation>
   </ref>
   <ref id="scirp.136564-ref14">
    <label>14</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Sluiter, R. (2009) Interpolation Methods for Climate Data: Literature Review. KNMI (Royal Netherlands Meteorological Institute), R&amp;D Information and Observation Technology. &gt;https://uaf-snap.org/wp-content/uploads/2020/08/Interpolation_methods_for_climate_data.pdf 
    </mixed-citation>
   </ref>
   <ref id="scirp.136564-ref15">
    <label>15</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     National Center for Atmospheric Research, Dennis, S. and National Center for Atmospheric Research Staff (2014) The Climate Data Guide: Statistical&amp;Diagnostic Methods Overview. &gt;https://climatedataguide.ucar.edu/climate-tools/statistical-diagnostic-methods-overview 
    </mixed-citation>
   </ref>
   <ref id="scirp.136564-ref16">
    <label>16</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Schober, P., Boer, C. and Schwarte, L.A. (2018) Correlation Coefficients: Appropriate Use and Interpretation. Anesthesia &amp; Analgesia, 126, 1763-1768. &gt;https://doi.org/10.1213/ane.0000000000002864 
    </mixed-citation>
   </ref>
   <ref id="scirp.136564-ref17">
    <label>17</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Keskar, N.S., Mudigere, D., Nocedal, J., Smelyanskiy, M. and Tang, P. (2017) On Large-Batch Training for Deep Learning: Generalization Gap and Sharp Minima. 
    </mixed-citation>
   </ref>
   <ref id="scirp.136564-ref18">
    <label>18</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Masters, D. and Luschi, C. (2018) Revisiting Small Batch Training for Deep Neural Networks. 
    </mixed-citation>
   </ref>
   <ref id="scirp.136564-ref19">
    <label>19</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Huilgol, P. (2024) Analytics Vidhya. Feature Transformation and Scaling Techniques to Boost Your Model Performance. &gt;https://www.analyticsvidhya.com/blog/2020/07/types-of-feature-transformation-and-scaling/ 
    </mixed-citation>
   </ref>
   <ref id="scirp.136564-ref20">
    <label>20</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     de Dieu Nyandwi, J. (2021) The Ultimate and Practical Guide on Feature Scaling. &gt;https://jeande.medium.com/the-ultimate-and-practical-guide-on-feature-scaling-d03fbe2cb25e 
    </mixed-citation>
   </ref>
   <ref id="scirp.136564-ref21">
    <label>21</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Scikit-Learn Developers (2023) Plot Different Scalers on the Same Data. &gt;https://scikit-learn.org/stable/auto_examples/preprocessing/plot_all_scaling.html#plot-all-scaling-quantile-transformer-section 
    </mixed-citation>
   </ref>
   <ref id="scirp.136564-ref22">
    <label>22</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Neptune Labs (2023) Fighting Overfitting with L1 or L2 Regularization: Which One Is Better? &gt;https://neptune.ai/blog/fighting-overfitting-with-l1-or-l2-regularization 
    </mixed-citation>
   </ref>
   <ref id="scirp.136564-ref23">
    <label>23</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Van Otten, N. (2023) L1 And L2 Regularization Explained, When to Use Them&amp;Practical How to Examples.&gt;https://spotintelligence.com/2023/05/26/l1-l2-regularization/ 
    </mixed-citation>
   </ref>
   <ref id="scirp.136564-ref24">
    <label>24</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Towards Data Science: Comparing Robustness of MAE, MSE, and RMSE. &gt;https://towardsdatascience.com/comparing-robustness-of-mae-mse-and-rmse-6d69da870828 
    </mixed-citation>
   </ref>
   <ref id="scirp.136564-ref25">
    <label>25</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Van Wyk, A. (2022) Encoding Cyclical Features for Deep Learning. &gt;https://www.kaggle.com/code/avanwyk/encoding-cyclical-features-for-deep-learning
    </mixed-citation>
   </ref>
  </ref-list>
 </back>
</article>