<?xml version="1.0" encoding="UTF-8"?><!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v3.0 20080202//EN" "http://dtd.nlm.nih.gov/publishing/3.0/journalpublishing3.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" dtd-version="3.0" xml:lang="en" article-type="research article">
 <front>
  <journal-meta>
   <journal-id journal-id-type="publisher-id">
    ijids
   </journal-id>
   <journal-title-group>
    <journal-title>
     International Journal of Internet and Distributed Systems
    </journal-title>
   </journal-title-group>
   <issn pub-type="epub">
    2327-7157
   </issn>
   <issn publication-format="print">
    2327-7165
   </issn>
   <publisher>
    <publisher-name>
     Scientific Research Publishing
    </publisher-name>
   </publisher>
  </journal-meta>
  <article-meta>
   <article-id pub-id-type="doi">
    10.4236/ijids.2024.62002
   </article-id>
   <article-id pub-id-type="publisher-id">
    ijids-137581
   </article-id>
   <article-categories>
    <subj-group subj-group-type="heading">
     <subject>
      Articles
     </subject>
    </subj-group>
    <subj-group subj-group-type="Discipline-v2">
     <subject>
      Computer Science 
     </subject>
     <subject>
       Communications
     </subject>
    </subj-group>
   </article-categories>
   <title-group>
    Deep Learning-Based Two-Step Approach for Intrusion Detection in Networks
   </title-group>
   <contrib-group>
    <contrib contrib-type="author" xlink:type="simple">
     <name name-style="western">
      <surname>
       Kamagaté Beman
      </surname>
      <given-names>
       Hamidja
      </given-names>
     </name> 
     <xref ref-type="aff" rid="aff1"> 
      <sup>1</sup>
     </xref>
    </contrib>
    <contrib contrib-type="author" xlink:type="simple">
     <name name-style="western">
      <surname>
       Kanga
      </surname>
      <given-names>
       Koffi
      </given-names>
     </name> 
     <xref ref-type="aff" rid="aff1"> 
      <sup>1</sup>
     </xref>
    </contrib>
    <contrib contrib-type="author" xlink:type="simple">
     <name name-style="western">
      <surname>
       Kouassi
      </surname>
      <given-names>
       Adless
      </given-names>
     </name> 
     <xref ref-type="aff" rid="aff1"> 
      <sup>1</sup>
     </xref>
    </contrib>
    <contrib contrib-type="author" xlink:type="simple">
     <name name-style="western">
      <surname>
       Olivier
      </surname>
      <given-names>
       Asseu
      </given-names>
     </name> 
     <xref ref-type="aff" rid="aff1"> 
      <sup>1</sup>
     </xref> 
     <xref ref-type="aff" rid="aff2"> 
      <sup>2</sup>
     </xref>
    </contrib>
    <contrib contrib-type="author" xlink:type="simple">
     <name name-style="western">
      <surname>
       Souleymane
      </surname>
      <given-names>
       Oumtanaga
      </given-names>
     </name> 
     <xref ref-type="aff" rid="aff3"> 
      <sup>3</sup>
     </xref>
    </contrib>
   </contrib-group> 
   <aff id="aff1">
    <addr-line>
     aEcole Supérieure Africaine des Technologies de l’Information et de la Communication (ESATIC), LASTIC Laboratory of ESATIC, Abidjan, Côte d’Ivoire
    </addr-line> 
   </aff> 
   <aff id="aff2">
    <addr-line>
     aUMRI STI, INPHB INPUMRI STI, INPHB INP-Houphouët Boigny, Yamoussoukro, Côte d’Ivoire
    </addr-line> 
   </aff> 
   <aff id="aff3">
    <addr-line>
     aUMRI MSN INPHB INP-Houphouët Boigny, Yamoussoukro, Côte d’Ivoire
    </addr-line> 
   </aff> 
   <pub-date pub-type="epub">
    <day>
     22
    </day> 
    <month>
     11
    </month>
    <year>
     2024
    </year>
   </pub-date> 
   <volume>
    06
   </volume> 
   <issue>
    02
   </issue>
   <fpage>
    25
   </fpage>
   <lpage>
    39
   </lpage>
   <history>
    <date date-type="received">
     <day>
      2,
     </day>
     <month>
      May
     </month>
     <year>
      2024
     </year>
    </date>
    <date date-type="published">
     <day>
      25,
     </day>
     <month>
      May
     </month>
     <year>
      2024
     </year> 
    </date> 
    <date date-type="accepted">
     <day>
      25,
     </day>
     <month>
      May
     </month>
     <year>
      2024
     </year> 
    </date>
   </history>
   <permissions>
    <copyright-statement>
     © Copyright 2014 by authors and Scientific Research Publishing Inc. 
    </copyright-statement>
    <copyright-year>
     2014
    </copyright-year>
    <license>
     <license-p>
      This work is licensed under the Creative Commons Attribution International License (CC BY). http://creativecommons.org/licenses/by/4.0/
     </license-p>
    </license>
   </permissions>
   <abstract>
    Intrusion Detection Systems (IDS) are essential for computer security, with various techniques developed over time. However, many of these methods suffer from high false positive rates. To address this, we propose an approach utilizing Recurrent Neural Networks (RNN). Our method starts by reducing the dataset’s dimensionality using a Deep Auto-Encoder (DAE), followed by intrusion detection through a Bidirectional Long Short-Term Memory (BiLSTM) network. The proposed DAE-BiLSTM model outperforms Random Forest, AdaBoost, and standard BiLSTM models, achieving an accuracy of 0.97, a recall of 0.95, and an AUC of 0.93. Although BiLSTM is slightly less effective than DAE-BiLSTM, both RNN-based models outperform AdaBoost and Random Forest. ROC curves show that DAE-BiLSTM is the most effective, demonstrating strong detection capabilities with a low false positive rate. While AdaBoost performs well, it is less effective than RNN models but still surpasses Random Forest.
   </abstract>
   <kwd-group> 
    <kwd>
     Cybersecurity
    </kwd> 
    <kwd>
      CICIDDS2017
    </kwd> 
    <kwd>
      Intrusion Detection
    </kwd> 
    <kwd>
      BiLSTM
    </kwd> 
    <kwd>
      Deep Auto-Encoder
    </kwd>
   </kwd-group>
  </article-meta>
 </front>
 <body>
  <sec id="s1">
   <title>1. Introduction</title>
   <p>The security of computer systems is a sensitive and worrying issue. The spectacular progress of information and communication technologies today offers inescapable facilities for file transfer, messaging and many other forms of information exchange. The development of computerized exchanges has unfortunately been accompanied by the development of malicious activities whose motives are as numerous as they are dangerous, and which evolve over time <xref ref-type="bibr" rid="scirp.137581-1">
     <a href="#ref1">[1]</a>
    </xref>. Taking advantage of the growing connectivity of information technology systems, particularly to the Internet, the possibilities for remote attacks are greater and more threatening.</p>
   <p>Intrusion Detection Systems (IDS) can be set up to detect any attempt to violate security mechanisms, by monitoring systems on an ongoing basis. Intrusion detection involves scanning network traffic, collecting all events, analyzing them and generating alarms if malicious attempts are identified <xref ref-type="bibr" rid="scirp.137581-2">
     [2]
    </xref>. Intrusion Detection Systems (IDS) are essential for cybersecurity. They detect threats before they cause damage, block intrusions by monitoring network activities in real-time, and protect sensitive data from unauthorized access, especially in critical sectors. By identifying vulnerabilities, they help strengthen system resilience and ensure compliance with security standards, facilitating audits. These functions make IDS essential tools for proactive cybersecurity. They are widely deployed across IT systems and have become integral to the design of security strategies. They are generally used to monitor access and information flow, with the aim of determining any malicious behavior, whether from inside or outside the information system, and making this information available to security administrators. As an option, intrusion detection systems can react to malicious behavior and take countermeasures.</p>
   <p>In recent years, a range of approaches has been developed to enhance intrusion detection systems, with a focus on both rule-based methods and machine learning techniques. Rule-based systems are relatively straightforward to implement and are effective at identifying known attacks while maintaining a low rate of false positives <xref ref-type="bibr" rid="scirp.137581-3">
     [3]
    </xref>. However, they have limitations when it comes to detecting unknown attacks. On the other hand, machine learning methods are capable of identifying novel threats, such as “zero-day” attacks, but they come with their own set of challenges. Although machine learning-based IDS can be effective in recognizing new threats <xref ref-type="bibr" rid="scirp.137581-4">
     [4]
    </xref>, they may also produce a high number of false positives and false negatives. This means they might incorrectly label legitimate traffic as malicious or fail to detect genuinely suspicious traffic, potentially undermining the overall reliability of network protection mechanisms.</p>
   <p>In this work, we propose a two-tier approach to tackle these challenges. The first tier utilizes deep autoencoders to compress the training dataset, focusing on extracting only the most relevant features. The second tier utilizes a Bidirectional Long Short-Term Memory (BiLSTM) network, a type of Recurrent Neural Network (RNN), to process the compressed dataset from the first tier and classify traffic as either normal or anomalous. The BiLSTM stands out for its bidirectional processing, capturing both past and future contexts to enrich the understanding of event relationships, essential for detecting complex intrusion patterns. Its long-term memory enables it to retain information over extended sequences, useful for spotting abnormal behaviors in network traffic. Finally, its adaptability to time series makes it an effective real-time model, responsive to emerging threats. These features enhance its accuracy and effectiveness in proactive intrusion detection <xref ref-type="bibr" rid="scirp.137581-5">
     [5]
    </xref>.</p>
   <p>The subsequent sections of this work are organized as follows: Section 2 provides a review of related work, Section 3 outlines the methodology and proposed framework, and Section 4 covers the experimentation, results, and discussion. The paper concludes with a summary and recommendations for future research.</p>
  </sec><sec id="s2">
   <title>2. Related Work</title>
   <p>For intrusion detection, the first approach used in the literature is the rule-based approach. It compares the signature of incoming flows to a range of signatures stored in a database that are already identified as malicious. If a signature matches one in the database, it indicates that the traffic may be malicious. This approach is widely developed in many research papers.</p>
   <p>In paper <xref ref-type="bibr" rid="scirp.137581-6">
     [6]
    </xref>, the authors introduce a signature-based intrusion detection system (IDS) designed to detect denial-of-service (DoS) and routing attacks in IoT networks. The major innovation of this IDS is its hybrid configuration, which combines centralized and distributed components. The detection module is installed on the main router, while lightweight modules are positioned near IoT devices to monitor and report traffic. The work <xref ref-type="bibr" rid="scirp.137581-7">
     [7]
    </xref> analyzes intrusion detection systems based on fuzzy logic, which are designed using various data mining techniques. Thanks to fuzzy logic, these systems can better account for the uncertainty and ambiguity inherent in intrusion detection data, thus providing more effective management of complex situations where attack patterns are not always clear or well-defined. In <xref ref-type="bibr" rid="scirp.137581-8">
     [8]
    </xref>, the authors performed an experimental study to evaluate the detection capabilities of signature-based intrusion detection systems in the context of web attacks. The results show that the predefined configuration rulesets of generic solutions like Snort, ModSecurity, and Nemesida Free offer lower-than-expected detection performance for known attacks, even when set to the most sensitive configurations. From these works, it can be concluded that although rule-based detection systems are effective in identifying known intrusions, their performance can be significantly impacted by the system configurations. Furthermore, these systems are unable to detect threats whose signatures are not already listed in their database.</p>
   <p>The second one, behavioral approach assumes that normal activity is different from intrusive activity. All that’s needed is a profile for normal activity, and a mechanism for comparing current activity with the established profile to detect significant deviations that will be considered as possible intrusions.</p>
   <p>Study <xref ref-type="bibr" rid="scirp.137581-9">
     [9]
    </xref> presents a behavioral approach to intrusion detection that combines Accelerated Particle Swarm Optimization (APSO) with a Support Vector Machine (SVM) to create an IDS model. The simulation results demonstrate a significant improvement in performance. Compared to other methods applied to the same dataset, the proposed model achieves superior detection accuracy.</p>
   <p>To address the challenges posed by the dynamic and constantly evolving nature of malicious attacks, which also occur in large volumes and require scalable solutions, study <xref ref-type="bibr" rid="scirp.137581-10">
     [10]
    </xref> focuses on using Deep Neural Networks to detect malware and malicious behavior at both the network and host levels. Experiments conducted on various datasets demonstrate a high level of accuracy in intrusion detection. Study in <xref ref-type="bibr" rid="scirp.137581-11">
     [11]
    </xref>.</p>
   <p>In work <xref ref-type="bibr" rid="scirp.137581-12">
     [12]
    </xref>, the authors emphasize the increasing difficulty in accurately detecting intrusions due to the evolving nature of cyber-attacks, which pose significant risks to data confidentiality, integrity, and availability. After evaluating several machine learning algorithms, which, despite being capable of detecting unknown attacks, produced many false positives, the study concludes that decision trees are the most effective among machine learning algorithm, delivering superior performance in intrusion detection. As a result, an optimized framework was developed based on this algorithm.</p>
   <p>Compared to the previous signature-based approach, this method can detect unknown attacks, but it tends to produce a higher number of false positive alarms and requires more time to reach convergence, especially when there are many features. To address these shortcomings, we suggest employing a two-step deep learning approach for intrusion detection. Initially, we utilize a stacked autoencoder to reduce the dimensionality of features. Subsequently, the resultant vector serves as input to a Bidirectional Long Short-Term Memory to detect whether the behavior is normal or abnormal.</p>
  </sec><sec id="s3">
   <title>3. Materials and Methods</title>
   <p>This section outlines our methodological approach to intrusion detection. It comprises two phases: dimensionality reduction and classification. Dimensionality reduction step help as to select less features than features in the original database. Then at second step we apply classification to vector of the less selected features.</p>
   <sec id="s3_1">
    <title>3.1. Dimensionality Reduction with Deep Auto-Encoder</title>
    <p>Instance classification encounters several problems, such as large number of features. In order to improve classification accuracy and reduce computation times, we need to reduce the dimensionality of the data. For that end, we use Deep Auto-Encoder (DAE) <xref ref-type="bibr" rid="scirp.137581-13">
      [13]
     </xref>. Contrary to the conventional Ordinary Auto-Encoder (OAE) with a single hidden layer, the Deep Auto-Encoder (DAE) depicted in <xref ref-type="fig" rid="fig1">
      Figure 1
     </xref> utilizes multiple hidden layers.</p>
    <p>The hidden layers link the input layer with the output layer. They are stacked one after the other to form the encoder with the input layer, and the decoder with the output layer. In order to extract the strongest and most important features for proper classification, the DAE goes through two main phases: encoder and decoder. The encoding phase takes place between the input layer and the hidden layers up to the bottleneck layer. In this phase, the input data is transformed into reduced-dimensional representation, while retaining relevant features.</p>
    <p>To train the autoencoder, we began by determining the number of hidden layers and neurons per layer, monitoring the risk of overfitting. Next, we selected the activation function, opting for one that offers both speed and stability. Regarding the cost function, in line with the literature <xref ref-type="bibr" rid="scirp.137581-13">
      [13]
     </xref> recommending mean squared error (MSE) for reconstructions, we adopted this approach. For the learning rate, we started with a low value, which we gradually adjusted to avoid overly slow convergence. We then determined the appropriate number of epochs and batch size to stabilize the model. The final step was to assess the quality of the reconstruction on a validation set, which led us to adjust parameters such as the number of layers, the learning rate, and the batch size.</p>
    <fig id="fig1" position="float">
     <label>Figure 1</label>
     <caption>
      <title>Figure 1. Deep Auto-Encoder (DAE) with hidden layers.</title>
     </caption>
     <graphic mimetype="image" position="float" xlink:type="simple" xlink:href="https://html.scirp.org/file/1120105-rId14.jpeg?20241122031030" />
    </fig>
    <p>The encoder evaluates (formula 1) vector h (latent or bottleneck vector) with size m from the input vector X which represents the original data.</p>
    <p>
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mi>
         h 
       </mi> 
       <mo>
         = 
       </mo> 
       <mi>
         f 
       </mi> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <msup> 
          <mi>
            W 
          </mi> 
          <mrow> 
           <mrow> 
            <mo>
              ( 
            </mo> 
            <mn>
              1 
            </mn> 
            <mo>
              ) 
            </mo> 
           </mrow> 
          </mrow> 
         </msup> 
         <mi>
           X 
         </mi> 
         <mo>
           + 
         </mo> 
         <msup> 
          <mi>
            b 
          </mi> 
          <mrow> 
           <mrow> 
            <mo>
              ( 
            </mo> 
            <mn>
              1 
            </mn> 
            <mo>
              ) 
            </mo> 
           </mrow> 
          </mrow> 
         </msup> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
      </mrow> 
     </math>(1)</p>
    <p>where: 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msup> 
        <mi>
          W 
        </mi> 
        <mrow> 
         <mrow> 
          <mo>
            ( 
          </mo> 
          <mn>
            1 
          </mn> 
          <mo>
            ) 
          </mo> 
         </mrow> 
        </mrow> 
       </msup> 
      </mrow> 
     </math> is a matrix 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mi>
         m 
       </mi> 
       <mo>
         × 
       </mo> 
       <mi>
         n 
       </mi> 
      </mrow> 
     </math>, 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msup> 
        <mi>
          b 
        </mi> 
        <mrow> 
         <mrow> 
          <mo>
            ( 
          </mo> 
          <mn>
            1 
          </mn> 
          <mo>
            ) 
          </mo> 
         </mrow> 
        </mrow> 
       </msup> 
      </mrow> 
     </math> is a bias vector of size n and 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mi>
         f 
       </mi> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mtext>
            
        </mtext> 
        <mo>
          ) 
        </mo> 
       </mrow> 
      </mrow> 
     </math> is an encoder activation function.</p>
    <p>After that, came the decoder phase. It takes place between the hidden layers (starting with the bottleneck layer) and the output layer which has the same number of nodes as the input layer. The hidden layers in this phase are symmetrical to the hidden layers in the previous phase, with one in common (the bottleneck layer). The DAE try to have an output vector 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mover accent="true"> 
       <mi>
         X 
       </mi> 
       <mo>
         ^ 
       </mo> 
      </mover> 
     </math> that is closet possible representation of the initial attributes (those of the input layer: X), by minimizing the loss function. In this work loss function is mean squared error loss function (formula (2)).</p>
    <p>
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mtext>
         Loss 
       </mtext> 
       <mo>
         = 
       </mo> 
       <mfrac> 
        <mn>
          1 
        </mn> 
        <mi>
          n 
        </mi> 
       </mfrac> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <mstyle displaystyle="true"> 
          <msubsup> 
           <mo>
             ∑ 
           </mo> 
           <mrow> 
            <mi>
              i 
            </mi> 
            <mo>
              = 
            </mo> 
            <mn>
              1 
            </mn> 
           </mrow> 
           <mi>
             n 
           </mi> 
          </msubsup> 
          <mrow> 
           <msup> 
            <mrow> 
             <mrow> 
              <mo>
                ( 
              </mo> 
              <mrow> 
               <msub> 
                <mi>
                  X 
                </mi> 
                <mi>
                  i 
                </mi> 
               </msub> 
               <mo>
                 − 
               </mo> 
               <msub> 
                <mover accent="true"> 
                 <mi>
                   X 
                 </mi> 
                 <mo>
                   ^ 
                 </mo> 
                </mover> 
                <mi>
                  i 
                </mi> 
               </msub> 
              </mrow> 
              <mo>
                ) 
              </mo> 
             </mrow> 
            </mrow> 
            <mn>
              2 
            </mn> 
           </msup> 
          </mrow> 
         </mstyle> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
      </mrow> 
     </math>(2)</p>
    <p>The decoder tries to build output vector from the latent vector h, the result of which is the vector 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mover accent="true"> 
       <mi>
         X 
       </mi> 
       <mo>
         ^ 
       </mo> 
      </mover> 
     </math>.</p>
    <p>
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mover accent="true"> 
        <mi>
          X 
        </mi> 
        <mo>
          ^ 
        </mo> 
       </mover> 
       <mo>
         = 
       </mo> 
       <mi>
         g 
       </mi> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <msup> 
          <mi>
            W 
          </mi> 
          <mrow> 
           <mrow> 
            <mo>
              ( 
            </mo> 
            <mn>
              2 
            </mn> 
            <mo>
              ) 
            </mo> 
           </mrow> 
          </mrow> 
         </msup> 
         <mi>
           h 
         </mi> 
         <mo>
           + 
         </mo> 
         <msup> 
          <mi>
            b 
          </mi> 
          <mrow> 
           <mrow> 
            <mo>
              ( 
            </mo> 
            <mn>
              2 
            </mn> 
            <mo>
              ) 
            </mo> 
           </mrow> 
          </mrow> 
         </msup> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
      </mrow> 
     </math>(3)</p>
    <p>where 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msup> 
        <mi>
          W 
        </mi> 
        <mrow> 
         <mrow> 
          <mo>
            ( 
          </mo> 
          <mn>
            2 
          </mn> 
          <mo>
            ) 
          </mo> 
         </mrow> 
        </mrow> 
       </msup> 
      </mrow> 
     </math> is a matrix 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mi>
         n 
       </mi> 
       <mo>
         × 
       </mo> 
       <mi>
         m 
       </mi> 
      </mrow> 
     </math>, 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msup> 
        <mi>
          b 
        </mi> 
        <mrow> 
         <mrow> 
          <mo>
            ( 
          </mo> 
          <mn>
            2 
          </mn> 
          <mo>
            ) 
          </mo> 
         </mrow> 
        </mrow> 
       </msup> 
      </mrow> 
     </math> is a bias vector of size n, 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mover accent="true"> 
       <mi>
         X 
       </mi> 
       <mo>
         ^ 
       </mo> 
      </mover> 
     </math> is a vector of size m, a reconstruction of the input vector X as close as possible to it.</p>
   </sec>
   <sec id="s3_2">
    <title>3.2. Data Classification with BiLSTM</title>
    <p>After remarkably reduced the number of attributes (features) in the original database, in this section, we propose to predict the classes of the new instances, for which we’ve chosen to apply Bidirectional Long Short-Term Memory (BiLSTM) <xref ref-type="bibr" rid="scirp.137581-14">
      [14]
     </xref>. BiLSTM architecture is an extension of recurrent neural networks (RNN), designed to capture temporal dependencies in data sequences in both past and future directions. It is composed with two Long Short-Term Memory (LSTM) <xref ref-type="bibr" rid="scirp.137581-15">
      [15]
     </xref>.</p>
    <p>A sequence of LSTM <xref ref-type="bibr" rid="scirp.137581-16">
      [16]
     </xref> is made up of cells connected across time steps t. Each cell’s output is regulated by a group of activation functions 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mrow> 
        <mo>
          { 
        </mo> 
        <mrow> 
         <mi>
           f 
         </mi> 
         <mrow> 
          <mo>
            ( 
          </mo> 
          <mi>
            t 
          </mi> 
          <mo>
            ) 
          </mo> 
         </mrow> 
         <mo>
           , 
         </mo> 
         <mi>
           i 
         </mi> 
         <mrow> 
          <mo>
            ( 
          </mo> 
          <mi>
            t 
          </mi> 
          <mo>
            ) 
          </mo> 
         </mrow> 
         <mo>
           , 
         </mo> 
         <mi>
           o 
         </mi> 
         <mrow> 
          <mo>
            ( 
          </mo> 
          <mi>
            t 
          </mi> 
          <mo>
            ) 
          </mo> 
         </mrow> 
        </mrow> 
        <mo>
          } 
        </mo> 
       </mrow> 
      </mrow> 
     </math>, referred to as gates <xref ref-type="bibr" rid="scirp.137581-17">
      [17]
     </xref>, as depicted in <xref ref-type="fig" rid="fig2">
      Figure 2
     </xref>. These gates yield values in 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msup> 
        <mi>
          ℝ 
        </mi> 
        <mi>
          T 
        </mi> 
       </msup> 
      </mrow> 
     </math>, where T represents the dimension (the number of cells in the sequence) of the LSTM. The function 
     <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mi>
         f 
       </mi> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mi>
          t 
        </mi> 
        <mo>
          ) 
        </mo> 
       </mrow> 
      </mrow> 
     </math>, often referred to as the forget gate, determines the degree to which the information from the preceding cell 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          c 
        </mi> 
        <mrow> 
         <mi>
           t 
         </mi> 
         <mo>
           − 
         </mo> 
         <mn>
           1 
         </mn> 
        </mrow> 
       </msub> 
      </mrow> 
     </math> is discarded <xref ref-type="bibr" rid="scirp.137581-18">
      [18]
     </xref>. The input gate 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mi>
         i 
       </mi> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mi>
          t 
        </mi> 
        <mo>
          ) 
        </mo> 
       </mrow> 
      </mrow> 
     </math> specifies the portion of information to be stored in cell 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          c 
        </mi> 
        <mi>
          t 
        </mi> 
       </msub> 
      </mrow> 
     </math>. Meanwhile, the output gate 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mi>
         o 
       </mi> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mi>
          t 
        </mi> 
        <mo>
          ) 
        </mo> 
       </mrow> 
      </mrow> 
     </math> governs the fraction of the internal state transmitted to the subsequent cell.</p>
    <fig id="fig2" position="float">
     <label>Figure 2</label>
     <caption>
      <title>Figure 2. Composition of the LSTM cell.</title>
     </caption>
     <graphic mimetype="image" position="float" xlink:type="simple" xlink:href="https://html.scirp.org/file/1120105-rId55.jpeg?20241122031031" />
    </fig>
    <p>The cell input 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mi>
         x 
       </mi> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mi>
          t 
        </mi> 
        <mo>
          ) 
        </mo> 
       </mrow> 
      </mrow> 
     </math> at time t concatenates with output of a cell 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mi>
         c 
       </mi> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <mi>
           t 
         </mi> 
         <mo>
           − 
         </mo> 
         <mn>
           1 
         </mn> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
      </mrow> 
     </math> at previous time step 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <mi>
           t 
         </mi> 
         <mo>
           − 
         </mo> 
         <mn>
           1 
         </mn> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
      </mrow> 
     </math> which is 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mi>
         h 
       </mi> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <mi>
           t 
         </mi> 
         <mo>
           − 
         </mo> 
         <mn>
           1 
         </mn> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
      </mrow> 
     </math>. The resultant vector traverses through the input node and passes through the gates responsible for input, forgetting, and output to give 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mi>
         h 
       </mi> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mi>
          t 
        </mi> 
        <mo>
          ) 
        </mo> 
       </mrow> 
      </mrow> 
     </math>. Evaluation of 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mi>
         h 
       </mi> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mi>
          t 
        </mi> 
        <mo>
          ) 
        </mo> 
       </mrow> 
      </mrow> 
     </math> (Equation (4.6)) for each cell can be find in many papers, for example in <xref ref-type="bibr" rid="scirp.137581-19">
      [19]
     </xref> and <xref ref-type="bibr" rid="scirp.137581-20">
      [20]
     </xref>.</p>
    <p>
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mi>
         f 
       </mi> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mi>
          t 
        </mi> 
        <mo>
          ) 
        </mo> 
       </mrow> 
       <mo>
         = 
       </mo> 
       <mi>
         σ 
       </mi> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <msub> 
          <mi>
            w 
          </mi> 
          <mrow> 
           <mi>
             f 
           </mi> 
           <mi>
             h 
           </mi> 
          </mrow> 
         </msub> 
         <mi>
           h 
         </mi> 
         <mrow> 
          <mo>
            ( 
          </mo> 
          <mrow> 
           <mi>
             t 
           </mi> 
           <mo>
             − 
           </mo> 
           <mn>
             1 
           </mn> 
          </mrow> 
          <mo>
            ) 
          </mo> 
         </mrow> 
         <mo>
           + 
         </mo> 
         <msub> 
          <mi>
            w 
          </mi> 
          <mrow> 
           <mi>
             f 
           </mi> 
           <mi>
             x 
           </mi> 
          </mrow> 
         </msub> 
         <mi>
           x 
         </mi> 
         <mrow> 
          <mo>
            ( 
          </mo> 
          <mi>
            t 
          </mi> 
          <mo>
            ) 
          </mo> 
         </mrow> 
         <mo>
           + 
         </mo> 
         <msub> 
          <mi>
            b 
          </mi> 
          <mi>
            f 
          </mi> 
         </msub> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
      </mrow> 
     </math> (4.1)</p>
    <p>
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mi>
         i 
       </mi> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mi>
          t 
        </mi> 
        <mo>
          ) 
        </mo> 
       </mrow> 
       <mo>
         = 
       </mo> 
       <mi>
         σ 
       </mi> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <msub> 
          <mi>
            w 
          </mi> 
          <mrow> 
           <mi>
             i 
           </mi> 
           <mi>
             h 
           </mi> 
          </mrow> 
         </msub> 
         <mi>
           h 
         </mi> 
         <mrow> 
          <mo>
            ( 
          </mo> 
          <mrow> 
           <mi>
             t 
           </mi> 
           <mo>
             − 
           </mo> 
           <mn>
             1 
           </mn> 
          </mrow> 
          <mo>
            ) 
          </mo> 
         </mrow> 
         <mo>
           + 
         </mo> 
         <msub> 
          <mi>
            w 
          </mi> 
          <mrow> 
           <mi>
             i 
           </mi> 
           <mi>
             x 
           </mi> 
          </mrow> 
         </msub> 
         <mi>
           x 
         </mi> 
         <mrow> 
          <mo>
            ( 
          </mo> 
          <mi>
            t 
          </mi> 
          <mo>
            ) 
          </mo> 
         </mrow> 
         <mo>
           + 
         </mo> 
         <msub> 
          <mi>
            b 
          </mi> 
          <mi>
            i 
          </mi> 
         </msub> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
      </mrow> 
     </math>(4.2)</p>
    <p>
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mover accent="true"> 
        <mi>
          c 
        </mi> 
        <mo>
          ˜ 
        </mo> 
       </mover> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mi>
          t 
        </mi> 
        <mo>
          ) 
        </mo> 
       </mrow> 
       <mo>
         = 
       </mo> 
       <mi>
         tanh 
       </mi> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <msub> 
          <mi>
            w 
          </mi> 
          <mrow> 
           <mover accent="true"> 
            <mi>
              c 
            </mi> 
            <mo>
              ˜ 
            </mo> 
           </mover> 
           <mi>
             h 
           </mi> 
          </mrow> 
         </msub> 
         <mi>
           h 
         </mi> 
         <mrow> 
          <mo>
            ( 
          </mo> 
          <mrow> 
           <mi>
             t 
           </mi> 
           <mo>
             − 
           </mo> 
           <mn>
             1 
           </mn> 
          </mrow> 
          <mo>
            ) 
          </mo> 
         </mrow> 
         <mo>
           + 
         </mo> 
         <msub> 
          <mi>
            w 
          </mi> 
          <mrow> 
           <mover accent="true"> 
            <mi>
              c 
            </mi> 
            <mo>
              ˜ 
            </mo> 
           </mover> 
           <mi>
             x 
           </mi> 
          </mrow> 
         </msub> 
         <mi>
           x 
         </mi> 
         <mrow> 
          <mo>
            ( 
          </mo> 
          <mi>
            t 
          </mi> 
          <mo>
            ) 
          </mo> 
         </mrow> 
         <mo>
           + 
         </mo> 
         <msub> 
          <mi>
            b 
          </mi> 
          <mover accent="true"> 
           <mi>
             c 
           </mi> 
           <mo>
             ˜ 
           </mo> 
          </mover> 
         </msub> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
      </mrow> 
     </math>(4.3)</p>
    <p>
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mi>
         c 
       </mi> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mi>
          t 
        </mi> 
        <mo>
          ) 
        </mo> 
       </mrow> 
       <mo>
         = 
       </mo> 
       <mi>
         f 
       </mi> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mi>
          t 
        </mi> 
        <mo>
          ) 
        </mo> 
       </mrow> 
       <mo>
         ⋅ 
       </mo> 
       <mi>
         c 
       </mi> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <mi>
           t 
         </mi> 
         <mo>
           − 
         </mo> 
         <mn>
           1 
         </mn> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
       <mo>
         + 
       </mo> 
       <mi>
         i 
       </mi> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mi>
          t 
        </mi> 
        <mo>
          ) 
        </mo> 
       </mrow> 
       <mo>
         ⋅ 
       </mo> 
       <mover accent="true"> 
        <mi>
          c 
        </mi> 
        <mo>
          ˜ 
        </mo> 
       </mover> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mi>
          t 
        </mi> 
        <mo>
          ) 
        </mo> 
       </mrow> 
      </mrow> 
     </math>(4.4)</p>
    <p>
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mi>
         o 
       </mi> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mi>
          t 
        </mi> 
        <mo>
          ) 
        </mo> 
       </mrow> 
       <mo>
         = 
       </mo> 
       <mi>
         σ 
       </mi> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <msub> 
          <mi>
            w 
          </mi> 
          <mrow> 
           <mi>
             o 
           </mi> 
           <mi>
             h 
           </mi> 
          </mrow> 
         </msub> 
         <mi>
           h 
         </mi> 
         <mrow> 
          <mo>
            ( 
          </mo> 
          <mrow> 
           <mi>
             t 
           </mi> 
           <mo>
             − 
           </mo> 
           <mn>
             1 
           </mn> 
          </mrow> 
          <mo>
            ) 
          </mo> 
         </mrow> 
         <mo>
           + 
         </mo> 
         <msub> 
          <mi>
            w 
          </mi> 
          <mrow> 
           <mi>
             o 
           </mi> 
           <mi>
             x 
           </mi> 
          </mrow> 
         </msub> 
         <mi>
           x 
         </mi> 
         <mrow> 
          <mo>
            ( 
          </mo> 
          <mi>
            t 
          </mi> 
          <mo>
            ) 
          </mo> 
         </mrow> 
         <mo>
           + 
         </mo> 
         <msub> 
          <mi>
            b 
          </mi> 
          <mi>
            o 
          </mi> 
         </msub> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
      </mrow> 
     </math>(4.5)</p>
    <p>
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mi>
         h 
       </mi> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mi>
          t 
        </mi> 
        <mo>
          ) 
        </mo> 
       </mrow> 
       <mo>
         = 
       </mo> 
       <mi>
         o 
       </mi> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mi>
          t 
        </mi> 
        <mo>
          ) 
        </mo> 
       </mrow> 
       <mo>
         ⋅ 
       </mo> 
       <mi>
         tanh 
       </mi> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <mi>
           c 
         </mi> 
         <mrow> 
          <mo>
            ( 
          </mo> 
          <mi>
            t 
          </mi> 
          <mo>
            ) 
          </mo> 
         </mrow> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
      </mrow> 
     </math>(4.6)</p>
    <p>It is important to notice that, there is many variants of LSTM cell depend on how function 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
        σ 
      </mi> 
     </math> and tanh are arranged and their number in the cell <xref ref-type="bibr" rid="scirp.137581-18">
      [18]
     </xref>. An LSTM layer operates unidirectionally, processing input sequentially from left to right, thus encoding dependencies solely on preceding elements in the sequence. To address this limitation, we employ an additional LSTM layer that operates in the opposite direction, enabling the detection of dependencies on subsequent elements in the text by processing input from right to left. This arrangement gives rise to a neural network known as a Bidirectional LSTM (BiLSTM) <xref ref-type="bibr" rid="scirp.137581-21">
      [21]
     </xref>. It enables the network to assign equal significance to both the initial and final segments of the sequence, leading to enhanced performance.</p>
    <fig id="fig3" position="float">
     <label>Figure 3</label>
     <caption>
      <title>Figure 3. Architecture of a bidirectional LSTM (BiLSTM) layer.</title>
     </caption>
     <graphic mimetype="image" position="float" xlink:type="simple" xlink:href="https://html.scirp.org/file/1120105-rId82.jpeg?20241122031031" />
    </fig>
    <p>Additionally, we illustrate an instance of a bidirectional LSTM layer in <xref ref-type="fig" rid="fig3">
      Figure 3
     </xref>. The result of a BiLSTM layer is the concatenation of the outputs in both directions, as show in Equation (5)</p>
    <p>
     <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mi>
         H 
       </mi> 
       <mo>
         = 
       </mo> 
       <mrow> 
        <mo>
          [ 
        </mo> 
        <mrow> 
         <mover accent="true"> 
          <mrow> 
           <msub> 
            <mi>
              h 
            </mi> 
            <mi>
              T 
            </mi> 
           </msub> 
          </mrow> 
          <mo stretchy="true">
            ⇀ 
          </mo> 
         </mover> 
         <mo>
           , 
         </mo> 
         <mover accent="true"> 
          <mrow> 
           <msub> 
            <mi>
              h 
            </mi> 
            <mi>
              T 
            </mi> 
           </msub> 
          </mrow> 
          <mo stretchy="true">
            ← 
          </mo> 
         </mover> 
        </mrow> 
        <mo>
          ] 
        </mo> 
       </mrow> 
      </mrow> 
     </math>(5)</p>
    <p>
     <xref ref-type="bibr" rid="scirp.137581-"></xref>where 
     <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mover accent="true"> 
        <mrow> 
         <msub> 
          <mi>
            h 
          </mi> 
          <mi>
            T 
          </mi> 
         </msub> 
        </mrow> 
        <mo stretchy="true">
          ⇀ 
        </mo> 
       </mover> 
      </mrow> 
     </math> and 
     <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mover accent="true"> 
        <mrow> 
         <msub> 
          <mi>
            h 
          </mi> 
          <mi>
            T 
          </mi> 
         </msub> 
        </mrow> 
        <mo stretchy="true">
          ← 
        </mo> 
       </mover> 
      </mrow> 
     </math> are define by Equation (4.6).</p>
   </sec>
   <sec id="s3_3">
    <title>3.3. Processing Framework</title>
    <p>The intrusion detection method we propose is divided into two steps. The first step involves dimensionality reduction using Deep Auto-Encoder (DAE), followed by a phase of behavior class prediction with BiLSTM recurrent neural networks. The dimensionality reduction process begins with data cleaning, which involves identifying, correcting, or removing erroneous data. Errors are indicated by missing values (NaN), infinite values (Inf), and empty cells within the dataset. To address these issues, we have developed an algorithm that replaces NaN values with 0, substitutes Inf values with the maximum value of the dataset, and fills empty cells with the mean value of the relevant attributes.</p>
    <p>Following data cleaning, we moved on to normalization. This preprocessing step adjusts the data values to ensure they fall within a consistent range, which helps machine learning algorithms function more efficiently and reliably. In this study, we applied Z-Score normalization) as our chosen technique. This method standardizes the data to have a mean of 0 and a standard deviation of 1, as specified in Equation (6).</p>
    <p>
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msup> 
        <mi>
          X 
        </mi> 
        <mo>
          ′ 
        </mo> 
       </msup> 
       <mo>
         = 
       </mo> 
       <mrow> 
        <mrow> 
         <mrow> 
          <mo>
            ( 
          </mo> 
          <mrow> 
           <mi>
             X 
           </mi> 
           <mo>
             − 
           </mo> 
           <mi>
             μ 
           </mi> 
          </mrow> 
          <mo>
            ) 
          </mo> 
         </mrow> 
        </mrow> 
        <mo>
          / 
        </mo> 
        <mi>
          σ 
        </mi> 
       </mrow> 
      </mrow> 
     </math> (6)</p>
    <p>where X represents the original data value, X’ is the normalized value, with 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
        μ 
      </mi> 
     </math> and 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
        σ 
      </mi> 
     </math> representing the mean and standard deviation, respectively. Z-Score normalization is the most appropriate method for autoencoders <xref ref-type="bibr" rid="scirp.137581-22">
      [22]
     </xref> as it provides a uniform and centered scale for the data, which enhances the efficiency of the autoencoder’s reconstruction and the model’s stability. The final step in this phase was to design and train the deep autoencoder architecture to extract reduced and representative features from the original data. <xref ref-type="bibr" rid="scirp.137581-#a1">
      Algorithm 1
     </xref> below outlines the sequence of steps involved in the dimensionality reduction process.</p>
    <p>
     <xref ref-type="bibr" rid="scirp.137581-"></xref>Algorithm 1. Dimensionality reduction.</p>
    <table class="MsoTableGrid custom-table" border="0" cellspacing="0" cellpadding="0"> 
     <tr> 
      <td class="aleft" width="6.96%"><p style="text-align:left">1</p></td> 
      <td class="aleft" width="93.04%"><p style="text-align:left">Input: D (original Dataset)</p></td> 
     </tr> 
     <tr> 
      <td class="aleft" width="6.96%"><p style="text-align:left">2</p></td> 
      <td class="aleft" width="93.04%"><p style="text-align:left"> Length = row number of D</p></td> 
     </tr> 
     <tr> 
      <td class="aleft" width="6.96%"><p style="text-align:left">3</p></td> 
      <td class="aleft" width="93.04%"><p style="text-align:left">Output: Compressed Dataset</p></td> 
     </tr> 
     <tr> 
      <td class="aleft" width="6.96%"><p style="text-align:left">4</p></td> 
      <td class="aleft" width="93.04%"><p style="text-align:left">For (i ← 1 to i ≤ Length)</p></td> 
     </tr> 
     <tr> 
      <td class="aleft" width="6.96%"><p style="text-align:left">5</p></td> 
      <td class="aleft" width="93.04%"><p style="text-align:left"> If 
        <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
          <msub> 
           <mi>
             D 
           </mi> 
           <mi>
             i 
           </mi> 
          </msub> 
          <mrow> 
           <mo>
             ( 
           </mo> 
           <mrow> 
            <msub> 
             <mi>
               x 
             </mi> 
             <mn>
               1 
             </mn> 
            </msub> 
            <mo>
              , 
            </mo> 
            <msub> 
             <mi>
               x 
             </mi> 
             <mn>
               2 
             </mn> 
            </msub> 
            <mo>
              , 
            </mo> 
            <mo>
              ⋯ 
            </mo> 
            <mo>
              , 
            </mo> 
            <msub> 
             <mi>
               x 
             </mi> 
             <mi>
               n 
             </mi> 
            </msub> 
           </mrow> 
           <mo>
             ) 
           </mo> 
          </mrow> 
         </mrow> 
        </math> contain inf, Nan or empty replace </p></td> 
     </tr> 
     <tr> 
      <td class="aleft" width="6.96%"><p style="text-align:left">6</p></td> 
      <td class="aleft" width="93.04%"><p style="text-align:left"> respectively with max value, zero and maximum value </p></td> 
     </tr> 
     <tr> 
      <td class="aleft" width="6.96%"><p style="text-align:left">7</p></td> 
      <td class="aleft" width="93.04%"><p style="text-align:left"> End if </p></td> 
     </tr> 
     <tr> 
      <td class="aleft" width="6.96%"><p style="text-align:left">8</p></td> 
      <td class="aleft" width="93.04%"><p style="text-align:left">End for </p></td> 
     </tr> 
     <tr> 
      <td class="aleft" width="6.96%"><p style="text-align:left">9</p></td> 
      <td class="aleft" width="93.04%"><p style="text-align:left">Update D</p></td> 
     </tr> 
     <tr> 
      <td class="aleft" width="6.96%"><p style="text-align:left">10</p></td> 
      <td class="aleft" width="93.04%"><p style="text-align:left">Apply Data normalization and update D (Equation (6))</p></td> 
     </tr> 
     <tr> 
      <td class="aleft" width="6.96%"><p style="text-align:left">11</p></td> 
      <td class="aleft" width="93.04%"><p style="text-align:left">Build _the _DAE (input data :D, Number of layers, activation function):</p></td> 
     </tr> 
     <tr> 
      <td class="aleft" width="6.96%"><p style="text-align:left">12</p></td> 
      <td class="aleft" width="93.04%"><p style="text-align:left"> For each layer in the encoder phase: Apply 
        <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
          <mi>
            h 
          </mi> 
          <mrow> 
           <mo>
             ( 
           </mo> 
           <mi>
             l 
           </mi> 
           <mo>
             ) 
           </mo> 
          </mrow> 
          <mo>
            = 
          </mo> 
          <mi>
            f 
          </mi> 
          <mrow> 
           <mo>
             ( 
           </mo> 
           <mrow> 
            <mi>
              W 
            </mi> 
            <mrow> 
             <mo>
               ( 
             </mo> 
             <mi>
               l 
             </mi> 
             <mo>
               ) 
             </mo> 
            </mrow> 
            <mo>
              ∗ 
            </mo> 
            <mi>
              h 
            </mi> 
            <mrow> 
             <mo>
               ( 
             </mo> 
             <mrow> 
              <mi>
                l 
              </mi> 
              <mo>
                − 
              </mo> 
              <mn>
                1 
              </mn> 
             </mrow> 
             <mo>
               ) 
             </mo> 
            </mrow> 
            <mo>
              + 
            </mo> 
            <mi>
              b 
            </mi> 
            <mrow> 
             <mo>
               ( 
             </mo> 
             <mi>
               l 
             </mi> 
             <mo>
               ) 
             </mo> 
            </mrow> 
           </mrow> 
           <mo>
             ) 
           </mo> 
          </mrow> 
         </mrow> 
        </math> </p></td> 
     </tr> 
     <tr> 
      <td class="aleft" width="6.96%"><p style="text-align:left">13</p></td> 
      <td class="aleft" width="93.04%"><p style="text-align:left"> For each layer l in the decoder phase: Apply 
        <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
          <mi>
            h 
          </mi> 
          <mrow> 
           <mo>
             ( 
           </mo> 
           <mi>
             l 
           </mi> 
           <mo>
             ) 
           </mo> 
          </mrow> 
          <mo>
            = 
          </mo> 
          <mi>
            f 
          </mi> 
          <mrow> 
           <mo>
             ( 
           </mo> 
           <mrow> 
            <mi>
              W 
            </mi> 
            <mrow> 
             <mo>
               ( 
             </mo> 
             <mi>
               l 
             </mi> 
             <mo>
               ) 
             </mo> 
            </mrow> 
            <mo>
              ∗ 
            </mo> 
            <mi>
              h 
            </mi> 
            <mrow> 
             <mo>
               ( 
             </mo> 
             <mrow> 
              <mi>
                l 
              </mi> 
              <mo>
                − 
              </mo> 
              <mn>
                1 
              </mn> 
             </mrow> 
             <mo>
               ) 
             </mo> 
            </mrow> 
            <mo>
              + 
            </mo> 
            <mi>
              b 
            </mi> 
            <mrow> 
             <mo>
               ( 
             </mo> 
             <mi>
               l 
             </mi> 
             <mo>
               ) 
             </mo> 
            </mrow> 
           </mrow> 
           <mo>
             ) 
           </mo> 
          </mrow> 
         </mrow> 
        </math> </p></td> 
     </tr> 
     <tr> 
      <td class="aleft" width="6.96%"><p style="text-align:left">14</p></td> 
      <td class="aleft" width="93.04%"><p style="text-align:left">Train_the_DAE (input vector X in D, encoder network, decoder network, Epoch, Batch size): </p></td> 
     </tr> 
     <tr> 
      <td class="aleft" width="6.96%"><p style="text-align:left">15</p></td> 
      <td class="aleft" width="93.04%"><p style="text-align:left"> Z = encoder(X) (Equation (1))</p></td> 
     </tr> 
     <tr> 
      <td class="aleft" width="6.96%"><p style="text-align:left">16</p></td> 
      <td class="aleft" width="93.04%"><p style="text-align:left"> X’ = decoder(Z) (Equation (3))</p></td> 
     </tr> 
     <tr> 
      <td class="aleft" width="6.96%"><p style="text-align:left">17</p></td> 
      <td class="aleft" width="93.04%"><p style="text-align:left"> Loss = MeanSquaredError (X, X’) (Equation (2))</p></td> 
     </tr> 
     <tr> 
      <td class="aleft" width="6.96%"><p style="text-align:left">18</p></td> 
      <td class="aleft" width="93.04%"><p style="text-align:left">Compressed dataset:</p></td> 
     </tr> 
     <tr> 
      <td class="aleft" width="6.96%"><p style="text-align:left">19</p></td> 
      <td class="aleft" width="93.04%"><p style="text-align:left"> Update D with latent vector Z</p></td> 
     </tr> 
    </table>
    <p>In line 12 and 13, 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mi>
         h 
       </mi> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mn>
          0 
        </mn> 
        <mo>
          ) 
        </mo> 
       </mrow> 
      </mrow> 
     </math> is the input data. 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mi>
         f 
       </mi> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mo>
          . 
        </mo> 
        <mo>
          ) 
        </mo> 
       </mrow> 
      </mrow> 
     </math> is activation function. 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mi>
         W 
       </mi> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mi>
          l 
        </mi> 
        <mo>
          ) 
        </mo> 
       </mrow> 
      </mrow> 
     </math> and 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mi>
         b 
       </mi> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mi>
          l 
        </mi> 
        <mo>
          ) 
        </mo> 
       </mrow> 
      </mrow> 
     </math> are the weights and biases of layer 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mo> 
       </mo> 
       <mi>
         l 
       </mi> 
      </mrow> 
     </math>.</p>
    <p>The compressed data from Algorithm 1 is subsequently utilized in the classification phase, which is based on a Bidirectional LSTM, to determine if the traffic is normal or indicative of an attack. The steps of this process are outlined in <xref ref-type="bibr" rid="scirp.137581-#a2">
      Algorithm 2
     </xref>.</p>
    <p>
     <xref ref-type="bibr" rid="scirp.137581-"></xref>Algorithm 2. Intrusion detection.</p>
    <table class="MsoTableGrid custom-table" border="0" cellspacing="0" cellpadding="0"> 
     <tr> 
      <td class="aleft" width="8.44%"><p style="text-align:left">1</p></td> 
      <td class="aleft" width="91.56%"><p style="text-align:left">Input: 
        <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
          <mi>
            D 
          </mi> 
          <mo>
            = 
          </mo> 
          <mrow> 
           <mo>
             [ 
           </mo> 
           <mrow> 
            <msub> 
             <mi>
               X 
             </mi> 
             <mn>
               1 
             </mn> 
            </msub> 
            <mo>
              , 
            </mo> 
            <msub> 
             <mi>
               X 
             </mi> 
             <mn>
               2 
             </mn> 
            </msub> 
            <mo>
              , 
            </mo> 
            <mo>
              ⋯ 
            </mo> 
            <mo>
              , 
            </mo> 
            <msub> 
             <mi>
               X 
             </mi> 
             <mi>
               n 
             </mi> 
            </msub> 
           </mrow> 
           <mo>
             ] 
           </mo> 
          </mrow> 
         </mrow> 
        </math> (Compressed data from Algorithm 1)</p></td> 
     </tr> 
     <tr> 
      <td class="aleft" width="8.44%"><p style="text-align:left">3</p></td> 
      <td class="aleft" width="91.56%"><p style="text-align:left">Output: Class of traffic (normal or abnormal)</p></td> 
     </tr> 
     <tr> 
      <td class="aleft" width="8.44%"><p style="text-align:left">4</p></td> 
      <td class="aleft" width="91.56%"><p style="text-align:left">Split Data into training and testing set: </p></td> 
     </tr> 
     <tr> 
      <td class="aleft" width="8.44%"><p style="text-align:left">5</p></td> 
      <td class="aleft" width="91.56%"><p style="text-align:left"> 
        <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
          <msub> 
           <mi>
             X 
           </mi> 
           <mrow> 
            <mi>
              t 
            </mi> 
            <mi>
              r 
            </mi> 
            <mi>
              a 
            </mi> 
            <mi>
              i 
            </mi> 
            <mi>
              n 
            </mi> 
           </mrow> 
          </msub> 
          <mo>
            = 
          </mo> 
          <msub> 
           <mi>
             X 
           </mi> 
           <mn>
             1 
           </mn> 
          </msub> 
          <mo>
            , 
          </mo> 
          <mtext> 
          </mtext> 
          <msub> 
           <mi>
             X 
           </mi> 
           <mn>
             2 
           </mn> 
          </msub> 
          <mo>
            , 
          </mo> 
          <mo>
            ⋯ 
          </mo> 
          <mo>
            , 
          </mo> 
          <msub> 
           <mi>
             X 
           </mi> 
           <mi>
             k 
           </mi> 
          </msub> 
         </mrow> 
        </math> </p></td> 
     </tr> 
     <tr> 
      <td class="aleft" width="8.44%"><p style="text-align:left">6</p></td> 
      <td class="aleft" width="91.56%"><p style="text-align:left"> 
        <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
          <msub> 
           <mi>
             X 
           </mi> 
           <mrow> 
            <mi>
              t 
            </mi> 
            <mi>
              e 
            </mi> 
            <mi>
              s 
            </mi> 
            <mi>
              t 
            </mi> 
           </mrow> 
          </msub> 
          <mo>
            = 
          </mo> 
          <msub> 
           <mi>
             X 
           </mi> 
           <mrow> 
            <mi>
              k 
            </mi> 
            <mo>
              + 
            </mo> 
            <mn>
              1 
            </mn> 
           </mrow> 
          </msub> 
          <mo>
            , 
          </mo> 
          <msub> 
           <mi>
             X 
           </mi> 
           <mrow> 
            <mi>
              k 
            </mi> 
            <mo>
              + 
            </mo> 
            <mn>
              2 
            </mn> 
           </mrow> 
          </msub> 
          <mo>
            , 
          </mo> 
          <mo>
            ⋯ 
          </mo> 
          <mo>
            , 
          </mo> 
          <msub> 
           <mi>
             X 
           </mi> 
           <mi>
             n 
           </mi> 
          </msub> 
         </mrow> 
        </math></p></td> 
     </tr> 
     <tr> 
      <td class="aleft" width="8.44%"><p style="text-align:left">7</p></td> 
      <td class="aleft" width="91.56%"><p style="text-align:left">Bulding the model: </p></td> 
     </tr> 
     <tr> 
      <td class="aleft" width="8.44%"><p style="text-align:left">8</p></td> 
      <td class="aleft" width="91.56%"><p style="text-align:left"> Forward LSTM: </p></td> 
     </tr> 
     <tr> 
      <td class="aleft" width="8.44%"><p style="text-align:left">9</p></td> 
      <td class="aleft" width="91.56%"><p style="text-align:left"> For each time step t from 1 to T</p></td> 
     </tr> 
     <tr> 
      <td class="aleft" width="8.44%"><p style="text-align:left">10</p></td> 
      <td class="aleft" width="91.56%"><p style="text-align:left"> For each X<sub>i</sub> in X<sub>train</sub> do: </p></td> 
     </tr> 
     <tr> 
      <td class="aleft" width="8.44%"><p style="text-align:left">11</p></td> 
      <td class="aleft" width="91.56%"><p style="text-align:left"> Calculate 
        <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
          <mi>
            f 
          </mi> 
          <mrow> 
           <mo>
             ( 
           </mo> 
           <mi>
             t 
           </mi> 
           <mo>
             ) 
           </mo> 
          </mrow> 
         </mrow> 
        </math> (Equation (4.1)), 
        <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
          <mi>
            i 
          </mi> 
          <mrow> 
           <mo>
             ( 
           </mo> 
           <mi>
             t 
           </mi> 
           <mo>
             ) 
           </mo> 
          </mrow> 
         </mrow> 
        </math> (Equation (4.2) and 
        <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
          <mover accent="true"> 
           <mi>
             c 
           </mi> 
           <mo>
             ˜ 
           </mo> 
          </mover> 
          <mrow> 
           <mo>
             ( 
           </mo> 
           <mi>
             t 
           </mi> 
           <mo>
             ) 
           </mo> 
          </mrow> 
         </mrow> 
        </math> (Equation (4.3))</p></td> 
     </tr> 
     <tr> 
      <td class="aleft" width="8.44%"><p style="text-align:left">12</p></td> 
      <td class="aleft" width="91.56%"><p style="text-align:left"> Update Cell state 
        <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
          <mi>
            c 
          </mi> 
          <mrow> 
           <mo>
             ( 
           </mo> 
           <mi>
             t 
           </mi> 
           <mo>
             ) 
           </mo> 
          </mrow> 
         </mrow> 
        </math> (Equation (4.4)) </p></td> 
     </tr> 
     <tr> 
      <td class="aleft" width="8.44%"><p style="text-align:left">13</p></td> 
      <td class="aleft" width="91.56%"><p style="text-align:left"> Calculate 
        <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
          <mi>
            o 
          </mi> 
          <mrow> 
           <mo>
             ( 
           </mo> 
           <mi>
             t 
           </mi> 
           <mo>
             ) 
           </mo> 
          </mrow> 
         </mrow> 
        </math>( Equation (4.4)) and 
        <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
          <mover accent="true"> 
           <mrow> 
            <mi>
              h 
            </mi> 
            <mrow> 
             <mo>
               ( 
             </mo> 
             <mi>
               t 
             </mi> 
             <mo>
               ) 
             </mo> 
            </mrow> 
           </mrow> 
           <mo stretchy="true">
             ⇀ 
           </mo> 
          </mover> 
         </mrow> 
        </math> Equation (4.5))</p></td> 
     </tr> 
     <tr> 
      <td class="aleft" width="8.44%"><p style="text-align:left">14</p></td> 
      <td class="aleft" width="91.56%"><p style="text-align:left"> Backward LSTM:</p></td> 
     </tr> 
     <tr> 
      <td class="aleft" width="8.44%"><p style="text-align:left">15</p></td> 
      <td class="aleft" width="91.56%"><p style="text-align:left"> For each time step t from T to 1</p></td> 
     </tr> 
     <tr> 
      <td class="aleft" width="8.44%"><p style="text-align:left">16</p></td> 
      <td class="aleft" width="91.56%"><p style="text-align:left"> For each X<sub>i</sub> in X<sub>train</sub> do</p></td> 
     </tr> 
     <tr> 
      <td class="aleft" width="8.44%"><p style="text-align:left">17</p></td> 
      <td class="aleft" width="91.56%"><p style="text-align:left"> Calculate 
        <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
          <mi>
            f 
          </mi> 
          <mrow> 
           <mo>
             ( 
           </mo> 
           <mi>
             t 
           </mi> 
           <mo>
             ) 
           </mo> 
          </mrow> 
         </mrow> 
        </math> (Equation (4.1)), 
        <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
          <mi>
            i 
          </mi> 
          <mrow> 
           <mo>
             ( 
           </mo> 
           <mi>
             t 
           </mi> 
           <mo>
             ) 
           </mo> 
          </mrow> 
         </mrow> 
        </math> (Equation (4.2) and 
        <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
          <mover accent="true"> 
           <mi>
             c 
           </mi> 
           <mo>
             ˜ 
           </mo> 
          </mover> 
          <mrow> 
           <mo>
             ( 
           </mo> 
           <mi>
             t 
           </mi> 
           <mo>
             ) 
           </mo> 
          </mrow> 
         </mrow> 
        </math> (Equation (4.3)) </p></td> 
     </tr> 
     <tr> 
      <td class="aleft" width="8.44%"><p style="text-align:left">18</p></td> 
      <td class="aleft" width="91.56%"><p style="text-align:left"> Update Cell state 
        <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
          <mtext> 
          </mtext> 
          <mi>
            c 
          </mi> 
          <mrow> 
           <mo>
             ( 
           </mo> 
           <mi>
             t 
           </mi> 
           <mo>
             ) 
           </mo> 
          </mrow> 
         </mrow> 
        </math> (Equation (4.4))</p></td> 
     </tr> 
     <tr> 
      <td class="aleft" width="8.44%"><p style="text-align:left">19</p></td> 
      <td class="aleft" width="91.56%"><p style="text-align:left"> Calculate 
        <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
          <mi>
            o 
          </mi> 
          <mrow> 
           <mo>
             ( 
           </mo> 
           <mi>
             t 
           </mi> 
           <mo>
             ) 
           </mo> 
          </mrow> 
         </mrow> 
        </math> (Equation (4.4)) and 
        <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
          <mover accent="true"> 
           <mrow> 
            <mi>
              h 
            </mi> 
            <mrow> 
             <mo>
               ( 
             </mo> 
             <mi>
               t 
             </mi> 
             <mo>
               ) 
             </mo> 
            </mrow> 
           </mrow> 
           <mo stretchy="true">
             ← 
           </mo> 
          </mover> 
         </mrow> 
        </math> Equation (4.5))</p></td> 
     </tr> 
     <tr> 
      <td class="aleft" width="8.44%"><p style="text-align:left">20</p></td> 
      <td class="aleft" width="91.56%"><p style="text-align:left"> Concatenation: </p></td> 
     </tr> 
     <tr> 
      <td class="aleft" width="8.44%"><p style="text-align:left">21</p></td> 
      <td class="aleft" width="91.56%"><p style="text-align:left"> For each time step t, concatenate the forward and backward LSTM hidden states</p></td> 
     </tr> 
     <tr> 
      <td class="aleft" width="8.44%"><p style="text-align:left">22</p></td> 
      <td class="aleft" width="91.56%"><p style="text-align:left"> 
        <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
          <mi>
            h 
          </mi> 
          <mrow> 
           <mo>
             ( 
           </mo> 
           <mi>
             t 
           </mi> 
           <mo>
             ) 
           </mo> 
          </mrow> 
          <mo>
            = 
          </mo> 
          <mrow> 
           <mo>
             [ 
           </mo> 
           <mrow> 
            <mover accent="true"> 
             <mrow> 
              <mi>
                h 
              </mi> 
              <mrow> 
               <mo>
                 ( 
               </mo> 
               <mi>
                 t 
               </mi> 
               <mo>
                 ) 
               </mo> 
              </mrow> 
             </mrow> 
             <mo stretchy="true">
               ⇀ 
             </mo> 
            </mover> 
            <mo>
              , 
            </mo> 
            <mover accent="true"> 
             <mrow> 
              <mi>
                h 
              </mi> 
              <mrow> 
               <mo>
                 ( 
               </mo> 
               <mi>
                 t 
               </mi> 
               <mo>
                 ) 
               </mo> 
              </mrow> 
             </mrow> 
             <mo stretchy="true">
               ← 
             </mo> 
            </mover> 
           </mrow> 
           <mo>
             ] 
           </mo> 
          </mrow> 
         </mrow> 
        </math> (Equation (5))</p></td> 
     </tr> 
     <tr> 
      <td class="aleft" width="8.44%"><p style="text-align:left">23</p></td> 
      <td class="aleft" width="91.56%"><p style="text-align:left">Output Layer: 
        <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
          <msub> 
           <mover accent="true"> 
            <mi>
              y 
            </mi> 
            <mo>
              ^ 
            </mo> 
           </mover> 
           <mi>
             t 
           </mi> 
          </msub> 
          <mo>
            = 
          </mo> 
          <mi>
            σ 
          </mi> 
          <mrow> 
           <mo>
             ( 
           </mo> 
           <mrow> 
            <mi>
              W 
            </mi> 
            <mi>
              h 
            </mi> 
            <mrow> 
             <mo>
               ( 
             </mo> 
             <mi>
               t 
             </mi> 
             <mo>
               ) 
             </mo> 
            </mrow> 
           </mrow> 
           <mo>
             ) 
           </mo> 
          </mrow> 
         </mrow> 
        </math> ( 
        <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
           σ 
         </mi> 
        </math> is an activation function, W is the weight matrix)</p></td> 
     </tr> 
     <tr> 
      <td class="aleft" width="8.44%"><p style="text-align:left">24</p></td> 
      <td class="aleft" width="91.56%"><p style="text-align:left">Training: </p></td> 
     </tr> 
     <tr> 
      <td class="aleft" width="8.44%"><p style="text-align:left">25</p></td> 
      <td class="aleft" width="91.56%"><p style="text-align:left"> Use the loss function on 
        <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
          <msub> 
           <mover accent="true"> 
            <mi>
              y 
            </mi> 
            <mo>
              ^ 
            </mo> 
           </mover> 
           <mi>
             t 
           </mi> 
          </msub> 
         </mrow> 
        </math> to compute the error </p></td> 
     </tr> 
     <tr> 
      <td class="aleft" width="8.44%"><p style="text-align:left">26</p></td> 
      <td class="aleft" width="91.56%"><p style="text-align:left"> Optimize the weight of 
        <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
          <msub> 
           <mover accent="true"> 
            <mi>
              y 
            </mi> 
            <mo>
              ^ 
            </mo> 
           </mover> 
           <mi>
             t 
           </mi> 
          </msub> 
         </mrow> 
        </math> to reduce loss</p></td> 
     </tr> 
     <tr> 
      <td class="aleft" width="8.44%"><p style="text-align:left">27</p></td> 
      <td class="aleft" width="91.56%"><p style="text-align:left">Evaluation:</p></td> 
     </tr> 
     <tr> 
      <td class="aleft" width="8.44%"><p style="text-align:left">28</p></td> 
      <td class="aleft" width="91.56%"><p style="text-align:left"> Test the model on the validation data X<sub>test</sub></p></td> 
     </tr> 
    </table>
   </sec>
  </sec><sec id="s4">
   <title>4. Experimentation, Results and Discussion</title>
   <sec id="s4_1">
    <title>4.1. Experimentation Framework AND Results</title>
    <p>The experiments were carried out on a PC equipped with a 16 GB RAM Core i7 processor running Windows 11 64-bit. For software, we utilized Python 3.7 within the Anaconda integrated development environment (IDE), which serves as the editor and includes various libraries such as Scikit-learn, TensorFlow, Keras, SciPy, NumPy, Pandas, and Matplotlib.</p>
    <p>Concerning the dataset, we use a most recent Canadian Institute for Cybersecurity (CIC) dataset called CIC2017 available at <xref ref-type="bibr" rid="scirp.137581-23">
      [23]
     </xref>. The CICIDS2017 dataset offers a wide range of attack scenarios, including Denial of Service (DoS), Distributed Denial of Service (DDoS), brute force, web attacks, infiltration, botnet, and others. This variety makes it ideal for testing and comparing different intrusion detection techniques. The dataset simulates realistic network traffic by combining both normal and malicious activities, closely resembling real-world conditions as opposed to fully synthetic datasets. It provides over 80 features extracted from network traffic, encompassing basic information like IP addresses and port numbers, as well as more advanced statistical metrics such as packet length and flow duration. Further details can be found at <xref ref-type="bibr" rid="scirp.137581-24">
      [24]
     </xref>.</p>
    <p>As mentioned in our proposal, we have chosen deep auto-encoder as module for dimensionality reduction. After carrying out several executions, to find the right optimal value for the number hidden layers, we return 3 hidden layers between the input layer which is a vector of size 80 and the bottleneck layer whose size is 25. The total number of trainable parameters is 22,141. The learning curve for our deep autoencoder was set to 20 epochs with a batch size of 320 (after several evaluation tests). <xref ref-type="table" rid="table1">
      Table 1
     </xref> summarizes the features of DAE.</p>
    <table-wrap id="table1">
     <label>
      <xref ref-type="table" rid="table1">
       Table 1
      </xref></label>
     <caption>
      <title>
       <xref ref-type="bibr" rid="scirp.137581-"></xref>Table 1. Characteristics deep auto-encoder.</title>
     </caption>
     <table class="MsoTableGrid custom-table" border="0" cellspacing="0" cellpadding="0"> 
      <tr> 
       <td class="custom-bottom-td acenter" width="17.09%" colspan="2"><p style="text-align:center">Steps</p></td> 
       <td class="custom-bottom-td acenter" width="17.09%"><p style="text-align:center">Layer type</p></td> 
       <td class="custom-bottom-td acenter" width="17.09%"><p style="text-align:center">Output size(features)</p></td> 
       <td class="custom-bottom-td acenter" width="17.09%"><p style="text-align:center">Number of parameters</p></td> 
      </tr> 
      <tr> 
       <td rowspan="5" class="custom-top-td acenter" width="17.09%"><p style="text-align:center">Encoding</p></td> 
       <td rowspan="4" class="custom-top-td acenter" width="17.09%"><p style="text-align:center"></p></td> 
       <td class="custom-top-td acenter" width="17.09%"><p style="text-align:center">Input</p></td> 
       <td class="custom-top-td acenter" width="17.09%"><p style="text-align:center">80</p></td> 
       <td class="custom-top-td acenter" width="17.09%"><p style="text-align:center">0</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="17.09%"><p style="text-align:center">Hidden_1</p></td> 
       <td class="acenter" width="17.09%"><p style="text-align:center">64</p></td> 
       <td class="acenter" width="17.09%"><p style="text-align:center">5184</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="17.09%"><p style="text-align:center">Hidden_2</p></td> 
       <td class="acenter" width="17.09%"><p style="text-align:center">50</p></td> 
       <td class="acenter" width="17.09%"><p style="text-align:center">3250</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="17.09%"><p style="text-align:center">Hidden_3</p></td> 
       <td class="acenter" width="17.09%"><p style="text-align:center">34</p></td> 
       <td class="acenter" width="17.09%"><p style="text-align:center">1734</p></td> 
      </tr> 
      <tr> 
       <td rowspan="5" class="acenter" width="17.09%"><p style="text-align:center">Decoding</p></td> 
       <td class="acenter" width="17.09%"><p style="text-align:center">Bottleneck</p></td> 
       <td class="acenter" width="17.09%"><p style="text-align:center">25</p></td> 
       <td class="acenter" width="17.09%"><p style="text-align:center">875</p></td> 
      </tr> 
      <tr> 
       <td rowspan="4" class="acenter" width="17.09%"><p style="text-align:center"></p></td> 
       <td class="acenter" width="17.09%"><p style="text-align:center">Hidden_4</p></td> 
       <td class="acenter" width="17.09%"><p style="text-align:center">34</p></td> 
       <td class="acenter" width="17.09%"><p style="text-align:center">884</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="17.09%"><p style="text-align:center">Hidden_5</p></td> 
       <td class="acenter" width="17.09%"><p style="text-align:center">50</p></td> 
       <td class="acenter" width="17.09%"><p style="text-align:center">1750</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="17.09%"><p style="text-align:center">Hidden_6</p></td> 
       <td class="acenter" width="17.09%"><p style="text-align:center">64</p></td> 
       <td class="acenter" width="17.09%"><p style="text-align:center">3264</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="17.09%"><p style="text-align:center">output</p></td> 
       <td class="acenter" width="17.09%"><p style="text-align:center">80</p></td> 
       <td class="acenter" width="17.09%"><p style="text-align:center">5200</p></td> 
      </tr> 
     </table>
    </table-wrap>
    <p>The proposed DAE enabled us to reduce the number of attributes from 80 to 25 (i.e. a compression factor of 1:3.2). <xref ref-type="fig" rid="fig4">
      Figure 4
     </xref> shows the graph of the loss function for the different number of epochs, and this in both train and test dataset. As the amount of loss is under 5 × 10<sup>−</sup><sup>4</sup> for epoch near 40. An output vector very close to the input vector is obtained from the 25 selected features. Then, BiLSTM help us to classify data instances with satisfactory accuracy.</p>
    <fig id="fig4" position="float">
     <label>Figure 4</label>
     <caption>
      <title>Figure 4. Training and test loss over 100 epochs.</title>
     </caption>
     <graphic mimetype="image" position="float" xlink:type="simple" xlink:href="https://html.scirp.org/file/1120105-rId151.jpeg?20241122031032" />
    </fig>
    <p>We also conducted another experiment to evaluate the performance of our proposal. For this, we used two metrics: accuracy and recall. Our proposal (DAE-BiLSTM) performs better than random forests, AdaBoost, or BiLSTM (see <xref ref-type="fig" rid="fig5">
      Figure 5
     </xref>). DAE-BiLSTM achieved an accuracy of 97% and a recall of 95%, respectively.</p>
    <fig id="fig5" position="float">
     <label>Figure 5</label>
     <caption>
      <title>Figure 5. Performance metrics (accuracy and recall) of Randon Forest, Adaboost, BiLSTM DAE-BiLSTM.</title>
     </caption>
     <graphic mimetype="image" position="float" xlink:type="simple" xlink:href="https://html.scirp.org/file/1120105-rId152.jpeg?20241122031032" />
    </fig>
    <p>Finally, we evaluated the ROC curves and AUC values of Random Forest, Adaboost, BiLSTM, and DAE-BiLSTM in an intrusion detection task. The results, shown in <xref ref-type="fig" rid="fig5">
      Figure 5
     </xref>, indicate that DAE-BiLSTM achieved the highest AUC, close to 0.93, outperforming all other models.</p>
    <fig id="fig6" position="float">
     <label>Figure 6</label>
     <caption>
      <title>Figure 6. ROC Curves for various intrusion detection model.</title>
     </caption>
     <graphic mimetype="image" position="float" xlink:type="simple" xlink:href="https://html.scirp.org/file/1120105-rId153.jpeg?20241122031032" />
    </fig>
   </sec>
   <sec id="s4_2">
    <title>4.2. Discussion</title>
    <p>Upon examining <xref ref-type="fig" rid="fig4">
      Figure 4
     </xref>, we observe that the training and test loss curves gradually decrease before stabilizing around the 40<sup>th</sup> epoch, reaching a loss rate significantly below 5 × 10<sup>−</sup><sup>4</sup>. At this point, the two curves converge, indicating that the model is learning effectively from the training data while generalizing properly to the test data. The progressive decrease in loss followed by stabilization at a low value reflects the model’s good performance. This convergence toward a low loss value, with no sign of divergence between the two curves, suggests that there is neither underfitting nor overfitting. The model successfully captures the characteristics of the data without overfitting the training set, which is an indicator of its generalization capability.</p>
    <p>With regard to the histogram in <xref ref-type="fig" rid="fig5">
      Figure 5
     </xref>, which represents the models’ performance on Accuracy and Recall, we note that the DAE-BiLSTM model is the most performant for this intrusion detection problem, with higher performance on both metrics. The BiLSTM follows closely, while Adaboost and Random Forest are less effective. This indicates that, for this complex problem, recurrent neural networks are a top choice. More specifically, the fact that DAE-BiLSTM has the highest Accuracy (0.97) and Recall (0.95) among the four models means that it is not only capable of making much more accurate predictions than the others, but it is also excellent at correctly identifying positive cases. Therefore, it is a model that generalizes well.</p>
    <p>The performance of the DAE-BiLSTM compared to other models can be explained by the fact that deep autoencoders improve data processing by reducing dimensionality, learning compact representations that highlight key features and filter out noise. This enhances the quality of input features for the BiLSTM and boosts detection accuracy. By focusing on essential characteristics, the DAE-BiLSTM reduces the risk of overfitting and improves generalization, enabling better identification of network attacks in new data. Lower input dimensionality also reduces computational costs, allowing faster training and more efficient resource use, while smoothing the optimization process and decreasing the likelihood of the model getting stuck in local minima, which potentially boosts accuracy. The refined features from the first step help the BiLSTM in the second step better distinguish between normal and malicious activities, reducing false positives and improving the overall reliability of the intrusion detection system.</p>
    <p>
     <xref ref-type="fig" rid="fig6">
      Figure 6
     </xref> also shows the DAE-BiLSTM’s superiority, as its Receiver Operating Characteristic (ROC) curve approaches the top-left corner, demonstrating outstanding intrusion detection accuracy with minimal errors. This model excels by combining the strengths of the Deep Autoencoder (DAE) for dimensionality reduction and the BiLSTM for capturing temporal patterns, resulting in exceptional performance. Even when used on its own, BiLSTM performs well, indicating that recurrent neural networks are effective for intrusion detection due to their ability to recognize temporal relationships in data. On the other hand, traditional models like Adaboost and Random Forest fall short, likely because they struggle to handle complex, high-dimensional data. Area Under the Curve (AUC) serves as a crucial metric for evaluating model performance, with higher values indicating better class separation. The DAE-BiLSTM leads the pack with an AUC of 0.93, followed by BiLSTM (0.86), Adaboost (0.76), and Random Forest (0.68). The success of DAE-BiLSTM lies in its dual approach: dimensionality reduction via the DAE, which filters noise and emphasizes important features, and temporal pattern recognition by the BiLSTM. This combination enhances the model’s generalization and reliability, making DAE-BiLSTM the optimal choice for intrusion detection tasks.</p>
   </sec>
  </sec><sec id="s5">
   <title>5. Conclusions</title>
   <p>In this study, we present a two-step approach for intrusion detection in computer systems, leveraging the advantages of deep autoencoders for dimensionality reduction and BiLSTM for capturing detailed temporal patterns in the data. In the first step, the deep autoencoder reduces the dimensionality of the input features, filtering out noise and retaining only the most relevant characteristics. In the second step, the refined feature vector serves as input to the BiLSTM, which further learns the temporal dependencies and patterns, enhancing the model’s ability to detect intrusions effectively.</p>
   <p>This combined approach has yielded promising results, outperforming several existing methods in the literature. The dimensionality reduction not only improves the quality of the input data but also accelerates the learning process, potentially leading to faster convergence during training. Future work will focus on validating this observation and exploring alternative deep learning architectures for the second stage to identify the most suitable model for generalization. By integrating dimensionality reduction with advanced temporal pattern recognition, the proposed method offers a robust and efficient solution for intrusion detection, demonstrating significant improvements in accuracy and reliability.</p>
  </sec>
 </body><back>
  <ref-list>
   <title>References</title>
   <ref id="scirp.137581-ref1">
    <label>1</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Rai, A., et al. (2020) A Review of Information Security: Issues and Techniques. International Journal for Research in Applied Science and Engineering Technology, 8, 953-960. &gt;https://doi.org/10.22214/ijraset.2020.5150
    </mixed-citation>
   </ref>
   <ref id="scirp.137581-ref2">
    <label>2</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Alanazi, H., Noor, R., Zaidan, B.B., et al. (2010) Intrusion Detection System: Overview. Journal of Computing, 2, 130-133. &gt;https://doi.org/10.48550/arXiv.1002.4047
    </mixed-citation>
   </ref>
   <ref id="scirp.137581-ref3">
    <label>3</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Salih, R., Den Hartog, J. and Smulders, E. (2020) Semantical Rule-Based False Positive Detection for IDS. &gt;https://pure.tue.nl/ws/portalfiles/portal/174214825/Salih_R..pdf
    </mixed-citation>
   </ref>
   <ref id="scirp.137581-ref4">
    <label>4</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Santhosh Kumar, S.V.N., Selvi, M. and Kannan, A. (2023) A Comprehensive Survey on Machine Learning‐Based Intrusion Detection Systems for Secure Communication in Internet of Things. Computational Intelligence and Neuroscience, 2023, Article ID: 8981988. &gt;https://doi.org/10.1155/2023/8981988
    </mixed-citation>
   </ref>
   <ref id="scirp.137581-ref5">
    <label>5</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Siami-Namini, S., Tavakoli, N. and Namin, A.S. (2019). The Performance of LSTM and BiLSTM in Forecasting Time Series. 2019 IEEE International Conference on Big Data (Big Data), Los Angeles, 9-12 December 2019, 3285-3292. &gt;https://doi.org/10.1109/bigdata47090.2019.9005997
    </mixed-citation>
   </ref>
   <ref id="scirp.137581-ref6">
    <label>6</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Ioulianou, P., Vasilakis, V., Moscholios, I., et al. (2018) A Signature-Based Intrusion Detection System for the Internet of Things. Information and Communication Technology Form, Graz, 11-13 July 2018. &gt;https://eprints.whiterose.ac.uk/133312/
    </mixed-citation>
   </ref>
   <ref id="scirp.137581-ref7">
    <label>7</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Masdari, M. and Khezri, H. (2020) A Survey and Taxonomy of the Fuzzy Signature-Based Intrusion Detection Systems. Applied Soft Computing, 92, Article ID: 106301. &gt;https://doi.org/10.1016/j.asoc.2020.106301
    </mixed-citation>
   </ref>
   <ref id="scirp.137581-ref8">
    <label>8</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Díaz-Verdejo, J., Muñoz-Calle, J., Estepa Alonso, A., Estepa Alonso, R. and Madinabeitia, G. (2022) On the Detection Capabilities of Signature-Based Intrusion Detection Systems in the Context of Web Attacks. Applied Sciences, 12, Article No. 852. &gt;https://doi.org/10.3390/app12020852
    </mixed-citation>
   </ref>
   <ref id="scirp.137581-ref9">
    <label>9</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Moukhafi, M., Bri, S. and El Yassini, K. (2018) Intrusion Detection System Based on a Behavioral Approach. In: Talbi, E.-G. and Nakib, A., Eds., Studies in Computational Intelligence, Springer International Publishing, 61-75. &gt;https://doi.org/10.1007/978-3-319-95104-1_4
    </mixed-citation>
   </ref>
   <ref id="scirp.137581-ref10">
    <label>10</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Vinayakumar, R., Alazab, M., Soman, K.P., Poornachandran, P., Al-Nemrat, A. and Venkatraman, S. (2019) Deep Learning Approach for Intelligent Intrusion Detection System. IEEE Access, 7, 41525-41550. &gt;https://doi.org/10.1109/access.2019.2895334
    </mixed-citation>
   </ref>
   <ref id="scirp.137581-ref11">
    <label>11</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Azam, Z., Islam, M.M. and Huda, M.N. (2023) Comparative Analysis of Intrusion Detection Systems and Machine Learning-Based Model Analysis through Decision Tree. IEEE Access, 11, 80348-80391. &gt;https://doi.org/10.1109/access.2023.3296444
    </mixed-citation>
   </ref>
   <ref id="scirp.137581-ref12">
    <label>12</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Azam, Z., Islam, M.M. and Huda, M.N. (2023) Comparative Analysis of Intrusion Detection Systems and Machine Learning Based Model Analysis through Decision Tree. IEEE Access, 11, 80348-80391.
    </mixed-citation>
   </ref>
   <ref id="scirp.137581-ref13">
    <label>13</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Mo, X., Pang, J. and Liu, Z. (2024) Deep Autoencoder Architecture with Outliers for Temporal Attributed Network Embedding. Expert Systems with Applications, 240, Article ID: 122596. &gt;https://doi.org/10.1016/j.eswa.2023.122596
    </mixed-citation>
   </ref>
   <ref id="scirp.137581-ref14">
    <label>14</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Yang, Y., Tu, S., Hashim Ali, R., Alasmary, H., Waqas, M. and Nouman Amjad, M. (2023) Intrusion Detection Based on Bidirectional Long Short-Term Memory with Attention Mechanism. Computers, Materials &amp; Continua, 74, 801-815. &gt;https://doi.org/10.32604/cmc.2023.031907
    </mixed-citation>
   </ref>
   <ref id="scirp.137581-ref15">
    <label>15</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Sherstinsky, A. (2020) Fundamentals of Recurrent Neural Network (RNN) and Long Short-Term Memory (LSTM) Network. Physica D: Nonlinear Phenomena, 404, Article ID: 132306. &gt;https://doi.org/10.1016/j.physd.2019.132306
    </mixed-citation>
   </ref>
   <ref id="scirp.137581-ref16">
    <label>16</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Tokpa, F.W.R., Kamagaté, B.H., Monsan, V. and Oumtanaga, S. (2023) Fake News Detection in Social Media: Hybrid Deep Learning Approaches. Journal of Advances in Information Technology, 14, 606-615. &gt;https://doi.org/10.12720/jait.14.3.606-615
    </mixed-citation>
   </ref>
   <ref id="scirp.137581-ref17">
    <label>17</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Yu, Y., Si, X., Hu, C. and Zhang, J. (2019) A Review of Recurrent Neural Networks: LSTM Cells and Network Architectures. Neural Computation, 31, 1235-1270. &gt;https://doi.org/10.1162/neco_a_01199
    </mixed-citation>
   </ref>
   <ref id="scirp.137581-ref18">
    <label>18</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Smagulova, K. and James, A.P. (2019) A Survey on LSTM Memristive Neural Network Architectures and Applications. The European Physical Journal Special Topics, 228, 2313-2324. &gt;https://doi.org/10.1140/epjst/e2019-900046-x
    </mixed-citation>
   </ref>
   <ref id="scirp.137581-ref19">
    <label>19</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Akandeh, A. and Salem, F.M. (2019) Slim LSTM NETWORKS: LSTM_6 and Lstm_C6. 2019 IEEE 62nd International Midwest Symposium on Circuits and Systems (MWSCAS), Dallas, 4-7 August 2019, 630-633. &gt;https://doi.org/10.1109/mwscas.2019.8884912
    </mixed-citation>
   </ref>
   <ref id="scirp.137581-ref20">
    <label>20</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Gill, K.S., Anand, V., Chauhan, R., Choudhary, A. and Gupta, R. (2023) CNN, LSTM, and Bi-LSTM Based Self-Attention Model Classification for User Review Sentiment Analysis. 2023 3rd International Conference on Smart Generation Computing, Communication and Networking (SMART GENCON), Bangalore, 29-31 December 2023, 1-6. &gt;https://doi.org/10.1109/smartgencon60755.2023.10442498
    </mixed-citation>
   </ref>
   <ref id="scirp.137581-ref21">
    <label>21</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Graves, A., Fernández, S. and Schmidhuber, J. (2005) Bidirectional LSTM Networks for Improved Phoneme Classification and Recognition. 15th International Conference, ICANN 2005, Warsaw, 11-15 September 2005, 799-804. &gt;https://doi.org/10.1007/11550907_126
    </mixed-citation>
   </ref>
   <ref id="scirp.137581-ref22">
    <label>22</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Liu, M., Zhu, T., Ye, J., Meng, Q., Sun, L. and Du, B. (2023) Spatio-Temporal Autoencoder for Traffic Flow Prediction. IEEE Transactions on Intelligent Transportation Systems, 24, 5516-5526. &gt;https://doi.org/10.1109/tits.2023.3243913
    </mixed-citation>
   </ref>
   <ref id="scirp.137581-ref23">
    <label>23</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Sharafaldin, I., Habibi Lashkari, A. and Ghorbani, A.A. (2018) Toward Generating a New Intrusion Detection Dataset and Intrusion Traffic Characterization. Proceedings of the 4th International Conference on Information Systems Security and Privacy, Funchal, 22-24 January 2018, 108-116. &gt;https://doi.org/10.5220/0006639801080116
    </mixed-citation>
   </ref>
   <ref id="scirp.137581-ref24">
    <label>24</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Intrusion Detection Evaluation Dataset (CIC-IDS2017). &gt;https://www.unb.ca/cic/datasets/ids-2017.html
    </mixed-citation>
   </ref>
  </ref-list>
 </back>
</article>