<?xml version="1.0" encoding="UTF-8"?><!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v3.0 20080202//EN" "http://dtd.nlm.nih.gov/publishing/3.0/journalpublishing3.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" dtd-version="3.0" xml:lang="en" article-type="research article">
 <front>
  <journal-meta>
   <journal-id journal-id-type="publisher-id">
    jis
   </journal-id>
   <journal-title-group>
    <journal-title>
     Journal of Information Security
    </journal-title>
   </journal-title-group>
   <issn pub-type="epub">
    2153-1234
   </issn>
   <issn publication-format="print">
    2153-1242
   </issn>
   <publisher>
    <publisher-name>
     Scientific Research Publishing
    </publisher-name>
   </publisher>
  </journal-meta>
  <article-meta>
   <article-id pub-id-type="doi">
    10.4236/jis.2025.164023
   </article-id>
   <article-id pub-id-type="publisher-id">
    jis-145026
   </article-id>
   <article-categories>
    <subj-group subj-group-type="heading">
     <subject>
      Articles
     </subject>
    </subj-group>
    <subj-group subj-group-type="Discipline-v2">
     <subject>
      Computer Science 
     </subject>
     <subject>
       Communications
     </subject>
    </subj-group>
   </article-categories>
   <title-group>
    Attack Detection and Alarming System on IOT Facilities Using Random Forest Enabled-Correlation Based Clustering (RF-CBC) Technique 
   </title-group>
   <contrib-group>
    <contrib contrib-type="author" xlink:type="simple">
     <name name-style="western">
      <surname>
       Adedayo David
      </surname>
      <given-names>
       Adeniyi
      </given-names>
     </name> 
     <xref ref-type="aff" rid="aff1"> 
      <sup>1</sup>
     </xref>
    </contrib>
    <contrib contrib-type="author" xlink:type="simple">
     <name name-style="western">
      <surname>
       Rhoda
      </surname>
      <given-names>
       Ajayi
      </given-names>
     </name> 
     <xref ref-type="aff" rid="aff2"> 
      <sup>2</sup>
     </xref>
    </contrib>
    <contrib contrib-type="author" xlink:type="simple">
     <name name-style="western">
      <surname>
       Josephine Olamatanmi
      </surname>
      <given-names>
       Mebawondu
      </given-names>
     </name> 
     <xref ref-type="aff" rid="aff3"> 
      <sup>3</sup>
     </xref>
    </contrib>
   </contrib-group> 
   <aff id="aff1">
    <addr-line>
     aDepartment of Mathematical and Computer Sciences, Faculty of Science, University of Medical Sciences, Ondo, Nigeria
    </addr-line> 
   </aff> 
   <aff id="aff2">
    <addr-line>
     aDepartment of Computer Science, University of New Haven, West Haven, USA
    </addr-line> 
   </aff> 
   <aff id="aff3">
    <addr-line>
     aDepartment of Computing, Afe Babalola University, Ado-Ekiti, Nigeria
    </addr-line> 
   </aff> 
   <pub-date pub-type="epub">
    <day>
     22
    </day> 
    <month>
     08
    </month>
    <year>
     2025
    </year>
   </pub-date> 
   <volume>
    16
   </volume> 
   <issue>
    04
   </issue>
   <fpage>
    447
   </fpage>
   <lpage>
    471
   </lpage>
   <history>
    <date date-type="received">
     <day>
      28,
     </day>
     <month>
      December
     </month>
     <year>
      2024
     </year>
    </date>
    <date date-type="published">
     <day>
      19,
     </day>
     <month>
      December
     </month>
     <year>
      2024
     </year> 
    </date> 
    <date date-type="accepted">
     <day>
      19,
     </day>
     <month>
      August
     </month>
     <year>
      2025
     </year> 
    </date>
   </history>
   <permissions>
    <copyright-statement>
     © Copyright 2014 by authors and Scientific Research Publishing Inc. 
    </copyright-statement>
    <copyright-year>
     2014
    </copyright-year>
    <license>
     <license-p>
      This work is licensed under the Creative Commons Attribution International License (CC BY). http://creativecommons.org/licenses/by/4.0/
     </license-p>
    </license>
   </permissions>
   <abstract>
    In the past decade, Internet Of Things (IOT) technology has become one of the fastest-growing and most widely used technologies and is rapidly becoming a basic feature of global civilization. However, the high connectivity and diversity of these IOT devices make them complex and vulnerable to both visible and invisible security threats that are capable of causing irrecoverable damage. To alleviate these challenges, this work presents a novel and analytical hybrid machine learning model that suitably combines the Random Forest with Correlation-Based Clustering techniques, in order to report and detect potential attacks on IOT facilities. This work also showcases the development of the Single Threshold Boxplot Outlier-Based feature scaling method (STBO). The (STBO) method is used to scale down the number of attributes in order to select the best feature at the pre-processing stage of the attack detection procedures. The implementation of the present system is accomplished with the aid of an in-house Python program using XAMP/Apache HTTP as the hosting server with MySQL application for database development and management. A comparative analysis of the present model alongside ANN, Traditional Random Forest, Naïve Bayes, and the traditional Clustering method shows that the proposed system outperformed the baseline methods studied, with precision rates and attack detection quality equal to or greater than 75% in most cases, and is therefore capable of providing a useful, faster, efficient, and accurate anomalous detection online and on a real-time basis consistently with low false positive and negative rates.
   </abstract>
   <kwd-group> 
    <kwd>
     Attack Detection
    </kwd> 
    <kwd>
      IOT
    </kwd> 
    <kwd>
      Outlier
    </kwd> 
    <kwd>
      Correlation
    </kwd> 
    <kwd>
      Random Forest
    </kwd> 
    <kwd>
      Clustering
    </kwd>
   </kwd-group>
  </article-meta>
 </front>
 <body>
  <sec id="s1">
   <title>1. Introduction</title>
   <p>In recent times, there has been an upsurge of interest in the usage of Internet of Things (IOT) worldwide. IOT extends the use of internet facilities by connecting millions of devices with the capability of interacting with one another through smart technologies such as smart cities, smart homes, smart hospitals, smart banking, smart schools, smart agriculture, etc.</p>
   <p>Internet Of a Thing (IOT) is becoming one of the fastest growing and most widely used technologies in the past decade in both the private and business domains, with the rapid growth of internet and network connectivity <xref ref-type="bibr" rid="scirp.145026-1">
     [1]
    </xref> <xref ref-type="bibr" rid="scirp.145026-2">
     [2]
    </xref>. The production and marketing of smart devices are increasing rapidly; more hardware devices such as sensors, actuators, microcontrollers, etc., and Internet Of a Thing (IOT) software are being introduced daily by manufacturers, purposely to gain a competitive market advantage. Currently, it is estimated that the number of IOT and connected devices is over 20 billion; it is expected that the number will reach up to 50 billion by the year 2025 <xref ref-type="bibr" rid="scirp.145026-3">
     [3]
    </xref>. However, many of these devices and products failed to take into consideration the issue of security during their design, and are vulnerable to security threats (both visible and invisible). The high number of these IOT devices, coupled with their diversity and complexity, represents a huge security risk such as Denial Of Service (DOS), Data Breaches (DB), tampering, spoofing, privilege escalation, and IOT botnets, to mention just a few, and therefore is capable of causing irrecoverable damage <xref ref-type="bibr" rid="scirp.145026-4">
     [4]
    </xref>.</p>
   <p>Security is one of the cornerstones of any information society, such as the Internet Of Things (IOT). The subject of information security has been given much attention in recent years; it has become a very important research and professional topic in the field of the Internet Of Things (IOT). Therefore, the need to ensure the security of the IOT network is of great significance to the success of IOT technology. In recent times, several studies on the topic of attack detection on IOT facilities have been carried out, some of which show great capability in protecting the IOT network. However, a good number of these traditional web attack detection technologies face several challenges, thus the need to research a more viable and progressive IOT attack detection system.</p>
   <p>In this work, a novel attack detection and alarming system for IoT networks based on hybrid Two-Phase attack detection techniques is developed by integrating Random Forest with the novel Correlation-Based Clustering (RF-CBC) machine learning techniques. Specifically, the proposed RF-CBC is capable of accepting the Uniform Resource Locator (URL) of potential users on the network traffic, analyzing the URL request in order to detect and report anomalous requests within the network. The contributions of this work are fourfold:</p>
   <p>First, a novel hybrid, two-phase machine learning algorithm called Random Forest Enabled-Correlation Distant Based Clustering (RF-CBC) is proposed. Attention is specifically focused on combining the Random Forest technique with the Correlation Based Clustering technique. In the first phase, the Random Forest is used to classify users into different clusters based on their access credentials, while the second phase uses the correlation-based clustering techniques to analyze the browsing pattern of each visitor to the facility in order to detect any abnormality. The Correlation clustering technique uses correlation statistics to determine the similarity between a given tuple and the other tuples in the clusters; classification is done based on the closest correlation to the given tuple instead of the popular Euclidean distance. This enables the designer of the attack detection and reporting system to have more varieties of such algorithms in order to be able to select the best-performing algorithm. The proposed correlation-based measuring technique is computationally efficient and accurate for scalable implementation. It is capable of handling large datasets with no assumption about data distribution. It also has the capability to show dissimilarity between the test sample and training sample. The proposed RF-CBC-based attack detection and alarming system is capable of providing usable, consistent, efficient, faster, and accurate anomalous detection online and in real time with low false positive and negative rates.</p>
   <p>Second, this study examines many existing intrusion detection systems on IOT devices with the aim of investigating the performance of the various algorithms used in developing the attack detection system; this is in order to arrive at a more viable and efficient way of realising the present system.</p>
   <p>Third, a novel feature selection model referred to as Single Threshold Boxplot Outlier Based Feature Selection (STBO) feature scaling method is proposed. This method is used to scale down the number of attributes in order to select the best features at the pre-processing stage of the attack detection process before applying the proposed RF-C model. This is important in order to overcome scalability and computational complexity problems common to many existing attack detection algorithms while showing capability in handling high dimensionality and noisy data, therefore improving the accuracy of the proposed machine learning algorithm to a large extent.</p>
   <p>Fourth, this present work proposed the construction of a specific attack detection and reporting system for Internet Of a Thing (IOT), facilitating the use of an experimental website developed using the Python programming language with XAMP/Apache HTTP being adopted as the hosting server and MySQL database management software for data acquisition, model extraction, and data management at the back end.</p>
   <p>The proposed attack detection and reporting system will accept a potential user’s browsing URL request, tokenize the URL request, store it in a data mart, and then pass the token to the feature learning model to analyze the URL request and transform it into a vector together with attached anomalous information. The proposed RF-CBC model first determines the access credentials of potential users, then classifies their URL information to determine the presence of any forms of attack, the result of which is used for final decision making and the identified attack is reported. The classifier is updated using the update module.</p>
   <p>This will assist IoT designers and administrators in planning an update of their IoT facilities to determine potential threats and facilitate the protection of IoT facilities against visible and invisible threats, as well as to enlighten the public at large.</p>
   <p>The proposed classical, hybrid-double phases attack detection algorithm, the RF-CBC algorithm that serves as the basis for the development of the attack detection and alarming system on the IOT facility, will be presented alongside the proposed Single Threshold Boxplot Outlier Based Feature Scaling (STBO) dimensionality reduction technique. A comparative analysis of the present model was done alongside four other machine learning algorithms, which included the ANN, Naïve Bayesian, Traditional Random Forest, and the traditional clustering algorithm. This is to demonstrate the superior performance of the proposed RF-CBC model and to justify the rationale behind the selection of the proposed RF-CBC model. The result of the experiment shows excellent performance of the designed system over the baseline method studied, with a precision rate and attack detection quality equal to or greater than 75% in most cases. The proposed attack detection and alarming system is capable of providing usable, consistent, efficient, faster, and accurate anomalous detection online and in real-time with a low false positive and negative rate.</p>
   <p>Finally, the experimental results were thoroughly presented, and the proposed system will be implemented online and in real-time on the web server of the University of Medical Sciences, Ondo City, Nigeria.</p>
  </sec><sec id="s2">
   <title>2. Review of Related Work</title>
   <p>This section examines many existing related works on intrusion detection systems in IoT facilities and the methods adopted, with the aim of investigating the performance of such algorithms used. This is done in order to arrive at more reliable and efficient ways of realizing the present system.</p>
   <p>Several scholars in the field of attack detection on IoT facilities have carried out several studies on the topic, some of which show high potential in protecting the IoT facilities. However, despite the promising results, challenges still persist in securing the IoT facilities. Scalability problems due to the high connectivity of IoT networks, the novelty of attack techniques, and the interpretability of many of the existing attack detection models remain a concern <xref ref-type="bibr" rid="scirp.145026-5">
     [5]
    </xref>-<xref ref-type="bibr" rid="scirp.145026-8">
     [8]
    </xref>.</p>
   <p>Maheswari et al. <xref ref-type="bibr" rid="scirp.145026-9">
     [9]
    </xref>, in their work, carried out research on a web attack detection system for IoT using ensemble classification. The results of their experiment show significant improvement in terms of accuracy when compared with the baseline methods used. However, their system was only able to detect some selected common types of attack, i.e., their system can only detect SQL injection and cross-site scripting, hence neglecting other types of attack. The present system is designed to take care of various types of attacks as they relate to IoT facilities.</p>
   <p>Yavuz, Unal, and Gul <xref ref-type="bibr" rid="scirp.145026-10">
     [10]
    </xref> proposed a deep learning-based machine learning approach for the detection of routine attacks. The results of their experiment show high accuracy and precision on their data set. However, their systems are limited by the number of attacks that can be detected. A system that could be used to detect multiple attack types is needed, hence the need for this present work.</p>
   <p>Learning Vector Quantization (LVQ) and K-NN were used by Naorum and Al-Sultani <xref ref-type="bibr" rid="scirp.145026-11">
     [11]
    </xref> for intrusion detection; their results record a good detection rate. The challenge with their approach is time complexity since the LVQ requires a long time to be trained, which requires the size of the classes to be equally likely, which is not the case with the present system. The present system is faster and more scalable in its operation.</p>
   <p>In the work of Jawhar and Mehrotra <xref ref-type="bibr" rid="scirp.145026-12">
     [12]
    </xref>, a hybrid intrusion detection system was proposed using fuzzy logic, neural network, and clustering algorithm with multi-layered perception to detect normal users and four attack types. The results of their experiment show a high detection rate of about 99.9%. However, their system is marred by the challenges of the distribution of the records in the training set not being close to equal between classes and the inability to detect multiple attack types.</p>
   <p>Kouassi, Monsan, and Adou <xref ref-type="bibr" rid="scirp.145026-13">
     [13]
    </xref>, in their work, explore the effectiveness of long-term memory neural networks (LSTMs) and Deep Neural Network (DNN) models for detecting attacks in IoT networks. The results of their experiment show that their models prove to be more effective for detecting attacks in IoT networks, particularly for sophisticated attacks.</p>
   <p>Sasia, Lashkaria, Lua, Xiong, and Iqbal <xref ref-type="bibr" rid="scirp.145026-14">
     [14]
    </xref> carried out a survey on various IoT attacks by categorizing attacks in the taxonomy according to various factors such as attack domains, attack threat type, attack executions, etc. These were accompanied by their respective remedies. Their study revealed several open research areas pertinent to the subject of IoT attacks.</p>
   <p>Siraparapu and Azad <xref ref-type="bibr" rid="scirp.145026-15">
     [15]
    </xref>, in their work, carried out a comprehensive review of systems for securing IoT devices in the digital era. It explores the role of IoT secure systems in Industry 4.0, optimizing manufacturing processes and supply chain management. It emphasizes the significance of IoT secure systems, discussing challenges, limitations, and benefits to organizations. Their analysis reveals emerging trends in IoT security standards and identifies critical gaps in current regulations, offering a forward-looking perspective on ensuring integrity and privacy across diverse domains.</p>
   <p>The present system is capable of overcoming the identified challenges of most of the existing systems studied, with the capability to handle multiple attack types due to its ability to monitor various clients to the website using their login credentials.</p>
   <p>Evidence from the available literature shows that the major challenges faced by many existing machine learning algorithms for prediction or classification are scalability problems, inability to handle noisy data, computational complexity, and low-dimensional attributes; these usually result in classification and predictive errors. These occur when the number of features and instances is too large <xref ref-type="bibr" rid="scirp.145026-16">
     [16]
    </xref>-<xref ref-type="bibr" rid="scirp.145026-19">
     [19]
    </xref>. As a result of this, many scholars in the field of machine learning have come up with a number of techniques for scaling down the number of attributes and data reduction <xref ref-type="bibr" rid="scirp.145026-20">
     [20]
    </xref>.</p>
   <p>Hegde et al. <xref ref-type="bibr" rid="scirp.145026-21">
     [21]
    </xref>, in their work, perform feature selection using a symmetrized feature selection algorithm in combination with stacked generalization-based metaheuristic techniques on chronic disease datasets from the Kaggle repository. The result of their experiment shows better accuracy than the baseline methods used. However, this method is applicable to small subsets of problems and configurations. The present system is capable of handling large datasets.</p>
   <p>Domingo and Hulten <xref ref-type="bibr" rid="scirp.145026-22">
     [22]
    </xref> used Hoeffding approaches to scale up machine learning algorithms. The method, which can be applied to choose among a set of discrete models or to estimate a continuous parameter, is capable of minimizing the time bound through the number of samples used, subject to the target limits on the loss of performance when using a subset of the data set. Their method was reported to produce very interesting results; however, the need to derive these bounds makes it difficult for many algorithms to use. The present system uses a single threshold bound, which makes it simple and easy to apply in many algorithms.</p>
   <p>Flores et al. <xref ref-type="bibr" rid="scirp.145026-23">
     [23]
    </xref> trained the ARMA model using a statistical analysis technique to eliminate irrelevant features; their results show a good level of accuracy but with the risk of eliminating useful information. The application of the single threshold plot box method is capable of overcoming these challenges.</p>
   <p>Sebban and Mock <xref ref-type="bibr" rid="scirp.145026-24">
     [24]
    </xref> carry out feature reduction using both filter and wrapper methods on a small subset of a dataset; this method performs well on their dataset. This method, however, is marred by the challenges of being computationally expensive with a high run time if used on a large dataset. These challenges are overcome by the present method of feature selection.</p>
  </sec><sec id="s3">
   <title>3. Methodology</title>
   <p>This section presents a description of a series of processes and methods utilised to achieve the objectives of the proposed system.</p>
   <sec id="s3_1">
    <title>3.1. Experimental Design</title>
    <p>Our proposed attack detection and alarming system for IoT facilities adopts a four-module architecture, viz:</p>
    <p>(a) The feature scaling and selection module: This module adopts the novel Single Threshold Boxplot Outlier Based Feature scaling method (STBO) to reduce the number of attributes in order to select the best and most useful feature for the proposed RF-CBC model at the pre-processing stage of the attack detection and reporting processes.</p>
    <p>(b) The profiling and classification module: This module accepts useful URL information and builds a user profile online and in real time in order to determine their different access credentials using random forest techniques. These are stored in a data mart/database created, leading to the creation of different clusters of users based on the profile of their access credentials.</p>
    <p>(c) The decision module: This module is designed to use the correlation statistic to analyze the user’s browsing pattern alongside their access credentials and the attack information in order to arrive at a decision for detecting the anomaly. This is achieved by computing the correlation between the current user’s URL information and the historical browsing information in the clusters stored in the data mart.</p>
    <p>(d) The update module: The update module is designed to fine-tune the update of the anomaly detection system; here, all raw URL requests and normalized data together with the intrusion detection result are stored in the database for further analysis in a feedback mechanism. This is to improve the robustness and reliability of the RF-CBC model as well as to facilitate further analysis by security experts.</p>
   </sec>
   <sec id="s3_2">
    <title>3.2. Data Collection</title>
    <p>In this study, a click log history of 15,374 anonymous users of the University of Medical Sciences (UNIMED) official website, who signed into their accounts over a period of 12 months from 12 May, 2023 to 3rd April, 2024, was selected randomly. The selected click log is made up of about 30,572 anonymous visitors to the UNIMED website server’s URL address database. <xref ref-type="fig" rid="fig1">
      Figure 1
     </xref> shows a sample of extracted users’ browsing history’s HTTP from the UNIMED official website located at <xref ref-type="bibr" rid="scirp.145026-http://unimed.edu.ng">
      http://unimed.edu.ng
     </xref>.</p>
    <fig id="fig1" position="float">
     <label>Figure 1</label>
     <caption>
      <title>Figure 1. Sample historical HTTP request to the UNIMED official website.</title>
     </caption>
     <graphic mimetype="image" position="float" xlink:type="simple" xlink:href="https://html.scirp.org/file/7801084-rId16.jpeg?20250822024410" />
    </fig>
    <p>We carried out a number of pre-processing operations on the raw user’s URL address database extracted by cleansing it, in order to eliminate irrelevant or noisy entries. After this, we developed a data mart of the log data and partitioned it into sessions <xref ref-type="bibr" rid="scirp.145026-25">
      [25]
     </xref>. After the session identification stage, we then applied the proposed Single Threshold Boxplot outlier-based feature scaling and feature selection algorithm to select the best feature for the proposed RF-CBC model. The proposed Random Forest algorithm was then used to group users into clusters based on similarities in their access.</p>
    <p>Credentials, browsing HTTP logs, and attack information, while the correlation-based Clusters algorithm was used for the final decision as to whether a user is an attacker or a normal user based on their access credentials and similarities in their search behaviour to that of the identified clusters and the attack information, the result of which is used to update the data mart for further analysis. The overall architecture of the entire RF-CBC attack detection and reporting system is shown in <xref ref-type="fig" rid="fig2">
      Figure 2
     </xref>.</p>
    <fig id="fig2" position="float">
     <label>Figure 2</label>
     <caption>
      <title>Figure 2. The overall architecture of the entire RF-CBC attack detection and reporting system.</title>
     </caption>
     <graphic mimetype="image" position="float" xlink:type="simple" xlink:href="https://html.scirp.org/file/7801084-rId17.jpeg?20250822024410" />
    </fig>
   </sec>
   <sec id="s3_3">
    <title>3.3. Data Pre-Processing Feature Selection</title>
    <p>Some of the problems with many machine learning algorithms are scalability and computational complexity. Noisy data and low-dimensional attributes, caused by too many features and instances, usually result in prediction or classification errors. To overcome these challenges, this work proposes a novel feature selection technique called the Single Threshold Boxplot Outlier Based Feature Scaling method (STBO) to scale down the number of attributes and select the best features for the proposed attack detection processes . In recent times, researchers in machine learning have proposed a number of feature selection techniques, such as Gini-index, χ<sup>2</sup> statistics, information gain, Wrapper method, and correlation technique, etc., . Some of these performed well on their data sets. Some of these techniques have been studied, and the pros and cons of each method have been thoroughly understood before proposing the current STBO technique.</p>
    <p>In this work, the click log history of anonymous visitors to the UNIMED official website was selected as described in Section 3.2. The terms that occur in the documents are presented as parameters, features, attributes, tuples, or variables. After session identification, the term count was taken. <xref ref-type="table" rid="table1">
      Table 1
     </xref> shows the term document matrix over all documents with the total term count. <xref ref-type="fig" rid="fig3">
      Figure 3
     </xref> shows the statistics of the occurrence of each attribute in the different documents extracted.</p>
    <fig id="fig3" position="float">
     <label>Figure 3</label>
     <caption>
      <title>Figure 3. The statistics of the occurrence of each attributes in the different documents extracted.</title>
     </caption>
     <graphic mimetype="image" position="float" xlink:type="simple" xlink:href="https://html.scirp.org/file/7801084-rId18.jpeg?20250822024412" />
    </fig>
    <table-wrap id="table1">
     <label>
      <xref ref-type="table" rid="table1">
       Table 1
      </xref></label>
     <caption>
      <title>
       <xref ref-type="bibr" rid="scirp.145026-"></xref>Table 1. Attributes documents matrix for 12 months with Total Term Count (TTC) over all documents using UNIMED users’ logs database.</title>
     </caption>
    </table-wrap>
    <fig id="fig4" position="float">
     <label>Figure 4</label>
     <caption>
      <title>In this work, we propose the Single Threshold Boxplot Outlier Based Feature scaling method of eliminating features that diverge significantly from the general pattern of the users’ click logs using the Boxplot, also called box-and-whisker plots. The Boxplot is a statistical method that uses five (5) summary statistics i.e. Minimum, Maximum, First Quartile (Q<sub>1</sub>), Second Quartile (Q<sub>2</sub>), and Third Quartile (Q<sub>3</sub>). Each quartile represents 25% of the data points <xref ref-type="bibr" rid="scirp.145026-27">
        [27]
       </xref>.We first arranged the data elements, i.e., The features’ total term count for each feature/tuple for all the documents was extracted in ascending order. We then calculate the Q<sub>1</sub>, Q<sub>2</sub> and Q<sub>3</sub>; we later calculate the interquartile range (IQR), i.e., IQR = (Q<sub>3</sub> − Q<sub>1</sub>). We also calculate the lower and upper bounds for the outlier, which will be included in the non-outlier zone.Lower Bound (LB) = Q<sub>1</sub> − 1.5 * IQRUpper Bound (UB) = Q<sub>3</sub> + 1.5 * IQR. <xref ref-type="bibr" rid="scirp.145026-24">
        [24]
       </xref>However, since our interest is to eliminate only irrelevant features, we eliminate only the data points that fall below the lower bound and consider them as outliers; hence, we consider only a single threshold, rather than the usual two thresholds.The attributes within the lower bound are to be used by the proposed RF-CBC algorithm.Algorithm listing for the Single Threshold Boxplot Outlier Based Feature Scaling method (STBO) is shown in <xref ref-type="bibr" rid="scirp.145026-#al1">
        Algorithm Listing 1
       </xref>.3.3.2. Algorithm ListingAlgorithm Listing 1. The single threshold boxplot outlier based feature scaling algorithm (STBO).<xref ref-type="bibr" rid="scirp.145026-"></xref><p class="imgGroupCss_v"><img class=" imgMarkCss lazy" data-original="https://html.scirp.org/file/7801084-rId20.jpeg?20250822024413" /></p></title>
     </caption>
     <graphic mimetype="image" position="float" xlink:type="simple" xlink:href="https://html.scirp.org/file/7801084-rId19.jpeg?20250822024412" />
    </fig>
    <p>This section presents the demonstration of the present STBO feature selection techniques using our collected data sets from the UNIMED official website, as shown in <xref ref-type="table" rid="table1">
      Table 1
     </xref>. We applied the STBO feature scaling algorithm to eliminate irrelevant attributes and retain only relevant attributes that fall above the lower bound for our attack detection system. We compute the total term count for all the selected attributes and then sort them in descending order of their TTC, as shown in <xref ref-type="table" rid="table2">
      Table 2
     </xref>. We compute Q<sub>1</sub>, Q<sub>2</sub>, Q<sub>3</sub>, the IQR, the lower and the upper bound. We then minimize the number of attributes by selecting attributes that fall above the lower bound from the data set extracted. This is done in order to increase the accuracy and efficiency of our RF-CBC machine learning algorithm. Given our data set according to <xref ref-type="table" rid="table1">
      Table 1
     </xref> and the sorted attributes according to their TTC, as shown in <xref ref-type="table" rid="table2">
      Table 2
     </xref>.</p>
    <table-wrap id="table2">
     <label>
      <xref ref-type="table" rid="table2">
       Table 2
      </xref></label>
     <caption>
      <title>
       <xref ref-type="bibr" rid="scirp.145026-"></xref>Table 2. The sorted attributes according to their TTC.</title>
     </caption>
    </table-wrap>
    <fig id="fig5" position="float">
     <label>Figure 5</label>
     <caption>
      <title>The Median count Q<sub>2</sub> = 1/2(N + 1)<sup>th</sup>=1/2(14 + 1)<sup>th</sup> = (15/2)<sup>th</sup> = 7.5<sup>th</sup>, we use 7<sup>th</sup> and the 8<sup>th</sup> element, i.e.=(515 + 540)/2 = 1055/2=527.5, therefore Q<sub>2</sub> = 527.5Finge count = 1/2(1 + median count), the result of which must be an integer = 1/2(1 + 7.5)<sup>th</sup> = 1/2(8.5)<sup>th</sup> = 4.25, fringe must be an integer therefore = 4.Q<sub>1</sub> = count from the beginning of the sorted attributes sum list, the number derived from the finge count, i.e., 4.i.e. Q<sub>1</sub> = 450Q<sub>3</sub> = count from the end of the sorted attributes sum list, the number derived from the finge count, i.e., 4.i.e. Q<sub>3</sub> = 561
       <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
  
         <mtext>
          
   IQR
  
         </mtext>
  
         <mo>
          
   =
  
         </mo>
  
         <msub> 
   
          <mi>
           
    Q
   
          </mi> 
   
          <mn>
           
    3
   
          </mn> 
  
         </msub> 
  
         <mo>
          
   −
  
         </mo>
  
         <msub> 
   
          <mi>
           
    Q
   
          </mi> 
   
          <mn>
           
    1
   
          </mn> 
  
         </msub> 
  
         <mo>
          
   =
  
         </mo>
  
         <mn>
          
   561
  
         </mn>
  
         <mo>
          
   −
  
         </mo>
  
         <mn>
          
   450
  
         </mn>
  
         <mo>
          
   =
  
         </mo>
  
         <mn>
          
   111
  
         </mn>
 
        </mrow>

       </math></title>
     </caption>
     <graphic mimetype="image" position="float" xlink:type="simple" xlink:href="https://html.scirp.org/file/7801084-rId21.jpeg?20250822024414" />
    </fig>
    <p>
     <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"> <mtable> 
       <mtr> 
        <mtd> 
         <mtext>
           Compute the Lower Bound 
         </mtext> 
         <mrow> 
          <mo>
            ( 
          </mo> 
          <mrow> 
           <mtext>
             LB 
           </mtext> 
          </mrow> 
          <mo>
            ) 
          </mo> 
         </mrow> 
         <mo>
           = 
         </mo> 
         <msub> 
          <mi>
            Q 
          </mi> 
          <mn>
            1 
          </mn> 
         </msub> 
         <mo>
           − 
         </mo> 
         <mn>
           1.5 
         </mn> 
         <mo>
           ∗ 
         </mo> 
         <mtext>
           IQR 
         </mtext> 
        </mtd> 
       </mtr> 
       <mtr> 
        <mtd> 
         <mo>
           = 
         </mo> 
         <mn>
           450 
         </mn> 
         <mo>
           − 
         </mo> 
         <mn>
           1.5 
         </mn> 
         <mo>
           ∗ 
         </mo> 
         <mn>
           111 
         </mn> 
        </mtd> 
       </mtr> 
       <mtr> 
        <mtd> 
         <mo>
           = 
         </mo> 
         <mn>
           450 
         </mn> 
         <mo>
           − 
         </mo> 
         <mn>
           166.5 
         </mn> 
        </mtd> 
       </mtr> 
       <mtr> 
        <mtd> 
         <mo>
           = 
         </mo> 
         <mn>
           283.5 
         </mn> 
        </mtd> 
       </mtr> 
      </mtable> 
     </math></p>
    <p>Any attributes TCT that fall below the lower bound of 283 are treated as outliers and eliminated.</p>
    <p>This section presents the evaluation of our proposed STBO feature selection technique in two different ways. First, we carried out feature selection on the extracted historical browsing pattern data from the University of Medical Sciences (UNIMED) website, using two baseline techniques which include Information Gain (IG), CFS, and the proposed STBO. <xref ref-type="table" rid="table3">
      Table 3
     </xref> shows the number of features selected by each technique. The experimental results show that the IG and the present STBO selected fewer features, with STBO selecting fewer features in comparison with the baseline methods studied.</p>
    <table-wrap id="table3">
     <label>
      <xref ref-type="table" rid="table3">
       Table 3
      </xref></label>
     <caption>
      <title>
       <xref ref-type="bibr" rid="scirp.145026-"></xref>Table 3. The number of features selected by CFS, IG and STBO feature selection techniques.</title>
     </caption>
     <table class="MsoTableGrid custom-table" border="0" cellspacing="0" cellpadding="0"> 
      <tr> 
       <td rowspan="2" class="acenter" width="31.43%"><p style="text-align:center">Total number of available attributes</p></td> 
       <td class="custom-bottom-td acenter" width="94.31%" colspan="3"><p style="text-align:center">Number of Attributes selected by each technique</p></td> 
      </tr> 
      <tr> 
       <td class="custom-bottom-td custom-top-td acenter" width="31.43%"><p style="text-align:center">CFS</p></td> 
       <td class="custom-bottom-td custom-top-td acenter" width="31.44%"><p style="text-align:center">IG</p></td> 
       <td class="custom-bottom-td custom-top-td acenter" width="31.44%"><p style="text-align:center">STBO</p></td> 
      </tr> 
      <tr> 
       <td class="custom-top-td acenter" width="31.43%"><p style="text-align:center">20</p></td> 
       <td class="custom-top-td acenter" width="31.43%"><p style="text-align:center">8</p></td> 
       <td class="custom-top-td acenter" width="31.44%"><p style="text-align:center">8</p></td> 
       <td class="custom-top-td acenter" width="31.44%"><p style="text-align:center">5</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="31.43%"><p style="text-align:center">25</p></td> 
       <td class="acenter" width="31.43%"><p style="text-align:center">14</p></td> 
       <td class="acenter" width="31.44%"><p style="text-align:center">12</p></td> 
       <td class="acenter" width="31.44%"><p style="text-align:center">7</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="31.43%"><p style="text-align:center">30</p></td> 
       <td class="acenter" width="31.43%"><p style="text-align:center">18</p></td> 
       <td class="acenter" width="31.44%"><p style="text-align:center">16</p></td> 
       <td class="acenter" width="31.44%"><p style="text-align:center">8</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="31.43%"><p style="text-align:center">60</p></td> 
       <td class="acenter" width="31.43%"><p style="text-align:center">25</p></td> 
       <td class="acenter" width="31.44%"><p style="text-align:center">19</p></td> 
       <td class="acenter" width="31.44%"><p style="text-align:center">10</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="31.43%"><p style="text-align:center">80</p></td> 
       <td class="acenter" width="31.43%"><p style="text-align:center">28</p></td> 
       <td class="acenter" width="31.44%"><p style="text-align:center">21</p></td> 
       <td class="acenter" width="31.44%"><p style="text-align:center">12</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="31.43%"><p style="text-align:center">90</p></td> 
       <td class="acenter" width="31.43%"><p style="text-align:center">32</p></td> 
       <td class="acenter" width="31.44%"><p style="text-align:center">23</p></td> 
       <td class="acenter" width="31.44%"><p style="text-align:center">15</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="31.43%"><p style="text-align:center">112</p></td> 
       <td class="acenter" width="31.43%"><p style="text-align:center">43</p></td> 
       <td class="acenter" width="31.44%"><p style="text-align:center">28</p></td> 
       <td class="acenter" width="31.44%"><p style="text-align:center">18</p></td> 
      </tr> 
     </table>
    </table-wrap>
    <p>To further evaluate the accuracy and effectiveness of the present method, we tested the degree of accuracy of our RF-CBC algorithm with the STBO, IG, and CSF feature selection algorithms. The accuracy of the RF-CBC was determined using the expression:</p>
    <p>
     <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mtext>
         Ac 
       </mtext> 
       <mo>
         = 
       </mo> 
       <mfrac> 
        <mrow> 
         <mtext>
           TN 
         </mtext> 
         <mo>
           + 
         </mo> 
         <mtext>
           TP 
         </mtext> 
        </mrow> 
        <mrow> 
         <mtext>
           TP 
         </mtext> 
         <mo>
           + 
         </mo> 
         <mtext>
           TN 
         </mtext> 
         <mo>
           + 
         </mo> 
         <mtext>
           FP 
         </mtext> 
         <mo>
           + 
         </mo> 
         <mtext>
           FN 
         </mtext> 
        </mrow> 
       </mfrac> 
      </mrow> 
     </math></p>
    <p>where:</p>
    <p>AC = Accuracy: This measures the proportion of correctly detected attacks and normal instances.</p>
    <p>TP = True Positive: This is the number of correctly detected attack instances,</p>
    <p>TN = True Negative: This is the number of correctly detected normal instances,</p>
    <p>FP = False Positive: This is the number of incorrectly detected attack instances, and</p>
    <p>FN = False Negative: This is the number of incorrectly detected normal instances.</p>
    <p>The result of the experiment, as shown in <xref ref-type="fig" rid="fig4">
      Figure 4
     </xref>, indicates that the proposed RF-CBC performed well, with the STBO technique significantly improving the accuracy of the RF-CBC algorithms in comparison with the IG and the CFS techniques at different levels of predictions and under the same experimental settings.</p>
    <fig id="fig6" position="float">
     <label>Figure 6</label>
     <caption>
      <title>Figure 4. The RF-CBC accuracy with the STBO, CFS and the IG feature selection methods at different level of prediction.</title>
     </caption>
     <graphic mimetype="image" position="float" xlink:type="simple" xlink:href="https://html.scirp.org/file/7801084-rId28.jpeg?20250822024416" />
    </fig>
   </sec>
   <sec id="s3_4">
    <title>3.4. The Proposed Random Forest Enabled-Correlation Based Clustering (RF-CBC) Technique</title>
    <p>Random forest is an ensemble decision tree in which each tree depends on a collection of random variables. It is a supervised machine learning algorithm widely used in classification and regression problems. Random forest builds decision trees using different samples, then takes the majority votes of the trees for classification or averaging in the case of regression. Leo Breiman was believed to have first come up with the idea of random forest <xref ref-type="bibr" rid="scirp.145026-28">
      [28]
     </xref>.</p>
    <p>Given an n-dimensional random vector 
     <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mi>
         X 
       </mi> 
       <mo>
         = 
       </mo> 
       <msup> 
        <mrow> 
         <mrow> 
          <mo>
            ( 
          </mo> 
          <mrow> 
           <msub> 
            <mi>
              X 
            </mi> 
            <mn>
              1 
            </mn> 
           </msub> 
           <mo>
             , 
           </mo> 
           <msub> 
            <mi>
              X 
            </mi> 
            <mn>
              2 
            </mn> 
           </msub> 
           <mo>
             , 
           </mo> 
           <msub> 
            <mi>
              X 
            </mi> 
            <mn>
              3 
            </mn> 
           </msub> 
           <mo>
             , 
           </mo> 
           <mo>
             ⋯ 
           </mo> 
           <mo>
             , 
           </mo> 
           <msub> 
            <mi>
              X 
            </mi> 
            <mi>
              n 
            </mi> 
           </msub> 
          </mrow> 
          <mo>
            ) 
          </mo> 
         </mrow> 
        </mrow> 
        <mtext>
          T 
        </mtext> 
       </msup> 
      </mrow> 
     </math> where 
     <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          X 
        </mi> 
        <mn>
          1 
        </mn> 
       </msub> 
       <mo>
         , 
       </mo> 
       <msub> 
        <mi>
          X 
        </mi> 
        <mn>
          2 
        </mn> 
       </msub> 
       <mo>
         , 
       </mo> 
       <msub> 
        <mi>
          X 
        </mi> 
        <mn>
          3 
        </mn> 
       </msub> 
       <mo>
         , 
       </mo> 
       <mo>
         ⋯ 
       </mo> 
       <mo>
         , 
       </mo> 
       <msub> 
        <mi>
          X 
        </mi> 
        <mi>
          n 
        </mi> 
       </msub> 
      </mrow> 
     </math> represent the real valued input with Y representing the predictive class (voting) <xref ref-type="bibr" rid="scirp.145026-29">
      [29]
     </xref>.</p>
    <p>Given an unknown Joint distribution 
     <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          P 
        </mi> 
        <mrow> 
         <mi>
           x 
         </mi> 
         <mi>
           y 
         </mi> 
        </mrow> 
       </msub> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <mi>
           X 
         </mi> 
         <mo>
           , 
         </mo> 
         <mi>
           Y 
         </mi> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
      </mrow> 
     </math></p>
    <p>
     <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mi>
         F 
       </mi> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mi>
          x 
        </mi> 
        <mo>
          ) 
        </mo> 
       </mrow> 
      </mrow> 
     </math> is a function to predict Y, which is determined by a loss function 
     <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mi>
         L 
       </mi> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <mi>
           Y 
         </mi> 
         <mo>
           , 
         </mo> 
         <mi>
           F 
         </mi> 
         <mrow> 
          <mo>
            ( 
          </mo> 
          <mi>
            X 
          </mi> 
          <mo>
            ) 
          </mo> 
         </mrow> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
      </mrow> 
     </math>, the loss function is defined to minimize the expected value of loss, which is denoted by Equation (1). <xref ref-type="bibr" rid="scirp.145026-29">
      [29]
     </xref>.</p>
    <p>
     <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          E 
        </mi> 
        <mrow> 
         <mi>
           x 
         </mi> 
         <mi>
           y 
         </mi> 
        </mrow> 
       </msub> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <mi>
           L 
         </mi> 
         <mrow> 
          <mo>
            ( 
          </mo> 
          <mrow> 
           <mi>
             Y 
           </mi> 
           <mo>
             , 
           </mo> 
           <mi>
             f 
           </mi> 
           <mrow> 
            <mo>
              ( 
            </mo> 
            <mi>
              x 
            </mi> 
            <mo>
              ) 
            </mo> 
           </mrow> 
          </mrow> 
          <mo>
            ) 
          </mo> 
         </mrow> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
      </mrow> 
     </math> (1)</p>
    <p>The subscript here denotes expectation with respect to joint distribution of X and Y</p>
    <p>
     <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mi>
         L 
       </mi> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <mi>
           Y 
         </mi> 
         <mo>
           , 
         </mo> 
         <mi>
           f 
         </mi> 
         <mrow> 
          <mo>
            ( 
          </mo> 
          <mi>
            x 
          </mi> 
          <mo>
            ) 
          </mo> 
         </mrow> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
      </mrow> 
     </math> measures how close 
     <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mi>
         f 
       </mi> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mi>
          x 
        </mi> 
        <mo>
          ) 
        </mo> 
       </mrow> 
      </mrow> 
     </math> is to Y, we choose zero-one-loss for classification, we have</p>
    <p>
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mi>
         L 
       </mi> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <mi>
           Y 
         </mi> 
         <mo>
           , 
         </mo> 
         <mi>
           f 
         </mi> 
         <mrow> 
          <mo>
            ( 
          </mo> 
          <mi>
            x 
          </mi> 
          <mo>
            ) 
          </mo> 
         </mrow> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
       <mo>
         = 
       </mo> 
       <mi>
         L 
       </mi> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <mi>
           Y 
         </mi> 
         <mo>
           ≠ 
         </mo> 
         <mi>
           f 
         </mi> 
         <mrow> 
          <mo>
            ( 
          </mo> 
          <mi>
            x 
          </mi> 
          <mo>
            ) 
          </mo> 
         </mrow> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
       <mo>
         = 
       </mo> 
       <mrow> 
        <mo>
          { 
        </mo> 
        <mtable columnalign="left"> 
         <mtr> 
          <mtd> 
           <mn>
             0 
           </mn> 
           <mo>
             , 
           </mo> 
           <mtext>
               
           </mtext> 
           <mtext>
               
           </mtext> 
           <mtext>
             if 
           </mtext> 
           <mtext>
               
           </mtext> 
           <mi>
             y 
           </mi> 
           <mo>
             = 
           </mo> 
           <mi>
             f 
           </mi> 
           <mrow> 
            <mo>
              ( 
            </mo> 
            <mi>
              x 
            </mi> 
            <mo>
              ) 
            </mo> 
           </mrow> 
          </mtd> 
         </mtr> 
         <mtr> 
          <mtd> 
           <mn>
             1 
           </mn> 
           <mo>
             , 
           </mo> 
           <mtext>
               
           </mtext> 
           <mtext>
               
           </mtext> 
           <mtext>
               
           </mtext> 
           <mtext>
               
           </mtext> 
           <mtext>
             otherwise 
           </mtext> 
          </mtd> 
         </mtr> 
        </mtable> 
       </mrow> 
      </mrow> 
     </math> (2)</p>
    <p>
     <xref ref-type="bibr" rid="scirp.145026-"></xref>If the set of possibilities value of Y is denoted by 
     <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"> <mi mathvariant="script">
        Y 
      </mi> 
     </math>, to maximise 
     <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          E 
        </mi> 
        <mrow> 
         <mi>
           x 
         </mi> 
         <mi>
           y 
         </mi> 
        </mrow> 
       </msub> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <mi>
           L 
         </mi> 
         <mrow> 
          <mo>
            ( 
          </mo> 
          <mrow> 
           <mi>
             Y 
           </mi> 
           <mo>
             , 
           </mo> 
           <mi>
             f 
           </mi> 
           <mrow> 
            <mo>
              ( 
            </mo> 
            <mi>
              x 
            </mi> 
            <mo>
              ) 
            </mo> 
           </mrow> 
          </mrow> 
          <mo>
            ) 
          </mo> 
         </mrow> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
      </mrow> 
     </math> for zero-one-loss we have</p>
    <p>
     <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mi>
         f 
       </mi> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mi>
          x 
        </mi> 
        <mo>
          ) 
        </mo> 
       </mrow> 
       <mo>
         = 
       </mo> 
       <mo> 
       </mo> 
       <mfrac> 
        <mrow> 
         <mi>
           arg 
         </mi> 
         <mi>
           max 
         </mi> 
         <mi>
           n 
         </mi> 
         <mrow> 
          <mo>
            ( 
          </mo> 
          <mrow> 
           <mrow> 
            <mrow> 
             <mi>
               Y 
             </mi> 
             <mo>
               = 
             </mo> 
             <mi>
               y 
             </mi> 
            </mrow> 
            <mo>
              | 
            </mo> 
           </mrow> 
           <mi>
             X 
           </mi> 
           <mo>
             = 
           </mo> 
           <mi>
             x 
           </mi> 
          </mrow> 
          <mo>
            ) 
          </mo> 
         </mrow> 
        </mrow> 
        <mrow> 
         <mi>
           Y 
         </mi> 
         <mo>
           ∈ 
         </mo> 
         <mi mathvariant="script">
           Y 
         </mi> 
        </mrow> 
       </mfrac> 
      </mrow> 
     </math> (3)</p>
    <p>This is also known as Bayes’ rule.</p>
    <p>In classification 
     <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mi>
         f 
       </mi> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mi>
          x 
        </mi> 
        <mo>
          ) 
        </mo> 
       </mrow> 
      </mrow> 
     </math> represent the most frequently predicted class popularly known as (“votting”) where 
     <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          h 
        </mi> 
        <mi>
          i 
        </mi> 
       </msub> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mi>
          x 
        </mi> 
        <mo>
          ) 
        </mo> 
       </mrow> 
       <mo>
         , 
       </mo> 
       <mo>
         ⋯ 
       </mo> 
       <mo>
         , 
       </mo> 
       <msub> 
        <mi>
          h 
        </mi> 
        <mi>
          j 
        </mi> 
       </msub> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mi>
          x 
        </mi> 
        <mo>
          ) 
        </mo> 
       </mrow> 
      </mrow> 
     </math> is the “base learning” combination that gives the “ensemble” prediction 
     <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mi>
         f 
       </mi> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mi>
          x 
        </mi> 
        <mo>
          ) 
        </mo> 
       </mrow> 
      </mrow> 
     </math> <xref ref-type="bibr" rid="scirp.145026-29">
      [29]
     </xref>.</p>
    <p>
     <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mi>
         f 
       </mi> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mi>
          x 
        </mi> 
        <mo>
          ) 
        </mo> 
       </mrow> 
       <mo>
         = 
       </mo> 
       <mo> 
       </mo> 
       <mfrac> 
        <mrow> 
         <mi>
           arg 
         </mi> 
         <mi>
           max 
         </mi> 
        </mrow> 
        <mrow> 
         <mi>
           Y 
         </mi> 
         <mo>
           ∈ 
         </mo> 
         <mi mathvariant="script">
           Y 
         </mi> 
        </mrow> 
       </mfrac> 
       <mstyle displaystyle="true"> 
        <msubsup> 
         <mo>
           ∑ 
         </mo> 
         <mrow> 
          <mi>
            j 
          </mi> 
          <mo>
            = 
          </mo> 
          <mn>
            1 
          </mn> 
         </mrow> 
         <mi>
           j 
         </mi> 
        </msubsup> 
        <mrow> 
         <mi>
           I 
         </mi> 
         <mrow> 
          <mo>
            ( 
          </mo> 
          <mrow> 
           <mi>
             y 
           </mi> 
           <mo>
             = 
           </mo> 
           <msubsup> 
            <mi>
              h 
            </mi> 
            <mi>
              j 
            </mi> 
            <mrow> 
             <mrow> 
              <mo>
                ( 
              </mo> 
              <mi>
                x 
              </mi> 
              <mo>
                ) 
              </mo> 
             </mrow> 
            </mrow> 
           </msubsup> 
          </mrow> 
          <mo>
            ) 
          </mo> 
         </mrow> 
        </mrow> 
       </mstyle> 
       <msub> 
        <mo> 
        </mo> 
        <mrow> 
         <mo> 
         </mo> 
         <mo> 
         </mo> 
         <mo> 
         </mo> 
         <mo> 
         </mo> 
         <mo> 
         </mo> 
         <mo> 
         </mo> 
         <mo> 
         </mo> 
         <mo> 
         </mo> 
        </mrow> 
       </msub> 
      </mrow> 
     </math> (4)</p>
    <p>
     <xref ref-type="bibr" rid="scirp.145026-"></xref>In Random forest, the J<sub>th</sub> based learner denotes a tree represented by 
     <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          h 
        </mi> 
        <mi>
          j 
        </mi> 
       </msub> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <mi>
           x 
         </mi> 
         <mo>
           , 
         </mo> 
         <msub> 
          <mi>
            θ 
          </mi> 
          <mi>
            j 
          </mi> 
         </msub> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
      </mrow> 
     </math> where 
     <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          θ 
        </mi> 
        <mi>
          j 
        </mi> 
       </msub> 
      </mrow> 
     </math> denotes collection of random variables and 
     <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          θ 
        </mi> 
        <mi>
          j 
        </mi> 
       </msub> 
      </mrow> 
     </math>’s are independent where 
     <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mi>
         j 
       </mi> 
       <mo>
         = 
       </mo> 
       <mn>
         1 
       </mn> 
       <mo>
         , 
       </mo> 
       <mo>
         ⋯ 
       </mo> 
       <mo>
         , 
       </mo> 
       <mi>
         j 
       </mi> 
      </mrow> 
     </math>. The random forest algorithm is used to classify each client into their different users’ categories based on their access credentials. Based on majority votes, these are stored in a data mart/database created, leading to the creation of different clusters of users based on the profile of their access credentials.</p>
    <p>In the early days of data mining/machine learning, humans adopted manual labeling of data. However, with the increase in the volume of data available these days, manual labeling of data has become difficult, tedious, and expensive. Hence, there is a need for automatic labeling. Clustering can be described as a method of grouping a set of data objects into different classes of similar objects <xref ref-type="bibr" rid="scirp.145026-30">
      [30]
     </xref>. Available literature shows that there are a number of clustering methods for machine learning, which include: K-modes, K-means, K-median, genetic K-means, intelligent K-means, etc. <xref ref-type="bibr" rid="scirp.145026-31">
      [31]
     </xref></p>
    <p>This present work adopts the K-modes clustering techniques. The K-Modes is a non-parametric algorithm capable of handling categorical data and optimizing a matching metric. The loss function (L<sub>o</sub>) is used without explicitly applying any distance metric. Here, similar trees are categorized into the same cluster based on majority voting. We select K-initial modes, then form k clusters by assigning all the data points to the cluster with the nearest mode (vote) using the matching metrics. We later recompute the modes of the clusters until the convergence criteria are met.</p>
    <p>The correlation statistic technique is used to analyze the users’ browsing patterns alongside their access credentials and the attack information to arrive at a decision for detecting the anomaly. This is achieved by computing the correlation between the client’s URL information and the historical browsing information in the clusters stored in the data mart.</p>
    <p>A review of different distance measurement techniques shows that a good number of the existing techniques are marred with various challenges ranging from inaccuracy, susceptibility to noisy data, computational complexity, to scalability challenges, etc.</p>
    <p>The present correlation Distance measurement technique is capable of overcoming these challenges while providing scalable, efficient, accurate, and simple distance measurement for any machine learning algorithm.</p>
    <p>The correlation distance measurement technique is used to measure the relation and association between phenomena. The correlation relationship between two tuples X and Y can be expressed using Equation (5) below <xref ref-type="bibr" rid="scirp.145026-29">
      [29]
     </xref>.</p>
    <p>
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mi>
         r 
       </mi> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <mi>
           x 
         </mi> 
         <mo>
           , 
         </mo> 
         <mi>
           y 
         </mi> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
       <mo>
         = 
       </mo> 
       <mfrac> 
        <mrow> 
         <msubsup> 
          <mstyle mathsize="140%" displaystyle="true"> 
           <mo>
             ∑ 
           </mo> 
          </mstyle> 
          <mrow> 
           <mi>
             i 
           </mi> 
           <mo>
             = 
           </mo> 
           <mn>
             1 
           </mn> 
          </mrow> 
          <mi>
            n 
          </mi> 
         </msubsup> 
         <mrow> 
          <mo>
            ( 
          </mo> 
          <mrow> 
           <msub> 
            <mi>
              x 
            </mi> 
            <mi>
              i 
            </mi> 
           </msub> 
           <mo>
             − 
           </mo> 
           <mover accent="true"> 
            <mi>
              x 
            </mi> 
            <mo>
              ¯ 
            </mo> 
           </mover> 
          </mrow> 
          <mo>
            ) 
          </mo> 
         </mrow> 
         <mrow> 
          <mo>
            ( 
          </mo> 
          <mrow> 
           <msub> 
            <mi>
              y 
            </mi> 
            <mi>
              i 
            </mi> 
           </msub> 
           <mo>
             − 
           </mo> 
           <mover accent="true"> 
            <mi>
              y 
            </mi> 
            <mo>
              ¯ 
            </mo> 
           </mover> 
          </mrow> 
          <mo>
            ) 
          </mo> 
         </mrow> 
        </mrow> 
        <mrow> 
         <msqrt> 
          <mrow> 
           <msubsup> 
            <mstyle mathsize="140%" displaystyle="true"> 
             <mo>
               ∑ 
             </mo> 
            </mstyle> 
            <mrow> 
             <mi>
               i 
             </mi> 
             <mo>
               = 
             </mo> 
             <mn>
               1 
             </mn> 
            </mrow> 
            <mi>
              n 
            </mi> 
           </msubsup> 
           <msup> 
            <mrow> 
             <mrow> 
              <mo>
                ( 
              </mo> 
              <mrow> 
               <msub> 
                <mi>
                  x 
                </mi> 
                <mi>
                  i 
                </mi> 
               </msub> 
               <mo>
                 − 
               </mo> 
               <mover accent="true"> 
                <mi>
                  x 
                </mi> 
                <mo>
                  ¯ 
                </mo> 
               </mover> 
              </mrow> 
              <mo>
                ) 
              </mo> 
             </mrow> 
            </mrow> 
            <mn>
              2 
            </mn> 
           </msup> 
           <msubsup> 
            <mstyle mathsize="140%" displaystyle="true"> 
             <mo>
               ∑ 
             </mo> 
            </mstyle> 
            <mrow> 
             <mi>
               i 
             </mi> 
             <mo>
               = 
             </mo> 
             <mn>
               1 
             </mn> 
            </mrow> 
            <mi>
              n 
            </mi> 
           </msubsup> 
           <msup> 
            <mrow> 
             <mrow> 
              <mo>
                ( 
              </mo> 
              <mrow> 
               <msub> 
                <mi>
                  y 
                </mi> 
                <mi>
                  i 
                </mi> 
               </msub> 
               <mo>
                 − 
               </mo> 
               <mover accent="true"> 
                <mi>
                  y 
                </mi> 
                <mo>
                  ¯ 
                </mo> 
               </mover> 
              </mrow> 
              <mo>
                ) 
              </mo> 
             </mrow> 
            </mrow> 
            <mn>
              2 
            </mn> 
           </msup> 
          </mrow> 
         </msqrt> 
        </mrow> 
       </mfrac> 
      </mrow> 
     </math> (5)</p>
    <p>where 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mover accent="true"> 
       <mi>
         x 
       </mi> 
       <mo>
         ¯ 
       </mo> 
      </mover> 
     </math> is the mean of Tuple, and 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mover accent="true"> 
       <mi>
         y 
       </mi> 
       <mo>
         ¯ 
       </mo> 
      </mover> 
     </math> is the mean of tuple y.</p>
    <p>X = Training tuple</p>
    <p>Y = test Tuple</p>
    <p>Correlation can have a value of 1, 0 or −1</p>
    <p>If the correlation is 1, then this implies that there is a perfect positive correlation between variable X and Y</p>
    <p>If the correlation is 0, that implies no correlation (the value of X and Y, don’t seem linked at all)</p>
    <p>If the correlation is −1, that implies there is a perfect negative correlation <xref ref-type="bibr" rid="scirp.145026-32">
      [32]
     </xref>.</p>
    <p>The random forest enabled Correlation based clustering (RF-CBC) algorithm listing is shown in <xref ref-type="bibr" rid="scirp.145026-#al2">
      Algorithm Listing 2
     </xref>.</p>
    <p>Algorithm Listing 2. The random forest enabled Correlation based clustering (RF-CBC) algorithm.</p>
    <fig id="fig7" position="float">
     <label>Figure 7</label>
     <caption>
      <title>3.4.4. Application of the Proposed RF-CBC Technique to Detect and Report Attack and AbnormalityThis section demonstrates the application of the present RF-CBR technique to detect and report abnormality on the experimental website, i.e., the UNIMED website, using our extracted user browsing history data set of the university website.Considering our experimental website i.e, the UNIMED website users’ credentials and click stream data are considered as a vector with attributes: user’s category, access type, allowed operation, (Click log<sub>1</sub>, Click log<sub>2</sub>, Click log<sub>3</sub> … Click log<sub>n</sub>), with clients/users represented by X<sub>1</sub>, X<sub>2</sub>, X<sub>3</sub>, X<sub>4</sub>, ⋯, X<sub>n</sub> as class labels, as presented in <xref ref-type="table" rid="table4">
        Table 4
       </xref>.Given an unknown user’s X<sub>5</sub>, determine the class of user X<sub>5</sub>. The random forest classifier will first be used to classify all the users into clusters based on similarities in their users’ credentials and their click logs. We then compute the correlation distance between users X<sub>5</sub> and all other vectors in each of the clusters by applying Equation (5), i.e.
       <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
  
         <mi>
          
   r
  
         </mi>
  
         <mrow>
   
          <mo>
           
    (
   
          </mo> 
   
          <mrow> 
    
           <mi>
            
     x
    
           </mi>
    
           <mo>
            
     ,
    
           </mo>
    
           <mi>
            
     y
    
           </mi>
   
          </mrow> 
   
          <mo>
           
    )
   
          </mo>
  
         </mrow>
  
         <mo>
          
   =
  
         </mo>
  
         <mfrac> 
   
          <mrow> 
    
           <msubsup> 
     
            <mstyle mathsize="140%" displaystyle="true"> 
             <mo>
               ∑ 
             </mo> 
            </mstyle> 
     
            <mrow> 
             <mi>
               i 
             </mi> 
             <mo>
               = 
             </mo> 
             <mn>
               1 
             </mn> 
            </mrow> 
     
            <mi>
              n 
            </mi> 
    
           </msubsup> 
    
           <mrow>
     
            <mo>
              ( 
            </mo> 
     
            <mrow> 
             <msub> 
              <mi>
                x 
              </mi> 
              <mi>
                i 
              </mi> 
             </msub> 
             <mo>
               − 
             </mo> 
             <mover accent="true"> 
              <mi>
                x 
              </mi> 
              <mo>
                ¯ 
              </mo> 
             </mover> 
            </mrow> 
     
            <mo>
              ) 
            </mo>
    
           </mrow>
    
           <mrow>
     
            <mo>
              ( 
            </mo> 
     
            <mrow> 
             <msub> 
              <mi>
                y 
              </mi> 
              <mi>
                i 
              </mi> 
             </msub> 
             <mo>
               − 
             </mo> 
             <mover accent="true"> 
              <mi>
                y 
              </mi> 
              <mo>
                ¯ 
              </mo> 
             </mover> 
            </mrow> 
     
            <mo>
              ) 
            </mo>
    
           </mrow>
   
          </mrow> 
   
          <mrow> 
    
           <msqrt> 
     
            <mrow> 
             <msubsup> 
              <mstyle mathsize="140%" displaystyle="true"> 
               <mo>
                 ∑ 
               </mo> 
              </mstyle> 
              <mrow> 
               <mi>
                 i 
               </mi> 
               <mo>
                 = 
               </mo> 
               <mn>
                 1 
               </mn> 
              </mrow> 
              <mi>
                n 
              </mi> 
             </msubsup> 
             <msup> 
              <mrow> 
               <mrow> 
                <mo>
                  ( 
                </mo> 
                <mrow> 
                 <msub> 
                  <mi>
                    x 
                  </mi> 
                  <mi>
                    i 
                  </mi> 
                 </msub> 
                 <mo>
                   − 
                 </mo> 
                 <mover accent="true"> 
                  <mi>
                    x 
                  </mi> 
                  <mo>
                    ¯ 
                  </mo> 
                 </mover> 
                </mrow> 
                <mo>
                  ) 
                </mo> 
               </mrow> 
              </mrow> 
              <mn>
                2 
              </mn> 
             </msup> 
             <msubsup> 
              <mstyle mathsize="140%" displaystyle="true"> 
               <mo>
                 ∑ 
               </mo> 
              </mstyle> 
              <mrow> 
               <mi>
                 i 
               </mi> 
               <mo>
                 = 
               </mo> 
               <mn>
                 1 
               </mn> 
              </mrow> 
              <mi>
                n 
              </mi> 
             </msubsup> 
             <msup> 
              <mrow> 
               <mrow> 
                <mo>
                  ( 
                </mo> 
                <mrow> 
                 <msub> 
                  <mi>
                    y 
                  </mi> 
                  <mi>
                    i 
                  </mi> 
                 </msub> 
                 <mo>
                   − 
                 </mo> 
                 <mover accent="true"> 
                  <mi>
                    y 
                  </mi> 
                  <mo>
                    ¯ 
                  </mo> 
                 </mover> 
                </mrow> 
                <mo>
                  ) 
                </mo> 
               </mrow> 
              </mrow> 
              <mn>
                2 
              </mn> 
             </msup> 
            </mrow> 
    
           </msqrt> 
   
          </mrow> 
  
         </mfrac> 
 
        </mrow>

       </math><xref ref-type="bibr" rid="scirp.145026-"></xref>Table 4. The UNIMED website data mart class label training tuples and user’s credentials click logs.
       <table class="MsoTableGrid custom-table" border="0" cellspacing="0" cellpadding="0"> 
 
        <tr> 
  
         <td class="custom-bottom-td acenter" width="8.88%"><p style="text-align:center">Users</p></td> 
  
         <td class="custom-bottom-td acenter" width="13.32%"><p style="text-align:center">Access type</p></td> 
  
         <td class="custom-bottom-td acenter" width="16.27%"><p style="text-align:center">Operation type</p></td> 
  
         <td class="custom-bottom-td acenter" width="12.96%"><p style="text-align:center">Log<sub>1</sub></p></td> 
  
         <td class="custom-bottom-td acenter" width="15.71%"><p style="text-align:center">Log<sub>2</sub></p></td> 
  
         <td class="custom-bottom-td acenter" width="15.71%"><p style="text-align:center">Log<sub>3</sub></p></td> 
  
         <td class="custom-bottom-td acenter" width="17.16%"><p style="text-align:center">Credentials type/Class Label</p></td> 
 
        </tr> 
 
        <tr> 
  
         <td class="custom-top-td acenter" width="8.88%"><p style="text-align:center">X<sub>1</sub></p></td> 
  
         <td class="custom-top-td acenter" width="13.32%"><p style="text-align:center">Full access</p></td> 
  
         <td class="custom-top-td acenter" width="16.27%"><p style="text-align:center">Unrestricted</p></td> 
  
         <td class="custom-top-td acenter" width="12.96%"><p style="text-align:center">Index</p></td> 
  
         <td class="custom-top-td acenter" width="15.71%"><p style="text-align:center">Staff profile</p></td> 
  
         <td class="custom-top-td acenter" width="15.71%"><p style="text-align:center">Exams and Record</p></td> 
  
         <td class="custom-top-td acenter" width="17.16%"><p style="text-align:center">Administrator</p></td> 
 
        </tr> 
 
        <tr> 
  
         <td class="acenter" width="8.88%"><p style="text-align:center">X<sub>2</sub></p></td> 
  
         <td class="acenter" width="13.32%"><p style="text-align:center">Privileged</p></td> 
  
         <td class="acenter" width="16.27%"><p style="text-align:center">Staff privileged</p></td> 
  
         <td class="acenter" width="12.96%"><p style="text-align:center">Index</p></td> 
  
         <td class="acenter" width="15.71%"><p style="text-align:center">Staff profile</p></td> 
  
         <td class="acenter" width="15.71%"><p style="text-align:center">Add courses</p></td> 
  
         <td class="acenter" width="17.16%"><p style="text-align:center">HOD</p></td> 
 
        </tr> 
 
        <tr> 
  
         <td class="acenter" width="8.88%"><p style="text-align:center">X<sub>3</sub></p></td> 
  
         <td class="acenter" width="13.32%"><p style="text-align:center">Restricted</p></td> 
  
         <td class="acenter" width="16.27%"><p style="text-align:center">Basic operation</p></td> 
  
         <td class="acenter" width="12.96%"><p style="text-align:center">Admission</p></td> 
  
         <td class="acenter" width="15.71%"><p style="text-align:center">faculty</p></td> 
  
         <td class="acenter" width="15.71%"><p style="text-align:center">Programmes</p></td> 
  
         <td class="acenter" width="17.16%"><p style="text-align:center">Visitor</p></td> 
 
        </tr> 
 
        <tr> 
  
         <td class="acenter" width="8.88%"><p style="text-align:center">X<sub>4</sub></p></td> 
  
         <td class="acenter" width="13.32%"><p style="text-align:center">Limited</p></td> 
  
         <td class="acenter" width="16.27%"><p style="text-align:center">Student operation</p></td> 
  
         <td class="acenter" width="12.96%"><p style="text-align:center">Registration</p></td> 
  
         <td class="acenter" width="15.71%"><p style="text-align:center">Courses</p></td> 
  
         <td class="acenter" width="15.71%"><p style="text-align:center">Fees payment</p></td> 
  
         <td class="acenter" width="17.16%"><p style="text-align:center">Student</p></td> 
 
        </tr> 
 
        <tr> 
  
         <td class="acenter" width="8.88%"><p style="text-align:center">X<sub>6</sub></p></td> 
  
         <td class="acenter" width="13.32%"><p style="text-align:center">Privileged</p></td> 
  
         <td class="acenter" width="16.27%"><p style="text-align:center">Staff privileged</p></td> 
  
         <td class="acenter" width="12.96%"><p style="text-align:center">Index</p></td> 
  
         <td class="acenter" width="15.71%"><p style="text-align:center">Staff profile</p></td> 
  
         <td class="acenter" width="15.71%"><p style="text-align:center">Approved courses</p></td> 
  
         <td class="acenter" width="17.16%"><p style="text-align:center">Tutor</p></td> 
 
        </tr> 
 
        <tr> 
  
         <td class="tbtextacenter" width="8.88%"><p style="text-align:center">…</p></td> 
  
         <td class="acenter" width="13.32%"><p style="text-align:center"></p></td> 
  
         <td class="acenter" width="16.27%"><p style="text-align:center"></p></td> 
  
         <td class="acenter" width="12.96%"><p style="text-align:center"></p></td> 
  
         <td class="acenter" width="15.71%"><p style="text-align:center"></p></td> 
  
         <td class="acenter" width="15.71%"><p style="text-align:center"></p></td> 
  
         <td class="acenter" width="17.16%"><p style="text-align:center"></p></td> 
 
        </tr> 
 
        <tr> 
  
         <td class="acenter" width="8.88%"><p style="text-align:center">X<sub>5</sub></p></td> 
  
         <td class="acenter" width="13.32%"><p style="text-align:center">Restricted</p></td> 
  
         <td class="acenter" width="16.27%"><p style="text-align:center">Basic operation</p></td> 
  
         <td class="acenter" width="12.96%"><p style="text-align:center">Staff profile</p></td> 
  
         <td class="acenter" width="15.71%"><p style="text-align:center">Exams and records</p></td> 
  
         <td class="acenter" width="15.71%"><p style="text-align:center">Upload Credentials</p></td> 
  
         <td class="acenter" width="17.16%"><p style="text-align:center">?</p></td> 
 
        </tr>

       </table>The data from our experimental website shown in <xref ref-type="table" rid="table4">
        Table 4
       </xref> are categorical in nature; this means they are non-numeric attributes. To be able to use them for numerical calculation, we adopt the nominal scale technique <xref ref-type="bibr" rid="scirp.145026-30">
        [30]
       </xref> <xref ref-type="bibr" rid="scirp.145026-33">
        [33]
       </xref>. The nominal scale involves assigning numbers to a category of data, such that no category is greater than or less than the other categories; it involves labeling data in a particular group according to the relevant attribute possessed, in no special or specific order or magnitude. It is strictly for identification purposes.Therefore, for a set of attributes, the users’ category, which consists of (Admin, HOD, Tutor, Student, Visitor), will be coded as 1 for Admin, 2 for HOD, 3 for Tutor, 4 for Student, and 5 for Visitor. Likewise, for the set of attributes, access type, which consists of (Full access, privileged, limited, and restricted), will also be coded as 1 for full access, 2 for privileged, 3 for limited, and 4 for restricted. For the set of attributes, operation type (unrestricted, staff privilege, basic operation, student privilege) will be coded as 1 for unrestricted, 2 for privilege, 3 for student privilege, and 4 for basic operation. For the attribute click log, which is made up of possible user clicks on the UNIMED website and includes: index (1), Staff portal (2), Admission (3), Student registration (4), Library (5), Exams portal (6), Faculty (7), Programmes (8), Add Courses (9), pay portal (10), Upload result (11), Upload credentials (12), etc., these are also labeled numerically in ascending order, such as 1, 2, 3, 4, …, n, where n represents the total number of possible clicks on the UNIMED website. The number in parentheses after each data tuple represents the nominal scale assigned to the tuples.The complete nominal scale representation of the given training tuple, as presented in <xref ref-type="table" rid="table4">
        Table 4
       </xref>, is shown in <xref ref-type="table" rid="table5">
        Table 5
       </xref>.</title>
     </caption>
     <graphic mimetype="image" position="float" xlink:type="simple" xlink:href="https://html.scirp.org/file/7801084-rId75.jpeg?20250822024420" />
    </fig>
    <table-wrap id="table4">
     <label>
      <xref ref-type="table" rid="table4">
       Table 4
      </xref></label>
     <caption>
      <title>
       <xref ref-type="bibr" rid="scirp.145026-"></xref>Table 5. The UNIMED website data mart class label training tuples and user’s credentials click logs.</title>
     </caption>
     <table class="MsoTableGrid custom-table" border="0" cellspacing="0" cellpadding="0"> 
      <tr> 
       <td class="custom-bottom-td acenter" width="11.51%"><p style="text-align:center">Users</p></td> 
       <td class="custom-bottom-td acenter" width="10.58%"><p style="text-align:center">Access type</p></td> 
       <td class="custom-bottom-td acenter" width="14.55%"><p style="text-align:center">Operation type</p></td> 
       <td class="custom-bottom-td acenter" width="8.36%"><p style="text-align:center">Log<sub>1</sub></p></td> 
       <td class="custom-bottom-td acenter" width="8.73%"><p style="text-align:center">Log<sub>2</sub></p></td> 
       <td class="custom-bottom-td acenter" width="11.37%"><p style="text-align:center">Log<sub>3</sub></p></td> 
       <td class="custom-bottom-td acenter" width="21.10%"><p style="text-align:center">Credentials type/Class Label</p></td> 
       <td class="custom-bottom-td acenter" width="13.79%"><p style="text-align:center">Status</p></td> 
      </tr> 
      <tr> 
       <td class="custom-top-td acenter" width="11.51%"><p style="text-align:center">X<sub>1</sub></p></td> 
       <td class="custom-top-td acenter" width="10.58%"><p style="text-align:center">1</p></td> 
       <td class="custom-top-td acenter" width="14.55%"><p style="text-align:center">1</p></td> 
       <td class="custom-top-td acenter" width="8.36%"><p style="text-align:center">1</p></td> 
       <td class="custom-top-td acenter" width="8.73%"><p style="text-align:center">2</p></td> 
       <td class="custom-top-td acenter" width="11.37%"><p style="text-align:center">6</p></td> 
       <td class="custom-top-td acenter" width="21.10%"><p style="text-align:center">Administrator</p></td> 
       <td class="custom-top-td acenter" width="13.79%"><p style="text-align:center">Normal</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="11.51%"><p style="text-align:center">X<sub>2</sub></p></td> 
       <td class="acenter" width="10.58%"><p style="text-align:center">2</p></td> 
       <td class="acenter" width="14.55%"><p style="text-align:center">2</p></td> 
       <td class="acenter" width="8.36%"><p style="text-align:center">1</p></td> 
       <td class="acenter" width="8.73%"><p style="text-align:center">2</p></td> 
       <td class="acenter" width="11.37%"><p style="text-align:center">9</p></td> 
       <td class="acenter" width="21.10%"><p style="text-align:center">HOD</p></td> 
       <td class="acenter" width="13.79%"><p style="text-align:center">Normal</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="11.51%"><p style="text-align:center">X<sub>3</sub></p></td> 
       <td class="acenter" width="10.58%"><p style="text-align:center">4</p></td> 
       <td class="acenter" width="14.55%"><p style="text-align:center">4</p></td> 
       <td class="acenter" width="8.36%"><p style="text-align:center">3</p></td> 
       <td class="acenter" width="8.73%"><p style="text-align:center">7</p></td> 
       <td class="acenter" width="11.37%"><p style="text-align:center">8</p></td> 
       <td class="acenter" width="21.10%"><p style="text-align:center">Visitor</p></td> 
       <td class="acenter" width="13.79%"><p style="text-align:center">Normal</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="11.51%"><p style="text-align:center">X<sub>4</sub></p></td> 
       <td class="acenter" width="10.58%"><p style="text-align:center">3</p></td> 
       <td class="acenter" width="14.55%"><p style="text-align:center">3</p></td> 
       <td class="acenter" width="8.36%"><p style="text-align:center">4</p></td> 
       <td class="acenter" width="8.73%"><p style="text-align:center">9</p></td> 
       <td class="acenter" width="11.37%"><p style="text-align:center">10</p></td> 
       <td class="acenter" width="21.10%"><p style="text-align:center">Student</p></td> 
       <td class="acenter" width="13.79%"><p style="text-align:center">Normal</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="11.51%"><p style="text-align:center">X<sub>6</sub></p></td> 
       <td class="acenter" width="10.58%"><p style="text-align:center">2</p></td> 
       <td class="acenter" width="14.55%"><p style="text-align:center">2</p></td> 
       <td class="acenter" width="8.36%"><p style="text-align:center">1</p></td> 
       <td class="acenter" width="8.73%"><p style="text-align:center"></p></td> 
       <td class="acenter" width="11.37%"><p style="text-align:center">9</p></td> 
       <td class="acenter" width="21.10%"><p style="text-align:center">Tutor</p></td> 
       <td class="acenter" width="13.79%"><p style="text-align:center">Normal</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="11.51%"><p style="text-align:center">X<sub>7</sub></p></td> 
       <td class="acenter" width="10.58%"><p style="text-align:center">4</p></td> 
       <td class="acenter" width="14.55%"><p style="text-align:center">4</p></td> 
       <td class="acenter" width="8.36%"><p style="text-align:center">1</p></td> 
       <td class="acenter" width="8.73%"><p style="text-align:center">6</p></td> 
       <td class="acenter" width="11.37%"><p style="text-align:center">9</p></td> 
       <td class="acenter" width="21.10%"><p style="text-align:center">Visitor</p></td> 
       <td class="acenter" width="13.79%"><p style="text-align:center">Attacker</p></td> 
      </tr> 
      <tr> 
       <td class="tbtextacenter" width="11.51%"><p style="text-align:center">…</p></td> 
       <td class="acenter" width="10.58%"><p style="text-align:center"></p></td> 
       <td class="acenter" width="14.55%"><p style="text-align:center"></p></td> 
       <td class="acenter" width="8.36%"><p style="text-align:center"></p></td> 
       <td class="acenter" width="8.73%"><p style="text-align:center"></p></td> 
       <td class="acenter" width="11.37%"><p style="text-align:center"></p></td> 
       <td class="acenter" width="21.10%"><p style="text-align:center"></p></td> 
       <td class="acenter" width="13.79%"><p style="text-align:center"></p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="11.51%"><p style="text-align:center">X<sub>5</sub></p></td> 
       <td class="acenter" width="10.58%"><p style="text-align:center">4</p></td> 
       <td class="acenter" width="14.55%"><p style="text-align:center">4</p></td> 
       <td class="acenter" width="8.36%"><p style="text-align:center">2</p></td> 
       <td class="acenter" width="8.73%"><p style="text-align:center">6</p></td> 
       <td class="acenter" width="11.37%"><p style="text-align:center">12</p></td> 
       <td class="acenter" width="21.10%"><p style="text-align:center">?</p></td> 
       <td class="acenter" width="13.79%"><p style="text-align:center">?</p></td> 
      </tr> 
     </table>
    </table-wrap>
    <p>The correlation between user X<sub>1</sub> and user x<sub>5</sub> can now be computed as follows:</p>
    <p>
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mi>
         r 
       </mi> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <mi>
           x 
         </mi> 
         <mo>
           , 
         </mo> 
         <mi>
           y 
         </mi> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
       <mo>
         = 
       </mo> 
       <mfrac> 
        <mrow> 
         <msubsup> 
          <mstyle mathsize="140%" displaystyle="true"> 
           <mo>
             ∑ 
           </mo> 
          </mstyle> 
          <mrow> 
           <mi>
             i 
           </mi> 
           <mo>
             = 
           </mo> 
           <mn>
             1 
           </mn> 
          </mrow> 
          <mi>
            n 
          </mi> 
         </msubsup> 
         <mrow> 
          <mo>
            ( 
          </mo> 
          <mrow> 
           <msub> 
            <mi>
              x 
            </mi> 
            <mi>
              i 
            </mi> 
           </msub> 
           <mo>
             − 
           </mo> 
           <mover accent="true"> 
            <mi>
              x 
            </mi> 
            <mo>
              ¯ 
            </mo> 
           </mover> 
          </mrow> 
          <mo>
            ) 
          </mo> 
         </mrow> 
         <mrow> 
          <mo>
            ( 
          </mo> 
          <mrow> 
           <msub> 
            <mi>
              y 
            </mi> 
            <mi>
              i 
            </mi> 
           </msub> 
           <mo>
             − 
           </mo> 
           <mover accent="true"> 
            <mi>
              y 
            </mi> 
            <mo>
              ¯ 
            </mo> 
           </mover> 
          </mrow> 
          <mo>
            ) 
          </mo> 
         </mrow> 
        </mrow> 
        <mrow> 
         <msqrt> 
          <mrow> 
           <msubsup> 
            <mstyle mathsize="140%" displaystyle="true"> 
             <mo>
               ∑ 
             </mo> 
            </mstyle> 
            <mrow> 
             <mi>
               i 
             </mi> 
             <mo>
               = 
             </mo> 
             <mn>
               1 
             </mn> 
            </mrow> 
            <mi>
              n 
            </mi> 
           </msubsup> 
           <msup> 
            <mrow> 
             <mrow> 
              <mo>
                ( 
              </mo> 
              <mrow> 
               <msub> 
                <mi>
                  x 
                </mi> 
                <mi>
                  i 
                </mi> 
               </msub> 
               <mo>
                 − 
               </mo> 
               <mover accent="true"> 
                <mi>
                  x 
                </mi> 
                <mo>
                  ¯ 
                </mo> 
               </mover> 
              </mrow> 
              <mo>
                ) 
              </mo> 
             </mrow> 
            </mrow> 
            <mn>
              2 
            </mn> 
           </msup> 
           <msubsup> 
            <mstyle mathsize="140%" displaystyle="true"> 
             <mo>
               ∑ 
             </mo> 
            </mstyle> 
            <mrow> 
             <mi>
               i 
             </mi> 
             <mo>
               = 
             </mo> 
             <mn>
               1 
             </mn> 
            </mrow> 
            <mi>
              n 
            </mi> 
           </msubsup> 
           <msup> 
            <mrow> 
             <mrow> 
              <mo>
                ( 
              </mo> 
              <mrow> 
               <msub> 
                <mi>
                  y 
                </mi> 
                <mi>
                  i 
                </mi> 
               </msub> 
               <mo>
                 − 
               </mo> 
               <mover accent="true"> 
                <mi>
                  y 
                </mi> 
                <mo>
                  ¯ 
                </mo> 
               </mover> 
              </mrow> 
              <mo>
                ) 
              </mo> 
             </mrow> 
            </mrow> 
            <mn>
              2 
            </mn> 
           </msup> 
          </mrow> 
         </msqrt> 
        </mrow> 
       </mfrac> 
      </mrow> 
     </math></p>
    <p>
     <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          X 
        </mi> 
        <mn>
          1 
        </mn> 
       </msub> 
       <mo>
         = 
       </mo> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <mn>
           1 
         </mn> 
         <mo>
           , 
         </mo> 
         <mn>
           1 
         </mn> 
         <mo>
           , 
         </mo> 
         <mn>
           1 
         </mn> 
         <mo>
           , 
         </mo> 
         <mn>
           2 
         </mn> 
         <mo>
           , 
         </mo> 
         <mn>
           6 
         </mn> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
      </mrow> 
     </math>, 
     <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          X 
        </mi> 
        <mn>
          5 
        </mn> 
       </msub> 
       <mo>
         = 
       </mo> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <mn>
           4 
         </mn> 
         <mo>
           , 
         </mo> 
         <mn>
           4 
         </mn> 
         <mo>
           , 
         </mo> 
         <mn>
           2 
         </mn> 
         <mo>
           , 
         </mo> 
         <mn>
           6 
         </mn> 
         <mo>
           , 
         </mo> 
         <mn>
           12 
         </mn> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
      </mrow> 
     </math></p>
    <p>The correlation between user X<sub>5</sub> and user X<sub>1</sub> is −1; this indicates that there is no relationship between user X<sub>1</sub> and user X<sub>5</sub>. The process is repeated between the unknown user X<sub>5</sub> and every other user. If the correlation is 1, then we categorize user X<sub>1</sub> into the cluster of user Y<sub>1</sub> or any other matching cluster. However, if the correlation matches that of the attacker’s cluster or if the correlation is −1 throughout, then user X<sub>1</sub> is suspected to be an attacker and is therefore reported as an attacker.</p>
   </sec>
  </sec><sec id="s4">
   <title>4. System Evaluation and Result Analysis</title>
   <p>We evaluate the proposed RF-CBC attack detection and reporting system using the University of Medical Sciences (UNIMED) website users’ browsing history data set. Being a real-world data set, the data set was pre-processed and relevant features were extracted using the proposed STBO feature scaling technique before applying the proposed RF-CBC to detect any abnormality on the website. To this effect, in-house software was developed using Python, with XAMP/Apache HTTP as the hosting server and MYSQL DBMS for data mart creation.</p>
   <sec id="s4_1">
    <title>4.1. System Evaluation</title>
    <p>As part of our experiment, we carried out a performance evaluation of our system through performance comparison of the present system with some baseline methods, which are the traditional random forest, the traditional clustering method, the Naïve Bayes, and the Artificial Neural Network techniques on the same datasets and environment. An experimental online attack detection and reporting system was developed to implement the RF-CBC model. In the developed attack detection system, the users enter their basic information, which is used to build their profile and credentials online and in real time. The attack detection report is triggered when the system observes abnormalities in the users’ browsing patterns. The source code for the developed application is available on request. The proposed system can be implemented online by uploading the present application to a web server, which can be invoked anytime online on any web browser.</p>
   </sec>
   <sec id="s4_2">
    <title>4.2. Data Set for Evaluation</title>
    <p>In order to evaluate our RF-CBC model, we used the extracted historical browsing history of the UNIMED website dataset. The click log history of 15,374 anonymous users of the UNIMED official website, who signed into their accounts over a period of 12 months from 12th May, 2023 to 30th April, 2024, was randomly selected, which is made up of about 30,572 sessions/SQL queries. The selected records were divided into five parts, four of which were used as training sets and the remaining part was used as the testing set. The class/clusters of the training part were considered known while those of the testing part were considered unknown; the training part is used to infer the unknown part.</p>
    <p>The software developed was used to run both the proposed RF-CBC and the baseline methods, which include: the Traditional Random Forest (TRF), the Naïve Bayesian (NBY), the Artificial Neural Network (ANN), and the Traditional Clustering (TCL) methods, for about sixty times. The results are used as a dataset for evaluation purposes. We calculate the Accuracy (ACC), True Positive (TP), True Negative (TN), False Positive (FP), False Negative (FN), Precision (P), and Recall (R).</p>
    <table-wrap id="table5">
     <label>
      <xref ref-type="table" rid="table5">
       Table 5
      </xref></label>
     <caption>
      <title>
       <xref ref-type="bibr" rid="scirp.145026-"></xref>Table 6. The confusion matrix for our attack detection system.</title>
     </caption>
     <table class="MsoTableGrid custom-table" border="0" cellspacing="0" cellpadding="0"> 
      <tr> 
       <td class="custom-bottom-td acenter" width="52.39%"><p style="text-align:center"></p></td> 
       <td class="custom-bottom-td acenter" width="52.39%"><p style="text-align:center">Relevant</p></td> 
       <td class="custom-bottom-td acenter" width="52.39%"><p style="text-align:center">Relevant</p></td> 
      </tr> 
      <tr> 
       <td class="custom-top-td acenter" width="52.39%"><p style="text-align:center">Predicted</p></td> 
       <td class="custom-top-td acenter" width="52.39%"><p style="text-align:center">TP</p></td> 
       <td class="custom-top-td acenter" width="52.39%"><p style="text-align:center">FP</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="52.39%"><p style="text-align:center">Not relevant</p></td> 
       <td class="acenter" width="52.39%"><p style="text-align:center">FN</p></td> 
       <td class="acenter" width="52.39%"><p style="text-align:center">TN</p></td> 
      </tr> 
     </table>
    </table-wrap>
    <p>where:</p>
    <p>AC = Accuracy: This measures the proportion of correctly detected attacks and normal instances.</p>
    <p>TP = True Positive: This is the number of correctly detected attack instances,</p>
    <p>TN = True Negative: This is the number of correctly detected normal instances,</p>
    <p>FP = False Positive: This is the number of incorrectly detected attack instances, and</p>
    <p>FN = False Negative: This is the number of incorrectly detected normal instances.</p>
    <p>We used the F1-Measure technique as presented in a confusion matrix shown in <xref ref-type="table" rid="table6">
      Table 6
     </xref>.</p>
    <p>As presented in the confusion matrix shown in <xref ref-type="table" rid="table6">
      Table 6
     </xref>, to evaluate the detection quality of the RF-CBC model, where F-Measure is the harmonic mean of precision and recall, i.e.</p>
    <p>
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mi>
         F 
       </mi> 
       <mo>
         = 
       </mo> 
       <mn>
         2 
       </mn> 
       <mo>
         ⋅ 
       </mo> 
       <mfrac> 
        <mrow> 
         <mtext>
           Precision 
         </mtext> 
         <mo>
           ⋅ 
         </mo> 
         <mtext>
           recall 
         </mtext> 
        </mrow> 
        <mrow> 
         <mtext>
           Precision 
         </mtext> 
         <mo>
           + 
         </mo> 
         <mtext>
           recall 
         </mtext> 
        </mrow> 
       </mfrac> 
      </mrow> 
     </math></p>
    <p>This is referred to as the F1- measure, since recall and precision are weighted.</p>
    <p>Precision is the number of correct predictions divided by the number of all returned predictions, i.e.</p>
    <p>
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mtext>
         Precision 
       </mtext> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mi>
          P 
        </mi> 
        <mo>
          ) 
        </mo> 
       </mrow> 
       <mo>
         = 
       </mo> 
       <mfrac> 
        <mrow> 
         <mtext>
           TP 
         </mtext> 
        </mrow> 
        <mrow> 
         <mtext>
           TP 
         </mtext> 
         <mo>
           + 
         </mo> 
         <mtext>
           FP 
         </mtext> 
        </mrow> 
       </mfrac> 
      </mrow> 
     </math></p>
    <p>Recall is the number of correct predictions divided by the number of all known interest supposed to be discovered, i.e.</p>
    <p>
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mtext>
         Recall 
       </mtext> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mi>
          R 
        </mi> 
        <mo>
          ) 
        </mo> 
       </mrow> 
       <mo>
         = 
       </mo> 
       <mfrac> 
        <mrow> 
         <mtext>
           TP 
         </mtext> 
        </mrow> 
        <mrow> 
         <mtext>
           TP 
         </mtext> 
         <mo>
           + 
         </mo> 
         <mtext>
           FN 
         </mtext> 
        </mrow> 
       </mfrac> 
      </mrow> 
     </math> <xref ref-type="bibr" rid="scirp.145026-17">
      [17]
     </xref> <xref ref-type="bibr" rid="scirp.145026-30">
      [30]
     </xref>.</p>
    <p>
     <xref ref-type="table" rid="table7">
      Table 7
     </xref> shows our experimental results on the UNIMED browsing history database with ACC, TP, TN, FP, TP, Precision, and Recall. <xref ref-type="fig" rid="fig5">
      Figure 5
     </xref> shows the result of our conducted experiment in F1-Measure using our browsing history click logs database.</p>
    <table-wrap id="table6">
     <label>
      <xref ref-type="table" rid="table6">
       Table 6
      </xref></label>
     <caption>
      <title>
       <xref ref-type="bibr" rid="scirp.145026-"></xref>Table 7. Experimental Results from our attack detection system</title>
     </caption>
     <table class="MsoTableGrid custom-table" border="0" cellspacing="0" cellpadding="0"> 
      <tr> 
       <td class="custom-bottom-td acenter" width="11.92%"><p style="text-align:center">Model</p></td> 
       <td class="custom-bottom-td acenter" width="9.41%"><p style="text-align:center">TP</p></td> 
       <td class="custom-bottom-td acenter" width="9.41%"><p style="text-align:center">TN</p></td> 
       <td class="custom-bottom-td acenter" width="9.41%"><p style="text-align:center">FP</p></td> 
       <td class="custom-bottom-td acenter" width="9.43%"><p style="text-align:center">FN</p></td> 
       <td class="custom-bottom-td acenter" width="16.80%"><p style="text-align:center">P</p></td> 
       <td class="custom-bottom-td acenter" width="16.80%"><p style="text-align:center">R</p></td> 
       <td class="custom-bottom-td acenter" width="16.82%"><p style="text-align:center">F1-measure</p></td> 
      </tr> 
      <tr> 
       <td class="custom-top-td acenter" width="11.92%"><p style="text-align:center">RF-CBC</p></td> 
       <td class="custom-top-td acenter" width="9.41%"><p style="text-align:center">8</p></td> 
       <td class="custom-top-td acenter" width="9.41%"><p style="text-align:center">2</p></td> 
       <td class="custom-top-td acenter" width="9.41%"><p style="text-align:center">1</p></td> 
       <td class="custom-top-td acenter" width="9.43%"><p style="text-align:center">0</p></td> 
       <td class="custom-top-td acenter" width="16.80%"><p style="text-align:center">0.888888889</p></td> 
       <td class="custom-top-td acenter" width="16.80%"><p style="text-align:center">1</p></td> 
       <td class="custom-top-td acenter" width="16.82%"><p style="text-align:center">0.941176471</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="11.92%"><p style="text-align:center">TRF</p></td> 
       <td class="acenter" width="9.41%"><p style="text-align:center">6</p></td> 
       <td class="acenter" width="9.41%"><p style="text-align:center">2</p></td> 
       <td class="acenter" width="9.41%"><p style="text-align:center">2</p></td> 
       <td class="acenter" width="9.43%"><p style="text-align:center">1</p></td> 
       <td class="acenter" width="16.80%"><p style="text-align:center">0.75</p></td> 
       <td class="acenter" width="16.80%"><p style="text-align:center">0.857142857</p></td> 
       <td class="acenter" width="16.82%"><p style="text-align:center">0.8</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="11.92%"><p style="text-align:center">ANN</p></td> 
       <td class="acenter" width="9.41%"><p style="text-align:center">6</p></td> 
       <td class="acenter" width="9.41%"><p style="text-align:center">3</p></td> 
       <td class="acenter" width="9.41%"><p style="text-align:center">1</p></td> 
       <td class="acenter" width="9.43%"><p style="text-align:center">1</p></td> 
       <td class="acenter" width="16.80%"><p style="text-align:center">0.857142857</p></td> 
       <td class="acenter" width="16.80%"><p style="text-align:center">0.857142857</p></td> 
       <td class="acenter" width="16.82%"><p style="text-align:center">0.857142857</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="11.92%"><p style="text-align:center">NBY</p></td> 
       <td class="acenter" width="9.41%"><p style="text-align:center">6</p></td> 
       <td class="acenter" width="9.41%"><p style="text-align:center">2</p></td> 
       <td class="acenter" width="9.41%"><p style="text-align:center">1</p></td> 
       <td class="acenter" width="9.43%"><p style="text-align:center">2</p></td> 
       <td class="acenter" width="16.80%"><p style="text-align:center">0.857142857</p></td> 
       <td class="acenter" width="16.80%"><p style="text-align:center">0.75</p></td> 
       <td class="acenter" width="16.82%"><p style="text-align:center">0.8</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="11.92%"><p style="text-align:center">TCL</p></td> 
       <td class="acenter" width="9.41%"><p style="text-align:center">5</p></td> 
       <td class="acenter" width="9.41%"><p style="text-align:center">3</p></td> 
       <td class="acenter" width="9.41%"><p style="text-align:center">2</p></td> 
       <td class="acenter" width="9.43%"><p style="text-align:center">1</p></td> 
       <td class="acenter" width="16.80%"><p style="text-align:center">0.714285714</p></td> 
       <td class="acenter" width="16.80%"><p style="text-align:center">0.833333333</p></td> 
       <td class="acenter" width="16.82%"><p style="text-align:center">0.769230769</p></td> 
      </tr> 
     </table>
    </table-wrap>
   </sec>
   <sec id="s4_3">
    <title>4.3. Presentation of Results and Discussion</title>
    <p>We recorded the results for both the TP, FP, TN, TF. We computed the Precision, the Recall, and the F1-Measure and compared the precision rate as shown in <xref ref-type="table" rid="table7">
      Table 7
     </xref>. <xref ref-type="table" rid="table6">
      Table 6
     </xref> shows the number of documents failing in each category using the confusion matrix. Our experimental result is shown in <xref ref-type="fig" rid="fig5">
      Figure 5
     </xref>. The result shows the excellent performance of the present RF-CBC model over the baseline methods studied. We established that the different algorithms studied generally have a peak point, though with significant differences in performance. The F1-measure value initially increases for each algorithm studied before their respective peak points, then gradually goes down after the peak point, meaning that the precision is nearly stable and recall increases before the peak, after which the precision decreases and recall is almost stable. <xref ref-type="fig" rid="fig5">
      Figure 5
     </xref> shows the F1-Measure of the proposed RF-CBC, the TRF, NBY, ANN, and the TCL algorithms at different lengths of prediction/detection using the UNIMED official website users’ click log dataset. The result shows the superiority of the present RF-CBC model over the baseline methods studied. The experimental result shows that the TRF, NBY, ANN, and TCL techniques recorded lower F1-Measure when run on our dataset, as shown in <xref ref-type="fig" rid="fig5">
      Figure 5
     </xref>; they performed poorly compared with the RF-CBC algorithm. The RF-CBC algorithm has the highest F1-Measure when used on our dataset. The experiment was carried out for over 60 different lengths of prediction using the same datasets and experimental settings. The Naïve Bayesian, the traditional Random Forest, and the traditional clustering algorithms perform a little better at short lengths of prediction below 15, but poorly at longer and worst at lengths of prediction longer than 25; this therefore results in a limited number of positive detections. However, the ANN performs a bit better than the TRF, NBY, and TCL techniques, but also only at shorter lengths and poorly at longer lengths of prediction and worst at lengths above 25, thereby resulting in a few positive detections. The proposed RF-CBC has over 80 F1-Measures for lengths of prediction between 1 and 20, after which it maintains about 75% F1-Measure at longer lengths of prediction as shown in <xref ref-type="fig" rid="fig5">
      Figure 5
     </xref>. The experimental result shows that the RF-CBC demonstrates higher potential in detecting and reporting abnormalities on our experimental website; hence, it remains the clear winner in this case.</p>
    <fig id="fig8" position="float">
     <label>Figure 8</label>
     <caption>
      <title>Figure 5. F1-Measures of RF-CBC, TRF, NBY, TCL and ANN at different length of prediction/detections.</title>
     </caption>
     <graphic mimetype="image" position="float" xlink:type="simple" xlink:href="https://html.scirp.org/file/7801084-rId90.jpeg?20250822024427" />
    </fig>
    <p>Furthermore, we authenticate the potential of the present RF-CBC model over the baseline method by comparing their respective execution speeds in seconds. We recorded the run time of each algorithm at different prediction lengths using our experimental datasets and under the same experimental settings. The experimental results indicate that the RF-CBC has the lowest run time and executes faster than the baseline methods studied in all cases.</p>
    <p>Though the NBY and the ANN also show low runtime compared to the TRF and the TCL technique, generally, the baseline methods’ runtime increases rapidly as the prediction length increases, as shown in <xref ref-type="fig" rid="fig6">
      Figure 6
     </xref>. The proposed RF-CBC recorded a lower runtime at all lengths; therefore, it achieves better results with large datasets and longer prediction lengths. The outstanding performance of our system may be due to the adoption of multiple techniques for the detection and reporting of intruders on the experimental website, such as the introduction of the STBO feature selection technique to select the best features that eliminate noisy data, user clustering based on their different credentials and types of allowed operation by the different categories of users on the website, the use of correlation distance measurement technique, etc. All these factors have made the RF-CBC the most appropriate method for this study.</p>
    <fig id="fig9" position="float">
     <label>Figure 9</label>
     <caption>
      <title>Figure 6. Run time of RF-CBC, ANN, TRF and TCL model at different length prediction/detection.</title>
     </caption>
     <graphic mimetype="image" position="float" xlink:type="simple" xlink:href="https://html.scirp.org/file/7801084-rId91.jpeg?20250822024427" />
    </fig>
    <p>Finally, the RF-CBC model shows potential for detecting and reporting many forms of attack on any IoT device. The model is specifically suitable for domains with large volumes of data and any number of relevant attributes. The RF-CBC also shows the capability to overcome some of the challenges of many existing attack detection systems, such as noisy data, scalability, poor distance measurement functions, and computational complexity. Therefore, it establishes a flexible, scalable, transparent, accurate, faster, computationally efficient, easy to understand, and easy to implement method of attack detection and reporting systems. Our experimental results show that the RF-CBC can outperform the traditional Random Forest, Naïve Bayesian, traditional Clustering, and the Artificial Neural Network techniques in a very difficult attack detection task and in a very large dataset online and in real time consistently.</p>
   </sec>
  </sec><sec id="s5">
   <title>5. Summary, Conclusion, and Recommendation</title>
   <p>The aims of this work are to develop an efficient, faster, scalable, robust, flexible, accurate, consistent, and easy-to-use attack detection and alarming system. This is achieved through the design and implementation of a novel attack detection algorithm, commonly referred to as the Random Forest enabled correlation-based clustering (RF-CBC) algorithm. To this effect, a click log history dataset of anonymous users of the UNIMED official website, who signed into their accounts over a period of twelve months, was randomly selected, and the extracted data was pre-processed. The data was cleansed to eliminate noisy or irrelevant entries, sessions were identified, and a data mart was created. A feature scaling technique called the Single Threshold Box Plot Outlier (STBO) algorithm was developed and used to select the best features before applying the proposed RF-CBC model for attack detection and reporting purposes.</p>
   <p>The results of the experiment conducted were presented and analysed. We also carried out a performance comparison of the developed system and four other baseline methods, which are the TRF, NBY, TCL, and ANN algorithms, to demonstrate the superiority of our system over the baseline methods studied. The results of this comparison show that the present system outperformed the baseline techniques in terms of speed and accuracy. This is aimed at assisting web developers and administrators to have a wider variety of algorithms from which the best performing can be selected, and to plan updates and improvements to the security of their websites and IOT facilities through the adoption of the present system, while also assisting IOT facilities users to be more secure and protected while conducting their legitimate business online.</p>
   <p>In conclusion, this work provides a basis for IoT facilities attack detection and an alarming system. The system collects the active users’ personal information and click stream information, builds a personalized profile and access credentials that will be used to monitor and determine the user’s activities and the types of operations he or she can perform on the site. The collected click stream and profile information are matched with similar clusters in the data mart in order to determine the user’s credentials, upon which the system determines whether the user is an attacker or a normal user. The results of our experiment show that our attack detection and reporting engine, powered by the RF-CBC algorithm, is capable of producing an efficient, faster, accurate, and scalable attack detection and reporting online and in real time consistently at any time while overcoming some of the challenges of the existing method studied.</p>
   <p>We are of the opinion that the current work can be improved further and therefore recommend that other scholars in the field explore alternative machine learning techniques and compare the results with the present work, to determine the most effective ways of solving similar problems in the future. Researchers could also further investigate the UNIMED official website users’ URL TP address database or similar websites on a continuous basis, so as to be abreast of new methods of attack in the near future.</p>
  </sec>
 </body><back>
  <ref-list>
   <title>References</title>
   <ref id="scirp.145026-ref1">
    <label>1</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Luo, C., Tan, Z., Min, G., Gan, J., Shi, W. and Tian, Z. (2021) A Novel Web Attack Detection System for Internet of Things via Ensemble Classification. IEEE Transactions on Industrial Informatics, 17, 5810-5818. &gt;https://doi.org/10.1109/tii.2020.3038761
    </mixed-citation>
   </ref>
   <ref id="scirp.145026-ref2">
    <label>2</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Andročec, D. and Vrček, N. (2018) Machine Learning for the Internet of Things Security: A Systematic Review. Proceedings of the 13th International Conference on Software Technologies, Porto, 563-570. &gt;https://doi.org/10.5220/0006841205630570
    </mixed-citation>
   </ref>
   <ref id="scirp.145026-ref3">
    <label>3</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Priya, V., Sumaiya Thaseen, I., Reddy Gadekallu, T., Aboudaif, M.K. and Abouel Nasr, E. (2021) Robust Attack Detection Approach for IIoT Using Ensemble Classifier. Computers, Materials &amp; Continua, 66, 2457-2470. &gt;https://doi.org/10.32604/cmc.2021.013852
    </mixed-citation>
   </ref>
   <ref id="scirp.145026-ref4">
    <label>4</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Alotaibi, B. and Alotaibi, M. (2020) A Stacked Deep Learning Approach for IoT Cyberattack Detection. Journal of Sensors, 2020, Article ID: 8828591. &gt;https://doi.org/10.1155/2020/8828591
    </mixed-citation>
   </ref>
   <ref id="scirp.145026-ref5">
    <label>5</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Chen, F., Deng, P., Wan, J., Zhang, D., Vasilakos, A.V. and Rong, X. (2015) Data Mining for the Internet of Things: Literature Review and Challenges. International Journal of Distributed Sensor Networks, 11, 1-14. &gt;https://doi.org/10.1155/2015/431047
    </mixed-citation>
   </ref>
   <ref id="scirp.145026-ref6">
    <label>6</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Mahdavinejad, M.S., Rezvan, M., Barekatain, M., Adibi, P., Barnaghi, P. and Sheth, A.P. (2018) Machine Learning for Internet of Things Data Analysis: A Survey. Digital Communications and Networks, 4, 161-175. &gt;https://doi.org/10.1016/j.dcan.2017.10.002
    </mixed-citation>
   </ref>
   <ref id="scirp.145026-ref7">
    <label>7</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Koay, A.M.Y., Ko, R.K.L., Hettema, H. and Radke, K. (2022) Machine Learning in Industrial Control System (ICS) Security: Current Landscape, Opportunities and Challenges. Journal of Intelligent Information Systems, 60, 377-405. &gt;https://doi.org/10.1007/s10844-022-00753-1
    </mixed-citation>
   </ref>
   <ref id="scirp.145026-ref8">
    <label>8</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Isaac, S., Ayodeji, D.K., Luqman, Y., Karma, S.M. and Aminu, J. (2024) Cyber Security Attack Detection Model Using Semi Supervised Learning. FUDMA Journal of Sciences (FJS), 8, 92-100. &gt;https://doi.org/10.33003/fjs-2024-0802-2343
    </mixed-citation>
   </ref>
   <ref id="scirp.145026-ref9">
    <label>9</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Maheswari, L.C.U., Srivalli, G., Shivani, G., Nikitha, G.S. and Kaveri, K. (2023) A Novel Web Attack Detection System for Internet of Things via Ensemble Classification. Turkish Journal of Computer and Mathematics Education, 14, 834-845.
    </mixed-citation>
   </ref>
   <ref id="scirp.145026-ref10">
    <label>10</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Yavuz, F.Y., Ünal, D. and Gül, E. (2018) Deep Learning for Detection of Routing Attacks in the Internet of Things. International Journal of Computational Intelligence Systems, 12, 39-58. &gt;https://doi.org/10.2991/ijcis.2018.25905181
    </mixed-citation>
   </ref>
   <ref id="scirp.145026-ref11">
    <label>11</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Al-Sultani, Z.N. (2012) Learning Vector Quantization (LVQ) and k-Nearest Neighbor for Intrusion Classification. World of Computer Science and Information Technology Journal, 2, 105-109.
    </mixed-citation>
   </ref>
   <ref id="scirp.145026-ref12">
    <label>12</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Jawhar, M. and Mehrotra, M. (2010) Design Network Intrusion Detection System Using Hybrid Fuzzy-Neural Network. International Journal of Computer Science and Security, 4, 285-294.
    </mixed-citation>
   </ref>
   <ref id="scirp.145026-ref13">
    <label>13</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Kouassi, B.M., Monsan, V. and Adou, K.J. (2024) Intelligent Detection and Identification of Attacks in IoT Networks Based on the Combination of DNN and LSTM Methods with a Set of Classifiers. Open Journal of Applied Sciences, 14, 2296-2319. &gt;https://doi.org/10.4236/ojapps.2024.148153
    </mixed-citation>
   </ref>
   <ref id="scirp.145026-ref14">
    <label>14</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Sasi, T., Lashkari, A.H., Lu, R., Xiong, P. and Iqbal, S. (2024) A Comprehensive Survey on IoT Attacks: Taxonomy, Detection Mechanisms and Challenges. Journal of Information and Intelligence, 2, 455-513. &gt;https://doi.org/10.1016/j.jiixd.2023.12.001
    </mixed-citation>
   </ref>
   <ref id="scirp.145026-ref15">
    <label>15</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Siraparapu, S.R. and Azad, S.M.A.K. (2024) Securing the IoT Landscape: A Comprehensive Review of Secure Systems in the Digital Era. e-Prime—Advances in Electrical Engineering, Electronics and Energy, 10, Article ID: 100798. &gt;https://doi.org/10.1016/j.prime.2024.100798
    </mixed-citation>
   </ref>
   <ref id="scirp.145026-ref16">
    <label>16</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Anbarasi, M., Anupriya, E. and Iyengar, S.N. (2010) Enhanced Prediction of Heart Disease with Feature Subset Selection using Genetic Algorithm. International Journal of Engineering Science and Technology, 2, 5370-5376.
    </mixed-citation>
   </ref>
   <ref id="scirp.145026-ref17">
    <label>17</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Adeniyi, A.D., Ajoge, N.S. and Sulaiman, U.I. (2021) Design and Realization of Pre-Ordered Feature Ranking Filtering (PFRF) Feature Selection Method for Machine Learning Algorithms. International Journal of Engineering and Technology Research, 21, 82-100.
    </mixed-citation>
   </ref>
   <ref id="scirp.145026-ref18">
    <label>18</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Kulkarni, S.A., Gurupur, V.P. and King, C. (2022) Impact Analysis of Stacked Machine Learning Algorithms Based Feature Selections for Deep Learning Algorithm Applied to Regression Analysis. SoutheastCon 2022, Mobile, 26 March-3 April 2022, 269-275. &gt;https://doi.org/10.1109/southeastcon48659.2022.9764105
    </mixed-citation>
   </ref>
   <ref id="scirp.145026-ref19">
    <label>19</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Mebawondu, O.J., Adetunmbi, A.O., Mebawondu, J.O. and Alowolodu, O.D. (2021) Feature Weighting and Classification Modeling for Network Intrusion Detection Using Machine Learning Algorithms. In: Misra, S. and Muhammad-Bello, B., Eds., Information and Communication Technology and Applications, Springer, 315-327. &gt;https://doi.org/10.1007/978-3-030-69143-1_25
    </mixed-citation>
   </ref>
   <ref id="scirp.145026-ref20">
    <label>20</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Mebawondu, O.J. (2024) Enhancing Intrusion Detection Systems with Efficient Deep Learning Techniques. 2024 IEEE 5th International Conference on Electro-Computing Technologies for Humanity (NIGERCON), Ado Ekiti, 26-28 November 2024, 1-5. &gt;https://doi.org/10.1109/nigercon62786.2024.10927178
    </mixed-citation>
   </ref>
   <ref id="scirp.145026-ref21">
    <label>21</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Hegde, S.K., Hegde, R., Hombalimath, V., Palanikkumar, D., Patwari, N. and Dankan Gowda, V. (2023) Symmetrized Feature Selection with Stacked Generalization Based Machine Learning Algorithm for the Early Diagnosis of Chronic Diseases. 2023 5th International Conference on Smart Systems and Inventive Technology (ICSSIT), Tirunelveli, 23-25 January 2023, 838-844. &gt;https://doi.org/10.1109/icssit55814.2023.10061062
    </mixed-citation>
   </ref>
   <ref id="scirp.145026-ref22">
    <label>22</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Domingos, P. and Hulten, G. (2002) Learning from Infinite Data in Finite Time. In: Dietterich, T.G., Becker, S. and Ghahramani, Z., Eds., Advances in Neural Information Processing Systems 14, The MIT Press, 673-680. &gt;https://doi.org/10.7551/mitpress/1120.003.0091
    </mixed-citation>
   </ref>
   <ref id="scirp.145026-ref23">
    <label>23</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Flores, J.J., Rodriguez, H. and Graff, M. (2010) Reducing the Search Space in Evolutive Design of ARIMA and ANN Models for Time Series Prediction. In: Sidorov, G., Hernández Aguirre, A. and Reyes García, C.A., Eds., Advances in Soft Computing, Springer, 325-336. &gt;https://doi.org/10.1007/978-3-642-16773-7_28
    </mixed-citation>
   </ref>
   <ref id="scirp.145026-ref24">
    <label>24</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Sebban, M. and Nock, R. (2002) A Hybrid Filter/Wrapper Approach of Feature Selection Using Information Theory. Pattern Recognition, 35, 835-846. &gt;https://doi.org/10.1016/s0031-3203(01)00084-x
    </mixed-citation>
   </ref>
   <ref id="scirp.145026-ref25">
    <label>25</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Antonelli, D., Baralis, E., Bruno, G., Cerquitelli, T., Chiusano, S. and Mahoto, N. (2013) Analysis of Diabetic Patients through Their Examination History. Expert Systems with Applications, 40, 4672-4678. &gt;https://doi.org/10.1016/j.eswa.2013.02.006
    </mixed-citation>
   </ref>
   <ref id="scirp.145026-ref26">
    <label>26</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     García-Pedrajas, N. and de Haro-García, A. (2012) Scaling up Data Mining Algorithms: Review and Taxonomy. Progress in Artificial Intelligence, 1, 71-87. &gt;https://doi.org/10.1007/s13748-011-0004-4
    </mixed-citation>
   </ref>
   <ref id="scirp.145026-ref27">
    <label>27</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Mazarei, A., Sousa, R., Mendes-Moreira, J., Molchanov, S. and Ferreira, H.M. (2024) Online Boxplot Derived Outlier Detection. International Journal of Data Science and Analytics, 19, 83-97. &gt;https://doi.org/10.1007/s41060-024-00559-0
    </mixed-citation>
   </ref>
   <ref id="scirp.145026-ref28">
    <label>28</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Breiman, L. (2001) Random Forests. Machine Learning, 45, 5-32. &gt;https://doi.org/10.1023/a:1010933404324
    </mixed-citation>
   </ref>
   <ref id="scirp.145026-ref29">
    <label>29</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Adele Cutler, D., Cutler, R. and Stevens, J.R. (2011) Ensemble Machine Learning: Methods and Applications. Springer.
    </mixed-citation>
   </ref>
   <ref id="scirp.145026-ref30">
    <label>30</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Han, J. and Kamber, M. (2006) Data Mining Concept and Techniques. 2nd Edition, Morgan Kaufmann Publishers, 285-350.
    </mixed-citation>
   </ref>
   <ref id="scirp.145026-ref31">
    <label>31</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Aggarwal, C.C. and Reddy, C.K. (2016) Data Clustering Algorithms and Applications. Chapman and Hall. &gt;https://doi.org/10.1201/9781315373515 
    </mixed-citation>
   </ref>
   <ref id="scirp.145026-ref32">
    <label>32</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Zaid, M.A. (2015) Correlation and Regression Analysis. The Statistical, Economic and Social Research and Training Centre for Islamic Countries (SESRIC) Kudüs Cad. &gt;https://www.sesric.org 
    </mixed-citation>
   </ref>
   <ref id="scirp.145026-ref33">
    <label>33</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Nworgu, B.G. (1991) Educational Research: Basic Issues and Methodology. Wisdom Publishes Ltd.
    </mixed-citation>
   </ref>
  </ref-list>
 </back>
</article>