<?xml version="1.0" encoding="UTF-8"?><!DOCTYPE article  PUBLIC "-//NLM//DTD Journal Publishing DTD v3.0 20080202//EN" "http://dtd.nlm.nih.gov/publishing/3.0/journalpublishing3.dtd"><article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" dtd-version="3.0" xml:lang="en" article-type="research article"><front><journal-meta><journal-id journal-id-type="publisher-id">JCC</journal-id><journal-title-group><journal-title>Journal of Computer and Communications</journal-title></journal-title-group><issn pub-type="epub">2327-5219</issn><publisher><publisher-name>Scientific Research Publishing</publisher-name></publisher></journal-meta><article-meta><article-id pub-id-type="doi">10.4236/jcc.2023.112009</article-id><article-id pub-id-type="publisher-id">JCC-123358</article-id><article-categories><subj-group subj-group-type="heading"><subject>Articles</subject></subj-group><subj-group subj-group-type="Discipline-v2"><subject>Computer Science&amp;Communications</subject></subj-group></article-categories><title-group><article-title>
 
 
  Machine Learning-Based Alarms Classification and Correlation in an SDH/WDM Optical Network to Improve Network Maintenance
 
</article-title></title-group><contrib-group><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Deussom</surname><given-names>Djomadji Eric Michel</given-names></name><xref ref-type="aff" rid="aff1"><sup>1</sup></xref><xref ref-type="corresp" rid="cor1"><sup>*</sup></xref></contrib><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Takembo</surname><given-names>Ntahkie Clovis</given-names></name><xref ref-type="aff" rid="aff1"><sup>1</sup></xref></contrib><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Tchapga</surname><given-names>Tchito Christian</given-names></name><xref ref-type="aff" rid="aff1"><sup>1</sup></xref></contrib><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Arabo</surname><given-names>Mamadou</given-names></name><xref ref-type="aff" rid="aff2"><sup>2</sup></xref></contrib><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Michael</surname><given-names>Ekonde Sone</given-names></name><xref ref-type="aff" rid="aff1"><sup>1</sup></xref></contrib></contrib-group><aff id="aff2"><addr-line>Division of Information and Communications Technology, The National School of Posts and Telecommunications and Information and Communication Technologies, University of Yaoundé I, Yaoundé, Cameroon</addr-line></aff><aff id="aff1"><addr-line>Department of Electrical and Electronic Engineering, College of Technology, University of Buea, Buea, Cameroon</addr-line></aff><pub-date pub-type="epub"><day>15</day><month>02</month><year>2023</year></pub-date><volume>11</volume><issue>02</issue><fpage>122</fpage><lpage>141</lpage><history><date date-type="received"><day>26,</day>	<month>January</month>	<year>2023</year></date><date date-type="rev-recd"><day>25,</day>	<month>February</month>	<year>2023</year>	</date><date date-type="accepted"><day>28,</day>	<month>February</month>	<year>2023</year></date></history><permissions><copyright-statement>&#169; Copyright  2014 by authors and Scientific Research Publishing Inc. </copyright-statement><copyright-year>2014</copyright-year><license><license-p>This work is licensed under the Creative Commons Attribution International License (CC BY). http://creativecommons.org/licenses/by/4.0/</license-p></license></permissions><abstract><p>
 
 
  The evolution of telecommunications has allowed the development of broadband services based mainly on fiber optic backbone networks. The operation and maintenance of these optical networks is made possible by using supervision platforms that generate alarms that can be archived in the form of log files. But analyzing the alarms in the log files is a laborious and difficult task for the engineers who need a degree of expertise. Identifying failures and their root cause can be time consuming and impact the quality of service, network availability and service level agreements signed between the operator and its customers. Therefore, it is more than important to study the different possibilities of alarms classification and to use machine learning algorithms for alarms correlation in order to quickly determine the root causes of problems faster. We conducted a research case study on one of the operators in Cameroon who held an optical backbone based on SDH and WDM technologies with data collected from 2016-03-28 to “2022-09-01” with 7201 rows and 18. In this paper, we will classify alarms according to different criteria and use 02 unsupervised learning algorithms namely the K-Means algorithm and the DBSCAN to establish correlations between alarms in order to identify root causes of problems and reduce the time to troubleshoot. To achieve this objective, log files were exploited in order to obtain the root causes of the alarms, and then K-Means algorithm and the DBSCAN were used firstly to evaluate their performance and their capability to identify the root cause of alarms in optical network.
 
</p></abstract><kwd-group><kwd>Optical Network</kwd><kwd> Alarms</kwd><kwd> Log Files</kwd><kwd> Root Cause Analysis</kwd><kwd> Machine Learning</kwd></kwd-group></article-meta></front><body><sec id="s1"><title>1. Introduction</title><p>The evolution of telecommunication networks, in particular optical networks based on Synchronous Digital Hierarchy (SDH) and Wavelength Division Multiplexing (WDM) technologies [<xref ref-type="bibr" rid="scirp.123358-ref1">1</xref>] [<xref ref-type="bibr" rid="scirp.123358-ref2">2</xref>] , has marked a major transformation of the world by enabling the massive sharing of information at high speed across the globe. Over the past decades, this hyper-connectivity has become deeply embedded in our daily lives and in most areas of activity. As a result, telecom operators must be able to ensure the smooth operation of these networks. To achieve this, it is essential to be able to efficiently identify the failures that occur and their origin.</p><p>Root cause analysis consists in finding the primary cause of a failure amongst a multitude of errors. Due to the complexity of networks, both in size and in technology, finding the root cause of a failure remains a difficult task. Today, the emergence of artificial intelligence with applications such as machine learning allows for a better root cause analysis. This can be possible by analyzing alarms logs. But to be able to implement these methods and end up with a maintenance aid tool, it is necessary to acquire a huge amount of data on the operating status of the optical network equipment. The acquisition of data will allow to train the machine learning algorithms [<xref ref-type="bibr" rid="scirp.123358-ref3">3</xref>] [<xref ref-type="bibr" rid="scirp.123358-ref4">4</xref>] [<xref ref-type="bibr" rid="scirp.123358-ref5">5</xref>] and to increase their performances.</p><p>It is within this framework that this article aims to implement a machine learning model that will facilitate the analysis of data from the log files of the optical network equipment from the Operation and Maintenance Centre (OMC) supervision servers of an operator. It will be a matter of recording the data extracted from the server, then performing a classification and a correlation of the alarms and determining the root causes. In fact, network maintenance almost always relies on alarm monitoring and management to identify existing problems and provide solutions. The proposed solutions must be implemented as quickly as possible in order to reduce the time of degradation or interruption of services. For optical networks, operators in the field sign Service Level Agreements (SLAs) with their customers and these SLAs must be respected otherwise financial compensation must be made by the operator to its customers. It is therefore important to find new approaches to help the rapid identification of network failures or various defects. Therefore, alarm classification and machine learning tools [<xref ref-type="bibr" rid="scirp.123358-ref6">6</xref>] [<xref ref-type="bibr" rid="scirp.123358-ref7">7</xref>] [<xref ref-type="bibr" rid="scirp.123358-ref8">8</xref>] are very useful to optimize network availability and increase the overall quality of service offered to customers. Several authors have focused on alarm classification approaches or on the use of machine learning to solve problems in networks [<xref ref-type="bibr" rid="scirp.123358-ref9">9</xref>] [<xref ref-type="bibr" rid="scirp.123358-ref10">10</xref>] [<xref ref-type="bibr" rid="scirp.123358-ref11">11</xref>] . We have for example E. DEUSSOM et al in ref. [<xref ref-type="bibr" rid="scirp.123358-ref9">9</xref>] who worked on fraud detection in mobile networks by analysing Call Data Records (CDRs) and traffic; moreover, B. BATCHAKUI et al in ref. [<xref ref-type="bibr" rid="scirp.123358-ref10">10</xref>] worked on the comparison of machine learning algorithms to improve the maintenance of Long Term Evolution (LTE) Time Division Duplexing (TDD) networks. M. Klemettinen, H. Mannila, and H. Toivonen in ref. [<xref ref-type="bibr" rid="scirp.123358-ref11">11</xref>] proposed a method for discovering recurring patterns of alarms in databases using a correlation system. They also present a tool with which network management experts can browse the large amounts of rules produced.</p><p>T. White and N. Ross in ref. [<xref ref-type="bibr" rid="scirp.123358-ref12">12</xref>] proposed architecture for an alarm correlation engine. They design a correlation engine directly into the Network Management System (NMS). The idea is to improve some aspects of the NMS to reduce the observed raw flow of alarms by sending only the most relevant information. Some methods define an alarm architecture using the Model-Based Reasoning principle. These methods introduce rule-based approaches to group alarms that refer to the same problem.</p><p>A. Bouillard, A. Junier, B. Ronot in ref. [<xref ref-type="bibr" rid="scirp.123358-ref13">13</xref>] proposed a method that creates several alarm dependency graphs, based on the sequence of alarm names, which allows a quick study of the alarm correlation problem through a powerful heuristic algorithm.</p><p>N. Amani, M. Fathi, M. Dehghan in ref. [<xref ref-type="bibr" rid="scirp.123358-ref14">14</xref>] proposed a new Case-Based Reasoning (CBR) method for alarm correlation in telecommunication networks. The proposed method has been simulated by developing three main modules: a module for generating faults and alarms, defining the network configuration, and filtering and correlating alarms using the CBR. One of the most important aspects of the results obtained was the speed of the system.</p><p>In this work, we conducted a research case study on one of the operators in Cameroon who hold an optical backbone based on SDH and WDM technology with data collected from 2016-03-28 to “2022-09-01” with 7201 rows and 18. The method used in this study uses machine learning with unsupervised learning techniques which are quite popular methods used in alarms classification and correlation. The purpose of this study is to propose an effective method that can be applied to detect root cause of problems based on alarm logs analysis. The rest of the paper is organized as follow: Materials and methods will be presented in the second section. The third section deals with the experimentation carried out and the results obtained; the fourth section presents the discussion of the results obtained. Finally, a general conclusion and perspectives are presented.</p></sec><sec id="s2"><title>2. Materials and Methods Used</title><p>In the context of the present work, the research was done by using real data collected on an optical network management platform in Cameroon using Huawei Network Cloud Engine management platform (NCE-Transmission). The Topology management and alarm management feature of the NCE-Tx was selected to monitor the network elements, network alarms and process the alarms based on their severity, types, sources, and impact on the network (see <xref ref-type="fig" rid="fig1">Figure 1</xref> and <xref ref-type="fig" rid="fig2">Figure 2</xref> which present the appearance the network entities through the topology management and the alarms on the alarm management platform of the NCE-Tx). Hence, the target data to be predicted are discrete data. Algorithms used for the classification of the collected data, as well as for the prediction of our target are presented and reviewed below: K-Means and DBSCAN.</p><p>In the context of this work, the target data to be exploited are unlabeled data. Thus, we will present the algorithms used for the classification of our data. These are the K-Means and DBSCAN algorithms.</p><sec id="s2_1"><title>2.1. K-Means Algorithm</title><p>K-means is an unsupervised non-hierarchical clustering algorithm. It allows to group in K distinct clusters the observations of the dataset. Thus, similar data will be found in the same cluster. Moreover, an observation can only be found in one cluster at a time (exclusivity of membership). The same observation cannot belong to two different clusters [<xref ref-type="bibr" rid="scirp.123358-ref15">15</xref>] . In order to group a dataset into K distinct clusters, the K-Means algorithm needs a way to compare the degree of similarity between the different observations. Thus, two data that are similar will have a reduced dissimilarity distance, while two different objects will have a larger separation distance. The K-means method is quite simple, starting with the selection of the number of clusters as many as K pieces, and then K pieces of data are randomly taken from the data set as a centroid to represent a cluster. All data are then calculated at the distance of the centroid and each data will be a member of a cluster represented by a centroid that has the closest distance to the data. Finally, the re-calculation of the centroid value obtained from the average value of each cluster [<xref ref-type="bibr" rid="scirp.123358-ref16">16</xref>] .</p></sec><sec id="s2_2"><title>2.2. DBSCAN Algorithm</title><p>DBSCAN (Density-Based Spatial Clustering of Applications with Noise) is a well-known unsupervised algorithm in clustering. DBSCAN is a simple algorithm that defines clusters using local density estimation [<xref ref-type="bibr" rid="scirp.123358-ref17">17</xref>] .</p><p>In order to train the system to be able to determine the different clusters and to propose the most optimal unsupervised classification, we followed an approach borrowed from “data science”. Indeed, from a set of data (dataset) collected from the databases of the operation and maintenance center of a cell phone operator during a period, we have executed on the said data, the machine learning algorithms presented in the previous paragraphs, in order to select the algorithm that will produce the most expected results. The approach consists of:</p><p>&#183; Processing the dataset;</p><p>&#183; Normalization/coding of the variables;</p><p>&#183; Determination of the model and its parameters;</p><p>&#183; Learning;</p><p>&#183; Test and interpretation of the results.</p></sec><sec id="s2_3"><title>2.3. Environment and Tools Used</title><p>To run the machine learning algorithms, we used the Python programming language, in particular version 3.8. It is the reference language used in the development of applications for artificial intelligence. It is easy to install, uncompiled, fast and light. The Python distribution used is Anaconda: it contains all the tools and libraries we need to do machine learning: Numpy, Matplotlib, Sklearn, Jupiter, Spider…etc.</p><sec id="s2_3_1"><title>2.3.1. Dataset Processing</title><p>The dataset is the set of examples that the machine must study. Our data were collected directly from the operator’s optical network monitoring platform system. Here we will focus on the alarm log file, whose entries are grouped into a single file. This is a plain text file in which each entry is recorded on one and only one line. The content of each line is however organized according to a nomenclature that can be configured by the user. Studying this nomenclature before the automatic analysis of the logs makes it possible to gain in efficiency (because the information is easily identifiable) and in time (because the analysis is then done on smaller chains). This information is among others the severity of the alarm, its identifier, the name of the failures, the location of the failure, the number of occurrences, etc… <xref ref-type="fig" rid="fig3">Figure 3</xref> is an extract of the data constituting our dataset.</p><p>We have a database with 7201 rows and 18 columns, 15 of which are of Type Object and 3 of Type Integer (Int64). The following python code gives the list of columns of the dataset as presented in <xref ref-type="fig" rid="fig4">Figure 4</xref>.</p><p>Before proceeding to the coding of the clustering algorithm of the solution to be applied with the help of the log files, we carried out some analyses with the tools offered by the ANACONDA environment, through its libraries Numpy and Matplotlib. These analyses allowed us to determine on the one hand, the columns or features to use, and the degree of correlation between the ALARM_ID parameter and the other parameters of the database. <xref ref-type="fig" rid="fig5">Figure 5</xref> presents the code used for the transformation of the values of the column “First occurred” in ms.</p><p>We transformed the data in order to make them more usable. These are:</p><p>o Transformation of the column “first occurred (ST)” into ms.</p><p>&#183; Assign the IDs to the elements of the “Location Info” column, from 0 to 4761;</p><p>&#183; Assign the IDs to the elements of the “Alarm Source” column, from 0 to 282;</p><p>&#183; Assign the IDs to the elements of the “Severity” column from 0 to 3.</p><p>The analysis of correlations through the heatmap (<xref ref-type="fig" rid="fig3">Figure 3</xref>) allows to visually representing correlations (or relationships) between variables.</p><p>sns.heatmap(corr, cmap = “RdBu”, vmin = −1, vmax = 1, annot = True)</p><p>From <xref ref-type="fig" rid="fig6">Figure 6</xref> below, we can observe weak correlations between Alarm ID and the other parameters:</p><p>o Number of occurrences;</p><p>o Date of the first occurrence;</p><p>o Location ID;</p><p>o Source ID.</p><p>We can see that the level of severity is not linked to other parameters. We will now classify these different alarms.</p></sec><sec id="s2_3_2"><title>2.3.2. Classification of Alarms According to Various Criteria</title><p><xref ref-type="fig" rid="fig7">Figure 7</xref> below presents the data header after exporting the logs.</p><p>v Exploration of the data</p><p>The data use for this paper are those collected from Huawei NCE Tx, it was collected from “2016-03-28 12:24:20” to “2022-09-01 12:22:58”.</p><p>v Alarm sources</p><p>The image below in <xref ref-type="fig" rid="fig8">Figure 8</xref> gives the top 10 sources of alarms; this information helps the maintenance engineer in the search for the root causes of the problems.</p><p>v Alarm severity levels</p><p><xref ref-type="fig" rid="fig9">Figure 9</xref> presents the command used for classification of alarms based on their severity. In the operation of the network, the knowledge of the severity level allows to define the priorities during the resolution of the problems. The critical severity problems strongly impact the services making them unavailable and are the first ones to be addressed for resolution. We have 4 levels of severity for alarms, they are the following:</p><p>• Minor;</p><p>• Critical;</p><p>• Major;</p><p>• Warning.</p><p>The knowledge on the number of occurrence of the alarms is also a form of classification of the alarms, the methods used for the resolutions of the most frequent alarms can be reused, moreover in terms of preventive maintenance; we can anticipate and prevent certain faults from occurring. <xref ref-type="fig" rid="fig1">Figure 1</xref>0 presents the top 10 recurring alarms.</p><p>v The most recurrent alarms in the dataset</p><p>By analyzing the most recurrent alarms in the dataset, it is possible to identify some common points of failure and find the root cause of a problem, it can be fiber cut, power problem or others issues. <xref ref-type="fig" rid="fig1">Figure 1</xref>0 presents the Python commands used to classify alarms with respect to their occurrence number while <xref ref-type="fig" rid="fig1">Figure 1</xref>1 presents a graphic with alarms type (name) with respect to the occurrence number.</p><p>v The top 10 most recurrent alarms in our dataset</p><p>The knowledge of the most recurrent alarms guides the maintenance engineer in the search for the root causes in the network.</p><p>v Top 10 critical alarms</p><p>v Top 10 major alarms</p><p>Note: It can be seen that the top 10 major alarms are similar to the top 10 critical alarms.</p><p><xref ref-type="fig" rid="fig1">Figure 1</xref>2 and <xref ref-type="fig" rid="fig1">Figure 1</xref>3 present occurrence of each alarm by severity degree, <xref ref-type="fig" rid="fig1">Figure 1</xref>2 is related to alarms with a critical severity while <xref ref-type="fig" rid="fig1">Figure 1</xref>3 presents alarms with major as alarm severity. From the previous paragraph, we have presented different ways of classifying alarms, depending on the source of the alarm, the level of severity of the alarm, the number of occurrences of the alarm and the geographical area affected by the alarm. It is now important to use two machine learning algorithms to establish correlations between these alarms with the ultimate goal of quickly identifying the root causes when problems occur and providing solutions in order to respect the Services Level Agreement between the optical network owner and the various customers.</p></sec><sec id="s2_3_3"><title>2.3.3. Learning and Creation of Prediction Models</title><p>After importing and cleaning the dataset, we need to start applying the different clustering algorithms for training. The data to be classified being unlabelled data, we are therefore facing a situation of Unsupervised Learning. Also, the algorithms retained for the research of our model are constituted of K-Means and DBSCAN.</p><p>Clustering with the K-Means algorithm</p><p>We did an 8 and 5 features approach to the dataset to find the underlying causes:</p><p>o “Alarm ID”,</p><p>o “Occurrences”,</p><p>o “First Occured in ms”,</p><p>o “Last Occured in ms”,</p><p>o “Location ID”,</p><p>o “Source ID”,</p><p>o “Severity ID”,</p><p>o “Cleared status”.</p><p>Choice of the number of clusters</p><p>To determine the optimal number of clusters, we use the “elbow” method.</p><p>This method reveals that there could be 5 underlying behaviors in the 8-feature datasets. <xref ref-type="fig" rid="fig1">Figure 1</xref>4 presents the result obtain after implementing the “elbow” function on the dataset and we get the coordinates of the cluster centers this is presented in <xref ref-type="fig" rid="fig1">Figure 1</xref>5.</p><p>The second method reveals that there could be 4 underlying behaviors in the 5-feature datasets (see <xref ref-type="fig" rid="fig1">Figure 1</xref>6) and we get the coordinates of the cluster centers which are presented in <xref ref-type="fig" rid="fig1">Figure 1</xref>7.</p><p>Clustering with the DBSCAN algorithm</p><p>We have this model with the configurations summarized in <xref ref-type="table" rid="table1">Table 1</xref>.</p><p>By using the python code, we obtained the result presented in <xref ref-type="table" rid="table2">Table 2</xref>.</p><table-wrap id="table1" ><label><xref ref-type="table" rid="table1">Table 1</xref></label><caption><title> Values to use for the model implementation</title></caption><table><tbody><thead><tr><th align="center" valign="middle" >Epsilon</th><th align="center" valign="middle" >Minpts</th></tr></thead><tr><td align="center" valign="middle" >1000</td><td align="center" valign="middle" >100</td></tr></tbody></table></table-wrap><table-wrap id="table2" ><label><xref ref-type="table" rid="table2">Table 2</xref></label><caption><title> Model results</title></caption><table><tbody><thead><tr><th align="center" valign="middle" >Number of data considered as noise</th><th align="center" valign="middle" >Number of clusters</th></tr></thead><tr><td align="center" valign="middle" >113</td><td align="center" valign="middle" >3</td></tr></tbody></table></table-wrap><p>To evaluate our two models, we will use the silhouette score. The silhouette coefficient or silhouette score is a measure used to calculate the quality of a clustering technique. Its value is between −1 and 1. To do this, we execute the python codes presented in <xref ref-type="fig" rid="fig1">Figure 1</xref>8 below.</p><p>It appears from this code that the K-Means score is 0.92 and the DBSCAN score is 0.87.</p></sec></sec></sec><sec id="s3"><title>3. Results and Discussion</title><sec id="s3_1"><title>3.1. Case of the K-Means Algorithm</title><p>After the implementation of the K-Means algorithm on the NCE-Tx dataset, we obtained the results presented below. It will be divided into two parts part A and part B (depending on the number of features used).</p><sec id="s3_1_1"><title>3.1.1. Part A—Approach with 8 Features</title><p>v Choice of the clusters number</p><p>To determine the optimal number of possible clusters, we opted for the elbow method; the result is presented in <xref ref-type="fig" rid="fig1">Figure 1</xref>4. <xref ref-type="fig" rid="fig1">Figure 1</xref>9 presents the coordinates of centroids, we have a total of 5. This method tells us that there could be 5 underlying behaviors in 8-featured datasets. <xref ref-type="fig" rid="fig2">Figure 2</xref>0 gives the coordinates of centroids.</p><p>v Behavior of centroids in the plane formed by Alarm ID and Location ID</p><p>Interpretation:</p><p>From <xref ref-type="fig" rid="fig2">Figure 2</xref>0, we found five (05) clusters grouping the different alarms of the optical network with similar behavior. Thus:</p><p>o The red color for cluster 1;</p><p>o Blue color for cluster 2;</p><p>o Cyan color for cluster 3;</p><p>o Green color for cluster 4;</p><p>o Yellow color for cluster 5;</p><p>o The black color for our centroids.</p><p>And we also observe, that there are noises constituting the elements of our cluster 1.</p><p>v Behavior of centroids in the plane formed by Alarm ID and Source ID</p><p>For the sake of representation in <xref ref-type="fig" rid="fig2">Figure 2</xref>1, we have limited ourselves to the Alarm IDs from 0 to 300. In this representation, we observe only one noise in the cluster distribution.</p><p>v Visualization</p><p>Interpretation: After observing <xref ref-type="fig" rid="fig2">Figure 2</xref>2, we can see that Location ID 70 which corresponds to “5-PQ1-40 (Prepaid Huawei)-PPI:1”, is strongly affected by Alarm IDs: 107 and 201, we can see it in <xref ref-type="fig" rid="fig2">Figure 2</xref>3, column 3.</p><p>So these behaviors would come from either: Cause 0 called cluster 0 or Cause 3 called cluster 3.</p><p>Interpretation: from <xref ref-type="fig" rid="fig2">Figure 2</xref>4 according to our distribution,</p><p>- We can see that the cause or cluster 0 which occurs frequently is responsible for almost all the alarms recorded.</p><p>- Considering the source ID equal to 70 which corresponds to “11-5-Limbe central”, the alarms of this source are mostly due to the cause/cluster 1.</p><p>Comments: By studying closely the behavior using these 8 features, we can see that the algorithm has clustered the alarms according to time. For example, cluster 0 takes into account the events that occurred from 2022.</p></sec><sec id="s3_1_2"><title>3.1.2. Part B—Approach with 5 Features</title><p>In this part, we will use the following 5 features:</p><p>o “Alarm ID”,</p><p>o “Location ID”,</p><p>o “Source ID”,</p><p>o “Severity ID”,</p><p>o “Occurrences”.</p><p>To determine the optimal number of possible clusters, we opted for the elbow method; we have the result presented in <xref ref-type="fig" rid="fig1">Figure 1</xref>6. This method reveals to us that there could be 4 underlying behaviors in 5-feature datasets, see <xref ref-type="fig" rid="fig2">Figure 2</xref>6 and <xref ref-type="fig" rid="fig2">Figure 2</xref>7. <xref ref-type="fig" rid="fig2">Figure 2</xref>5 gives the centroids coordinates.</p><p>Interpretation: From <xref ref-type="fig" rid="fig2">Figure 2</xref>7, with this model the behaviors around location ID 70 are all from cluster 0. The source ID 70 refers to the network entity of Limbe central.</p><p>Interpretation: By analyzing <xref ref-type="fig" rid="fig2">Figure 2</xref>8 and by considering the source ID equal to 70 which corresponds to “11-5-Limbe Central”, which in fact is the name of the network entity OSN installed in Limbe Central equipment room, the alarms of this source have the same behavior.</p></sec></sec><sec id="s3_2"><title>3.2. Case of the DBSCAN Algorithm</title><p>After implementing the DBSCAN algorithm on the data set from the NCE-Tx, we obtained the results presented below.</p><p>The optimal ε is equal to 1000 for a better partitioning of our dataset. This can be seen in the <xref ref-type="fig" rid="fig2">Figure 2</xref>9 below using the Scikit-Learn library.</p><p>v Visualization</p><p>Interpretation: For the sake of representation in <xref ref-type="fig" rid="fig3">Figure 3</xref>0, we have limited ourselves to Alarm IDs from 0 to 1000. In this representation, we can see for example that the alarms of the source ID 90 all have the same behavior or are all from the same cluster.</p><p>At the end of the experimentation that we have carried out in the previous paragraphs, the machine learning model that we propose is the K-Means model. Indeed, this model is retained because its quality of clustering technique is superior to that of DBSCAN. <xref ref-type="table" rid="table3">Table 3</xref> presents the results of the evaluation of the algorithms.</p><table-wrap id="table3" ><label><xref ref-type="table" rid="table3">Table 3</xref></label><caption><title> Results of the evaluation of the algorithms</title></caption><table><tbody><thead><tr><th align="center" valign="middle" >Points de comparaison</th><th align="center" valign="middle" >K-means</th><th align="center" valign="middle" >DBSCAN</th></tr></thead><tr><td align="center" valign="middle" >Nombre de clusters</td><td align="center" valign="middle" >4</td><td align="center" valign="middle" >3</td></tr><tr><td align="center" valign="middle"  rowspan="2"  >Variable</td><td align="center" valign="middle"  rowspan="2"  >K = 4</td><td align="center" valign="middle" >Epsilon = 1000</td></tr><tr><td align="center" valign="middle" >Minpts = 100</td></tr><tr><td align="center" valign="middle" >Nombre de donn&#233;es bruits</td><td align="center" valign="middle" >/</td><td align="center" valign="middle" >113</td></tr><tr><td align="center" valign="middle" >Coefficient de silhouette</td><td align="center" valign="middle" >0.92</td><td align="center" valign="middle" >0.87</td></tr></tbody></table></table-wrap><p>For x = 5 features, x being the number of features selected in the log file.</p><p>From the classification done using K-Means:</p><p>&#183; Cluster 0 is close to the SWDL alarms because it is the alarm that is closest to the first centroid.</p><p>&#183; Cluster 1 groups the alarms around the software (expired license, end of free period) because it is the alarm that is closest to the second centroid.</p><p>&#183; Cluster 2 is close to the PROTECTION_SUBNET_RISK alarm because it is the alarm closest to the third centroid.</p><p>&#183; Cluster 3 is close to the PORT_MODULE_OFFLINE, ALM_GFP, LASER_ MOD_ERR_EX, LCAS_PLCT, LCAS_TLCT, LCAS_PLCR, LCAS_FOPR alarms because they are the ones closest to the fourth centroid.</p><p>In view of these results, we can deduce the different root causes of the alarms grouped in different clusters using the maintenance guide of the vendor who in this case is Huawei.</p><p>The results obtained from our experimentation allow us to determine the root causes of the alarm clusters through the unsupervised classification of the optical network alarms. The unsupervised system proposed solutions for pre-processing the alarms to allow the network administrator to determine the causes in a log file. The work carried out and explored in the state of the art chapter shows us that the exploitation of AI in the maintenance of telecommunication networks has brought great efficiency. However, these works only facilitate the analysis of a log file and do not give the causes of the alarms. With the advent of expert systems, the problem has begun to be addressed, but has been hampered by the fact that these systems become obsolete very quickly in the face of a dynamic and extensive environment. The methodical learning offered by machine learning models therefore brings a step forward towards the intelligent processing of mobile network maintenance work.</p></sec></sec><sec id="s4"><title>4. Conclusions</title><p>The present work focused on the classification and correlation of optical network alarms based on Machine Learning algorithms. We applied two Unsupervised Algorithms, on log files of alarms from an optical network operator supervision platform in order to be able to determine the root causes of failures. Our goal was to obtain a machine learning model that will facilitate the analysis of alarm data from the optical layer supervision’s platform, record this data and obtain the root causes. This is based on the real alarms of the studied network.</p><p>To achieve this goal we have shown the development process of two unsupervised classification algorithms to group alarms from the information present on the log files of the optical network supervision platform.</p><p>The results obtained show the interest of the proposed approach and its capacity to facilitate the analysis of log files; this with the aim of obtaining the root causes. These results are therefore globally satisfactory. We propose as perspectives to add a functionality of failure prediction using a supervised model, to implement this model on the network and add in vendor management platform machine learning based modules even if they are licenses based to help network operators to easily detects root causes of problems and solve them.</p></sec><sec id="s5"><title>Conflicts of Interest</title><p>The authors declare no conflicts of interest regarding the publication of this paper.</p></sec><sec id="s6"><title>Cite this paper</title><p>Michel, D.D.E., Clovis, T.N., Christian, T.T., Mamadou, A. and Sone, M.E. (2023) Machine Learning-Based Alarms Classification and Correlation in an SDH/WDM Optical Network to Improve Network Maintenance. Journal of Computer and Communications, 11, 122-141. https://doi.org/10.4236/jcc.2023.112009</p></sec></body><back><ref-list><title>References</title><ref id="scirp.123358-ref1"><label>1</label><mixed-citation publication-type="other" xlink:type="simple">Baraketi, S. (2015) Ingénierie des réseaux optiques SDH et WDM et étude multicouche IP/MPLS sur OTN sur DWDM. Réseaux et télécommunicationsm. Universite Toulouse III Paul Sabatier, Franais.</mixed-citation></ref><ref id="scirp.123358-ref2"><label>2</label><mixed-citation publication-type="other" xlink:type="simple">Marwa, M.L. and Samira, M.M. (2017) Etude des Reseaux D’Acces Optique Exploitant le Multiplexage en Longueurs D’onde.</mixed-citation></ref><ref id="scirp.123358-ref3"><label>3</label><mixed-citation publication-type="other" xlink:type="simple">Expert. https://www.expert.ai/blog/machine-learning-definition</mixed-citation></ref><ref id="scirp.123358-ref4"><label>4</label><mixed-citation publication-type="other" xlink:type="simple">Talend. https://www.talend.com/fr/resources/what-is-machine-learning</mixed-citation></ref><ref id="scirp.123358-ref5"><label>5</label><mixed-citation publication-type="other" xlink:type="simple">Müller, A.C. and Guido, S. (2016) Introduction to Machine Learning with Python: A Guide for Data Scientists. O’Reilly Media Inc., Sebastopol, 1-319.</mixed-citation></ref><ref id="scirp.123358-ref6"><label>6</label><mixed-citation publication-type="other" xlink:type="simple">Berthier, E. (2019) Une Introduction au Machine Learning.</mixed-citation></ref><ref id="scirp.123358-ref7"><label>7</label><mixed-citation publication-type="other" xlink:type="simple">Chloé-Agathe (2019) Introduction au Machine Learning. 9, 10.</mixed-citation></ref><ref id="scirp.123358-ref8"><label>8</label><mixed-citation publication-type="other" xlink:type="simple">Teahouse. https://teahouse.fifty-five.com/fr/petit-guide-du-machine-learning-partie-4-lapprentissage-par-renforcement</mixed-citation></ref><ref id="scirp.123358-ref9"><label>9</label><mixed-citation publication-type="other" xlink:type="simple">Deussom, E., Matemtsap, M.B., Tchagna, K.A., et al. (2022) Machine Learning-Based Approach for Designing and Implementing a Collaborative Fraud Detection Model through CDR and Traffic Analysis. Transactions on Machine Learning and Artificial Intelligence, 10, 46-58.</mixed-citation></ref><ref id="scirp.123358-ref10"><label>10</label><mixed-citation publication-type="other" xlink:type="simple">Bernabe, B., et al. (2022) Comparing Machine Learning Algorithms for Improving the Maintenance of LTE Networks Based on Alarms Analysis. Journal of Computer and Communications, 10, 125-137. https://doi.org/10.4236/jcc.2022.1012010</mixed-citation></ref><ref id="scirp.123358-ref11"><label>11</label><mixed-citation publication-type="other" xlink:type="simple">Klemettinen, M., Mannila, H. and Toivonen, H. (1999) Rule Discovery in Telecommunication Alarm Data. Network and Systems Management Journal, 7, 395-423. https://doi.org/10.1023/A:1018787815779</mixed-citation></ref><ref id="scirp.123358-ref12"><label>12</label><mixed-citation publication-type="other" xlink:type="simple">White, T. and Ross, N. (1996) Fault Diagnosis and Network Entities in a Next Generation Network Management System. In Conference Reports: Expert Systems Applications in Artificial Intelligence, Paris, 1996, 517-522.</mixed-citation></ref><ref id="scirp.123358-ref13"><label>13</label><mixed-citation publication-type="other" xlink:type="simple">Bouillard, A. and Ronot, B. (2013) Alarms Correlation in Telecommunication Networks. https://hal.inria.fr/hal-00838969</mixed-citation></ref><ref id="scirp.123358-ref14"><label>14</label><mixed-citation publication-type="other" xlink:type="simple">Amani, N., et al. (2005) A Case-Based Reasoning Method for Alarm Filtering and Correlation in Telecommunication Networks. Canadian Conference on Electrical and Computer Engineering, Saskatoon, 1-4 May 2005, 2182-2186. https://ieeexplore.ieee.org/document/1557421</mixed-citation></ref><ref id="scirp.123358-ref15"><label>15</label><mixed-citation publication-type="other" xlink:type="simple">Benzaki, Y. (2018) Tout ce que vous voulez savoir sur l’algorithme K-Means. https://mrmint.fr/algorithme-k-means</mixed-citation></ref><ref id="scirp.123358-ref16"><label>16</label><mixed-citation publication-type="other" xlink:type="simple">Adinugroho, S. and Sari, Y.A. (2018) Implementasi Data Mining Menggunakan Weka. Universitas Brawijaya Press, Kota Malang.</mixed-citation></ref><ref id="scirp.123358-ref17"><label>17</label><mixed-citation publication-type="other" xlink:type="simple">Machine Learning &amp; Clustering: Focus sur l’Algorithme DBSCAN. https://datascientest.com/machine-learning-clustering-dbscan</mixed-citation></ref></ref-list></back></article>