<?xml version="1.0" encoding="UTF-8"?><!DOCTYPE article  PUBLIC "-//NLM//DTD Journal Publishing DTD v3.0 20080202//EN" "http://dtd.nlm.nih.gov/publishing/3.0/journalpublishing3.dtd"><article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" dtd-version="3.0" xml:lang="en" article-type="research article"><front><journal-meta><journal-id journal-id-type="publisher-id">OJOp</journal-id><journal-title-group><journal-title>Open Journal of Optimization</journal-title></journal-title-group><issn pub-type="epub">2325-7105</issn><publisher><publisher-name>Scientific Research Publishing</publisher-name></publisher></journal-meta><article-meta><article-id pub-id-type="doi">10.4236/ojop.2021.103006</article-id><article-id pub-id-type="publisher-id">OJOp-111650</article-id><article-categories><subj-group subj-group-type="heading"><subject>Articles</subject></subj-group><subj-group subj-group-type="Discipline-v2"><subject>Computer Science&amp;Communications</subject><subject> Engineering</subject><subject> Physics&amp;Mathematics</subject></subj-group></article-categories><title-group><article-title>
 
 
  Model for Anticipating Failures by Omission in Calculation Grids
 
</article-title></title-group><contrib-group><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Ramadane</surname><given-names>Adamou Yougouda</given-names></name><xref ref-type="aff" rid="aff1"><sup>1</sup></xref><xref ref-type="corresp" rid="cor1"><sup>*</sup></xref></contrib><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Marcellin</surname><given-names>Nkenlifack</given-names></name><xref ref-type="aff" rid="aff2"><sup>2</sup></xref></contrib><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Vivient</surname><given-names>Corneille Kamla</given-names></name><xref ref-type="aff" rid="aff1"><sup>1</sup></xref></contrib><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Laurent</surname><given-names>Bitjoka</given-names></name><xref ref-type="aff" rid="aff1"><sup>1</sup></xref></contrib></contrib-group><aff id="aff1"><addr-line>ENSAI, University of Ngaoundere, Ngaoundere, Cameroon</addr-line></aff><aff id="aff2"><addr-line>URIFIA, Department of Mathematics and Computer Science, Faculty of Science, University of Dschang, Dschang, Cameroon</addr-line></aff><pub-date pub-type="epub"><day>12</day><month>08</month><year>2021</year></pub-date><volume>10</volume><issue>03</issue><fpage>71</fpage><lpage>87</lpage><history><date date-type="received"><day>15,</day>	<month>June</month>	<year>2021</year></date><date date-type="rev-recd"><day>29,</day>	<month>August</month>	<year>2021</year>	</date><date date-type="accepted"><day>1,</day>	<month>September</month>	<year>2021</year></date></history><permissions><copyright-statement>&#169; Copyright  2014 by authors and Scientific Research Publishing Inc. </copyright-statement><copyright-year>2014</copyright-year><license><license-p>This work is licensed under the Creative Commons Attribution International License (CC BY). http://creativecommons.org/licenses/by/4.0/</license-p></license></permissions><abstract><p>
 
 
  Computer grids are infrastructures in which heterogeneous and distributed resources offer very high computing or storage performance. If they offer extreme computing performance, they are also subject to the appearance of many failures related to this type of architecture. While performing tasks, if the response time of a node in the system incomprehensibly exceeds the requirements of the specifications, the node experiences an omission failure. The task running in the failed node will be unavailable until the node resumes normal activity. Waiting not being a possible solution, many fault tolerance methods have been proposed. Despite this large number of fault tolerance methods on offer, computer grids are still prone to many failures by omission. In this work, a numerical study of the failures by omission which occur in the calculation grids during the execution of the tasks was carried out and a model allowing anticipating its failures was proposed with the formalism PDEVS (Parallel Discret EVent system Specification).
 
</p></abstract><kwd-group><kwd>Calculation Grids</kwd><kwd> Fault Tolerance</kwd><kwd> Failures by Omissions</kwd><kwd> PDEVS</kwd><kwd> DEVS</kwd><kwd> Modeling</kwd><kwd> Simulation</kwd><kwd> DEVSJAVA</kwd></kwd-group></article-meta></front><body><sec id="s1"><title>1. Introduction</title><p>The need to have more and more computing or storage power for projects has pushed humans to create increasingly efficient infrastructures [<xref ref-type="bibr" rid="scirp.111650-ref1">1</xref>]. Since the appearance of the first computers, passing by supercomputers, then clusters to computer grids, the objective has always been the search for ever higher performance. Computer networks are today the most efficient infrastructures in terms of power [<xref ref-type="bibr" rid="scirp.111650-ref2">2</xref>]. IT grid is a virtual infrastructure made up of a set of potentially shared, distributed, heterogeneous, delocalized and autonomous IT resources. This distributed, heterogeneous and delocalized architecture that computer grids possess is at the origin of many failures, including failures by omission [<xref ref-type="bibr" rid="scirp.111650-ref3">3</xref>]. A failure known as “by omission” is a failure in which the component temporarily ceases its activity and then resumes its normal activity [<xref ref-type="bibr" rid="scirp.111650-ref4">4</xref>]. The origins of failures by omission are numerous, we can cite: The increase in processor temperature because when the processor temperature approaches the maximum temperature allowed on the processor’s Integrated Heat Distributor (IHS) is also referred to as the chassis temperature. Congestion when the RAM occupancy rate approaches 100%. Many fault tolerance methods have been implemented so that the computer network can function despite the presence of faults [<xref ref-type="bibr" rid="scirp.111650-ref5">5</xref>] [<xref ref-type="bibr" rid="scirp.111650-ref6">6</xref>] [<xref ref-type="bibr" rid="scirp.111650-ref7">7</xref>]. There are several fault tolerance methods that are proposed in the literature [<xref ref-type="bibr" rid="scirp.111650-ref6">6</xref>]:</p><p>1) Fault tolerance based on the replication method. Different works have presented several replication methods: active, passive, semi-active, coordinator/cohort and adaptive. All of these methods consist of creating a set of replicas in different nodes. These replicas communicate with each other and with the system. When a failure occurs, one of the different replicas is chosen after consensus to return the result of the final processing [<xref ref-type="bibr" rid="scirp.111650-ref8">8</xref>]. But the fact that each of the replicas uses the resources and decreases the overall power of the grid, is not a suitable method for computing grids which require the maximum of resources.</p><p>2) Recovery-based fault tolerance. Two categories of restoration-based fault tolerance methods are proposed in the literature: backward recovery which consists of making system backups and, in the event of node failure, restoring from the last backup. And recovery by pursuit, which consists in the event of a failure of a node, reconstructs the failed job from a consensus between the other jobs [<xref ref-type="bibr" rid="scirp.111650-ref4">4</xref>].</p><p>Despite the evolution of fault tolerance methods, there are still many failures that occur in the performance of tasks in computer grids [<xref ref-type="bibr" rid="scirp.111650-ref9">9</xref>] [<xref ref-type="bibr" rid="scirp.111650-ref10">10</xref>]. For us, therefore, it is a question of answering the following question “Can we not propose a model that can anticipate failures by omission?”. The goal of this article is to propose a model which makes it possible to anticipate failures by omission in the calculation grids during the execution of the tasks. The plan for this article is as follows. In Section 2, we will present the corrected and failing models of a computing grid and in Section 3, we will present the models that allow anticipating failures by omission in computing grids.</p></sec><sec id="s2"><title>2. Modeling and Simulation of a Computer Grid</title><p>To model and simulate omission failures in computational grids, we use the PDEVS formalism to represent 100 nodes. In the following we present a calculation grid and the formalism used</p><sec id="s2_1"><title>2.1. Description of a Computer Grid</title><p>Ian Foster and Carl Kesselman were the first to use the word “grid” in their book “The Grid: Blueprint for a New Computing infrastructure” published in 2003 [<xref ref-type="bibr" rid="scirp.111650-ref11">11</xref>]. They compare a computer network to an electrical network in terms of the availability of resources. In an electrical network, all you have to do is plug an outlet into an outlet and electricity is supposed to flow. In a computer grid, it is enough to connect to the network to have the resources of the grid. a computer grid is defined as a set of shared, distributed, heterogeneous, delocalized and autonomous hardware and software resources in which large-scale scientific problems can be solved [<xref ref-type="bibr" rid="scirp.111650-ref6">6</xref>] [<xref ref-type="bibr" rid="scirp.111650-ref12">12</xref>]. The life cycle of a job in a grid IT begins when the user obtains a certificate from a certificate authority trusted by the grid. The user then registers in a virtual grid organization and obtains their user certificate [<xref ref-type="bibr" rid="scirp.111650-ref13">13</xref>]. The user can then connect to the platform from a particular machine acting as a user interface from which he can launch a job. This job is submitted from the user interface to the job management system (WMS) through its “tache” input port. The WMS searches for the calculation elements (CE) or nodes available to process the job. The system if job finds the grid at rest, ie in a “attente” state, it goes into an “active” state during which it processes the job. If the job finds the system in an “active” state, the job is just waiting for its turn to be processed. If a fault by omission is committed during the execution of the job. For example, jobs are sent to a node faster than it can process them, the node gets congested and an omission failure occurs on that node and its state goes to omission until the node decongests itself [<xref ref-type="bibr" rid="scirp.111650-ref14">14</xref>]. Once the processing of the parts of the job is finished, the WMS reconstructs the processed task and the processed job is returned by the “error” port [<xref ref-type="bibr" rid="scirp.111650-ref12">12</xref>].</p></sec><sec id="s2_2"><title>2.2. The PDEVS Formalism</title><p>To model and simulate failures by omission in computing grids, we used the PDEVS formalism. A computing grid being a very complex autonomous system seen as a single computer but which is a connection of several computers which share their resources in order to provide a very high power [<xref ref-type="bibr" rid="scirp.111650-ref15">15</xref>]. This modeling cannot be done with standard modeling formalisms because the latter are less suited to the modeling of complex and autonomous systems [<xref ref-type="bibr" rid="scirp.111650-ref9">9</xref>]. The modeling was done with the PDEVS modeling formalism. The PDEVS formalism allows modeling of causal and deterministic systems [<xref ref-type="bibr" rid="scirp.111650-ref10">10</xref>]. Its role is to provide a simple technique of parallelization of calculations [<xref ref-type="bibr" rid="scirp.111650-ref13">13</xref>]. A PDEVS atomic model is based on continuous time, inputs, outputs, states and functions [<xref ref-type="bibr" rid="scirp.111650-ref16">16</xref>]. More complex models are built by connecting several atomic models in a hierarchical fashion. The interactions are ensured by the input and output ports of the models, which favors modularity [<xref ref-type="bibr" rid="scirp.111650-ref17">17</xref>]. A computing grid is a set of interconnected nodes with the same goal. To model a computer grid, we therefore start by proposing the model of a node as an atomic model. Then we propose a grid model as being a coupled model which interconnects atomic models.</p><sec id="s2_2_1"><title>2.2.1. Model Description with Formal Specifications</title><p>The following model presents the formal specifications of an atomic model with PDEVS [<xref ref-type="bibr" rid="scirp.111650-ref18">18</xref>], which is an extension of the DEVS formalism [<xref ref-type="bibr" rid="scirp.111650-ref19">19</xref>]</p><p>M A = 〈 X , Y , S , δ i n t , δ e x t , δ c o n , λ , t a 〉</p><p>With:</p><p>&#183; the set of input ports and values:</p><p>X = { ( p , v ) | p ∈ P e , v ∈ V x }</p><p>where Pe and Vx are two finite sets representing the set of ports and input values;</p><p>&#183; the set of output ports and values:</p><p>Y = { ( p , v ) | p ∈ P s , v ∈ V y }</p><p>where Ps and Vy are two finite sets representing the set of output ports and the values carried by the events generated at the output;</p><p>&#183; the set of sequential states:</p><p>S = { s i | i ∈ R + }</p><p>&#183; the internal state transition function</p><p>δ i n t ( S ) → S</p><p>&#183; the external state transition function</p><p>δ e x t ( Q &#215; X ) → S</p><p>where, X<sup>b</sup> all input bags belonging to X, and Q the total state set, Q = { ( s , e ) | s ∈ S , 0 ≤ e ≤ t a ( s ) } , e is the time elapsed since last transition;</p><p>&#183; the confluent transition function</p><p>δ c o n ( S &#215; X ) → S</p><p>&#183; the output function:</p><p>λ ( S ) → Y</p><p>&#183; the time advance function:</p><p>t a ( S ) → R 0 + ∪ ∞</p><p>The following model presents the formal specifications of a model coupled with PDEVS</p><p>M C = 〈 X , Y , D , E O C , I C , E I C , S e l e c t 〉</p><p>With:</p><p>&#183; the set of input ports and values:</p><p>X = { ( p , v ) | p ∈ P e , v ∈ V x }</p><p>where Pe and Vx are two finite sets representing the set of ports and input values;</p><p>&#183; the set of output ports and values:</p><p>Y = { ( p , v ) | p ∈ P s , v ∈ V y }</p><p>where Ps and Vy are two finite sets representing the set of output ports and the values carried by the events generated at the output;</p><p>&#183; the set of the component names:</p><p>D = { m i | i ∈ R + }</p><p>&#183; external output couplings connect component outputs to external Outputs:</p><p>E O C = { ( ( m i , p s j ) , ( M C , p s k ) ) | i , j , s , k ∈ R + }</p><p>&#183; internal couplings connect component outputs to component inputs:</p><p>I C = { ( ( m i , p s j ) , ( m k , p e l ) ) | i , j , s , k ∈ R + }</p><p>&#183; external input couplings connect external inputs to component inputs</p><p>E I C = { ( ( M C , p e i ) , ( m j , p e k ) ) | e , i , j , k ∈ R + }</p><p>&#183; The selection function allowing to resolve the model activation conflict:</p><p>S e l e c t ( D ) → m i [ i ]</p></sec><sec id="s2_2_2"><title>2.2.2. Graphic Model</title><p>The model shown in <xref ref-type="fig" rid="fig1">Figure 1</xref> shows a description of the graphical model with PDEVS.</p><p>M A = 〈 X , Y , S , δ i n t , δ e x t , δ c o n , λ , t a 〉</p></sec></sec><sec id="s2_3"><title>2.3. Simulation</title><p>For the simulations, we use the devs-suite simulator version 4.0.0 developed by the arizona center for modeling and integrative simulation using DEVSJAVA as a language [<xref ref-type="bibr" rid="scirp.111650-ref20">20</xref>]. To simulate the model below, 100 nodes were used to model a calculation grid with the parameters listed in <xref ref-type="table" rid="table1">Table 1</xref>:</p><table-wrap id="table1" ><label><xref ref-type="table" rid="table1">Table 1</xref></label><caption><title> Characteristics of the nodes used</title></caption><table><tbody><thead><tr><th align="center" valign="middle" >Node</th><th align="center" valign="middle" >CPU</th><th align="center" valign="middle" >Cores</th><th align="center" valign="middle" >Tcase</th><th align="center" valign="middle" >RAM size</th></tr></thead><tr><td align="center" valign="middle" >1</td><td align="center" valign="middle" >2 &#215; Intel Xeon Gold 6130</td><td align="center" valign="middle" >16 cores/CPU</td><td align="center" valign="middle" >87˚C</td><td align="center" valign="middle" >192 GiB</td></tr><tr><td align="center" valign="middle" >3</td><td align="center" valign="middle" >2 &#215; Intel Core i5-3320M</td><td align="center" valign="middle" >4 cores/CPU</td><td align="center" valign="middle" >73˚C</td><td align="center" valign="middle" >12GiB</td></tr><tr><td align="center" valign="middle" >1</td><td align="center" valign="middle" >2 &#215; POWER8NVL 1.0</td><td align="center" valign="middle" >10 cores/CPU</td><td align="center" valign="middle" >74˚C</td><td align="center" valign="middle" >128 GiB</td></tr><tr><td align="center" valign="middle" >1</td><td align="center" valign="middle" >2 &#215; Intel Xeon Gold 5218</td><td align="center" valign="middle" >16 cores/CPU</td><td align="center" valign="middle" >87˚C</td><td align="center" valign="middle" >1.5 TiB</td></tr><tr><td align="center" valign="middle" >3</td><td align="center" valign="middle" >Intel Xeon Gold 6130</td><td align="center" valign="middle" >16 cores/CPU</td><td align="center" valign="middle" >87˚C</td><td align="center" valign="middle" >768 GiB</td></tr><tr><td align="center" valign="middle" >1</td><td align="center" valign="middle" >2 &#215; Intel Xeon E5-2630 v4</td><td align="center" valign="middle" >10 cores/CPU</td><td align="center" valign="middle" >74˚C</td><td align="center" valign="middle" >256 GiB</td></tr><tr><td align="center" valign="middle" >2</td><td align="center" valign="middle" >2 &#215; AMD EPYC 7301</td><td align="center" valign="middle" >16 cores/CPU</td><td align="center" valign="middle" >65˚C</td><td align="center" valign="middle" >128 GiB</td></tr><tr><td align="center" valign="middle" >4</td><td align="center" valign="middle" >2 &#215; Intel Xeon E5-2680 v4</td><td align="center" valign="middle" >14 cores/CPU</td><td align="center" valign="middle" >86˚C</td><td align="center" valign="middle" >768 GiB</td></tr><tr><td align="center" valign="middle" >1</td><td align="center" valign="middle" >4 &#215; Intel Xeon Gold 6126</td><td align="center" valign="middle" >12 cores/CPU</td><td align="center" valign="middle" >86˚C</td><td align="center" valign="middle" >192 GiB</td></tr><tr><td align="center" valign="middle" >3</td><td align="center" valign="middle" >2 &#215; Intel Xeon E5-2630L</td><td align="center" valign="middle" >6 cores/CPU</td><td align="center" valign="middle" >69.8˚C</td><td align="center" valign="middle" >32 GiB</td></tr><tr><td align="center" valign="middle" >1</td><td align="center" valign="middle" >2 &#215; Intel Xeon E5-2698 v4</td><td align="center" valign="middle" >20 cores/CPU</td><td align="center" valign="middle" >90˚C</td><td align="center" valign="middle" >512 GiB</td></tr><tr><td align="center" valign="middle" >5</td><td align="center" valign="middle" >Intel Xeon E5-2620</td><td align="center" valign="middle" >6 cores/CPU</td><td align="center" valign="middle" >77.4˚C</td><td align="center" valign="middle" >32 GiB</td></tr><tr><td align="center" valign="middle" >5</td><td align="center" valign="middle" >AMD EPYC 7642</td><td align="center" valign="middle" >48 cores/CPU</td><td align="center" valign="middle" >66˚C</td><td align="center" valign="middle" >512 GiB</td></tr><tr><td align="center" valign="middle" >1</td><td align="center" valign="middle" >2 &#215; Intel Xeon E5-2620 v4</td><td align="center" valign="middle" >8 cores/CPU</td><td align="center" valign="middle" >74˚C</td><td align="center" valign="middle" >64 GiB</td></tr><tr><td align="center" valign="middle" >2</td><td align="center" valign="middle" >2 &#215; Intel Xeon E5-2630</td><td align="center" valign="middle" >6 cores/CPU</td><td align="center" valign="middle" >77.4˚C</td><td align="center" valign="middle" >32 GiB</td></tr><tr><td align="center" valign="middle" >6</td><td align="center" valign="middle" >ThunderX2 99xx</td><td align="center" valign="middle" >32 cores/CPU</td><td align="center" valign="middle" >74˚C</td><td align="center" valign="middle" >256 GiB</td></tr><tr><td align="center" valign="middle" >1</td><td align="center" valign="middle" >2 &#215; AMD Opteron 250</td><td align="center" valign="middle" >1 core/CPU</td><td align="center" valign="middle" >65˚C</td><td align="center" valign="middle" >2 GiB</td></tr><tr><td align="center" valign="middle" >1</td><td align="center" valign="middle" >2 &#215; Intel Xeon E5-2630</td><td align="center" valign="middle" >6 cores/CPU</td><td align="center" valign="middle" >77.4˚C</td><td align="center" valign="middle" >32 GiB</td></tr><tr><td align="center" valign="middle" >1</td><td align="center" valign="middle" >2 &#215; Intel Xeon Silver 4110</td><td align="center" valign="middle" >8 cores/CPU</td><td align="center" valign="middle" >77˚C</td><td align="center" valign="middle" >128 GiB</td></tr><tr><td align="center" valign="middle" >1</td><td align="center" valign="middle" >8 &#215; Intel Xeon E5-2630 v3</td><td align="center" valign="middle" >8 cores/CPU</td><td align="center" valign="middle" >72.1˚C</td><td align="center" valign="middle" >128 GiB</td></tr><tr><td align="center" valign="middle" >1</td><td align="center" valign="middle" >2 &#215; Intel Xeon E5-2620 v3</td><td align="center" valign="middle" >6 cores/CPU</td><td align="center" valign="middle" >72.6˚C</td><td align="center" valign="middle" >64 GiB</td></tr><tr><td align="center" valign="middle" >1</td><td align="center" valign="middle" >2 &#215; Intel Xeon E5-2650</td><td align="center" valign="middle" >8 cores/CPU</td><td align="center" valign="middle" >77.4˚C</td><td align="center" valign="middle" >256 GiB</td></tr><tr><td align="center" valign="middle" >10</td><td align="center" valign="middle" >Intel Xeon Gold 5218R</td><td align="center" valign="middle" >20 cores/CPU</td><td align="center" valign="middle" >87˚C</td><td align="center" valign="middle" >96 GiB</td></tr><tr><td align="center" valign="middle" >1</td><td align="center" valign="middle" >2 &#215; Intel Xeon E5-2650</td><td align="center" valign="middle" >8 cores/CPU</td><td align="center" valign="middle" >77.4˚C</td><td align="center" valign="middle" >64 GiB</td></tr><tr><td align="center" valign="middle" >2</td><td align="center" valign="middle" >Intel Xeon E5-2650 v4</td><td align="center" valign="middle" >12 cores/CPU</td><td align="center" valign="middle" >80˚C</td><td align="center" valign="middle" >128 GiB</td></tr><tr><td align="center" valign="middle" >1</td><td align="center" valign="middle" >2 &#215; Intel Xeon E5-2603 v3</td><td align="center" valign="middle" >6 cores/CPU</td><td align="center" valign="middle" >72.8˚C</td><td align="center" valign="middle" >64 GiB</td></tr><tr><td align="center" valign="middle" >2</td><td align="center" valign="middle" >Intel Xeon E5-2630 v3</td><td align="center" valign="middle" >8 cores/CPU</td><td align="center" valign="middle" >72.1˚C</td><td align="center" valign="middle" >128 GiB</td></tr><tr><td align="center" valign="middle" >6</td><td align="center" valign="middle" >Intel Xeon E5-2630 v3</td><td align="center" valign="middle" >8 cores/CPU</td><td align="center" valign="middle" >72.1˚C</td><td align="center" valign="middle" >128 GiB</td></tr><tr><td align="center" valign="middle" >1</td><td align="center" valign="middle" >Intel Xeon Gold 5220</td><td align="center" valign="middle" >18 cores/CPU</td><td align="center" valign="middle" >87˚C</td><td align="center" valign="middle" >96 GiB</td></tr><tr><td align="center" valign="middle" >4</td><td align="center" valign="middle" >2 &#215; AMD EPYC 7452</td><td align="center" valign="middle" >32 cores/CPU</td><td align="center" valign="middle" >65˚C</td><td align="center" valign="middle" >128 GiB</td></tr><tr><td align="center" valign="middle" >5</td><td align="center" valign="middle" >AMD EPYC 7351</td><td align="center" valign="middle" >16 cores/CPU</td><td align="center" valign="middle" >66˚C</td><td align="center" valign="middle" >128 GiB</td></tr><tr><td align="center" valign="middle" >1</td><td align="center" valign="middle" >2 &#215; Intel Xeon Gold 6130</td><td align="center" valign="middle" >16 cores/CPU</td><td align="center" valign="middle" >87˚C</td><td align="center" valign="middle" >192 GiB</td></tr><tr><td align="center" valign="middle" >1</td><td align="center" valign="middle" >2 &#215; Intel Xeon E5-2660</td><td align="center" valign="middle" >8 cores/CPU</td><td align="center" valign="middle" >73˚C</td><td align="center" valign="middle" >64 GiB</td></tr><tr><td align="center" valign="middle" >7</td><td align="center" valign="middle" >2 &#215; Intel Xeon E5-2630L v4</td><td align="center" valign="middle" >10 cores/CPU</td><td align="center" valign="middle" >62˚C</td><td align="center" valign="middle" >128 GiB</td></tr><tr><td align="center" valign="middle" >1</td><td align="center" valign="middle" >2 &#215; Intel Xeon E5-2660 v2</td><td align="center" valign="middle" >10 cores/CPU</td><td align="center" valign="middle" >75˚C</td><td align="center" valign="middle" >128 GiB</td></tr><tr><td align="center" valign="middle" >1</td><td align="center" valign="middle" >3 &#215; Intel Xeon X5570</td><td align="center" valign="middle" >4 cores/CPU</td><td align="center" valign="middle" >75˚C</td><td align="center" valign="middle" >24 GiB</td></tr><tr><td align="center" valign="middle" >1</td><td align="center" valign="middle" >2 &#215; AMD Opteron 6164 HE</td><td align="center" valign="middle" >12 cores/CPU</td><td align="center" valign="middle" >65˚C</td><td align="center" valign="middle" >48 GiB</td></tr><tr><td align="center" valign="middle" >1</td><td align="center" valign="middle" >2 &#215; Intel Xeon E5-2630 v3</td><td align="center" valign="middle" >8 cores/CPU</td><td align="center" valign="middle" >72.1˚C</td><td align="center" valign="middle" >128 GiB</td></tr><tr><td align="center" valign="middle" >8</td><td align="center" valign="middle" >Intel Xeon E5-2630 v3</td><td align="center" valign="middle" >8 cores/CPU</td><td align="center" valign="middle" >72.1˚C</td><td align="center" valign="middle" >128 GiB</td></tr><tr><td align="center" valign="middle" >1</td><td align="center" valign="middle" >2 &#215; Intel Xeon X5670</td><td align="center" valign="middle" >6 cores/CPU</td><td align="center" valign="middle" >81.3˚C</td><td align="center" valign="middle" >96 GiB</td></tr></tbody></table></table-wrap><p>The 100 nodes are distributed as in the table above. For this modelization, it is considered that the distribution of tasks is perfectly balanced [<xref ref-type="bibr" rid="scirp.111650-ref16">16</xref>] that is to say:</p><p>Time total = NbFlops total ( Noeud1 ) Flops/s ( Noeud1 ) = ⋯ = NbFlops total ( Noeudn ) Flops/s ( Noeudn )</p><p>With: NbFlops the number of operations performed on all the processors of a node. Flops/s the computing power of a node and Time the execution time of each node of the grid and the tasks are submitted to the system in an identical time interval.</p></sec></sec><sec id="s3"><title>3. Results and Discussion</title><sec id="s3_1"><title>3.1. Modeling a Correct Node</title><p>A good node is a node that cannot receive failures. The specifications of such a node are shown in <xref ref-type="fig" rid="fig2">Figure 2</xref>.</p><p>The model starts in an initial “attente” state. when it receives a job via the “tache” port which is an external transition (continuous arrows), the node goes into an “active” state during which the job is processed. At the end of processing, the processed job is returned by the “resultat” port. If several jobs have been sent to the node, the jobs are processed one after the other. Once a job has been processed, the node continues to process the next job thanks to an internal transition (interrupted arrows). If during the processing of a job, another job is sent by the “tache” port, a confluence function will select which of the jobs will be executed between the internal and external transition. The representation of <xref ref-type="fig" rid="fig3">Figure 3</xref> is a formal description of the model of a correct node with the PDEVS formalism</p><p>The model having the above formal specifications of a correct node has been simulated and the results of this simulation are shown in the <xref ref-type="fig" rid="fig4">Figure 4</xref>.</p><p>The simulation of <xref ref-type="fig" rid="fig4">Figure 4</xref> is structured as follows: graph a represents the time of the last event, that is to say the time of the state being processed. Graph b represents the execution time of the next event. The graph c-1 represents the evolution of the states of the system during the execution of the jobs and the graph c-2 represents the state of the system before the arrival of the jobs. The graphs d and e represent the two input ports “tache” and “faute”, when a job is submitted for processing to the grid, it enters by the port “tache” and when a fault is committed it enters by the “faute” Port. The graphs f and g respectively represent the output ports “resultat” and “erreur”. When a job is successfully processed, the result is returned through the “resultat” port and when the system encounters a failure, the result is returned through the “erreur” port.</p><p>The job that are sent to the node are run for 100 seconds. The first tasks that are sent to the node are the first to be executed. The system boots into an “attente” state. When a transition occurs, the node executes the job for the defined execution time. When processing is complete, the corresponding output port takes the corresponding value. In this node, no failure ever occurs so no matter what port entered, the out port “resultat” will always receive an out value on each transition. In the rest of this work, we model a node that can receive failures by omission.</p></sec><sec id="s3_2"><title>3.2. Modeling a Failed Node</title><p>A node may receive failures by omission from time to time. The specifications of such a node capable of receiving failures by omission are shown in <xref ref-type="fig" rid="fig5">Figure 5</xref> and <xref ref-type="fig" rid="fig6">Figure 6</xref>.</p><p>A failed node is a good node until a failure occurs, in which case a failure by omission. in this model, the node is in an “active” state until an omission failure occurs and the node enters an omission state. The node remains for a random period which can range from less than a second to infinity. After this period during which the node has remained in an “omission” state, it resumes its normal service and passes to the “active” state. The simulation of <xref ref-type="fig" rid="fig7">Figure 7</xref> represents that of a model failing by omission.</p><p>Jobs sent to the node run for 100 seconds. The first jobs sent to the node are the first to be executed. The system starts up in an initial “attente” state. When a transition occurs, the node executes the task for the defined execution time. When processing is complete, the corresponding output port takes on the corresponding value. In this node, when failure by omission occurs in the node (red band), either by overheating of the processor or by overload of the RAM, the node goes into an “omission” state for a random period. In this simulation the failure by omission lasts 90 seconds. After this period, the node resumed its activity as it should. Task 4 will therefore run for 190 seconds and there is now a 90 second lag on all subsequent processed jobs.</p></sec><sec id="s3_3"><title>3.3. Modeling of a Calculation Grid</title><p>A computing grid is a set of heterogeneous, interconnected nodes considered as a single computer to provide great computing power. The modeling of a computing grid is therefore, with the PDEVS formalism, a coupled model where the various atomic models which constitute it are the various nodes. The formal specifications of a grid are presented in <xref ref-type="fig" rid="fig8">Figure 8</xref>.</p><p>In this modelization, the model N represents the calculation grid. Jobs can be submitted through two ports to the grid, “stain” and “fault”. The submitted jobs arrive at the “WMSI” task manager, which is responsible for finding available nodes and sending parts of the job to them. Once the jobs have been processed by the various nodes, the parts of the job are sent back to the “WMSO” task manager to be reconstructed, restored and sent back to the output port. The simulation of this model is shown in <xref ref-type="fig" rid="fig9">Figure 9</xref>.</p><p>In this simulation of a computing grid, every 150 seconds the system makes sure that the nodes are still all connected to the network (green bands) by exchanging messages while performing a system backup. In the event of a failure (red band), a fault tolerance technique is applied to the next message exchange. In the case of our simulation during the processing of job5, a failure occurs. failure detection occurs at t = 450 sec. The fault tolerance technique is applied and the job is processed. In this simulation, the restart recovery caused by node failure is general, that is to say all nodes are restored. Of which at t = 500, the system restarts where the failure occurred.</p></sec><sec id="s3_4"><title>3.4. Proposed Correction Model</title><p>In this part we propose a model which makes it possible to anticipate failures by omission in the calculation grid. Unlike the current method of performing system backups at a set frequency and restoring the last backup in the event of a failure, this model looks for signs that can lead to omission failure in every node. The principle is to observe the components of each node using the node’s sensors and exchanged messages. For example, if the temperature sensor of the processor signals a temperature which is approaching due to the temperature of the processor chassis, the failure is reported to the task manager who will no longer send jobs to the node likely to have a failure, the model is therefore represented in <xref ref-type="fig" rid="fig1">Figure 1</xref>0.</p><p>The formal description of this model is presented in <xref ref-type="fig" rid="fig1">Figure 1</xref>1.</p><p>The above model presents a model which allows failures by omission to be anticipated. Once the node begins to receive jobs, it goes into the “active” state. During job processing, if the node begins to receive too many jobs beyond what it can handle, the RAM occupancy rate increases to saturation causing failure by omission. Or if, during job processing, the processor temperature rises to approach the temperature of the processor chassis, the node temporarily stops operating indefinitely and causes an omission failure. These values being measured continuously, the principle is to send information on the temperature of the processor and the rate of occupation of the RAM to the task manager which will determine whether it should always forward jobs to this node or not based on the state of RAM and processor. In general, a safe range during normal workload is between 40˚C and 65˚C (or 104˚F - 149˚F) [<xref ref-type="bibr" rid="scirp.111650-ref17">17</xref>]. The Intel Xeon E5-2630L v4 processor has a chassis temperature of Tcase = 62˚C. In this model, if the processor temperature reaches 62˚C, the node reports an “omitted state”. Or if the RAM occupancy rate reaches 90% of the usable RAM, the node reports the failure and the grid task manager will no longer send tasks to the node (<xref ref-type="fig" rid="fig1">Figure 1</xref>2).</p><p>This simulation is that of a calculation grid which makes it possible to anticipate failures by omission. In this simulation, three nodes are likely to be bad and the rest are good nodes. If one of the nodes goes down, its processor temperature rises 15 degrees after each task. In this simulation, if the temperature of the Tcase processor chassis minus the T processor temperature is less than 10 degrees (Tcase-T ≤ 10) then a failure is reported in the node and the message is transmitted to the task manager which will ignore thereafter. Here this case occurs after the processing of job4, a node is reported as bad and the task manager</p><p>ignores it and the system continues to execute the jobs. Unlike other fault tolerance protocols which wait for the failure to occur before tolerating it, this model allows you to anticipate the failure by probing the various components of the node and determining if there may be a failure. This detection time and recovery of the system are almost zero with this model. But this model is only effective if the failure is normal, ie failures which are not linked to the attacks.</p></sec></sec><sec id="s4"><title>4. Conclusions</title><p>Failures by omission in computational grids have been studied numerically in the present work. The numerical methodology is based on the PDEVS modeling formalism in the devs-suite simulator version 4.0.0. The main results are summarized as follows:</p><p>1) We first present a correct node model which is a node which cannot receive failures and a node model which can receive failures by omission.</p><p>2) Next, we present a computational grid model which is a set of interconnected nodes. This grid which is a set of correct nodes and nodes that can receive failures by omission.</p><p>3) Finally, we propose a model which makes it possible to anticipate failures by omission in the calculation grids. This model is based on the behavior of certain components of the node to anticipate failures. Using this model will help limit omission failures that are caused by the hardware components of the node.</p></sec><sec id="s5"><title>Conflicts of Interest</title><p>The authors declare no conflicts of interest regarding the publication of this paper.</p></sec><sec id="s6"><title>Cite this paper</title><p>Yougouda, R.A., Nkenlifack, M., Kamla, V.C. and Bitjoka, L. (2021) Model for Anticipating Failures by Omission in Calculation Grids. Open Journal of Optimization, 10, 71-87. https://doi.org/10.4236/ojop.2021.103006</p></sec></body><back><ref-list><title>References</title><ref id="scirp.111650-ref1"><label>1</label><mixed-citation publication-type="other" xlink:type="simple">Kalfadj, F., Belabbas, Y. and Meddeber, M. (2009) Conception d’un Simulateur de Grilles Orient&amp;#233; Gestion d’ &amp;#201; quilibrage. Proceedings of the 2nd Conf&amp;#233;rence Internationale sur l’Informatique et ses Applications, Saida, 3-4 May 2009, 1-11.</mixed-citation></ref><ref id="scirp.111650-ref2"><label>2</label><mixed-citation publication-type="other" xlink:type="simple">Barchet-Estefanel, L.A. (2005) LaPIe: Communications Collectives adapt&amp;#233;es aux Grilles de Calcul. Institut National Polytechnique De Grenoble, Grenoble.</mixed-citation></ref><ref id="scirp.111650-ref3"><label>3</label><mixed-citation publication-type="other" xlink:type="simple">Ta&amp;#239;ani, F., Killijian, M.O. and Fabre, C.J. (2006) Intergiciels pour la tol&amp;#233;rance aux fautes. RSTI-TSI, 25, 599-630. https://doi.org/10.3166/tsi.25.599-630</mixed-citation></ref><ref id="scirp.111650-ref4"><label>4</label><mixed-citation publication-type="other" xlink:type="simple">Monnet, S. (2006) Gestion des donn&amp;#233;es dans les grilles de calcul: Support pour la tol&amp;#233;rance aux fautes et la coh&amp;#233;rence des donn&amp;#233;es. Universit&amp;#233; Rennes 1, Rennes.</mixed-citation></ref><ref id="scirp.111650-ref5"><label>5</label><mixed-citation publication-type="other" xlink:type="simple">Rebbah, M. (2015) Tol&amp;#233;rance aux fautes dans les grilles de calcul. Universit&amp;#233; des Sciences et de la Technologie d’Oran, Oran.</mixed-citation></ref><ref id="scirp.111650-ref6"><label>6</label><mixed-citation publication-type="other" xlink:type="simple">Ndiaye, N.M. (2013) Techniques de gestion des d&amp;#233;faillances dans les grilles informatiques tol&amp;#233;rantes aux fautes. Universit&amp;#233; Pierre et Marie Curie, Paris.</mixed-citation></ref><ref id="scirp.111650-ref7"><label>7</label><mixed-citation publication-type="other" xlink:type="simple">Akoka, J. and Comyn-Wattiau, I. (2006) Encyclop&amp;#233;die de l’informatique et des syst&amp;#232;mes d’information. Vuibert, Paris.</mixed-citation></ref><ref id="scirp.111650-ref8"><label>8</label><mixed-citation publication-type="other" xlink:type="simple">Krakoviak, S. (2004) Tolerences aux fautes, Universit&amp;#233; Joseph Fourier.</mixed-citation></ref><ref id="scirp.111650-ref9"><label>9</label><mixed-citation publication-type="other" xlink:type="simple">Ramat, &amp;#201;., Duboz, R. and Quesnel, G. (2012) Theorie de la mod&amp;#233;lisation et de la simulation, fondements formels et operationnels de record/VLE, INRA-Cirad- ULCO/LISIC.</mixed-citation></ref><ref id="scirp.111650-ref10"><label>10</label><mixed-citation publication-type="other" xlink:type="simple">Zeigler, B.P., Kim, T.G. and Praehofer, H. (2000) Theory of Modeling and Simulation. Academic Press, Orlando.</mixed-citation></ref><ref id="scirp.111650-ref11"><label>11</label><mixed-citation publication-type="other" xlink:type="simple">Foster, I. and Kesselman, C. (2003) The Grid: Blueprint for a New Computing Infrastructure. Morgan Kaufmann, Burlington.</mixed-citation></ref><ref id="scirp.111650-ref12"><label>12</label><mixed-citation publication-type="other" xlink:type="simple">Avizienis, A., Laprie, J.-C. and Randell, B. (2001) Fundamental Concepts of Dependability. Newcastle University, Newcastle.</mixed-citation></ref><ref id="scirp.111650-ref13"><label>13</label><mixed-citation publication-type="other" xlink:type="simple">Zeigler, B., Gon Kim, T. and Praehofer, H. (2000) Theory of Modeling and Simulation. 2nd Edition, Academic Press, New York.</mixed-citation></ref><ref id="scirp.111650-ref14"><label>14</label><mixed-citation publication-type="other" xlink:type="simple">Perez, J. (2009) Apprentissage Arti&amp;#226;ciel pour l’ordonnancement des taches dans les grilles de calcul. Universit&amp;#233; Paris-Sud, Paris.</mixed-citation></ref><ref id="scirp.111650-ref15"><label>15</label><mixed-citation publication-type="other" xlink:type="simple">Zeigler, B.P. (2005) Introduction to DEVS Modeling and Simulation with JAVA: Developing Component-Based Simulation Models. Arizona Center for Integrative Modeling and Simulation, Tucson.</mixed-citation></ref><ref id="scirp.111650-ref16"><label>16</label><mixed-citation publication-type="other" xlink:type="simple">Kadri, W. (2012) Placement dynamique des taches d&amp;#233;pendantes dans une grille de calcul. Universit&amp;#233; d’oran, Oran.</mixed-citation></ref><ref id="scirp.111650-ref17"><label>17</label><mixed-citation publication-type="other" xlink:type="simple">Villinger, S. (2021) Comment surveiller la temp&amp;#233;rature de votre processeur. https://www.avast.com/fr-fr/c-how-to-check-cpu-temperature</mixed-citation></ref><ref id="scirp.111650-ref18"><label>18</label><mixed-citation publication-type="other" xlink:type="simple">Ouedraogo, D., Wendsida Igo, S., Sawadogo, G.L., Compaore, A., Zeghmati, B. and Chesneau, X. (2020) Modeling and Numerical Simulation of Heat Transfers in a Metallic Pressure Cooker Isolated with Kapok Wool. Modeling and Numerical Simulation of Material Science, 10, 15-30. https://doi.org/10.4236/mnsms.2020.102002</mixed-citation></ref><ref id="scirp.111650-ref19"><label>19</label><mixed-citation publication-type="other" xlink:type="simple">Guessoum, Z., Briot, J.P., Faci, N. and Marin, O. (2004) Un m&amp;#233;canisme de r&amp;#233;plication adaptative pour des SMA tol&amp;#233;rants aux pannes. JFSMA.</mixed-citation></ref><ref id="scirp.111650-ref20"><label>20</label><mixed-citation publication-type="other" xlink:type="simple">Zeigler, B. (1976) Theory of Modeling and Simulation. Academic Press, London.</mixed-citation></ref></ref-list></back></article>