<?xml version="1.0" encoding="UTF-8"?><!DOCTYPE article  PUBLIC "-//NLM//DTD Journal Publishing DTD v3.0 20080202//EN" "http://dtd.nlm.nih.gov/publishing/3.0/journalpublishing3.dtd"><article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" dtd-version="3.0" xml:lang="en" article-type="research article"><front><journal-meta><journal-id journal-id-type="publisher-id">OALibJ</journal-id><journal-title-group><journal-title>Open Access Library Journal</journal-title></journal-title-group><issn pub-type="epub">2333-9705</issn><publisher><publisher-name>Scientific Research Publishing</publisher-name></publisher></journal-meta><article-meta><article-id pub-id-type="doi">10.4236/oalib.1108497</article-id><article-id pub-id-type="publisher-id">OALibJ-116209</article-id><article-categories><subj-group subj-group-type="heading"><subject>Articles</subject></subj-group><subj-group subj-group-type="Discipline-v2"><subject>Biomedical&amp;Life Sciences</subject><subject> Business&amp;Economics</subject><subject> Chemistry&amp;Materials Science</subject><subject> Computer Science&amp;Communications</subject><subject> Earth&amp;Environmental Sciences</subject><subject> Engineering</subject><subject> Medicine&amp;Healthcare</subject><subject> Physics&amp;Mathematics</subject><subject> Social Sciences&amp;Humanities</subject></subj-group></article-categories><title-group><article-title>
 
 
  Predicting the Perceived Employee Tendency of Leaving an Organization Using SVM and Naive Bayes Techniques
 
</article-title></title-group><contrib-group><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Ijeoma</surname><given-names>Lilian Emmanuel-Okereke</given-names></name><xref ref-type="aff" rid="aff1"><sup>1</sup></xref></contrib><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Sylvanus</surname><given-names>Okwudili Anigbogu</given-names></name><xref ref-type="aff" rid="aff1"><sup>1</sup></xref></contrib></contrib-group><aff id="aff1"><addr-line>Department of Computer Science, Nnamdi Azikiwe University, Awka, Nigeria</addr-line></aff><pub-date pub-type="epub"><day>04</day><month>03</month><year>2022</year></pub-date><volume>09</volume><issue>03</issue><fpage>1</fpage><lpage>15</lpage><history><date date-type="received"><day>17,</day>	<month>February</month>	<year>2022</year></date><date date-type="rev-recd"><day>25,</day>	<month>March</month>	<year>2022</year>	</date><date date-type="accepted"><day>28,</day>	<month>March</month>	<year>2022</year></date></history><permissions><copyright-statement>&#169; Copyright  2014 by authors and Scientific Research Publishing Inc. </copyright-statement><copyright-year>2014</copyright-year><license><license-p>This work is licensed under the Creative Commons Attribution International License (CC BY). http://creativecommons.org/licenses/by/4.0/</license-p></license></permissions><abstract><p>
 
 
  There are several experienced and highly skilled employees considered to be assets in every organization; a good and flexible working environment is required to retain them. The perceived exit of well-skilled and highly experienced employees may result in financial losses, poor sales with customers’ dissatisfaction and produced a low turnover. It also led to low production and output. The existing methods of operation lack the merit in producing accurate and reliable results. It could not generalize well with testing datasets and results in the problem of model over-fitting. Little or no work has been done in the area of predicting the perceived employee tendency of leaving an organization using Support Vector Machine (SVM) and Naive Bayes (NB) algorithm. The implemention was done using Python (Spyder IDE) in ANACONDA. In this paper, a model which is capable of predicting the perceived employee tendency of exiting an organization was developed using the support vector machine and the Naive Bayesian machine learning algorithm. The adopted techniques improved the prediction accuracy and generalized well with testing datasets in overcoming the problem of over-fitting. It also reduced the sudden occurrence of experienced and skilled employees leaving an organization. We adopted the SVM and NB to effectively handle overlapping and reduce data misclassification errors that can work well with a limited number of the dataset. The proposed NB model was trained, successfully tested and evaluated using the same dataset in comparison with the SVM technique. The experimental results of NB model produced 100% prediction accuracy with a 0.0000 RMSE error value in comparison with the SVM which gave a 97.00% success rate and 0.0258 RMSE value.
 
</p></abstract><kwd-group><kwd>Classification</kwd><kwd> Job Exit</kwd><kwd> Machine Learning</kwd><kwd> Naive Bayes</kwd><kwd> Support Vector Machine</kwd></kwd-group></article-meta></front><body><sec id="s1"><title>1. Introduction</title><p>Every business industry or organization that deals with skilled workers of human capital development focuses on profit maximization, turnover and cost minimization. There are profit-driven industries in the world today that suffer the backlog of financial losses caused by the exit of skilled and highly experienced workers [<xref ref-type="bibr" rid="scirp.116209-ref1">1</xref>]. The existing methods of operation are not accurate and reliable, and could not generalize well with testing datasets which results in the problem of model over-fitting.</p><p>Job exit is the process of an employee leaving an organization or industry which at times is beyond the control of that organization [<xref ref-type="bibr" rid="scirp.116209-ref2">2</xref>]. The act of highly experienced employees leaving an organization voluntarily and involuntarily can be controlled if we have a system that predicts the occurrence before it happens [<xref ref-type="bibr" rid="scirp.116209-ref3">3</xref>]. The voluntary exit of an employee from an organization is a crucial and important issue that results in a decline in human capital and financial loss when the best staff often leaves without prior knowledge. There are several factors responsible for or causes the perceived employee tendency of leaving an organization, namely: the organizational factors and individual factors [<xref ref-type="bibr" rid="scirp.116209-ref4">4</xref>]. The organizational factors include low salary, workload not commeasurable to salary and overtime pay, too many requirements for advancement, lack of appreciation for a job well done, etc [<xref ref-type="bibr" rid="scirp.116209-ref5">5</xref>]. The individual factors include frequent late-night meetings, family obligations, work conflicting with personal life responsibilities, personal relationships. There are other factors identified by [<xref ref-type="bibr" rid="scirp.116209-ref6">6</xref>] as new rules and organizational policies, lack of monetary benefits, extensive workload and stress, lack of leadership qualities, the relationship between manager and promotions, not being involved for staff training. Skilled and highly experienced employees are considered assets to any organization; therefore, a good and flexible working environment is recommended to retain them [<xref ref-type="bibr" rid="scirp.116209-ref7">7</xref>].</p><p>It is difficult and competitive to have qualified and highly experienced staff in fulfilling the needs of any organization around the globe. The success and efficiency of every organization depend on its capacity to retain skilled and well-experienced employees. The exit of well-skilled and highly experienced employees may result in losses in millions and billions of Naira, loss of revenue, poor sales and customers’ dissatisfaction with a low turnover in meeting up with its goal and objectives in an organization. The unskilled employees are prone to make more errors which may give rise to low production, output and little or no work has been done in the area of predicting the perceived employee tendency of leaving an organization. And most of the existing systems in the application suffered from data misclassification errors with overlapping patterns. It produced more False Positive (FP) cases than the True Positive (TP) and more False Positive (FP) classification than False Negative (FN) cases caused by model over-fitting.</p><p>The aim is to build a model capable of predicting the perceived Employee tendency of exiting an Organization using Support Vector Machine (SVM) and Naive Bayesian (NB) technique. This model is employed to help reduce overlapping and misclassification errors which affected the performance accuracy of the existing system techniques and could not generalize well with the testing dataset.</p><p>This paper is divided into different sections as followings: Section 1 contains the introduction, Section 2 presents a brief review of previous approaches relating to the study area and the gap in exploring the proposed model; Section 3 introduces materials and methods employed for developing the model; Section 4 focuses on the results and detailed discussion of results; Section 5 presents the conclusion to the paper.</p></sec><sec id="s2"><title>2. Related Work</title><p>Saradhi &amp; Palshikar [<xref ref-type="bibr" rid="scirp.116209-ref8">8</xref>] stated reasons for an employee exiting an organization to have a better offer or career growth relating to better salary, promotions, staff training and work environment. Ramamurthy et al. [<xref ref-type="bibr" rid="scirp.116209-ref9">9</xref>] developed a model in order to predict those employees qualified to be trained for a particular skill that suits their job function. Singh et al. [<xref ref-type="bibr" rid="scirp.116209-ref10">10</xref>] proposed a study on employee attrition using the C5.0 type of decision tree technique but suffers from model over-fitting and misclassification errors. Maharjan [<xref ref-type="bibr" rid="scirp.116209-ref11">11</xref>] developed a model to predict employee churn using SVM, Logistic regression and decision tree classifiers. An Extra tree class was invoked from Python SKlearn library to compute the score for all features. The information gain was used to filter relevant features, analyzed and compared without any order or rank. It uses internal mechanism in ranking the features without relying upon user calculated and ranked values. The dataset was imbalanced and employee attributes are less significant compared to non employee attributes. The dataset was divided into the ratio 70:30:70 of training and testing set to overcome the problem of over-fitting for the purpose of learning using a Stratified K-fold cross-validation test. A k value was defined based on what the entire dataset will get and divided in forming that number of K-folds. It provided a uniform data distribution with majority and minority across testing and training items. The LR performed better than the SVM. The accuracy rate was below average and could not be extended to work with handle cloud-based platform. Jayad et al. [<xref ref-type="bibr" rid="scirp.116209-ref12">12</xref>] proposed the use of Naive Bayes (NB) classification model as a technique in machine learning to predict employee performance drives the success of every organization. The implementation was done using WEKA and correctly classified and predicted target variable as required with 95.48% accuracy rate in 0.01 seconds and update performance score of 96.77% metrics of accuracy. The accuracy of NB increased along with the number of instances. The confusion matrix recorded more correctly predicted values (TP + TN) than wrongly predicted values (FN + FP). The model flags up error message when instances are below ten (10) and accuracy level could not be computed.</p><p>Yahia et al. [<xref ref-type="bibr" rid="scirp.116209-ref13">13</xref>] adopted a deep machine leaning technique predict employee attrition support system. The deep driven machine learning approach was employed to detect key employee attributes with feature extraction technique with two different dataset that influences staff attrition. A small size of human resourced dataset of about 450 responses and a large sized kaggle HR dataset of 15,000 samples are used to train and test the model. The voting classifier performed better and produced 99% in terms of accuracy with real life dataset compared to other methods. The model was not suitable to work with imbalanced dataset with companies that have high turnover. Kamath et al. [<xref ref-type="bibr" rid="scirp.116209-ref14">14</xref>] employed the combination of Random Forest (RF), SVM, DT and Logistic Regression (LR) machine learning techniques for human resource attrition status, management and forecasting. The dataset was divided into training, testing and validation set in the ratio of 70%, 15% and 15% respectively. The results revels that employee attrition depends mainly on employee satisfaction as compared to other features and attributes. The RF was the best in performance with r-square value of 0.9773 while other like DT, SVM and LR recorded 0.8473, 0.8315 and 0.2299 respectively. The model could not work with large and unstructured dataset. Alshehhi et al. [<xref ref-type="bibr" rid="scirp.116209-ref15">15</xref>] combined DT and RF classifiers in machine learning to predict employee retention rates in an organization. The FR classifier was employed to forecast employee characteristics with retention rates using a training data of 13-years and testing dataset of 14-years.The dataset was divided into two to avoid model overfitting. It was trained to predict the occurrence of employee retention across each year, categories of department and training and used to determine if the organization losing an experienced Staff or not because of training and retraining. The RF classifier outperformed the DT technique in terms of accuracy and error rate. The RF and DT techniques could not work with large volume of dataset.</p><p>Senanayake et al. [<xref ref-type="bibr" rid="scirp.116209-ref16">16</xref>] employed the RF learning algorithm in ML to predict employee resignation in Swedish armed forces. The RF model was train to learn and predict employee that are due for resign and recommend possible recruitment policies that can be used to replace such retiring employees using a sizable dataset. The RF model produced 89.067% accuracy in comparison with the zero-guess that gave 84.533%. The dataset was quite small to achieve high accuracy rate as required.</p></sec><sec id="s3"><title>3. Materials and Methods</title><p>In this paper, we are focusing on the use of SVM and Naive Bayes (NB) techniques to handle the problem of outliers efficiently with better accuracy rate and effectively handle overlapping classifications. The gamma, C set to 1.0 and random state variables are employed in the SVM class to have a better performance rate. We are adopting the Gaussian Naive bayes type of classifier because it is highly scalable with number of data points, predictors and not sensitivity to irrelevant data features.</p><sec id="s3_1"><title>3.1. Data Source</title><p>The dataset (<xref ref-type="table" rid="table1">Table 1</xref>) used was obtained from a well-structured self study questionnaire distributed and collected through survey as a primary source containing five hundred and fourteen (514) items with attributes: timestamp, promoted, job satisfaction, work hour per day, training and working experience, job security, changed jobs, and employee exit as target. The dataset was divided into 80%</p><p>training 80 100 % &#215; 514 = 412 items and 20% testing 20 100 % &#215; 514 = 102 set for predicting the perceived employee tendency of leaving an organization.</p></sec><sec id="s3_2"><title>3.2. Data Preprocessing</title><p>The pre-processing stage is necessary for the training and reduces threshold value. It was adopted to help manipulate data and improve model performance because in gathering data sometimes poses difficulties and may result into out-of-range, missing, noisy and false data values. This involves data cleaning, instance selection, data normalization, transformation, feature extraction and selection. The preprocessing produces training data as output which can effectively be interpreted by models.</p></sec><sec id="s3_3"><title>3.3. Classification</title><p>The classification system is adopted as a supervised learning process of determining or predicting data classes referred to as target, labels or categories. Classification is a predictive task or modeling of estimating a mapping function from</p><table-wrap id="table1" ><label><xref ref-type="table" rid="table1">Table 1</xref></label><caption><title> Employee dataset</title></caption><table><tbody><thead><tr><th align="center" valign="middle" ></th><th align="center" valign="middle" >Gender</th><th align="center" valign="middle" >Experience</th><th align="center" valign="middle" >Job security</th><th align="center" valign="middle" >Working hours</th><th align="center" valign="middle" >- - - -</th><th align="center" valign="middle" >Target</th></tr></thead><tr><td align="center" valign="middle" >0</td><td align="center" valign="middle" >Female</td><td align="center" valign="middle" >8</td><td align="center" valign="middle" >Yes</td><td align="center" valign="middle" >4</td><td align="center" valign="middle" >- - - -</td><td align="center" valign="middle" >Yes</td></tr><tr><td align="center" valign="middle" >1</td><td align="center" valign="middle" >Male</td><td align="center" valign="middle" >25</td><td align="center" valign="middle" >Yes</td><td align="center" valign="middle" >8</td><td align="center" valign="middle" >- - - -</td><td align="center" valign="middle" >No</td></tr><tr><td align="center" valign="middle" >2</td><td align="center" valign="middle" >Female</td><td align="center" valign="middle" >7</td><td align="center" valign="middle" >Yes</td><td align="center" valign="middle" >8</td><td align="center" valign="middle" >- - - -</td><td align="center" valign="middle" >No</td></tr><tr><td align="center" valign="middle" >3</td><td align="center" valign="middle" >Female</td><td align="center" valign="middle" >15</td><td align="center" valign="middle" >Yes</td><td align="center" valign="middle" >7</td><td align="center" valign="middle" >- - - -</td><td align="center" valign="middle" >No</td></tr><tr><td align="center" valign="middle" >4</td><td align="center" valign="middle" >Female</td><td align="center" valign="middle" >13</td><td align="center" valign="middle" >Yes</td><td align="center" valign="middle" >6</td><td align="center" valign="middle" >- - - -</td><td align="center" valign="middle" >No</td></tr><tr><td align="center" valign="middle" >- - -</td><td align="center" valign="middle" >- - - -</td><td align="center" valign="middle" >- - -</td><td align="center" valign="middle" >- - -</td><td align="center" valign="middle" ></td><td align="center" valign="middle" >- - - -</td><td align="center" valign="middle" >- - -</td></tr><tr><td align="center" valign="middle" >510</td><td align="center" valign="middle" >Female</td><td align="center" valign="middle" >30</td><td align="center" valign="middle" >Yes</td><td align="center" valign="middle" >8</td><td align="center" valign="middle" >- - - -</td><td align="center" valign="middle" >No</td></tr><tr><td align="center" valign="middle" >511</td><td align="center" valign="middle" >Female</td><td align="center" valign="middle" >3</td><td align="center" valign="middle" >Yes</td><td align="center" valign="middle" >8</td><td align="center" valign="middle" >- - - -</td><td align="center" valign="middle" >No</td></tr><tr><td align="center" valign="middle" >512</td><td align="center" valign="middle" >Female</td><td align="center" valign="middle" >4</td><td align="center" valign="middle" >No</td><td align="center" valign="middle" >10</td><td align="center" valign="middle" >- - - -</td><td align="center" valign="middle" >Yes</td></tr><tr><td align="center" valign="middle" >513</td><td align="center" valign="middle" >Female</td><td align="center" valign="middle" >15</td><td align="center" valign="middle" >Yes</td><td align="center" valign="middle" >9</td><td align="center" valign="middle" >- - - -</td><td align="center" valign="middle" >No</td></tr><tr><td align="center" valign="middle" >514</td><td align="center" valign="middle" >Female</td><td align="center" valign="middle" >15</td><td align="center" valign="middle" >Yes</td><td align="center" valign="middle" >9</td><td align="center" valign="middle" >- - - -</td><td align="center" valign="middle" >No</td></tr></tbody></table></table-wrap><p>input variables represented as “X” variable to a discrete output variable represented with “y”. It depend mainly the area of application and the nature of available dataset [<xref ref-type="bibr" rid="scirp.116209-ref17">17</xref>]. The NB classifier is work based on Bayes theorem under simple assumption and the attributes are conditionally independent.</p></sec><sec id="s3_4"><title>3.4. Feature Extraction</title><p>The feature selection process was adopted to determine the correlation between variable or attribute pars based on the level of correlation using a score value. The higher the score value the higher the correlation between attribute pairs [<xref ref-type="bibr" rid="scirp.116209-ref18">18</xref>].</p></sec><sec id="s3_5"><title>3.5. Support Vector Machines (SVM) Classifier</title><p>The SVM is one of the simplest and more preferred machine learning techniques used by data professionals because it tendency of producing better and high accuracy with less computational error [<xref ref-type="bibr" rid="scirp.116209-ref19">19</xref>]. The SVM uses two main concepts namely: hypothesis space and the loss functioning finding an “optimal” hyper-plane as a solution to any learning problem [<xref ref-type="bibr" rid="scirp.116209-ref20">20</xref>]. The SVM is memory efficient and uses subsets of training data points in the support vectors called decision function. The simplest formulation of SVM is the linear one, where the hyper-plane lies on the space of the input data [<xref ref-type="bibr" rid="scirp.116209-ref21">21</xref>]. The SVM estimator was defined on the training dataset and tested to effectively predict the target variable. A SVM classifier was invoked from the sklearn.svm library in python and SVM model created. The gamma variable set to be scalable, c=1,0 and random states set to 101 with the Python script: svc=SVC(gamma=‘scale’, C=1.0, rando_state=101). The model was trained with training dataset with svm.fit(X_train,y_test) and predicted using the testing dataset[svc.predict(X_test)]. The visualization was done using mat_plot_lib library in python. A SVM classifier was created with the pre-processed training data to make predictions about employees exit.</p></sec><sec id="s3_6"><title>3.6. Naive Bayes Classifier</title><p>The Naive Bayes (NB) technique is one of the most popular known supervised machine learning algorithms that uses Bayes theorem. The Gaussian NB classification algorithm works with the principles of conditional probability as given by Bayes theorem. The Bayes theorem gives the conditional probability of an event “H” given whether event “D” has occurred. The Bayesian theorem basically computes the conditional probability of the occurrence of an event based on prior knowledge of conditions that might be related to the event [<xref ref-type="bibr" rid="scirp.116209-ref22">22</xref>]. It provides update to probability of hypothesis (H) for some given instance of data (D) which can be expressed in Equation (3.1) as follows:</p><p>P ( H / D ) = P ( D / H ) P ( H ) P ( D ) (3.1)</p><p>where P(D/H) is the probability of hypothesis and P(D) dataset features/parameters.</p><p>The character or feature variables are encoded using label encoder at preprocessing stage and feature scaling technique employed for the training and testing dataset of the independent variables in producing better classification report.</p><p>The D is given as:</p><p>D = ( d 1 , d 2 , d 3 , ⋯ d n ) (3.2)</p><p>where d<sub>1</sub>, d<sub>2</sub>, d<sub>3</sub>, ..., d<sub>n</sub> represents the features mapped into the outlook.</p></sec><sec id="s3_7"><title>3.7. Performance Evaluation</title><p>The prediction accuracy, confusion matrix, classification report and ROC curve are employed to evaluate the performance of SVM and NB classifiers. The Classification accuracy is the ratio of correctly classified data points to the total no. of points in the dataset which ranges from 0% - 100%.</p><p>Classification accuracy = Number of correct classifications Total number of classifications = TP + TN TP + TN + FP + FN (3.3)</p><p>Precision: is a metrics used to measure the positive classifications represented as follows?</p><p>Precision = TP TP + FP (3.4)</p><p>Recall is a metric used to measure the false negative classifications represented as:</p><p>Recall = TP TP + FN (3.5)</p><p>F1-score: takes into consideration the true positive and false positive regardless of false negative and false positive classifications. The F1-score is sensitive to which class is positive and negative as given below in Equation (3.6):</p><p>F1-score = 2 ∗ Precision ∗ Recall Precision + Recall = 2 ∗ TP 2 ∗ TP + FP + FN (3.6)</p><p>The RMSE is a diagnostic tool employed to evaluate the quality of model predictions. It shows how far the model predictions fall from measured true values using the Euclidean distance. The RMSE computes residual and mean of each data point with the square of the same mean. The RMSE can be expressed in Equation (3.7) as:</p><p>RMSE = ∑ i = 1 N ‖ Y ( i ) − y ( i ) ‖ 2 N ︷ (3.7)</p><p>where N is the number of points, Y(i): the i-th measurement and y ︷ is the corresponding predictions</p></sec></sec><sec id="s4"><title>4. Results and Discussion</title><p>The results of SVM and NB classifiers are obtained through the use of seaborn heatmap, clustering graph, Charts and tables. The heatmap was employed in visualizing the correlations between target variable and other attributes or variables of the dataset. Clustering graph to group the nodes of exiting and not exiting employees into two different clusters represented with red and green colors. The Bar plots to show the categorical data with heights proportional to the value it represented and tables as a useful structural representation of organizing data into rows and columns. The design and implementation was done with some varying finetuned hyper-parameter values to have a better classification result. The prediction and classification accuracy of both model are visualized and discussed using confusion matrix, ROC and classification report as given bellow.</p><p><xref ref-type="fig" rid="fig1">Figure 1</xref> is the heat map or correlation matrix used to measure the relationship between variables. The matrix depicts a linear correlation between all possible</p><p>pairs of employee experience, job security, working hours, no. of changed jobs, promotions, job exit and etc. There is a strong relationship as shown in the main diagonal and other pairs.</p><p><xref ref-type="fig" rid="fig2">Figure 2</xref> depicts the number of those employees perceived to be leaving represented with red and those not exiting using blue color obtained from the proposed system dataset. The exiting employees as obtained from the dataset gave 245 items and those not leaving produced 269 items as visualized.</p><p><xref ref-type="fig" rid="fig3">Figure 3</xref> shows the two different clusters employees using the NB classifier obtained from the proposed system dataset. The employees perceived to exit the organization are grouped into one cluster as represented with red color and those staying with another cluster with blue color.</p><p><xref ref-type="fig" rid="fig4">Figure 4</xref> depicts the employee years of working experience as visualized and arranged in ascending order. It ranges from 1 to 39 as obtained from the proposed system dataset for decision making. The employees with 39-years of experience are very few and the least in number which requires good treatment in retaining them and compared to those with 10-years of experience with the highest number of staff as shown in the Bar chart.</p><p><xref ref-type="fig" rid="fig5">Figure 5</xref> shows the employee working hours per day that ranges from 3, 4, 5 hours to a maximum of 45 hours including overtime as obtained from the dataset. Those employees that work 8 hours per day are the highest compare to those spend 45, 25, 15, 14 hour and so as show in the Bar plot.</p><p><xref ref-type="fig" rid="fig6">Figure 6</xref> depicts the confusion matrix of proposed SVM classifier with leading diagonal elements or values showing the total number of correctly predicted values</p><p>that are equal to the actual or true values above and below the main diagonal cell values or off-diagonal elements shows the wrongly predicted values. The higher the diagonal values the better the prediction accuracy. From the confusion matrix: The total No. of correct predictions = TP + TN = 76 + 75 = 151 and wrong predictions = FP+ FN = 4 + 0 = 4.</p><p><xref ref-type="table" rid="table2">Table 2</xref> depicts the classification report of SVM containing the precession, recall and f1-score accuracy of exiting and not exiting employees. The precision accuracy score those employees not leaving the organization produced 0.95 and those exiting gave 1.00. The recall score for exiting employees recorded 0.95 and those not leaving to be 1.00 and f1-score 0.97 for both employees either exiting or not leaving.</p><p><xref ref-type="fig" rid="fig7">Figure 7</xref> shows the confusion matrix of the proposed NB classifier at testing stage with the correct predictions displayed at the secondary diagonal and wrongly predicted values recorded above and below the main diagonal called the off-diagonal elements. The total No. of correct predictions = TP + TN = 76 + 79 = 155 and wrongly predicted = FP + FN = 0 + 0 = 0 shown in <xref ref-type="fig" rid="fig7">Figure 7</xref> where TP is true positive, FP false positive, FN false negative and TN true negative.</p><table-wrap id="table2" ><label><xref ref-type="table" rid="table2">Table 2</xref></label><caption><title> The classification report of SVM</title></caption><table><tbody><thead><tr><th align="center" valign="middle" ></th><th align="center" valign="middle" >Precision</th><th align="center" valign="middle" >Recall</th><th align="center" valign="middle" >F1-score</th><th align="center" valign="middle" >Support</th></tr></thead><tr><td align="center" valign="middle" >0</td><td align="center" valign="middle" >0.95</td><td align="center" valign="middle" >1.00</td><td align="center" valign="middle" >0.97</td><td align="center" valign="middle" >76</td></tr><tr><td align="center" valign="middle" >1</td><td align="center" valign="middle" >1.00</td><td align="center" valign="middle" >0.95</td><td align="center" valign="middle" >0.97</td><td align="center" valign="middle" >79</td></tr><tr><td align="center" valign="middle" >Accuracy</td><td align="center" valign="middle" ></td><td align="center" valign="middle" >0.97</td><td align="center" valign="middle" >0.97</td><td align="center" valign="middle" >155</td></tr><tr><td align="center" valign="middle" >Macro avg</td><td align="center" valign="middle" >0.97</td><td align="center" valign="middle" >0.97</td><td align="center" valign="middle" >0.97</td><td align="center" valign="middle" >155</td></tr><tr><td align="center" valign="middle" >Weighted avg</td><td align="center" valign="middle" >0.98</td><td align="center" valign="middle" >0.97</td><td align="center" valign="middle" >0.97</td><td align="center" valign="middle" >155</td></tr></tbody></table></table-wrap><table-wrap id="table3" ><label><xref ref-type="table" rid="table3">Table 3</xref></label><caption><title> The classification report of NB</title></caption><table><tbody><thead><tr><th align="center" valign="middle" ></th><th align="center" valign="middle" >Precision</th><th align="center" valign="middle" >Recall</th><th align="center" valign="middle" >F1-score</th><th align="center" valign="middle" >Support</th></tr></thead><tr><td align="center" valign="middle" >0</td><td align="center" valign="middle" >1.00</td><td align="center" valign="middle" >1.00</td><td align="center" valign="middle" >1.00</td><td align="center" valign="middle" >76</td></tr><tr><td align="center" valign="middle" >1</td><td align="center" valign="middle" >1.00</td><td align="center" valign="middle" >1.00</td><td align="center" valign="middle" >1.00</td><td align="center" valign="middle" >79</td></tr><tr><td align="center" valign="middle" >Accuracy</td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" >1.00</td><td align="center" valign="middle" >155</td></tr><tr><td align="center" valign="middle" >Macro avg</td><td align="center" valign="middle" >1.00</td><td align="center" valign="middle" >1.00</td><td align="center" valign="middle" >1.00</td><td align="center" valign="middle" >155</td></tr><tr><td align="center" valign="middle" >Weighted avg</td><td align="center" valign="middle" >1.00</td><td align="center" valign="middle" >1.00</td><td align="center" valign="middle" >1.00</td><td align="center" valign="middle" >155</td></tr></tbody></table></table-wrap><p><xref ref-type="table" rid="table3">Table 3</xref> shows the classification report of NB classifier with precession, recall and f1-score classification accuracy of 1.00 for exiting and not exiting employees in an organization. The precision accuracy score gage 1.00, recall 1.00 and f1-score to be 1.00. There is a diplomatic tire as recorded for the precision, recall and f1-score values from the classification report.</p><p><xref ref-type="fig" rid="fig8">Figure 8</xref> is the Receiver Operating Characteristic (ROC) graph of SVM and NB classifiers showing the trade-off between sensitivity or true positive rate and specificity (1-FPR). The NB ROC curve is closer to top-left corner of the graph and is perfect and performed better than SVM model. The SVM curve is a bit away from the top X- and Y-axis with respect to the number of False Positive Rate (FPR) and True Positive Rate (TPR) as shown in the ROC graph. The proposed gave points lying along the diagonal (True Positive Rate = False Positive Rate) as expected.</p><table-wrap id="table4" ><label><xref ref-type="table" rid="table4">Table 4</xref></label><caption><title> Training and validation test accuracy</title></caption><table><tbody><thead><tr><th align="center" valign="middle" >Model</th><th align="center" valign="middle" >Prediction Accuracy (%)</th><th align="center" valign="middle" >RMSE</th></tr></thead><tr><td align="center" valign="middle" >SVM</td><td align="center" valign="middle" >97.0%</td><td align="center" valign="middle" >0.0258</td></tr><tr><td align="center" valign="middle" >NB</td><td align="center" valign="middle" >100%</td><td align="center" valign="middle" >0.0000</td></tr></tbody></table></table-wrap><p><xref ref-type="table" rid="table4">Table 4</xref> shows the prediction accuracy measured in percentage and RMSE of SVM and NB classifiers. The result of NB classifier is recorded 100% to be faster with no RMSE value compared to the SVM techniques that produced 97.0% accuracy with 0.0258 RMSE value of testing dataset.</p></sec><sec id="s5"><title>5. Conclusion</title><p>The prediction of the accuracy of NB was higher compared to SVM in terms of prediction accuracy and RMSE for all fine-tuned hyper-parameter values in determining the anticipated exit of skilled and highly experienced employees from an organization. This will help industries detect and prevent the possible occurrence of experienced employees’ exit that may pose a danger to their throughput and can serve as a benchmark to other researchers because the model is scalable. The SVM prediction accuracy was high but recorded with a small error rate compared to the NB with a 100% accuracy rate with zero RMSE margin. The use of Python programming language simplified the implementation task because it has several machine learning inbuilt libraries and classes with deployable tools which can be achieved through a few lines of codes been optimized to achieve its best performance level.</p></sec><sec id="s6"><title>Conflicts of Interest</title><p>The authors declare no conflicts of interest.</p></sec><sec id="s7"><title>Cite this paper</title><p>Emmanuel-Okereke, I.L. and Anigbogu, S.O. (2022) Predicting the Perceived Employee Tendency of Leaving an Organization Using SVM and Naive Bayes Techniques. Open Access Library Journal, 9: e8497. https://doi.org/10.4236/oalib.1108497</p></sec></body><back><ref-list><title>References</title><ref id="scirp.116209-ref1"><label>1</label><mixed-citation publication-type="other" xlink:type="simple">Khera, S. and Divya, N. (2018) Predictive Modelling of Employee Turnover in Indian IT Industry Using Machine Learning Techniques. Vision, 23, 12-21.  
https://doi.org/10.1177/0972262918821221</mixed-citation></ref><ref id="scirp.116209-ref2"><label>2</label><mixed-citation publication-type="other" xlink:type="simple">Pandiyan, P., Kannadasan, K. and Vinoth, R. (2013) Prospective Control in an Organization through Two Grade Systems. 4D International Journal of Multidisciplinary Research and Development, 1, 22-25.</mixed-citation></ref><ref id="scirp.116209-ref3"><label>3</label><mixed-citation publication-type="other" xlink:type="simple">Kalaivani, J., Vinoth, R. and Elangovan, S.R. (2014) Survival Time to Trace the threshold Grade Level in an Organization. International Journal of Multidisciplinary Research and Development, 1, 22-25.</mixed-citation></ref><ref id="scirp.116209-ref4"><label>4</label><mixed-citation publication-type="other" xlink:type="simple">Morrell, K., Loan-Clarke, J. and Wilkingson, A. (2004) The Role of Shocks in Employee Turnover. British Journal of Management, 15, 335-349.  
https://doi.org/10.1111/j.1467-8551.2004.00423.x</mixed-citation></ref><ref id="scirp.116209-ref5"><label>5</label><mixed-citation publication-type="other" xlink:type="simple">Kannadasan, K., Pandiyan, P., Vinoth, R. and Saminathan, R. (2013) Time to Recruitment in an Organization through Three Parameter Generalized Exponential Model. Journal of Reliability and Statistical Studies, 6, 21-28.</mixed-citation></ref><ref id="scirp.116209-ref6"><label>6</label><mixed-citation publication-type="other" xlink:type="simple">Morrell, K. (2005) Towards a Typology of Nursing Turnover: The Role of Shocks in Nurses’ Decision to Leave. Journal of Advanced Nursing, 49, 315-322.  
https://doi.org/10.1111/j.1365-2648.2004.03290.x</mixed-citation></ref><ref id="scirp.116209-ref7"><label>7</label><mixed-citation publication-type="other" xlink:type="simple">Kuwaiti, A.A., Raman, V., Subbarayalu, A.V., Palanivel, R.M. and Prabaharan, S. (2018) Predicting the Exit Time of Employees in an Organization Using Statistical Model. International Journal of Scientific and Technology Research (IJSTR), 5, 213-217.</mixed-citation></ref><ref id="scirp.116209-ref8"><label>8</label><mixed-citation publication-type="other" xlink:type="simple">Saradhi, V.V. and Palshikar, G.K. (2011) Employee Churn Prediction. Expert Systems with Applications, 38, 19-30. https://doi.org/10.1016/j.eswa.2010.07.134</mixed-citation></ref><ref id="scirp.116209-ref9"><label>9</label><mixed-citation publication-type="other" xlink:type="simple">Ramamurthy, K.N., Singh, M., Davis, M., Kevern, J.A., Klein, U. and Peran, M. (2015) Identifying Employees for Re-Skilling Using an Analytics-Based Approach. 2015 IEEE International Conference on Data Mining Workshop (ICDMW), Atlantic City, 14-17 November 2015, 345-354. https://doi.org/10.1109/ICDMW.2015.206</mixed-citation></ref><ref id="scirp.116209-ref10"><label>10</label><mixed-citation publication-type="other" xlink:type="simple">Singh, M., Varshney, K.R., Wang, J., Mojsilovic, A., Gill, A.R., Faur, P.I. and Ezry, R. (2012) An Analytics Approach for Proactively Combating Voluntary Attrition of Employees. 2012 IEEE 12th International Conference on Data Mining Workshops (ICDMW), Brussels, 10-13 December 2012, 317-323.  
https://doi.org/10.1109/ICDMW.2012.136</mixed-citation></ref><ref id="scirp.116209-ref11"><label>11</label><mixed-citation publication-type="other" xlink:type="simple">Maharjan, R. (2011) Employee Churn Prediction Using Logistic Regression and Support Vector Machine Support Vector Machine. SJSU Scholar Works: Master Degree Work, 1-72.</mixed-citation></ref><ref id="scirp.116209-ref12"><label>12</label><mixed-citation publication-type="other" xlink:type="simple">Jayadi, R., Firmantyo, H.M., Dzaka, M.T.J., Sualdy, M. and Putra, A.M. (2019) Employee Performance Prediction Using Naive Bayes. International Journal of Advanced Trends in Computer Science and Engineering, 6, 3031-3035.  
https://doi.org/10.30534/ijatcse/2019/59862019</mixed-citation></ref><ref id="scirp.116209-ref13"><label>13</label><mixed-citation publication-type="other" xlink:type="simple">Yahia, N.B., Hlel, J. and Colomo-Palacios, R. (2017) From Big Data to Deep Data to Su- pport People Analytics for Employee Attrition Prediction. IEEE Access, 34, 1-12.</mixed-citation></ref><ref id="scirp.116209-ref14"><label>14</label><mixed-citation publication-type="other" xlink:type="simple">Kamath, R.S., Jamsandekar, S.S. and Naik, P.G. (2019) Machine Learning Approach for Employee Attrition Analysis. International Journal of Trend in Scientific Research and Development (IJTSRD), 5, 62-67. https://doi.org/10.31142/ijtsrd23065</mixed-citation></ref><ref id="scirp.116209-ref15"><label>15</label><mixed-citation publication-type="other" xlink:type="simple">Alshehhi, K., Zawbaa, S.B. and Tariq, M.U. (2021) Employee Retention Prediction in Corporate Organizations Using Machine Learning Methods. Academic of Entrepreneurship Journal, 27, 1-16.</mixed-citation></ref><ref id="scirp.116209-ref16"><label>16</label><mixed-citation publication-type="other" xlink:type="simple">Senanayake, D., Muthugama, L., Mendis, L. and Madushanka, T. (2015) Customer Ch- urn Prediction: A Cognitive Approach. Internation Journal of Computer, Electrical, Automation, Control and Information Engineering, 9, 23-43.</mixed-citation></ref><ref id="scirp.116209-ref17"><label>17</label><mixed-citation publication-type="other" xlink:type="simple">Foley, A.E. (2019) Using Machine Learning to Predict Employee Resignation in the Swe- dish Armed Forces. Kth Royal Institute of Technology School of Industrial Engineering and Management, Stockholm, 1-85.</mixed-citation></ref><ref id="scirp.116209-ref18"><label>18</label><mixed-citation publication-type="other" xlink:type="simple">Panjasuchat, M. and Limpiyakorn, Y. (2020) Applying Reinforcement Learning for Cus- tomer Churn Prediction. Journal of Physics: Conference Series, 1619, 12015.  
https://doi.org/10.1088/1742-6596/1619/1/012016</mixed-citation></ref><ref id="scirp.116209-ref19"><label>19</label><mixed-citation publication-type="other" xlink:type="simple">Zhang, X., Zhu, J., Xu, S. and Wan, Y. (2012) Predicting Customer Churn through Interpersonal Influence. Knowledge-Based Systems, 28, 97-104.  
https://doi.org/10.1016/j.knosys.2011.12.005</mixed-citation></ref><ref id="scirp.116209-ref20"><label>20</label><mixed-citation publication-type="other" xlink:type="simple">Bryant, P.C. and Allen, D.G. (2013) Compensation, Benefits and Employee Turnover: HR Strategies for Retaining Top Talent. Compensation and Benefits Review, 45, 171- 175. https://doi.org/10.1177/0886368713494342</mixed-citation></ref><ref id="scirp.116209-ref21"><label>21</label><mixed-citation publication-type="other" xlink:type="simple">Byun, H. and Lee, S.W. (2003) A Survey on Pattern Recognition Applications of Support Vector Machines. International Journal of Pattern Recognition and Artificial Intelligence (IJPRAI), 17, 459-486. https://doi.org/10.1142/S0218001403002460</mixed-citation></ref><ref id="scirp.116209-ref22"><label>22</label><mixed-citation publication-type="other" xlink:type="simple">Ramakrishnan, R., Bhattacharya, S. and Dhanya, P. (2018) Predict Employee Attrition by Using Predictive Analytics. Benchmarking: An International Journal, 26, 2-18.  
https://doi.org/10.1108/BIJ-03-2018-0083</mixed-citation></ref></ref-list></back></article>