<?xml version="1.0" encoding="UTF-8"?><!DOCTYPE article  PUBLIC "-//NLM//DTD Journal Publishing DTD v3.0 20080202//EN" "http://dtd.nlm.nih.gov/publishing/3.0/journalpublishing3.dtd"><article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" dtd-version="3.0" xml:lang="en" article-type="research article"><front><journal-meta><journal-id journal-id-type="publisher-id">OJBM</journal-id><journal-title-group><journal-title>Open Journal of Business and Management</journal-title></journal-title-group><issn pub-type="epub">2329-3284</issn><publisher><publisher-name>Scientific Research Publishing</publisher-name></publisher></journal-meta><article-meta><article-id pub-id-type="doi">10.4236/ojbm.2020.84102</article-id><article-id pub-id-type="publisher-id">OJBM-101579</article-id><article-categories><subj-group subj-group-type="heading"><subject>Articles</subject></subj-group><subj-group subj-group-type="Discipline-v2"><subject>Business&amp;Economics</subject></subj-group></article-categories><title-group><article-title>
 
 
  Detection of Fraud Patterns in Accounting Accounts Using Data Mining Techniques
 
</article-title></title-group><contrib-group><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Alexander</surname><given-names>Báez Hernández</given-names></name><xref ref-type="aff" rid="aff1"><sup>1</sup></xref></contrib><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Debrayan</surname><given-names>Bravo Hidalgo</given-names></name><xref ref-type="aff" rid="aff2"><sup>2</sup></xref></contrib></contrib-group><aff id="aff1"><addr-line>AdministrativeSciencesDepartment, Universidad Central del Ecuador, Quito, Ecuador</addr-line></aff><aff id="aff2"><addr-line>Department of Electrical Engineering and Electronics, Universidad San Francisco de Quito (USFQ), Quito, Ecuador</addr-line></aff><pub-date pub-type="epub"><day>04</day><month>06</month><year>2020</year></pub-date><volume>08</volume><issue>04</issue><fpage>1609</fpage><lpage>1618</lpage><history><date date-type="received"><day>4,</day>	<month>June</month>	<year>2020</year></date><date date-type="rev-recd"><day>17,</day>	<month>July</month>	<year>2020</year>	</date><date date-type="accepted"><day>20,</day>	<month>July</month>	<year>2020</year></date></history><permissions><copyright-statement>&#169; Copyright  2014 by authors and Scientific Research Publishing Inc. </copyright-statement><copyright-year>2014</copyright-year><license><license-p>This work is licensed under the Creative Commons Attribution International License (CC BY). http://creativecommons.org/licenses/by/4.0/</license-p></license></permissions><abstract><p>
 
 
  Accounting databases with fraudulent transactions inside was used to detect fraud patterns by data mining tool. The object was accomplished by the following method: first, inside data
  ,
   fraudulent transactions according 
  to 
  three
   
  fraud pattern
  s
   w
  ere
   settled, over it the algorithm
  s
  , Euclidian distance and local outlier factors w
  ere
   run using Rapidminer program. As 
  a 
  result the fraud patterns w
  ere 
  show
  n
   in different ways according to the specific graphics contributed by the program. In conclusion, clusters grouping by Euclidian distance with k Means algorithm (k = 4) allowed an adequate visualization of the values
  ’
   distribution, as consequence was detected the first and third fraud patterns. The application of the outlier’s detection algorithm (LOF) detected the three fraud patterns in a clear way as a consequence of the insolate outliers in different graphics, show
  n
   by Rapidminer program, with different variables correlated.
 
</p></abstract><kwd-group><kwd>Recognition</kwd><kwd> Financial Fraud</kwd><kwd> Tool</kwd><kwd> Data Mining</kwd></kwd-group></article-meta></front><body><sec id="s1"><title>1. Introduction</title><p>Accounting data analysis has become a necessary process for fraud detecting. It is a currency problem that has increased the interest in accounting researching (Seo, Choi, Choi, Lee, &amp; Lee, 2009). Fraud Detection and control have raised relevance in the audit process (Westhausen, 2017) where the informatics program is taking more presence as a tool capable to perceive the deception in accounting data (Gottschalk &amp; Solli-Saether, 2010).</p><p>Data mining is one of the most important current models of modern intelligent business analysis and decision support tools. This relevance is accepted by the transcendental professional accounting organizations. The American Institute of Certified Public Accountants (AICPA) has identified data mining as one of the top ten technologies going forward, and the Institute of Internal Auditors (IIA) has listed data mining as one of the top four research priorities (Koh &amp; Low, 2004). Additionally, Chartered Global Management Accountants (CGMA) has reported that N 50% of corporate leaders rank big data and data mining among the top ten corporate priorities that are central to the data-driven business era (CGMA, 2013). Data mining has been defined as the process of identifying valid, potentially novel, and ultimately understandable patterns in the data (Pujari, 2001). It is also known as the process of extracting knowledge from massive amounts of data (Han et al., 2006) to increase the efficiency of decisions in a particular discipline. The key approach to data mining, therefore, is to leverage an organization's data assets for financial or non-financial benefits. Therefore, data mining has been applied to almost all commercial and non-commercial disciplines, including accounting.</p><p>Data mining has given the possibility to analyze big data where some elements have been considered in untypical way or behavior in its input or manage as a support of traditional making decisions (Mraović, 2008). It has turned in strong tool to analyze the accounting data (Debreceny &amp; Gray, 2010; Joudaki et al., 2015; Qin, 2014; Wu, Ou, Lin, Chang, &amp; Yen, 2012). Business knowledge, data interpretation and logical reason to apply analytic method to build the data are embraced with data mining analysis (Jackson, 2002). A lot of algorithms for different analyses are possible in data mining (Wu et al., 2008). A broad scope of application in accounting science was resumed by Amani &amp; Fadlalla (2017) in the following statements:</p><p>1) Data classification and mapping to predefine qualitative attributes in itself</p><p>2) Data clustering in clusters with specifical signification</p><p>3) Focus in predictive approach to find future numerical values or a classification of no numerical values</p><p>4) Outliers detection, indeed values that are out of range</p><p>5) Optimization, the best solution choice for a means set</p><p>6) Visualization, adequate image for the best data comprehension</p><p>7) Regression, estimation of a dependent variable from a set of independent variables.</p><p>Outliers’ detection was done for two different algorithms:</p><p>1) Outliers detection considered the distance it locates higher atypical values in a data set no labeled to get subset of it. This subset is denominated outliers solution set (Angiulli, Basta, &amp; Pizzuti, 2006).</p><p>2) Detecting outlier in a data set considered a Local Outlier Factor (LOF). El LOF is focused on local density where the location is expressed for more close K neighbors whose distance is used to estimate the density. When it compares the local density from object with the local density of their neighbors it possible to identify areas with similar density and point with minor density that its neighbors. It is outliers (Breunig, Kriegel, Ng, &amp; Sander, 2000). First data classification used a common clustering suggested by Abonyi &amp; Feil (2007) and K middle algorithmic considered of frequency utilization for this purpose, according to Likas, Vlassis, &amp; Verbeek (2003). Detecting of fraud pattern in accounting data with outlier data mining algorithmic was the objective.</p></sec><sec id="s2"><title>2. Methods</title><p>Data base was made with real assent of a month from enterprise accounting bank. Inside of it was settling transactions from three-fraud pattern.</p><p>1) A little Withdraws (credit) of 0.05 cents many times.</p><p>2) Withdraws (credit), a constant amount with regular period of time.</p><p>3) A high quantitative withdraw (credit), more than average account entries.</p><p>RapidMiner program was used for Data Analysis according to the argument from (Amer &amp; Goldstein, 2012; Hofmann &amp; Klinkenberg, 2013; Jungermann, 2009). In the process, RapidMiner program reads first the database from Excel and selects (filter) the real values to avoid null values in attributes analyzed in the algorithmics as first step. Continued was applied the algorithmic process, Euclidian distance and local outlier factors as is showed in <xref ref-type="fig" rid="fig1">Figure 1</xref> and <xref ref-type="fig" rid="fig2">Figure 2</xref>.</p><p>The outliers were determined by operator using LOF algorithmic and local density in it (Breunig et al., 2000).</p></sec><sec id="s3"><title>3. Result and Discussion</title><sec id="s3_1"><title>3.1. Bibliometrics Analysis of Detection of Fraud Pattern in Accounting Accounts Using Data Mining Techniques</title><p>Using as search criteria “detection of fraud and data mining” in the title, abstract, and the keywords of the scientific contributions in the Scopus academic directory. 921 results were detected between Conference Paper, Article and Review. <xref ref-type="fig" rid="fig3">Figure 3</xref> shows the quantitative evolution that this line of research has undergone in recent years. The trend of the curve is increasing, in this figure. It is evident that the international scientific community considers fraud detection techniques through data mining as an effective and practical method. This figure was made using the bibliometric analysis tools provided by Scupus.</p><p>The bibliographic information of each of the 921 detected documents was imported from the Scopus platform, in (.ris) format. This process allowed this data to be processed in the VOSviewer science bibliometric analysis software. This computational tool allows building and visualizing bibliometric networks of scientific activity in journals, researchers or publications. As well as generating maps of scientific activity based on citations, bibliographic coupling, co-citations or authorship relationships, among others. The thermos research limitations density map shows the concentration of the most recent terms in a given research area. <xref ref-type="fig" rid="fig4">Figure 4</xref> shows a term density map that shows which are the most frequent terms and their relationship in research in this area of knowledge. This type of figures allows a very clear idea of current research trends and trends in the subject in question.</p></sec><sec id="s3_2"><title>3.2. Detection of Fraud Pattern in Accounting Accounts Using Data Mining Techniques</title><p>&#183; Clustering</p><p>Cluster groups are observable in <xref ref-type="fig" rid="fig5">Figure 5</xref>. They are maxims values. In red color is perceived an outlier value that means a fraud pattern. Three outliers grouped in dark blue color the account entries attached at 0.05 as value. The high level of operation was inferred by dark color around the point as evidence of fraud pattern. The calculated centroids for this clustering distribution are presented in <xref ref-type="table" rid="table1">Table 1</xref>.</p><p>&#183; Detection by amount</p><p><xref ref-type="fig" rid="fig6">Figure 6</xref> was made with the same data that was used to obtain <xref ref-type="fig" rid="fig5">Figure 5</xref>. A lot of transaction (high frequency) between 0 - 5000.00 was showed. The difference with <xref ref-type="fig" rid="fig5">Figure 5</xref> is that values around 0.05 aren’t separate in a group. Indeed, it does not detect the one pattern fraud. Analogous condition occurred with fraud Pattern three, where the outlier values aren’t insolated because they are inside of a range of high values transaction. Withdraw frequency distribution doesn’t let looking the fraud patterns. This is only possible with data maning tecniques.</p><p>True cluster separated in a group high withdraw including a value from the pattern three. These values are considered as suspicious, out of range. False cluster grouped in dark blue color a lot of small transaction around cero, leading toward one pattern fraud (<xref ref-type="fig" rid="fig7">Figure 7</xref>).</p><p>Blue dark color is representative of a cluster that group small values around cero. It is an indication of fraud pattern one. The values designed with color that trend to dark red are relevant because they are high withdrawn. It must be considered in the representative sample in account audit universe. It, also, included a value representative of the pattern fraud three (<xref ref-type="fig" rid="fig8">Figure 8</xref>).</p><p>These figures exposed the most precise information. In dark blue color are a lot of small account entries per day around cero, evidence of the fraud pattern 1. In spectrum of red color appear per day transactions with high values, suspicions of</p><p>irregular entries settle in accounts. The matter is that available information per day gives a temporality sequence in the way that the fraud happened to building architecture used in its committed. It is visible in <xref ref-type="fig" rid="fig9">Figure 9</xref> clusters of transactions in dark color blue concentrate in specifics days. This indicated a suspicious regular performance in the accounting operations associated with time. It’s a way that detected the second pattern as a paradox: an irregular proceeding committed as shady regular accounting record correlated with time period.</p><p>Research limitations:</p><p>&#183; Bibliographic references you will consider are from the Scopus database.</p><p>&#183; Only manuscripts published in the English language are considered.</p><p>&#183; Database was made only with real assent of a month from enterprise accounting bank.</p></sec></sec><sec id="s4"><title>4. Conclusion</title><p>The application of Data Mining using the detection of Outliers (both Distance Algorithm and Local Outliers Factor) made it possible to detect (fraudulent) patterns injected into an accounting database that simulated three possible types of fraud:</p><table-wrap id="table1" ><label><xref ref-type="table" rid="table1">Table 1</xref></label><caption><title> Centroids calculation for withdrawing</title></caption><table><tbody><thead><tr><th align="center" valign="middle" >Attribute</th><th align="center" valign="middle" >Cluster 0</th><th align="center" valign="middle" >Cluster 1</th><th align="center" valign="middle" >Cluster 2</th><th align="center" valign="middle" >Cluster 3</th></tr></thead><tr><td align="center" valign="middle" >WITHDRAW</td><td align="center" valign="middle" >12420.83</td><td align="center" valign="middle" >28537.636</td><td align="center" valign="middle" >5679.66</td><td align="center" valign="middle" >882.63</td></tr></tbody></table></table-wrap><p>1) A little Withdraws (credit) of 0.05 cents many times.</p><p>2) Withdraws (credit), a constant amount with a regular period of time.</p><p>3) A high quantitative withdraw (credit), more than average account entries.</p><p>The detection of Outliers offered the following possibilities:</p><p>The clusters grouping indicated, in a simple way, the possible movements’ distribution for auditors. The use of the k Means algorithm (k = 4) allowed an adequate visualization of the values’ distribution. It has detected the pattern third in cluster one, and first pattern in cluster three (<xref ref-type="fig" rid="fig5">Figure 5</xref>. Clustering Withdraw). The application of the outlier's detection algorithm (LOF) detected, <xref ref-type="fig" rid="fig6">Figure 6</xref> and <xref ref-type="fig" rid="fig7">Figure 7</xref>, the first and third pattern in a clear way as a consequence of the insolate outliers (true and false) or transactions density (high and low) Daily Outliers distribution, <xref ref-type="fig" rid="fig5">Figure 5</xref>, detected all the patterns. It’s the only that remarked the second pattern as an expression of a paradoxid result: a regular transaction correlated with time variable is an irregular procedure in accounting data. It’s the core idea used to detect the second pattern.</p></sec><sec id="s5"><title>Conflicts of Interest</title><p>The authors declare no conflicts of interest regarding the publication of this paper.</p></sec><sec id="s6"><title>Cite this paper</title><p>Hern&#225;ndez, A. B., &amp; Hidalgo, D. B. (2020). Detection of Fraud Patterns in Accounting Accounts Using Data Mining Techniques. Open Journal of Business and Management, 8, 1609-1618. https://doi.org/10.4236/ojbm.2020.84102</p></sec></body><back><ref-list><title>References</title><ref id="scirp.101579-ref1"><label>1</label><mixed-citation publication-type="other" xlink:type="simple">Abonyi, J., &amp; Feil, B. (2007). Cluster Analysis for Data Mining and System Identification. Basel: Birkhauser.</mixed-citation></ref><ref id="scirp.101579-ref2"><label>2</label><mixed-citation publication-type="other" xlink:type="simple">Amani, F. A., &amp; Fadlalla, A. M. (2017). Data Mining Applications in Accounting: A Review of the Literature and Organizing Framework. International Journal of Accounting Information Systems, 24, 32-58. https://doi.org/10.1016/j.accinf.2016.12.004</mixed-citation></ref><ref id="scirp.101579-ref3"><label>3</label><mixed-citation publication-type="other" xlink:type="simple">Amer, M., &amp; Goldstein, M. (2012). Nearest-Neighbor and Clustering Based Anomaly Detection Algorithms for RapidMiner. Proc. of the 3rd RapidMiner Community Meeting and Conference (RCOMM 2012) (pp. 1-12).</mixed-citation></ref><ref id="scirp.101579-ref4"><label>4</label><mixed-citation publication-type="other" xlink:type="simple">Angiulli, F., Basta, S., &amp; Pizzuti, C. (2006). Distance-Based Detection and Prediction of Outliers. IEEE Transactions on Knowledge and Data Engineering, 18, 145-160.  
https://doi.org/10.1109/TKDE.2006.29</mixed-citation></ref><ref id="scirp.101579-ref5"><label>5</label><mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Breunig</surname><given-names> M. M.</given-names></name>,<name name-style="western"><surname> Kriegel</surname><given-names> H.-P.</given-names></name>,<name name-style="western"><surname> Ng</surname><given-names> R. T.</given-names></name>,<name name-style="western"><surname> &amp; Sander</surname><given-names> J. </given-names></name>,<etal>et al</etal>. (<year>2000</year>)<article-title>. LOF: Identifying Density-Based Local Outliers</article-title><source> ACM SIGMOD Record</source><volume> 29</volume>,<fpage> 93</fpage>-<lpage>104</lpage>.<pub-id pub-id-type="doi"></pub-id></mixed-citation></ref><ref id="scirp.101579-ref6"><label>6</label><mixed-citation publication-type="other" xlink:type="simple">Chartered Global Management Accountant (CGMA) (2013). Report from Insight to Impact: Unlocking Opportunities in Big Data.  
http://www.cgma.org/Resources/Reports/DownloadableDocuments/From_insight_to_impact-unlocking_the_opportunities_in_big_data.pdf</mixed-citation></ref><ref id="scirp.101579-ref7"><label>7</label><mixed-citation publication-type="other" xlink:type="simple">Debreceny, R. S., &amp; Gray, G. L. (2010). Data Mining Journal Entries for Fraud Detection: An Exploratory Study. International Journal of Accounting Information Systems, 11, 157-181. https://doi.org/10.1016/j.accinf.2010.08.001</mixed-citation></ref><ref id="scirp.101579-ref8"><label>8</label><mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Gottschalk</surname><given-names> P.</given-names></name>,<name name-style="western"><surname> &amp; Solli-Saether</surname><given-names> H. </given-names></name>,<etal>et al</etal>. (<year>2010</year>)<article-title>. Computer Information Systems in Financial Crime Investigations</article-title><source> Journal of Computer Information Systems</source><volume> 50</volume>,<fpage> 41</fpage>-<lpage>49</lpage>.<pub-id pub-id-type="doi"></pub-id></mixed-citation></ref><ref id="scirp.101579-ref9"><label>9</label><mixed-citation publication-type="other" xlink:type="simple">Han, J., Kamber, M., &amp; Pei, J. (2006). Data Mining: Concepts and Techniques. Burlington, MA: Morgan Kaufmann.</mixed-citation></ref><ref id="scirp.101579-ref10"><label>10</label><mixed-citation publication-type="other" xlink:type="simple">Hofmann, M., &amp; Klinkenberg, R. (2013). RapidMiner: Data Mining Use Cases and Business Analytics Applications. Boca Raton, FL: CRC Press.</mixed-citation></ref><ref id="scirp.101579-ref11"><label>11</label><mixed-citation publication-type="other" xlink:type="simple">Jackson, J. (2002). Data Mining; A Conceptual Overview. Communications of the Association for Information Systems, 8, 19. https://doi.org/10.17705/1CAIS.00819</mixed-citation></ref><ref id="scirp.101579-ref12"><label>12</label><mixed-citation publication-type="other" xlink:type="simple">Joudaki, H., Rashidian, A., Minaei-Bidgoli, B., Mahmoodi, M., Geraili, B., Nasiri, M., &amp; Arab, M. (2015). Using Data Mining to Detect Health Care Fraud and Abuse: A Review of Literature. Global Journal of Health Science, 7, 194.  
https://doi.org/10.5539/gjhs.v7n1p194</mixed-citation></ref><ref id="scirp.101579-ref13"><label>13</label><mixed-citation publication-type="other" xlink:type="simple">Jungermann, F. (2009). Information Extraction with Rapidminer. The Proceedings of the GSCL Symposium Sprachtechnologie und eHumanities, TU Dortmund.</mixed-citation></ref><ref id="scirp.101579-ref14"><label>14</label><mixed-citation publication-type="other" xlink:type="simple">Koh, H. C., &amp; Low, C. K. (2004). Going Concern Prediction Using Data Mining Techniques. Managerial Auditing Journal, 19, 462-476.  
https://doi.org/10.1108/02686900410524436</mixed-citation></ref><ref id="scirp.101579-ref15"><label>15</label><mixed-citation publication-type="other" xlink:type="simple">Likas, A., Vlassis, N., &amp; Verbeek, J. J. (2003). The Global k-Means Clustering Algorithm. Pattern Recognition, 36, 451-461. https://doi.org/10.1016/S0031-3203(02)00060-2</mixed-citation></ref><ref id="scirp.101579-ref16"><label>16</label><mixed-citation publication-type="other" xlink:type="simple">Mraovic, B. (2008). Relevance of Data Mining for Accounting: Social Implications. Social Responsibility Journal, 4, 439-455. https://doi.org/10.1108/17471110810909858</mixed-citation></ref><ref id="scirp.101579-ref17"><label>17</label><mixed-citation publication-type="other" xlink:type="simple">Pujari, A. K. (2001). Data Mining Techniques. Bangalore: Universities Press.</mixed-citation></ref><ref id="scirp.101579-ref18"><label>18</label><mixed-citation publication-type="other" xlink:type="simple">Qin, B. Z. (2014). Study on the Financial Audit Risk Prevention Based on Data Mining Technology. Applied Mechanics and Materials, 608-609, 351-354.  
https://doi.org/10.4028/www.scientific.net/AMM.608-609.351</mixed-citation></ref><ref id="scirp.101579-ref19"><label>19</label><mixed-citation publication-type="other" xlink:type="simple">Seo, K., Choi, J., Choi, Y., Lee, D., &amp; Lee, S. (2009). Research about Extracting and Analyzing Accounting Data of Company to Detect Financial Fraud. IEEE International Conference on Intelligence and Security Informatics (pp. 200-202).  
https://doi.org/10.1109/ISI.2009.5137302</mixed-citation></ref><ref id="scirp.101579-ref20"><label>20</label><mixed-citation publication-type="other" xlink:type="simple">Westhausen, H.-U. (2017). The Escalating Relevance of Internal Auditing as Anti-Fraud Control. Journal of Financial Crime, 24, 322-328.  
https://doi.org/10.1108/JFC-06-2016-0041</mixed-citation></ref><ref id="scirp.101579-ref21"><label>21</label><mixed-citation publication-type="other" xlink:type="simple">Wu, R. S., Ou, C. S., Lin, H. Y., Chang, S. I., &amp; Yen, D. C. (2012). Using Data Mining Technique to Enhance Tax Evasion Detection Performance. Expert Systems with Applications, 39, 8769-8777. https://doi.org/10.1016/j.eswa.2012.01.204</mixed-citation></ref><ref id="scirp.101579-ref22"><label>22</label><mixed-citation publication-type="other" xlink:type="simple">Wu, X., Kumar, V., Quinlan, J. R., Ghosh, J., Yang, Q., Motoda, H., Philip, S. Y. et al. (2008). Top 10 Algorithms in Data Mining. Knowledge and Information Systems, 14, 1-37. https://doi.org/10.1007/s10115-007-0114-2</mixed-citation></ref></ref-list></back></article>