<?xml version="1.0" encoding="UTF-8"?><!DOCTYPE article  PUBLIC "-//NLM//DTD Journal Publishing DTD v3.0 20080202//EN" "http://dtd.nlm.nih.gov/publishing/3.0/journalpublishing3.dtd"><article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" dtd-version="3.0" xml:lang="en" article-type="research article"><front><journal-meta><journal-id journal-id-type="publisher-id">IIM</journal-id><journal-title-group><journal-title>Intelligent Information Management</journal-title></journal-title-group><issn pub-type="epub">2160-5912</issn><publisher><publisher-name>Scientific Research Publishing</publisher-name></publisher></journal-meta><article-meta><article-id pub-id-type="doi">10.4236/iim.2014.62005</article-id><article-id pub-id-type="publisher-id">IIM-43830</article-id><article-categories><subj-group subj-group-type="heading"><subject>Articles</subject></subj-group><subj-group subj-group-type="Discipline-v2"><subject>Computer Science&amp;Communications</subject></subj-group></article-categories><title-group><article-title>
 
 
  Extracting Significant Patterns for Oral Cancer Detection Using Apriori Algorithm
 
</article-title></title-group><contrib-group><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>eha</surname><given-names>Sharma</given-names></name><xref ref-type="aff" rid="aff1"><sup>1</sup></xref><xref ref-type="corresp" rid="cor1"><sup>*</sup></xref></contrib><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Hari</surname><given-names>Om</given-names></name><xref ref-type="aff" rid="aff2"><sup>2</sup></xref><xref ref-type="corresp" rid="cor1"><sup>*</sup></xref></contrib></contrib-group><aff id="aff1"><addr-line>Padmashree Dr. D. Y. Patil Institute of Master of Computer Applications, Pune, India</addr-line></aff><aff id="aff2"><addr-line>Indian School of Mines, Dhanbad, India</addr-line></aff><author-notes><corresp id="cor1">* E-mail:<email>nvsharma@rediffmail.com(ES)</email>;<email>hariom4india@gmail.com(HO)</email>;</corresp></author-notes><pub-date pub-type="epub"><day>17</day><month>03</month><year>2014</year></pub-date><volume>06</volume><issue>02</issue><fpage>30</fpage><lpage>37</lpage><history><date date-type="received"><day>13</day>	<month>January</month>	<year>2014</year></date><date date-type="rev-recd"><day>12</day>	<month>February</month>	<year>2014</year>	</date><date date-type="accepted"><day>10</day>	<month>March</month>	<year>2014</year></date></history><permissions><copyright-statement>&#169; Copyright  2014 by authors and Scientific Research Publishing Inc. </copyright-statement><copyright-year>2014</copyright-year><license><license-p>This work is licensed under the Creative Commons Attribution International License (CC BY). http://creativecommons.org/licenses/by/4.0/</license-p></license></permissions><abstract><p>
 
 
   Presently, no effective tool exists for early diagnosis and treatment of oral cancer. Here, we describe an approach for cancer detection and prevention based on analysis using association rule mining. The data analyzed are pertaining to clinical symptoms, history of addiction, co-morbid condition and survivability of the cancer patients. The extracted rules are useful in taking clinical judgments and making right decisions related to the disease. The results shown here are promising and show the potential use of this approach toward eventual development of diagnostic assay and treatment with sufficient support and confidence suitable for detection of early-stage oral cancer. 
 
</p></abstract><kwd-group><kwd>Data Mining; Association Rule Mining; Apriori; Oral Cancer; WEKA</kwd></kwd-group></article-meta></front><body><sec id="s1"><title>1. Introduction</title><p>In this paper, we have adopted Fayyad et al.’s definition of knowledge discovery and data mining. Knowledge discovery is the “non-trivial process of identifying valid, novel, potentially useful and ultimately understandable patterns in data” [<xref ref-type="bibr" rid="scirp.43830-ref1">1</xref>] . Data mining is one of the steps in the process of knowledge discovery, consists of applying data analysis and discovery (learning) algorithms that produce a particular enumeration of patterns (or models) over the data. [<xref ref-type="bibr" rid="scirp.43830-ref2">2</xref>] [<xref ref-type="bibr" rid="scirp.43830-ref3">3</xref>] . Data mining could also be characterized as the procedure of finding useful patterns or meaning in raw data which subsequently can be used to develop a predictive model [<xref ref-type="bibr" rid="scirp.43830-ref3">3</xref>] -[<xref ref-type="bibr" rid="scirp.43830-ref5">5</xref>] . It is variously been called as KDD (knowledge discovery in databases), knowledge discovery, knowledge extraction, information discovery, information harvesting, data archeology and data pattern processing [<xref ref-type="bibr" rid="scirp.43830-ref6">6</xref>] . Knowledge discovery involves the additional steps of target data set selection, data preprocessing, and data reduction (reducing the number of variables), which occur prior to data mining. It also involves the additional steps of information interpretation and consolidation of the information extracted during the data mining process. These extracted patterns will provide useful knowledge to decision makers [<xref ref-type="bibr" rid="scirp.43830-ref7">7</xref>] .</p><p>Application of data mining are many—it has been utilized seriously and also widely by marketers, for direct marketing and cross-selling or up-selling; by financial institutions, for credit scoring and fraud detection; by manufacturers, for quality control and maintenance scheduling and by retailers, for market segmentation and store layout [<xref ref-type="bibr" rid="scirp.43830-ref8">8</xref>] . Data mining is becoming increasingly popular, if not increasingly essential in healthcare industry as well. It is so because modern medicine and various healthcare transactions generate almost daily, huge amounts of heterogeneous data that is too complex and voluminous to be processed and analyzed by traditional methods. For example, medical data may contain SPECT images, signals like ECG, clinical information like temperature, cholesterol levels, etc., as well as the physician’s interpretation. Those who deal with such data understand that there is a widening gap between data collection and data comprehension. Computerized techniques are needed to help humans address this problem [<xref ref-type="bibr" rid="scirp.43830-ref9">9</xref>] . Data mining provides the methodology and technology to transform these mounds of data into useful information for decision making, with the intention of offer valuable quality services at reasonable costs, which is a main concern envisage by the healthcare organizations (hospitals, medical centers). Data mining applications can incredibly profit all stake holders of the healthcare industry such as hospitals, clinics, physicians, and patients, for example, by identifying effective treatments and best practices.</p><p>This paper presents an application of data mining in early detection and prevention of oral cancer and discusses how the generated patterns can be effectively used by physicians. The World Health Organization’s Global Burden of Disease statistics distinguished malignancy or cancer as the second largest global cause of death, after cardiovascular disease [<xref ref-type="bibr" rid="scirp.43830-ref10">10</xref>] . Cancer is the fastest growing segment of the disease burden; global cancer deaths are anticipated to increase from 7.1 million in 2002 to 11.5 million in 2030 [<xref ref-type="bibr" rid="scirp.43830-ref11">11</xref>] . Advances in prevention, diagnostics and treatment of cancer have contributed to the improved prognosis for cancer patients: one third of cancers are preventable and another third are curable through early detection and effective therapy [<xref ref-type="bibr" rid="scirp.43830-ref12">12</xref>] .</p><p>The objective of this article is to explore relevant literature and present the same in Section 2, cover the information about oral cancer in Section 3, examine the data mining methodology and then appropriate algorithm to implement in Section 4. Experimental results are presented in Section 5 and finally Section 6 highlights the conclusions and offers some future directions. At the end, acknowledgement and references are mentioned.</p></sec><sec id="s2"><title>2. Literature Review</title><p>Kaladhar et al. [<xref ref-type="bibr" rid="scirp.43830-ref13">13</xref>] predict oral cancer survivability using the CART, Random Forest, LMT, and Na&#239;ve Bayesian classification algorithms, which classify the cancer survival using 10 fold cross validation and training dataset. Among these algorithms, the Random Forest technique classifies more accurately the cancer survival dataset as compared to other methods. Singh et al. [<xref ref-type="bibr" rid="scirp.43830-ref14">14</xref>] have applied the apriori algorithm with transaction reduction on the data of cancer symptoms by considering five different types of cancer to find the symptoms that help the cancer to spread and also the cancer type that spreads faster. Srikant et al. [<xref ref-type="bibr" rid="scirp.43830-ref15">15</xref>] have considered the problem of integrating constraints in the form of boolean expression that appoint the presence or absence of items in rules. Nahar et al. [<xref ref-type="bibr" rid="scirp.43830-ref16">16</xref>] discuss the significant prevention factors for a particular type of cancer. To find out the prevention factors, they have first constructed a prevention factor dataset through an extensive literature review and then three association rule mining algorithms: Apriori, Predictive apriori, and Tertius algorithms have been applied on that data to discover most of the significant prevention factors against a specific type of cancer. Experimental results illustrate that the Apriori is the most useful association rule-mining algorithm for discovery of the prevention factors.</p><p>Swami et al. [<xref ref-type="bibr" rid="scirp.43830-ref17">17</xref>] discuss the multidimensional association rules and the model for smoking habits in order to take some preventive measures to reduce the various habits of smoking in youths. Milovic et al. [<xref ref-type="bibr" rid="scirp.43830-ref18">18</xref>] discuss the applicability of data mining in healthcare and explain how the patterns can be used by physicians to determine diagnoses, prognoses, and apply for patients in healthcare organizations. A detailed survey on various methods adopted by the researchers for identification and classification of oral cancer detection at an earlier stage has been given in [<xref ref-type="bibr" rid="scirp.43830-ref19">19</xref>] . Chuang et al. [<xref ref-type="bibr" rid="scirp.43830-ref20">20</xref>] consider DNA repair genes by choosing a single nucleotide polymorphisms (SNPs) dataset with 238 samples of oral cancer and control patients for disease prediction. They report that the performance of the holdout cross validation is much better than cross validation and the best classification accuracy is 64.2%. Gadewal et al. [<xref ref-type="bibr" rid="scirp.43830-ref21">21</xref>] have enlarged the oral cancer gene database to 374 genes by adding 132 gene entries to enable fast retrieval of updated information.</p></sec><sec id="s3"><title>3. Oral Cancer</title><p>Oral malignancy is a heterogeneous assembly of tumors rolling out from diverse parts of the oral cavity, with distinctive predisposing factors, prevalence, and treatment outcomes. Oral tumor is one of the ten most incessant diseases worldwide with a yearly occurrence of over 300,000 cases, of which 62% arise in developing nations [<xref ref-type="bibr" rid="scirp.43830-ref22">22</xref>] [<xref ref-type="bibr" rid="scirp.43830-ref23">23</xref>] . There is a huge contrast in the rate of oral tumor in diverse regions of the worlds. The age-adjusted rates of oral tumor differ from over 20 for every 100,000 population in India, to 10 for every 100,000 population in the U.S., and less than 2 for every 100,000 population in the Middle East [<xref ref-type="bibr" rid="scirp.43830-ref24">24</xref>] [<xref ref-type="bibr" rid="scirp.43830-ref25">25</xref>] . In comparison with the U.S. population, where oral cavity malignancy represents only about of 3% of malignancies, it accounts for over 30% of all growths in India. It has been estimated that 83,000 new oral cancer cases occur every year iIn India [<xref ref-type="bibr" rid="scirp.43830-ref26">26</xref>] [<xref ref-type="bibr" rid="scirp.43830-ref27">27</xref>] . The variation in incidence and pattern of oral cancer is due to regional differences in the prevalence of risk factors. But as oral cancer has well-defined risk factors, these may be modified—giving real hope for primary prevention.</p><p>The clinician's issue is separating malignant lesions from a nearly infinite amount of other poorly characterized, questionable, and crudely comprehended lesions that also occur in the oral cavity. Most oral lesions are benign, yet many have a manifestation that may be effectively befuddled with threatening lesions and some are now considered pre-malignant because they have been statistically correlated with subsequently cancerous changes [<xref ref-type="bibr" rid="scirp.43830-ref28">28</xref>] . On the other hand, some malignant lesions seen in an early stage may mistaken for a benign. Early carcinomas are presumably asymptotic and ensuing signs are regularly misjudged in light of the fact that they imitate numerous benevolent lesions and the distress is negligible. Professional consultation is thus often delayed, increasing the chance for local spread and regional metastases. Stress must be placed on gaining access to high risk individuals for periodic oral examinations and educational efforts to increase the skill of primary health care providers in recognizing this problem. Squamous cell carcinoma accounts for 90% of the total number of malignant oral lesions. Therefore, the problem of oral cancer is primarily that of pathogenesis, diagnosis, and management of squamous cell carcinoma originating from oral muscular surface [<xref ref-type="bibr" rid="scirp.43830-ref29">29</xref>] . The aim of this work is to apply the association rule mining on the data pertaining to clinical symptoms, history of addiction, co-morbid condition and survivability in order to evaluate the clinical features, diagnosis, and treatment of oral cancer patients.</p></sec><sec id="s4"><title>4. Association Rule Mining</title><p>Data mining technique, association rule mining is applied to search the hidden relationships among the attributes. It identifies strong rules discovered in databases using different measures of interestingness. Thus, an association rule is a pattern that states when X occurs, Y occurs with certain probability. In this paper, we adopt the standard definition of association rules [<xref ref-type="bibr" rid="scirp.43830-ref30">30</xref>] -[<xref ref-type="bibr" rid="scirp.43830-ref36">36</xref>] .</p>Apriori Algorithm<p>The apriori is a classic algorithm for frequent item set mining and association rule learning over the transactional databases [<xref ref-type="bibr" rid="scirp.43830-ref37">37</xref>] . It proceeds by identifying the frequent individual items in the database and extending them to larger and larger item sets as long as those item sets appear sufficiently often in the database. The frequent item sets determined by a apriori can be used to determine association rules, which highlight general trends in the database [<xref ref-type="bibr" rid="scirp.43830-ref38">38</xref>] . Association rules mining using apriori algorithm uses a “bottom up” approach, breadth-first search and a hash tree structure to count the candidate item sets efficiently. A two-step apriori algorithm is explained with the help of flowchart as shown in  <xref ref-type="fig" rid="fig1">Figure 1</xref> and the algorithm is mentioned below:</p><p>Apriori algorithm: Candidate Generation and Test Approach Step 1: Initially, scan database (DB) once to get frequent 1-itemset.</p><p>Step 2: Generate length (k + 1) candidate item sets from length k frequent item sets.</p><p>Step 3: Test candidates against DB.</p><p>Step 4: Terminate, if no frequent or candidate set can be Generated.</p><p>To select interesting rules from the set of all possible rules generated, constraints on various measures of significance and interest can be used. The best-known constraints are minimum thresholds on support and confidence.</p><p>Support: The rule holds with support supp in T (the transaction data set) if supp % of transactions contain <inline-formula><inline-graphic xlink:href="tmlimages\2-8701282x\0c26ea6a-afea-4076-bc5a-e8bb1db906dd.png" xlink:type="simple"/></inline-formula> [<xref ref-type="bibr" rid="scirp.43830-ref39">39</xref>] .</p><p><inline-formula><inline-graphic xlink:href="tmlimages\2-8701282x\6dce0092-5c66-4a22-bc78-df41095cd02a.png" xlink:type="simple"/></inline-formula>.</p><p>Confidence: The rule holds with confidence conf in T if conf % of transactions that contain X also contain Y [<xref ref-type="bibr" rid="scirp.43830-ref39">39</xref>] [<xref ref-type="bibr" rid="scirp.43830-ref40">40</xref>] .</p><p><img src="htmlimages\2-8701282x\b15a1fa4-12f7-4ae1-a35a-bc58f1bda394.png" /></p><p>Lift: It is the probability of the observed support to that expected, if X and Y were independent [<xref ref-type="bibr" rid="scirp.43830-ref41">41</xref>].</p><p><img src="htmlimages\2-8701282x\2a93e6f1-a960-43cb-bce0-b7bf22af32b8.png" /></p><p>Leverage: It measures the difference of X and Y appearing together in the dataset and what would be expected if X and Y were statistically dependent [<xref ref-type="bibr" rid="scirp.43830-ref42">42</xref>] .</p><p><img src="htmlimages\2-8701282x\098eb55f-8783-4f7d-a968-1afc7245d285.png" /></p><p>Conviction: It is the probability of the expected frequency that X occurs without Y (that is to say, the frequency that the rule makes an incorrect prediction) [<xref ref-type="bibr" rid="scirp.43830-ref43">43</xref>] .</p><p><img src="htmlimages\2-8701282x\1db25dcc-23e9-4d14-96a9-164cb5c63cd0.png" /></p></sec><sec id="s5"><title>5. Experimental Results</title><p>The database for this work is created by collecting data through a retrospective chart review and the entire process is presented in [<xref ref-type="bibr" rid="scirp.43830-ref44">44</xref>] . There are total 33 variables and 1025 records of patients were created for the analysis. A data mining tool—WEKA 3.7.9 [<xref ref-type="bibr" rid="scirp.43830-ref45">45</xref>] has been used to explore the behaviour of the apriori algorithm for extracting the significant patterns for early detection of oral cancer. The oral cancer data is initially stored in MS Excel sheet, then converted into comma separated values (.csv file) and subsequently to attribute relation file format (.arff file), which is the acceptable format to WEKA tool. Minimum support defined by the tool for the generated rule is 0.1 (103 instances) and minimum confidence is 0.9. Association rules for early detection of the oral cancer patients on the basis clinical symptoms, history of addiction and co-morbid condition are mentioned below and the same is presented in the graphical form in Figures 2-4:</p><p>Rule 1. Clinical-Symptom = Ulcer (577) ==&gt; Survival = Dead (498) &lt; conf: (0.98) &gt; lift: (2.31) lev: (0.24) [<xref ref-type="bibr" rid="scirp.43830-ref246">246</xref>] conv:(28.27).</p><p>Rule 2. Clinical-Symptom = Burning-Sensation (287) ==&gt; Survival = alive (286) &lt; conf: (1) &gt; lift: (2.27) lev: (0.16) [<xref ref-type="bibr" rid="scirp.43830-ref160">160</xref>] conv: (80.64).</p><p>Rule Details: Rule 1 and Rule 2 suggest ulcer as the clinical symptom may indicate oral cancer with more certainty in comparison to other clinical symptoms like burning sensation, loosening of tooth and mass, which subsequently lead to high mortality.</p><p>Rule 3. History-of-Addiction = Tobacco-Smoking and History-of-Addiction1 = Alcohol (131) ==&gt; Survival = Dead (131) &lt; conf: (1) &gt; lift: (2.6) lev: (0.08) [<xref ref-type="bibr" rid="scirp.43830-ref80">80</xref>] conv: (80.64).</p><p>Rule Details: Rule 3 proposes that history of addiction like tobacco-smoking or tobacco-chewing and alcohol accounts for most oral cancers. Heavy smokers who use tobacco for a long time are most at risk. The risk is even</p><p>higher for tobacco users who drink alcohol heavily. In fact, three out of four oral cancers occur in people who use alcohol, tobacco, or both alcohol and tobacco.</p><p>Rule 4. Co-Morbid-Condition = Hypertension (261) ==&gt; Survival = Dead (261) &lt; conf: (1) &gt; lift: (1.78) lev: (0.11) [<xref ref-type="bibr" rid="scirp.43830-ref114">114</xref>] conv: (114.33).</p><p>Rule Details: Rule 4 puts forward that co-morbid condition like hypertension may also be a reason for oral cancer and subsequently for high mortality.</p><p>Rule 5. Clinical-Symptom = Ulcer and History-of-Addiction = Tobacco-Smoking (131) ==&gt; Survival = Dead (131) &lt; conf: (1) &gt; lift: (1.78) lev: (0.06) [<xref ref-type="bibr" rid="scirp.43830-ref57">57</xref>] conv: (57.38).</p><p>Rule Details: Rule 5 suggests that if clinical-Symptom is ulcer and history of addiction is tobacco-smoking, chances of patients suffering from oral cancer is high which may lead to mortality.</p><p>Rule 6. History-of-Addiction = Tobacco-Chewing and Co-Morbid-Condition = Hypertension (130) ==&gt; Survival = Dead (130) &lt; conf: (1) &gt; lift: (1.78) lev: (0.06) [<xref ref-type="bibr" rid="scirp.43830-ref56">56</xref>] conv: (56.95).</p><p>Rule 7. History-of-Addiction = Tobacco-Smoking and Co-Morbid-Condition = Hypertension (131) ==&gt; Survival = Dead (131) &lt; conf: (1) &gt; lift: (1.78) lev: (0.06) [<xref ref-type="bibr" rid="scirp.43830-ref57">57</xref>] conv: (57.38).</p><p>Rule Details: Rule 6 and Rule 7 hint that history of addiction like tobacco-smoking and tobacco-chewing along with co-morbid condition like hypertension increases the probability of oral cancer which may be the reason for high mortality.</p><p>The significant patterns generated using apriori algorithm can be summarized as follows:</p><p>If Clinical Symptoms = Ulcer.</p><p>If History of Addiction = Tobacco-Chewing/Smoking or Alcohol.</p><p>If Co-Morbid Condition = Hypertension.</p><p>ThenOral Cancer is suspected which has to be confirmed through biopsy and other diagnostic procedure.</p></sec><sec id="s6"><title>6. Conclusion and Future Work</title><p>The data mining technique which is adopted for the research work is association rule mining. The algorithm used for implementing association rule mining is a popular apriori algorithm. Apriori has been used to extract the association among various valuable data pertaining to clinical symptoms, history of addiction and co-morbid condition. The rules generated would certainly assist the practitioners in early discovery of oral cancer and consequently help in prevention of the disease. The experimental results demonstrate that all the generated rules hold the highest confidence level, thereby, making them very useful for early detection and prevention of oral cancer. In future, we intend to extend this research work by attempting to extract significant patterns and useful rules through the association rule mining algorithm using various attributes like predisposing factors, gross examination, tumor site, tumor size, neck node, etc. and use it more effectively for early detection and prevention of oral cancer.</p></sec><sec id="s7"><title>Acknowledgements</title><p>The authors would like to thank Dr. Vijay Sharma, MS, ENT, for his valuable contribution in understanding the occurrence and diagnosis of Oral Cancer. The authors devote their sincere thanks to the management and staff of Indian School of Mines, for their constant support and motivation.</p></sec></body><back><ref-list><title>References</title><ref id="scirp.43830-ref1"><label>1</label><mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Fayyad</surname><given-names> U. M.</given-names></name>,<name name-style="western"><surname> Piatetsky-Shapiro</surname><given-names> G. and Smyth</given-names></name>,<name name-style="western"><surname> P. </surname><given-names>  </given-names></name>,<etal>et al</etal>. (<year>1996</year>)<article-title>Data Mining to Knowledge Discovery in Databases</article-title><source> AI Magazine</source><volume> 17</volume>,<fpage> 37</fpage>-<lpage>54</lpage>.<pub-id pub-id-type="doi"></pub-id></mixed-citation></ref><ref id="scirp.43830-ref2"><label>2</label><mixed-citation publication-type="other" xlink:type="simple">Han, J., Kamber, M. and Pei, J. (2011) Data Mining: Concepts and Techniques. 3rd Edison, Morgan Kaufmann Publishers, Burlington.</mixed-citation></ref><ref id="scirp.43830-ref3"><label>3</label><mixed-citation publication-type="other" xlink:type="simple">Khosla, R. and Dillon, T. (1997) Knowledge Discovery, Data Mining and Hybrid Systems. Engineering Intelligent Hybrid Multi-Agent Systems, Kluwer Academic Publishers, Norwell, 143-177.</mixed-citation></ref><ref id="scirp.43830-ref4"><label>4</label><mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Kincade</surname><given-names> K. </given-names></name>,<etal>et al</etal>. (<year>1998</year>)<article-title>Data Mining: Digging for Healthcare Gold</article-title><source> Insurance &amp; Technology</source><volume> 23</volume>,<fpage> IM2</fpage>-<lpage>IM7</lpage>.<pub-id pub-id-type="doi"></pub-id></mixed-citation></ref><ref id="scirp.43830-ref5"><label>5</label><mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Milley</surname><given-names> A. </given-names></name>,<etal>et al</etal>. (<year>2000</year>)<article-title>Healthcare and Data Mining</article-title><source> Health Management Technology</source><volume> 21</volume>,<fpage> 44</fpage>-<lpage>47</lpage>.<pub-id pub-id-type="doi"></pub-id></mixed-citation></ref><ref id="scirp.43830-ref6"><label>6</label><mixed-citation publication-type="other" xlink:type="simple">Fayyad, U.M., Piatetsky-Shapiro, G. and Smyth, P. (1996) Data Mining to Knowledge Discovery: An Overview. Advances in Knowledge Discovery and Data Mining, AAAI Press/MIT Press, 1-36.</mixed-citation></ref><ref id="scirp.43830-ref7"><label>7</label><mixed-citation publication-type="other" xlink:type="simple">Houston, A.L., Chen, H., Hubbard, S.M., Schatz, B.R. , Ng, T.D., Sewell, R.R. and Tolle, K.M. (1999) Medical Data Mining on the Internet: Research on a Cancer Information System. Artificial Intelligence Review, 13, 437-466. http://dx.doi.org/10.1023/A:1006548623067</mixed-citation></ref><ref id="scirp.43830-ref8"><label>8</label><mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Koh</surname><given-names> H.C. and Tan</given-names></name>,<name name-style="western"><surname> G. </surname><given-names>  </given-names></name>,<etal>et al</etal>. (<year>2005</year>)<article-title>Data Mining Applications in Healthcare</article-title><source> Journal of Healthcare Information Management</source><volume> 19</volume>,<fpage> 64</fpage>-<lpage>72</lpage>.<pub-id pub-id-type="doi"></pub-id></mixed-citation></ref><ref id="scirp.43830-ref9"><label>9</label><mixed-citation publication-type="other" xlink:type="simple">Cios, K.J. (2001) Medical Data Mining and Knowledge Discovery. Studies in Fuzziness and Soft Computing, 60, 502.</mixed-citation></ref><ref id="scirp.43830-ref10"><label>10</label><mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Mathers</surname><given-names> C.D. and Loncar</given-names></name>,<name name-style="western"><surname> D. </surname><given-names>  </given-names></name>,<etal>et al</etal>. (<year>2006</year>)<article-title>Projections of Global Mortality and Burden of Disease</article-title><source> PLOS Medicine</source><volume> 3</volume>,<fpage> 2002</fpage>-<lpage>2030</lpage>.<pub-id pub-id-type="doi"></pub-id></mixed-citation></ref><ref id="scirp.43830-ref11"><label>11</label><mixed-citation publication-type="other" xlink:type="simple">World Health Organization (2007) Department of Measurement and Health Information Systems: World Health Statistics 2007, World Health Organization, Geneva.</mixed-citation></ref><ref id="scirp.43830-ref12"><label>12</label><mixed-citation publication-type="other" xlink:type="simple">World Health Organization (2002) Department of Management of Noncommunicable Diseases: National Cancer Control Programmes, World Health Organization, Geneva.</mixed-citation></ref><ref id="scirp.43830-ref13"><label>13</label><mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Kaladhar</surname><given-names> D.S.V.G.K.</given-names></name>,<name name-style="western"><surname> Chandana</surname><given-names> B. and Kumar</given-names></name>,<name name-style="western"><surname> P.B. </surname><given-names>  </given-names></name>,<etal>et al</etal>. (<year>2011</year>)<article-title>Predicting Cancer Survivability Using Classification Algorithms</article-title><source> International Journal of Research and Reviews in Computer Science (IJRRCS)</source><volume> 2</volume>,<fpage> 340</fpage>-<lpage>343</lpage>.<pub-id pub-id-type="doi"></pub-id></mixed-citation></ref><ref id="scirp.43830-ref14"><label>14</label><mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Singh</surname><given-names> S.</given-names></name>,<name name-style="western"><surname> Yadav</surname><given-names> M. and Gupta</given-names></name>,<name name-style="western"><surname> H. </surname><given-names>  </given-names></name>,<etal>et al</etal>. (<year>2012</year>)<article-title>Finding the Chances and Prediction of Cancer through Apriori Algorithm with Transaction Reduction</article-title><source> International Journal of Advanced Computer Research</source><volume> 2</volume>,<fpage> 23</fpage>-<lpage>28</lpage>.<pub-id pub-id-type="doi"></pub-id></mixed-citation></ref><ref id="scirp.43830-ref15"><label>15</label><mixed-citation publication-type="other" xlink:type="simple">Srikant, R., Vu, Q. and Agrawal, R. (1997) Mining Association Rules with Item Constraints. Proceeding KDD9, New Port Beach, 14-17 August 1997, 67-73.</mixed-citation></ref><ref id="scirp.43830-ref16"><label>16</label><mixed-citation publication-type="other" xlink:type="simple">Nahar, J., Kevin, S.T., Ali, A.B.M.S. and Chen, Y.P. (2009) Significant Cancer Prevention Factor Extraction: An Association Rule Discovery Approach. Journal of Medical Systems, 35, 353-367. DOI 10.1007/s10916-009-9372-8</mixed-citation></ref><ref id="scirp.43830-ref17"><label>17</label><mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Swami</surname><given-names> S.</given-names></name>,<name name-style="western"><surname> Thakur</surname><given-names> R.S. and Chandel</given-names></name>,<name name-style="western"><surname> R.S. </surname><given-names>  </given-names></name>,<etal>et al</etal>. (<year>2011</year>)<article-title>Multi-Dimensional Association Rules Extraction in Smoking Habits Database</article-title><source> International Journal of Advanced Networking and Applications</source><volume> 3</volume>,<fpage> 1176</fpage>-<lpage>1179</lpage>.<pub-id pub-id-type="doi"></pub-id></mixed-citation></ref><ref id="scirp.43830-ref18"><label>18</label><mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Milovic</surname><given-names> B. and Milovic</given-names></name>,<name name-style="western"><surname> M. </surname><given-names>  </given-names></name>,<etal>et al</etal>. (<year>2012</year>)<article-title>Prediction and Decision Making in Health Care Using Data Mining</article-title><source> International Journal of Public Health Science</source><volume> 1</volume>,<fpage> 69</fpage>-<lpage>78</lpage>.<pub-id pub-id-type="doi"></pub-id></mixed-citation></ref><ref id="scirp.43830-ref19"><label>19</label><mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Anuradha</surname><given-names> K. and Sankaranarayanan</given-names></name>,<name name-style="western"><surname> K. </surname><given-names>  </given-names></name>,<etal>et al</etal>. (<year>2012</year>)<article-title>Identification of Suspicious Regions to Detect Oral Cancers at an Earlier Stage—A Literature SURVEY</article-title><source> International Journal of Advances in Engineering &amp; Technology</source><volume> 3</volume>,<fpage> 84</fpage>-<lpage>91</lpage>.<pub-id pub-id-type="doi"></pub-id></mixed-citation></ref><ref id="scirp.43830-ref20"><label>20</label><mixed-citation publication-type="other" xlink:type="simple">Chuang, L.Y., Wu, K.C., Chang, H.W. and Yang, C.H. (2011) Support Vector Machine-Based Prediction for Oral Cancer Using Four SNPs in DNA Repair Genes. Proceedings of the International Multi Conference of Engineers and Computer Scientists, Hong Kong,16-18 March 2011, 16-18.</mixed-citation></ref><ref id="scirp.43830-ref21"><label>21</label><mixed-citation publication-type="other" xlink:type="simple">Gadewal, N.S. and Zingde, S.M. (2011) Database and Interaction Network of Genes Involved in Oral Cancer: Version II. Bioinformation, 6, 169-170. http://dx.doi.org/10.6026/97320630006169</mixed-citation></ref><ref id="scirp.43830-ref22"><label>22</label><mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Perkin</surname><given-names> D.M. and Lara</given-names></name>,<name name-style="western"><surname> E. </surname><given-names>  </given-names></name>,<etal>et al</etal>. (<year>1980</year>)<article-title>Estimates of the World Wide Frequency of Sixteen Major Cancers</article-title><source> International Journal of Cancer</source><volume> 41</volume>,<fpage> 184</fpage>-<lpage>197</lpage>.<pub-id pub-id-type="doi"></pub-id></mixed-citation></ref><ref id="scirp.43830-ref23"><label>23</label><mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Elango</surname><given-names> J.K.</given-names></name>,<name name-style="western"><surname> Gangadharan</surname><given-names> P.</given-names></name>,<name name-style="western"><surname> Sumithra</surname><given-names> S. and Kuriakose</given-names></name>,<name name-style="western"><surname> M.A. </surname><given-names>  </given-names></name>,<etal>et al</etal>. (<year>2006</year>)<article-title>Trends of Head and Neck Cancers in Urban And Rural India</article-title><source> Asian Pacific Journal of Cancer Prevention</source><volume> 7</volume>,<fpage> 108</fpage>-<lpage>112</lpage>.<pub-id pub-id-type="doi"></pub-id></mixed-citation></ref><ref id="scirp.43830-ref24"><label>24</label><mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Sankaranarayan</surname><given-names> R.</given-names></name>,<name name-style="western"><surname> Masuyer</surname><given-names> E.</given-names></name>,<name name-style="western"><surname> Swaminathan</surname><given-names> R.</given-names></name>,<name name-style="western"><surname> Ferley</surname><given-names> J. and Whelan</given-names></name>,<name name-style="western"><surname> S. </surname><given-names>  </given-names></name>,<etal>et al</etal>. (<year>1998</year>)<article-title>Head and Neck Cancer: A Global Perspective on Epidemiology and Prognosis</article-title><source> Anticancer Research</source><volume> 18</volume>,<fpage> 4779</fpage>-<lpage>4786</lpage>.<pub-id pub-id-type="doi"></pub-id></mixed-citation></ref><ref id="scirp.43830-ref25"><label>25</label><mixed-citation publication-type="other" xlink:type="simple">Sankaranarayanan, R., Ramadas, K. and Thomas, K. (2005) Effect of Screening on Oral Cancer Mortality in Kerala, India: A Cluster-Randomised Controlled Trial. The Lancet, 365, 1927-1933. http://dx.doi.org/10.1016/S0140-6736(05)66658-5</mixed-citation></ref><ref id="scirp.43830-ref26"><label>26</label><mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Manoharan</surname><given-names> N.</given-names></name>,<name name-style="western"><surname> Tyagi</surname><given-names> B.B. and Raina</given-names></name>,<name name-style="western"><surname> V. </surname><given-names>  </given-names></name>,<etal>et al</etal>. (<year>2004</year>)<article-title>Cancer incidences in rural Delhi</article-title><source> Asian Pacific Journal of Cancer Prevention</source><volume> 11</volume>,<fpage> 73</fpage>-<lpage>78</lpage>.<pub-id pub-id-type="doi"></pub-id></mixed-citation></ref><ref id="scirp.43830-ref27"><label>27</label><mixed-citation publication-type="other" xlink:type="simple">Agrawal, M., Pandey, S., Jain, S. and Maitin, S. (2012) Oral Cancer Awareness of the General Public in Gorakhpur City, India. Asian Pacific Journal of Cancer Prevention, 13, 5195-5199. http://dx.doi.org/10.7314/APJCP.2012.13.10.5195</mixed-citation></ref><ref id="scirp.43830-ref28"><label>28</label><mixed-citation publication-type="other" xlink:type="simple">American Cancer Society (1996) Cancer Facts and figures, Atlanta (GA), The Society.</mixed-citation></ref><ref id="scirp.43830-ref29"><label>29</label><mixed-citation publication-type="other" xlink:type="simple">Cancer Research Capign (1990) Oral Cancer. Fact Sheet, 14.</mixed-citation></ref><ref id="scirp.43830-ref30"><label>30</label><mixed-citation publication-type="other" xlink:type="simple">Agrawal, R., Imielinski, T. and Swami, A. (1993) Mining Association Rules between Sets of Items in Large Databases. Proceedings of the 1993 ACM SIGMOD International Conference on Management of Data, Washington DC, May 1993, 207-216.</mixed-citation></ref><ref id="scirp.43830-ref31"><label>31</label><mixed-citation publication-type="other" xlink:type="simple">An, J., Chen, Y.P.P. and Chen, H. (2005) DDR: An Index Method for Large Time Series Datasets. Information Systems, 30, 333-348. http://dx.doi.org/10.1016/j.is.2004.05.001</mixed-citation></ref><ref id="scirp.43830-ref32"><label>32</label><mixed-citation publication-type="other" xlink:type="simple">Chen, Y.P.P. and Chen, F. (2008) Targets for Drug Discovery Using Bioinformatics. Expert Opinion on Therapeutic Targets, 12, 383-389. http://dx.doi.org/10.1517/14728222.12.4.383</mixed-citation></ref><ref id="scirp.43830-ref33"><label>33</label><mixed-citation publication-type="other" xlink:type="simple">Lau, R.Y.K., Tang, M., Wong, O., Milliner, S.W. and Chen, Y.P.P. (2006) An Evolutionary Learning Approach for Adaptive Negotiation Agents. International Journal of Intelligent Systems, 21, 41-72. http://dx.doi.org/10.1002/int.20120</mixed-citation></ref><ref id="scirp.43830-ref34"><label>34</label><mixed-citation publication-type="other" xlink:type="simple">Ordonez, C. (2006) Association Rule Discovery with the Train and Test Approach for Heart Disease Prediction. IEEE Transaction on Information Technology. Biomed, 10, 334-343. http://dx.doi.org/10.1109/TITB.2006.864475</mixed-citation></ref><ref id="scirp.43830-ref35"><label>35</label><mixed-citation publication-type="other" xlink:type="simple">Ordonez, C. and Omiecinski, E. (1999) Discovering Association Rules Based on Image Content. IEEE Advances in Digital Libraries Conference (ADL’99), Baltimore, 19-21 May 1999, 38-49.</mixed-citation></ref><ref id="scirp.43830-ref36"><label>36</label><mixed-citation publication-type="other" xlink:type="simple">Ordonez, C., Santana, C.A. and Braal, L. (2000) Discovering Interesting Association Rules in Medical Data. ACM DMKD Workshop, Dallas, 14 May 2000, 78-85.</mixed-citation></ref><ref id="scirp.43830-ref37"><label>37</label><mixed-citation publication-type="other" xlink:type="simple">Agrawal, R. and Srikant, R. (1994) Fast Algorithms for Mining Association Rules in Large Databases. Proceedings of the 20th International Conference on Very Large Data Bases, VLDB, Santiago de Chile, 12-15 September 1994, 487-499.</mixed-citation></ref><ref id="scirp.43830-ref38"><label>38</label><mixed-citation publication-type="other" xlink:type="simple">Zaki, M.J. (2004) Mining Non-Redundant Association Rules. Data Mining and Knowledge Discovery, 9, 223-248. http://dx.doi.org/10.1023/B:DAMI.0000040429.96086.c7</mixed-citation></ref><ref id="scirp.43830-ref39"><label>39</label><mixed-citation publication-type="other" xlink:type="simple">Agrawal, R., Imielinski, T. and Swami, A. (1993) Mining Association Rules between Sets of Items in Large Databases. Proceedings of the 1993 ACM SIGMOD International Conference on Management of Data, Washington DC, 26-28 May 1993, 207-216.</mixed-citation></ref><ref id="scirp.43830-ref40"><label>40</label><mixed-citation publication-type="other" xlink:type="simple">Hipp, J., Güntzer, U. and Nakhaeizadeh, G. (2000) Algorithms for Association Rule Mining—A General Survey and Comparison. ACM SIGKDD Explorations Newsletter, 2, 58-64. http://dx.doi.org/10.1145/360402.360421</mixed-citation></ref><ref id="scirp.43830-ref41"><label>41</label><mixed-citation publication-type="other" xlink:type="simple">Brin, S., Motwani, R., Ullman, J.D. and Tsur, S. (1997) Dynamic Itemset Counting and Implication Rules for Market Basket Data. Proceedings of the ACM SIGMOD International Conference on Management of Data (SIGMOD 1997), Tucson, 13-15 May 1997, 255-264.</mixed-citation></ref><ref id="scirp.43830-ref42"><label>42</label><mixed-citation publication-type="other" xlink:type="simple">Piatetsky-Shapiro, G. (1991) Discovery, Analysis, and Presentation of Strong Rules. Knowledge Discovery in Databases. AAAI/MIT Press, Cambridge, 248, 255-264.</mixed-citation></ref><ref id="scirp.43830-ref43"><label>43</label><mixed-citation publication-type="other" xlink:type="simple">Brin, S., Motwani, R., Ullman, J.D. and Tsur, S. (1997) Dynamic Itemset Counting and Implication Rules for Market Basket Data. Proceedings of the ACM SIGMOD International Conference on Management of Data (SIGMOD 1997), Tucson, 13-15 May 1997, 265-276.</mixed-citation></ref><ref id="scirp.43830-ref44"><label>44</label><mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Sharma</surname><given-names> N. and Om</given-names></name>,<name name-style="western"><surname> H. </surname><given-names>  </given-names></name>,<etal>et al</etal>. (<year>2012</year>)<article-title>Framework for Early Detection and Prevention of Oral Cancer Using Data Mining</article-title><source> International Journal of Advances in Engineering &amp; Technology</source><volume> 4</volume>,<fpage> 302</fpage>-<lpage>310</lpage>.<pub-id pub-id-type="doi"></pub-id></mixed-citation></ref><ref id="scirp.43830-ref45"><label>45</label><mixed-citation publication-type="other" xlink:type="simple">Witten, I.H. and Frank, E. (2005) Data Mining: Practical Machine Learning Tool and Techniques. 2nd Edition, Elsevier, Amsterdam.</mixed-citation></ref></ref-list></back></article>