<?xml version="1.0" encoding="UTF-8"?><!DOCTYPE article  PUBLIC "-//NLM//DTD Journal Publishing DTD v3.0 20080202//EN" "http://dtd.nlm.nih.gov/publishing/3.0/journalpublishing3.dtd"><article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" dtd-version="3.0" xml:lang="en" article-type="research article"><front><journal-meta><journal-id journal-id-type="publisher-id">JBiSE</journal-id><journal-title-group><journal-title>Journal of Biomedical Science and Engineering</journal-title></journal-title-group><issn pub-type="epub">1937-6871</issn><publisher><publisher-name>Scientific Research Publishing</publisher-name></publisher></journal-meta><article-meta><article-id pub-id-type="doi">10.4236/jbise.2016.91002</article-id><article-id pub-id-type="publisher-id">JBiSE-62911</article-id><article-categories><subj-group subj-group-type="heading"><subject>Articles</subject></subj-group><subj-group subj-group-type="Discipline-v2"><subject>Biomedical&amp;Life Sciences</subject></subj-group></article-categories><title-group><article-title>
 
 
  Application of Word Embedding to Drug Repositioning
 
</article-title></title-group><contrib-group><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>uc</surname><given-names>Luu Ngo</given-names></name><xref ref-type="aff" rid="aff1"><sup>1</sup></xref></contrib><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Naoki</surname><given-names>Yamamoto</given-names></name><xref ref-type="aff" rid="aff1"><sup>1</sup></xref></contrib><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Vu</surname><given-names>Anh Tran</given-names></name><xref ref-type="aff" rid="aff1"><sup>1</sup></xref></contrib><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Ngoc</surname><given-names>Giang Nguyen</given-names></name><xref ref-type="aff" rid="aff1"><sup>1</sup></xref></contrib><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Dau</surname><given-names>Phan</given-names></name><xref ref-type="aff" rid="aff1"><sup>1</sup></xref></contrib><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Favorisen</surname><given-names>Rosyking Lumbanraja</given-names></name><xref ref-type="aff" rid="aff1"><sup>1</sup></xref></contrib><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Mamoru</surname><given-names>Kubo</given-names></name><xref ref-type="aff" rid="aff2"><sup>2</sup></xref></contrib><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Kenji</surname><given-names>Satou</given-names></name><xref ref-type="aff" rid="aff2"><sup>2</sup></xref><xref ref-type="corresp" rid="cor1"><sup>*</sup></xref></contrib></contrib-group><aff id="aff2"><addr-line>Institute of Science and Engineering, Kanazawa University, Kanazawa, Japan</addr-line></aff><aff id="aff1"><addr-line>Graduate School of Natural Science and Technology, Kanazawa University, Kanazawa, Japan</addr-line></aff><author-notes><corresp id="cor1">* E-mail:<email>ken@t.kanazawa-u.ac.jp(KS)</email>;</corresp></author-notes><pub-date pub-type="epub"><day>20</day><month>01</month><year>2016</year></pub-date><volume>09</volume><issue>01</issue><fpage>7</fpage><lpage>16</lpage><history><date date-type="received"><day>2</day>	<month>December</month>	<year>2015</year></date><date date-type="rev-recd"><day>accepted</day>	<month>18</month>	<year>January</year>	</date><date date-type="accepted"><day>21</day>	<month>January</month>	<year>2016</year></date></history><permissions><copyright-statement>&#169; Copyright  2014 by authors and Scientific Research Publishing Inc. </copyright-statement><copyright-year>2014</copyright-year><license><license-p>This work is licensed under the Creative Commons Attribution International License (CC BY). http://creativecommons.org/licenses/by/4.0/</license-p></license></permissions><abstract><p>
 
 
  As a key technology of rapid and low-cost drug development, drug repositioning is getting popular. In this study, a text mining approach to the discovery of unknown drug-disease relation was tested. Using a word embedding algorithm, senses of over 1.7 million words were well represented in sufficiently short feature vectors. Through various analysis including clustering and classification, feasibility of our approach was tested. Finally, our trained classification model achieved 87.6% accuracy in the prediction of drug-disease relation in cancer treatment and succeeded in discovering novel drug-disease relations that were actually reported in recent studies.
 
</p></abstract><kwd-group><kwd>Distributed Representation of Word Sense</kwd><kwd> Discovery of Drug-Disease Relation</kwd><kwd> Word Analogy</kwd></kwd-group></article-meta></front><body><sec id="s1"><title>1. Introduction</title><p>To develop an effective and highly-demanded drug, hundreds million dollars and 10 or more years for R &amp; D and clinical trial are typically required. Structure-based drug design (SBDD) is actively studied to reduce the cost and time by in-silico screening of candidate chemicals [<xref ref-type="bibr" rid="scirp.62911-ref1">1</xref>] [<xref ref-type="bibr" rid="scirp.62911-ref2">2</xref>] ; however, still it requires long time for tests on animals and human. Against such a background, a concept of drug repositioning (or drug repurposing, reprofiling, etc.) is attracting much interest and expectation from academic researchers and pharmaceutical companies [<xref ref-type="bibr" rid="scirp.62911-ref3">3</xref>] . One of the famous examples of drug repositioning is a treatment of multiple myeloma by thalidomide that was initially developed for relieving nausea and vomiting in pregnancy. Since drug repositioning means reuse of approved drugs for another purpose, their safety and method of production have already been confirmed.</p><p>Besides biomedical experiments, computational methods are developed for drug repositioning. Most of them adopt network-based algorithms and combination of various databases including gene expression and pathway data [<xref ref-type="bibr" rid="scirp.62911-ref4">4</xref>] . On the other hand, it is also suggested that text mining has much potential for drug repositioning. In biomedical text mining, named entities (genes, proteins, etc.) are recognized and the relations among them are extracted (e.g. “Gefitinib” “EGFR”). Additionally, biomedical ontologies or WordNet [<xref ref-type="bibr" rid="scirp.62911-ref5">5</xref>] are utilized for the sources of semantic information. In this study, we applied word embedding, implemented as word2vec [<xref ref-type="bibr" rid="scirp.62911-ref6">6</xref>] - [<xref ref-type="bibr" rid="scirp.62911-ref8">8</xref>] , for efficient representation of semantic information of words in a sufficiently large subset of PubMed abstracts. Through the clustering and classification experiments especially on anti-cancer drugs and cancer-re- lated diseases, it is suggested that the word vectors, generated by word embedding for drugs and diseases, are representing rich semantic information and promising for drug repositioning.</p></sec><sec id="s2"><title>2. Materials and Methods</title><p>In this section, data and algorithms are described. Overview of processing pipeline is shown in <xref ref-type="fig" rid="fig1">Figure 1</xref>.</p><sec id="s2_1"><title>2.1. Raw Corpus as a Set of Sentences</title><p>As a raw corpus, we used a subset of PubMed abstracts downloaded in October 2013, filtered by the keyword “cancer”. From 3,099,076 abstracts, 14,847,050 sentences were extracted.</p></sec><sec id="s2_2"><title>2.2. Parsing</title><p>Enju [<xref ref-type="bibr" rid="scirp.62911-ref9">9</xref>] was used for POS recognition of words and conversion into base forms. Since the sentences were extracted from biomedical abstracts, “-genia” option was specified. As a result, part-of-speech (POS) and base form are recognized for each words.</p><fig id="fig1"  position="float"><label><xref ref-type="fig" rid="fig1">Figure 1</xref></label><caption><title> Overview of processing pipeline. Box colors indicate: light blue for corpus, light green for databases, yellow for word vectors, and pink for concatenated word vectors</title></caption><graphic mimetype="image"   position="float"  xlink:type="simple"  xlink:href="http://html.scirp.org/file/2-9102253x6.png"/></fig></sec><sec id="s2_3"><title>2.3. Post Processing of Parsing Result</title><p>So that word2vec can differently treat the same word with different POS categories, they were attached right after the base form of words (e.g. “care” -&gt;“care(V)”). For readability, nouns are kept as is. To simplify the input for word2vec, we removed all words except nouns, adjectives, adverbs, and verbs.</p></sec><sec id="s2_4"><title>2.4. Named Entity Recognition and Conversion into Single Words</title><p>Biological terms typically consist of two or more words. In addition, they have many synonyms. Since word2vec basically treats a sentence as a sequence of words, it is needed to recognize biological synonyms, aggregate them into primary terms, and convert them into single words (e.g. “yolk sac tumor” -&gt;“endodermal sinus tumor” -&gt;“endodermal_sinus_tumor”). In this study, primary names and synonyms of drugs, diseases, and genes were extracted from PharmGKB [<xref ref-type="bibr" rid="scirp.62911-ref10">10</xref>] and used for recognition and aggregation (genes are used only for showing distribution of word vectors). For each converted single words, prefixes indicating their semantic categories were attached for later processing (e.g. “endodermal_sinus_tumor” -&gt;“disease::endodermal_sinus_tumor”).</p><p>Related to the conversion above, we need to consider about the existence of original single words. Firstly, if a synonym word is aggregated into primary word, the original word disappears and is not used for word embedding. Secondly, if a multi-word term is converted into a single word, all the original single words in the multi-word term disappear. Thirdly, if two multi-word term occur in a sentence with overlapping, it is impossible to replace both of them at once. To avoid there problems, a sentence is converted into the sentences which containing at most one converted word per sentence. For example, if a sentence contains two terms to be converted, three sentences including original one are generated. After all conversion, 14,847,050 sentences are expanded to 45,264,480.</p></sec><sec id="s2_5"><title>2.5. Word Embedding</title><p>In the field of natural language processing and text mining, computational representation of a linguistic unit (e.g. documents, paragraphs, sentences, terms, and words) is essential. The simplest one for document is bag-of-word model, which represents each document as a vector of word frequencies in it. In case of word representation, only the neighboring words in the same sentence are counted. For better analysis, stop-words are removed and raw frequencies are modified by term weighting such as tf-idf. After that, these vectors are used to evaluate the characteristics of the units and similarities between them (vector space model).</p><p>One of the serious problems in such a representation and analysis is high dimensionality and sparseness of vectors. For instance, 10 millions of sentences may contain one million of different words, then the dimension of a vector is also one million. In addition, since frequency of word follows Zipf’s law, most of the one million of words only occur a few times, which makes the vectors quite sparse. Though there exist traditional algorithms for dimension reduction or compression like Principal Component Analysis (PCA) and Latent Semantic Analysis (LSA), this problem is not fully solved.</p><p>Word embedding for distributed representation of word sense is a new approach to this problem. Based on neural network algorithm, reasonably short numerical vectors (e.g. 100 dimensions) are calculated for all words in a set of sentences. Through the application studies, it is proved that the vector space constructed by word embedding represents word senses and distances (similarities) between them quite well. Additionally, in this space of word sense, word analogy works well in some domains. For example, given three words “man”, “woman”, and “king”, word analogy could predict “queen” by calculating vector(“man”) − vector(“woman”) + vector(“king”) and searching for the nearest word vector vector(“queen”). Though word analogy might allow wide variety of applications, the most desired one is discovery of unknown relations.</p><p>In this study, we used word2vec software, a de facto standard implementation of word embedding algorithm, with the following parameters by default.</p><p> vector size = 200</p><p> window size = 8</p><p> minimum count of words to be embedded = 1 (i.e. all words)</p><p> model = continuous bag of words</p><p>As a result, 1,772,186 words were embedded into word vectors (2303 for drugs, 3069 for diseases, 8703 for genes, and 1,758,111 for others).</p></sec><sec id="s2_6"><title>2.6. Word Classes</title><p>For the evaluation of clustering results, ATC codes [<xref ref-type="bibr" rid="scirp.62911-ref11">11</xref>] and MeSH tree numbers [<xref ref-type="bibr" rid="scirp.62911-ref12">12</xref>] were attached to drug and disease names, respectively. ATC codes were extracted from DrugBank [<xref ref-type="bibr" rid="scirp.62911-ref13">13</xref>] . Due to the incompleteness of data annotation, only 1253 drugs out of 2303 and 2745 diseases out of 3069 have such classification.</p></sec><sec id="s2_7"><title>2.7. Drug-Disease Relations</title><p>For the evaluation of difference vectors between drugs and diseases, relations between drugs and diseases occurring in the corpus were extracted from CTD [<xref ref-type="bibr" rid="scirp.62911-ref14">14</xref>] . Only the 12,462 relations with therapeutic evidences were adopted for obtaining trustable results. In the set of relations, the mapping from drugs to diseases is many-to- many. For example, “drug::gefitinib” is related to 17 different diseases, and “disease::lung_neoplasm” is mapped from 60 different drugs.</p><p>In order to conduct detailed analysis on cancer-related drugs and diseases, 12,462 extracted drug-disease relations were further filtered so that both of drug and disease names in each relation are attached to an ATC code and a MeSH tree number beginning with “L” (Antineoplastic and immunomodulating agents) and “C04” (Neoplasms), respectively. As a result, 1097 relations consist of 104 anti-cancer drugs and 107 cancer-related diseases were extracted for detailed analysis.</p></sec><sec id="s2_8"><title>2.8. Clustering</title><p>For visual evaluation of word vector quality, we performed hierarchical clustering with cosine distance and Ward’s method [<xref ref-type="bibr" rid="scirp.62911-ref15">15</xref>] . Before the clustering, 2303 drugs and 3069 diseases occurring in the corpus were reduced to 1282 and 1051, respectively, since other drugs and diseases did not occur in CTD.</p></sec><sec id="s2_9"><title>2.9. Classification</title><p>Support Vector Machine (SVM) was adopted for learning and predicting possible relations between drugs and diseases. As an implementation, ksvm function included in kernlab package for R software was used with default parameters.</p></sec></sec><sec id="s3"><title>3. Experimental Results</title><sec id="s3_1"><title>3.1. Distribution of Word Vectors</title><p><xref ref-type="fig" rid="fig2">Figure 2</xref> illustrates the 3D plot of vectors corresponding to 2303 drugs, 3069 diseases, and 8703 for genes. For visualization, the dimension of vector was reduced from 200 to 3 by PCA. In the left panel of the figure, it is</p><fig id="fig2"  position="float"><label><xref ref-type="fig" rid="fig2">Figure 2</xref></label><caption><title> Distribution of word vectors visualized through PCA and 3D plot. Left panel: blue, red, green colors indicate word vectors for drugs, diseases, and genes. Right panel: color gradation from light blue to light pink indicates the frequency of words (from rare to frequent)</title></caption><graphic mimetype="image"   position="float"  xlink:type="simple"  xlink:href="http://html.scirp.org/file/2-9102253x7.png"/></fig><p>shown that the distributions of word vectors in three categories are clearly separated. In the right panel, it is also shown that the frequent words have clear separation, whereas it is relatively difficult to discriminate the categories of rare words.</p></sec><sec id="s3_2"><title>3.2. Cluster Analysis</title><p><xref ref-type="fig" rid="fig3">Figure 3</xref> and <xref ref-type="fig" rid="fig4">Figure 4</xref> show the result of hierarchical clustering for drugs and diseases, respectively. In the right panels of them, entire pictures of clustering results for 1282 drugs and 1051 diseases are shown. In the right panel of <xref ref-type="fig" rid="fig3">Figure 3</xref>, it can be seen that most of the anti-cancer drugs are condensed in the ninth cluster from the top (left panel for more detail). It indicates that the word vectors for drugs well represent the characteristics of corresponding drugs. Also in <xref ref-type="fig" rid="fig4">Figure 4</xref>, we can see that the seventh cluster from the top contains a number of cancer-related diseases, however, also in sixth and ninth clusters. The difference between these results might come from the fact that diseases can be classified from different perspectives (tissues, mechanism, etc.).</p></sec><sec id="s3_3"><title>3.3. Applicability of Simple Word Analogy</title><p>Though word analogy is quite attractive, it does not always works well. To evaluate the applicability of word analogy to the discovery of new relation between drug and disease, we checked whether most of the displacement vectors between confirmed drug-disease pairs (i.e. correct relations) are similar in length and parallel to each other or not. Unfortunately, as shown in <xref ref-type="fig" rid="fig5">Figure 5</xref>, the displacement vectors have wide range of lengths and directions. It indicates that the simple application of word analogy to drug repositioning cannot achieve high performance.</p></sec><sec id="s3_4"><title>3.4. Classification of Correct and Incorrect Drug-Disease Relations</title><p>Instead of simple application of word analogy, we constructed a classification model using SVM. For all combinations of 104 anti-cancer drugs and 107 cancer-related diseases (i.e. 11,128 drug-disease pairs), drug vectors and disease vectors were concatenated and binary class labels (i.e. positive or negative) were added according to 1097 correct drug-disease relations extracted from CTD. Due to the imbalance of two classes, 1097 out of 10,031 negative examples were randomly selected so that the numbers of positive and negative examples are balanced.</p><p>The result of performance evaluation is shown in <xref ref-type="table" rid="table1">Table 1</xref>. Each accuracy is an average of 100 times 10-fold cross-validation with different subsets of negative examples. In this table, it was revealed that the performance was not so affected by vector size and window size, and the best accuracy was 87.6%. For exploratory use of the classification model to discover candidate drug-disease pairs, it means sufficiently high performance.</p><p>Finally, we tested all combinations of 2199 drugs not used in training and 107 cancer-related diseases (in total, 235,293 drug-disease pairs). In case of the classification model trained by 11,128 examples, only 64 test examples were predicted as positive, and all the drugs in the examples were anti-cancer drugs (but not included in 104 anti-cancer drugs used for training). By controlling the degree of class imbalance in training data, it is possible to predict a pair of non-anti-cancer drug and cancer-related disease as positive. For example, using the classification model trained by 1097 positive and 8776 negative examples (degree of imbalance is 1:8), 10 times training and test by 235,293 drug-disease pairs discovered the following candidate drugs for repositioning to cancer treatment, where the numbers indicate how many times they were discovered in 10 times training and test.</p><table-wrap id="table1" ><label><xref ref-type="table" rid="table1">Table 1</xref></label><caption><title> Accuracy of classifying correct and incorrect drug-disease relations by SVM</title></caption><table><tbody><thead><tr><th align="center" valign="middle"  rowspan="2"  >Vector size</th><th align="center" valign="middle"  colspan="4"  >Accuracy</th></tr></thead><tr><td align="center" valign="middle" >Window size = 2</td><td align="center" valign="middle" >Window size = 3</td><td align="center" valign="middle" >Window size = 4</td><td align="center" valign="middle" >Window size = 8</td></tr><tr><td align="center" valign="middle" >50</td><td align="center" valign="middle" >0.872</td><td align="center" valign="middle" >0.873</td><td align="center" valign="middle" >0.875</td><td align="center" valign="middle" >0.872</td></tr><tr><td align="center" valign="middle" >75</td><td align="center" valign="middle" >0.873</td><td align="center" valign="middle" >0.874</td><td align="center" valign="middle" >0.874</td><td align="center" valign="middle" >0.874</td></tr><tr><td align="center" valign="middle" >100</td><td align="center" valign="middle" >0.874</td><td align="center" valign="middle" >0.874</td><td align="center" valign="middle" >0.874</td><td align="center" valign="middle" >0.876</td></tr><tr><td align="center" valign="middle" >200</td><td align="center" valign="middle" >0.874</td><td align="center" valign="middle" >0.874</td><td align="center" valign="middle" >0.874</td><td align="center" valign="middle" >0.874</td></tr></tbody></table></table-wrap><fig id="fig3"  position="float"><label><xref ref-type="fig" rid="fig3">Figure 3</xref></label><caption><title> Result of hierarchical clustering on drugs. Red, green, blue, yellow colors for characters indicate that the drugs are classified in ATC codes as “L01: Antineoplastic Agents”, “L02: Endocrine Therapy”, “L03: Immunostimulants”, and “L04: Immunosuppressants”, respectively</title></caption><graphic mimetype="image"   position="float"  xlink:type="simple"  xlink:href="http://html.scirp.org/file/2-9102253x8.png"/></fig><fig id="fig4"  position="float"><label><xref ref-type="fig" rid="fig4">Figure 4</xref></label><caption><title> Result of hierarchical clustering on diseases. Green, yellow, red, dodger blue, pink, blue, and spring green colors for characters indicate that the diseases are classified in MeSH tree numbers as “C04.182: Cysts”, “C04.445: Hamartoma”, “C04.557: Neoplasms by Histologic Type”, “C04.588: Neoplasms by Site”, “C04.697: Neoplastic Processes”, “C04.730: Paraneoplastic Syndromes”, and “C04.834: Precancerous Conditions”</title></caption><graphic mimetype="image"   position="float"  xlink:type="simple"  xlink:href="http://html.scirp.org/file/2-9102253x9.png"/></fig><fig id="fig5"  position="float"><label><xref ref-type="fig" rid="fig5">Figure 5</xref></label><caption><title> Distribution of displacement vectors for cancer-related drug-disease relations in CTD. Blue and red points represent anti-cancer drugs and cancer-related diseases</title></caption><graphic mimetype="image"   position="float"  xlink:type="simple"  xlink:href="http://html.scirp.org/file/2-9102253x10.png"/></fig><p>drug::urokinase (10), drug::photodynamic_therapy (10), drug::oxygen (10), drug::nonoxynol-9 (10),</p><p>drug::nitroglycerin (10), drug::nitrogen (10), drug::l-phenylalanine (10), drug::l-methionine (10),</p><p>drug::l-glutamine (10), drug::l-cysteine (10), drug::glutathione (10), drug::glucose (10),</p><p>drug::epoxide (10), drug::enzyme (10), drug::collagenase (10), drug::bisphosphonate (10),</p><p>drug::amino_acid (10), drug::amide (10), drug::clarithromycin (9), drug::vitamin (8), drug::l-proline (7),</p><p>drug::vitamin_e (6), drug::xanthophylls (4), drug::phospholipid (4), drug::palifermin (4), drug::ether (4),</p><p>drug::ethacrynic_acid (4), drug::denosumab (4), drug::egfr_inhibitor (2), drug::pyruvic_acid (1).</p><p>Besides too general names like “drug::enzyme” and “drug::amide”, it is notable that the above list includes approved anti-cancer drugs (e.g. “drug::denosumab”), anti-cancer drugs under investigation (e.g. “drug:: clarithromycin”, “drug::bisphosphonate”, and “drug::xanthophyll”), and drugs potentially promote cancer (e.g. “drug::urokinase” and “drug::collagenase”). Especially, it should be emphasized that repositioning of clarithro- mycin to anti-cancer agent has been reported in 2015 [<xref ref-type="bibr" rid="scirp.62911-ref16">16</xref>] , despite the fact that the corpus was downloaded in 2013. Though further screening based on expert’s knowledge is necessary, this result demonstrate that the classification of concatenated word vector is a promising approach to in-silico screening of drug-disease relations for drug repositioning.</p></sec></sec><sec id="s4"><title>4. Discussion and Conclusion</title><p>One of the reasons why word embedding by word2vec becomes popular is its functionality of word analogy [<xref ref-type="bibr" rid="scirp.62911-ref6">6</xref>] . For example, if a sufficient amount of corpus is converted into word vectors and used in the analogy, it could predict the fourth word “California” from three given words “Chicago”, “Illinois”, and “Stockton”. Since a state for a city is unique, it works well: it readily means that the analogy easily fails for one-to-many relationship (e.g. predicting “Stockton” from “Illinois”, “Chicago”, and “California”). About drug-disease relationship, at first we expected that one drug is used for basically one disease. However, as shown in <xref ref-type="fig" rid="fig5">Figure 5</xref>, it was one-to-many from both sides of drug-disease relation. For another problem like gene-protein relationship, accuracy of word analogy might be high since only one protein is produced from one gene, ignoring alternative splicing.</p><p>Although word analogy was not available, word2vec provided significant advantage in the text mining from a large number of biomedical texts in this study. It efficiently encoded more than 1.7 million words into quite short vectors (e.g. 200 dimensions). If we use traditional word frequency and vector space model, one vector for a word is a vector of 1.7 million features with extremely high sparsity. Due to the efficiency of encoding, we could process the whole corpus in reasonable memory space and computation time. Furthermore, the word vectors generated by word2vec seem to well reflect the semantic space of biomedical words. <xref ref-type="fig" rid="fig2">Figure 2</xref> illustrates that the words in different semantic categories are well separated in case of sufficiently high frequency of occurrences. Also, the results of clustering shown in <xref ref-type="fig" rid="fig3">Figure 3</xref> and <xref ref-type="fig" rid="fig4">Figure 4</xref> indicate that similarities among words in the same category are also fine. They might be promising results for further application of word embedding in biomedical text mining.</p><p>In this study, it was revealed that word embedding is effective for representing sense of all words in a large number of cancer-related PubMed abstracts. Furthermore, concatenation of word vectors of drugs and diseases well represents their relations and could be used for finding candidate drugs for repositioning by classification. For better performance of classification, various feature selection and over-sampling algorithms [<xref ref-type="bibr" rid="scirp.62911-ref17">17</xref>] will be tested in the future work.</p></sec><sec id="s5"><title>Acknowledgements</title><p>In this research, the super-computing resource was provided by Human Genome Center, the Institute of Medical Science, the University of Tokyo. Additional computation time was provided by the super computer system in Research Organization of Information and Systems (ROIS), National Institute of Genetics (NIG). This work was supported by JSPS KAKENHI Grant Number 26330328.</p></sec><sec id="s6"><title>Cite this paper</title><p>Duc LuuNgo,NaokiYamamoto,Vu AnhTran,Ngoc GiangNguyen,DauPhan,Favorisen RosykingLumbanraja,MamoruKubo,KenjiSatou, (2016) Application of Word Embedding to Drug Repositioning. Journal of Biomedical Science and Engineering,09,7-16. doi: 10.4236/jbise.2016.91002</p></sec></body><back><ref-list><title>References</title><ref id="scirp.62911-ref1"><label>1</label><mixed-citation publication-type="other" xlink:type="simple">Ferreira, L.G., dos Santos, R.N., Oliva, G. and Andricopulo, A.D. (2015) Molecular Docking and Structure-Based Drug Design Strategies. Molecules, 20, 13384-13421. &lt;BR/&gt;http://dx.doi.org/10.3390/molecules200713384</mixed-citation></ref><ref id="scirp.62911-ref2"><label>2</label><mixed-citation publication-type="other" xlink:type="simple">Bajorath, J. (2015) Computer-Aided Drug Discovery. F1000Research, 4, 630. &lt;BR/&gt;http://dx.doi.org/10.12688/f1000research.6653.1</mixed-citation></ref><ref id="scirp.62911-ref3"><label>3</label><mixed-citation publication-type="other" xlink:type="simple">Ashburn, T.T. and Thor, K.B. (2004) Drug Repositioning: Identifying and Developing New Uses for Existing Drugs. Nature Reviews Drug Discovery, 3, 673-683. &lt;BR/&gt;http://dx.doi.org/10.12688/f1000research.6653.1</mixed-citation></ref><ref id="scirp.62911-ref4"><label>4</label><mixed-citation publication-type="other" xlink:type="simple">Emig, D., Ivliev, A., Pustovalova, O., Lancashire, L., Bureeva, S., Nikolsky, Y. and Bessarabova, M. (2013) Drug Target Prediction and Repositioning Using an Integrated Network-Based Approach. PLoS ONE, 8, e60618.&lt;BR/&gt;http://dx.doi.org/10.1371/journal.pone.0060618</mixed-citation></ref><ref id="scirp.62911-ref5"><label>5</label><mixed-citation publication-type="other" xlink:type="simple">Fellbaum, C. (1998) WordNet: An Electronic Lexical Database. MIT, Cambridge, MA.</mixed-citation></ref><ref id="scirp.62911-ref6"><label>6</label><mixed-citation publication-type="other" xlink:type="simple">Mikolov, T., Chen, K., Corrado, G. and Dean, J. (2013) Efficient Estimation of Word Representations in Vector Space. Proceedings of Workshop at ICLR. arXiv:1301.3781v1</mixed-citation></ref><ref id="scirp.62911-ref7"><label>7</label><mixed-citation publication-type="other" xlink:type="simple">Mikolov, T., Sutskever, I., Chen, K., Corrado, G. and Dean, J. (2013) Distributed Representations of Words and Phrases and Their Compositionality. Proceedings of NIPS. arXiv:1301.3781v3</mixed-citation></ref><ref id="scirp.62911-ref8"><label>8</label><mixed-citation publication-type="other" xlink:type="simple">Mikolov, T., Yih, W.T. and Zweig, G. (2013) Linguistic Regularities in Continuous Space Word Representations. Proceedings of NAACL HLT, 746-751.</mixed-citation></ref><ref id="scirp.62911-ref9"><label>9</label><mixed-citation publication-type="other" xlink:type="simple">Miyao, Y. and Tsujii, J. (2008) Feature Forest Models for Probabilistic HPSG Parsing. Computational Linguistics, 34, 35-80. &lt;BR/&gt;http://dx.doi.org/10.1162/coli.2008.34.1.35</mixed-citation></ref><ref id="scirp.62911-ref10"><label>10</label><mixed-citation publication-type="other" xlink:type="simple">Whirl-Carrillo, M., McDonagh, E.M., Hebert, J.M., Gong, L., Sangkuhl, K., Thorn, C.F., Altman, R.B. and Klein, T.E. (2012) Pharmacogenomics Knowledge for Personalized Medicine. Clinical Pharmacology &amp; Therapeutics, 92, 414-417. &lt;BR/&gt;http://dx.doi.org/10.1038/clpt.2012.96</mixed-citation></ref><ref id="scirp.62911-ref11"><label>11</label><mixed-citation publication-type="other" xlink:type="simple">WHO Collaborating Centre for Drug Statistics Methodology (2015) ATC Classi-fication Index with DDDs. WHO Collaborating Centre, Oslo.</mixed-citation></ref><ref id="scirp.62911-ref12"><label>12</label><mixed-citation publication-type="other" xlink:type="simple">Lipscomb, C.E. (2000) Medical Subject Headings (MeSH). Bulletin of the Medical Library Association, 88, 265.</mixed-citation></ref><ref id="scirp.62911-ref13"><label>13</label><mixed-citation publication-type="other" xlink:type="simple">Wishart, D.S., Knox, C, Guo, A.C., Shrivastava, S., Hassanali, M., Stothard, P., Chang, Z. and Woolsey, J. (2006) DrugBank: A Comprehensive Resource for in Silico Drug Discovery and Exploration. Nucleic Acids Research, 34, D668-D672. &lt;BR/&gt;http://dx.doi.org/10.1093/nar/gkj067</mixed-citation></ref><ref id="scirp.62911-ref14"><label>14</label><mixed-citation publication-type="other" xlink:type="simple">Davis, A.P., Grondin, C.J., Lennon-Hopkins, K., Saraceni-Richards, C., Sciaky, D., King, B.L., Wiegers, T.C. and Mattingly, C.J. (2015) The Comparative Toxicogenomics Database’s 10th Year Anniversary: Update 2015. Nucleic Acids Research, 43, D914-D920. &lt;BR/&gt;http://dx.doi.org/10.1093/nar/gku935</mixed-citation></ref><ref id="scirp.62911-ref15"><label>15</label><mixed-citation publication-type="other" xlink:type="simple">Xu, R. and Wunsch, D.I.I. (2005) Survey of Clustering Algorithms. IEEE Transactions on Neural Networks, 16, 645- 678. &lt;BR/&gt;http://dx.doi.org/10.1109/TNN.2005.845141</mixed-citation></ref><ref id="scirp.62911-ref16"><label>16</label><mixed-citation publication-type="other" xlink:type="simple">Pantziarka, P., Bouche, G., Meheus, L., Sukhatme, V. and Sukhatme, V.P. (2015) Repurposing Drugs in Oncology (ReDO)-Clarithromycin as an Anti-Cancer Agent. eCancer Medical Science, 9, 513. &lt;BR/&gt;http://dx.doi.org/10.3332/ecancer.2015.521</mixed-citation></ref><ref id="scirp.62911-ref17"><label>17</label><mixed-citation publication-type="other" xlink:type="simple">Dang, X.T., Hirose, O., Bui, D.H., Saethang, T., Tran, V.A., Nguyen, T.L.A., Le, T.T.K., Kubo, M., Yamada, Y. and Satou, K. (2013) A Novel Over-Sampling Method and Its Application to Cancer Classification from Gene Expression Data. Chem-Bio Informatics Journal, 13, 19-29. &lt;BR/&gt;http://dx.doi.org/10.1273/cbij.13.19</mixed-citation></ref></ref-list></back></article>