<?xml version="1.0" encoding="UTF-8"?><!DOCTYPE article  PUBLIC "-//NLM//DTD Journal Publishing DTD v3.0 20080202//EN" "http://dtd.nlm.nih.gov/publishing/3.0/journalpublishing3.dtd"><article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" dtd-version="3.0" xml:lang="en" article-type="research article"><front><journal-meta><journal-id journal-id-type="publisher-id">NS</journal-id><journal-title-group><journal-title>Natural Science</journal-title></journal-title-group><issn pub-type="epub">2150-4091</issn><publisher><publisher-name>Scientific Research Publishing</publisher-name></publisher></journal-meta><article-meta><article-id pub-id-type="doi">10.4236/ns.2021.139034</article-id><article-id pub-id-type="publisher-id">NS-111957</article-id><article-categories><subj-group subj-group-type="heading"><subject>Articles</subject></subj-group><subj-group subj-group-type="Discipline-v2"><subject>Biomedical&amp;Life Sciences</subject><subject> Chemistry&amp;Materials Science</subject><subject> Earth&amp;Environmental Sciences</subject><subject> Medicine&amp;Healthcare</subject><subject> Physics&amp;Mathematics</subject></subj-group></article-categories><title-group><article-title>
 
 
  Exploring Local Chemical Space in De Novo Molecular Generation Using Multi-Agent Deep Reinforcement Learning
 
</article-title></title-group><contrib-group><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Wei</surname><given-names>Hu</given-names></name><xref ref-type="aff" rid="aff1"><sub>1</sub></xref></contrib></contrib-group><aff id="aff1"><label>1</label><addr-line>Department of Computer Science, Houghton College, Houghton, NY, USA</addr-line></aff><pub-date pub-type="epub"><day>02</day><month>09</month><year>2021</year></pub-date><volume>13</volume><issue>09</issue><fpage>412</fpage><lpage>424</lpage><history><date date-type="received"><day>8,</day>	<month>August</month>	<year>2021</year></date><date date-type="rev-recd"><day>13,</day>	<month>September</month>	<year>2021</year>	</date><date date-type="accepted"><day>16,</day>	<month>September</month>	<year>2021</year></date></history><permissions><copyright-statement>&#169; Copyright  2014 by authors and Scientific Research Publishing Inc. </copyright-statement><copyright-year>2014</copyright-year><license><license-p>This work is licensed under the Creative Commons Attribution International License (CC BY). http://creativecommons.org/licenses/by/4.0/</license-p></license></permissions><abstract><p>
 
 
  Single-agent reinforcement learning (RL) is commonly used to learn how to play computer games, in which the agent makes one move before making the next in a sequential decision process. Recently single agent was also employed in the design of molecules and drugs. While a single agent is a good fit for computer games, it has limitations when used in mol-ecule design. Its sequential learning makes it impossible to modify or improve the previous steps while working on the current step. In this paper, we proposed to apply the mul-ti-agent RL approach to the research of molecules, which can optimize all sites of a mole-cule simultaneously. To elucidate the validity of our approach, we chose one chemical compound Favipiravir to explore its local chemical space. Favipiravir is a broad-spectrum inhibitor of viral RNA polymerase, and is one of the compounds that are currently being used in SARS-CoV-2 (COVID-19) clinical trials. Our experiments revealed the collabora-tive learning of a team of deep RL agents as well as the learning of its individual learning agent in the exploration of Favipiravir. In particular, our multi-agents not only discovered the molecules near Favipiravir in chemical space, but also the learnability of each site in the string representation of Favipiravir, critical information for us to understand the underline mechanism that supports machine learning of molecules.
 
</p></abstract><kwd-group><kwd>Multi-Agent Reinforcement Learning</kwd><kwd> Actor-Critic</kwd><kwd> Molecule Design</kwd><kwd> SARS-CoV-2</kwd><kwd> COVID-19</kwd></kwd-group></article-meta></front><body><sec id="s1"><title>1. INTRODUCTION</title><p>The COVID-19 pandemic demonstrated an urgent need for speedy methods to design drug and vaccine, and to ensure the effectiveness and safety of these products. However, achieving this goal is not an easy task as the number of drug-like molecules is estimated to be between 10<sup>30</sup> and 10<sup>60</sup>. In recent years, artificial intelligence (AI) has been introduced to guide more intelligent search of drug-like molecules. Recent applications of machine learning techniques including supervised, unsupervised, and reinforcement learning (RL) have shown their success in this challenging area.</p><p>One unique feature of RL is that it doesn’t learn from static data stored in files like supervised and unsupervised learning, instead it learns from dynamic data generated by the interaction between an agent and the environment. Many deep supervised and unsupervised learning techniques even require large datasets. This requirement can pose challenges in molecule design, and for a molecule of interest there may not be a corresponding training dataset existed. The RL approach, however, does not require predefined training datasets, instead it only needs to use a reward function. This has the potential to eventually lead to unexpected new molecules that no human has thought about so far [<xref ref-type="bibr" rid="scirp.111957-ref1">1</xref>]. Many current machine learning techniques are unable to effectively control the properties of the generated molecules, but RL methods can treat the design of a molecule as a computer game so an agent can learn to generate molecules with desired properties, considering desired properties as its goal in a game.</p><p>Several recent papers showed the feasibility of using reinforcement learning for molecular design [2 - 5]. However, most of current research used single-agent RL approach that works on a molecule at one site at a time sequentially, which does not allow for any change or improvement of the previous steps while working on the current step. In this paper, we proposed to leverage multi-agent RL for exploring chemical space, which has the advantage of optimizing all sites of a molecule concurrently by a team of RL agents.</p><p>The RNA polymerase inhibitor Favipiravir is currently in clinical trials as a treatment for infection with severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2). Favipiravir was first used against SARS-CoV-2 in Wuhan, China, then in other countries. In June 2020, Favipiravir received the approval in India for mild and moderate COVID-19 infections. As of the 23rd of July, 2020, there are 32 studies registered on clinicaltrials.gov to assess the effect of this drug in the treatment of COVID-19 [<xref ref-type="bibr" rid="scirp.111957-ref6">6</xref>]. The purpose of our work was to use multi-agent RL approach to discover the local chemical space of Favipiravir.</p><p>There are several related studies on chemical space using machine learning, including genetic algorithms [<xref ref-type="bibr" rid="scirp.111957-ref7">7</xref>], recurrent neural networks [<xref ref-type="bibr" rid="scirp.111957-ref8">8</xref>], and reinforcement learning [<xref ref-type="bibr" rid="scirp.111957-ref1">1</xref>]. In particular, RL agents can explore chemical space out of their own motivation, which may discover unexpected new molecules [<xref ref-type="bibr" rid="scirp.111957-ref1">1</xref>].</p></sec><sec id="s2"><title>2. METHODS</title><p>Reinforcement learning as a field of machine learning adopts a trial and error learning approach. There are three important components in RL, state, action, and reward, which define the interaction between the agent and the environment. The goal of RL is for an agent to learn an optimal policy that can achieve maximal long-term reward. There are two main categories of RL, valued-based and policy-based. In the value-based, an agent learns a value function that represents long-term rewards, and then uses this function to define a policy, which means learning a policy indirectly from a value function. In the policy-based, an agent directly optimizes its policy. Actor-critic algorithms (<xref ref-type="fig" rid="fig1">Figure 1</xref>) incorporate both ideas: they have an actor to take actions (policy-based), and then use a critic to evaluate the actions (value-based) [9 - 11].</p><p>Our work employed a multi-agent actor-critic algorithm in which a group of actor-critic agents learn from others, and each agent needs to consider its own state as well as the states of its neighboring agents. Multi-agent learning can be divided into three major categories: fully cooperative, fully competitive, and mixed cooperative–competitive. In this paper, we used a cooperative approach, implying all the agents collaborate to maximize a common long-term reward. Our idea is simple: we not only need history to better position us for future learning, but also use our current learning outcome to change or improve our previous learning. A single-agent can only learn how to move forward based on history.</p><p>There are several challenges of multi-agent learning. In single-agent learning, the agent only interacts with the environment, while in multi-agent learning, one agent needs to interacts with the environment as well as other agents. In multi-agent learning, one agent not only has to consider the state change of the environment caused by its own actions, but also those from other agents. From one agent’s view, the other agents may be considered to be part of the environment, which becomes non-stationary from one agent’s</p><p>perspective. As the number of agents increases, the size of the joint action space of all agents grows exponentially. In this paper, for any one agent, we used a local joint action space of its neighboring agents.</p></sec><sec id="s3"><title>3. RESULTS</title><p>A common task in molecule design is to search the local chemical space around known molecules or drugs. In this work, the known drug was Favipiravir. Before a molecule can be processed by computers, it needs a representation so computers can understand its chemical information. There are several ways to encode molecules including graphs and strings. Molecular graph representation uses nodes and edges in a graph to reveal the atoms and bonds that make up the molecule, whereas string presentation uses characters for the same purpose. Another representation is using chemical descriptors to create chemical fingerprints, which are vectors encoding physicochemical or structural properties. Hashed fingerprints use a hash function to hash the vectors into vectors of fixed size usually consisting of 512, 1024, or 2048 bits [<xref ref-type="bibr" rid="scirp.111957-ref12">12</xref>].</p><p>In this paper, we used SELFIES (Self-Referencing Embedded Strings) string representation (version 1.0.3) which is an improvement over SMILES (Simplified Molecular Input Line Entry System) strings since an arbitrary SMILES string could represent a chemically infeasible or invalid molecular structure, whereas all possible SELFIES strings represent only valid molecules. Using SELFIES strings therefore can eliminate the common post processing step in using SMILES strings, a known deficiency of SMILES strings [<xref ref-type="bibr" rid="scirp.111957-ref13">13</xref>].</p><p>The alphabet of SELFIES is a collection of symbols or characters used to encode chemical structures of molecules, which is made of these elements: {N-1expl, #P, S+1expl, #P+1expl, =S-1expl, =C+1expl, S, =P+1expl, #S+1expl, =O+1expl, #P-1expl, =P, Cl, =O, C, S-1expl, P, Expl=Ring1, =S+1expl, =P-1expl, O-1expl, C-1expl, Ring3, Branch1_1, #C+1expl, Branch1_3, #O+1expl, Ring1, #N, Ring2, =C-1expl, =N, =N-1expl, Branch3_1, Br, Branch2_1, Expl=Ring3, =N+1expl, #S-1expl, O, Branch2_2, Branch3_2, P-1expl, =C, =S, #S, P+1expl, Branch2_3, N+1expl, C+1expl, O+1expl, #N+1expl, #C-1expl, Branch1_2, #C, I, Expl=Ring2, Hexpl, N, F, Branch3_3}.</p><p>For example, an Favipiravir molecule is encoded as a SELFIES string of length 21: [C][=C] [Branch1_1][P][N][=C][Branch1_1][Branch2_1][C][Branch1_2][C][=O][N][Ring1][Branch1_3][C][Branch1_2][C][=O][N][F] (for clarity, square bracket is used to enclose each symbol), and its SMILES representation is C1=C(N=C(C(=O)N1)C(=O)N)F (no square bracket is used).</p><p>This section of results has two parts: multi-agents and single-agent, so we could see the difference of the two. At the same time, we employed different reward functions in each part to get better understanding of how reward functions can affect the learning of RL agents in each part.</p><sec id="s3_1"><title>3.1. Multi-Agents</title><p>In each episode, a team of 21 actor-critic agents were given a random SELFIES string of length 21, and each site of this string was assigned an agent to gain maximal reward based on a given reward function, which is defined in Sections 3.1 and 3.2 respectively. The SELFIES string for Favipiravir was the target representation. The learning representation was the string that the 21 agents were updating from a random SELFIES string, used as a circular linear list when defining left and right neighboring states for a given state. The state of each agent was the character at the assigned site, and the joint state of each agent was the collection of three states, left state, state, right state, which corresponded to three agents, left agent, current agent, right agent. The action space of each agent was the alphabet of SELFIES of size 61. During training each agent j (j from 1 to 21) took one action from the alphabet, and used this action (character) to update the current character at site j. One training episode consisted of 21 updates, with one update from each agent.</p><sec id="s3_1_1"><title>3.1.1. Character and Molecule Similarity Based Multi-Agents</title><p>The range of molecule similarity between learning and target molecules in <xref ref-type="table" rid="table1">Table 1</xref> is (0, 1) and that of character similarity is (0, 21). To visualize the correlation between character similarity and molecule similarity in the reinforcement learning process, we scaled the molecule similarity from (0, 1) to (0, 21). Three experiments were conducted and the average of the collaborative learning outcomes from the multiple actor-critic agents were collected (<xref ref-type="fig" rid="fig2">Figure 2</xref>). To smooth the collaborative learning curves, a moving average of window size 50 was applied. The molecule similarity increased as the character similarity (<xref ref-type="fig" rid="fig2">Figure 2</xref>). Each agent’s cumulative rewards based on the reward function introduced in <xref ref-type="table" rid="table1">Table 1</xref> are shown in <xref ref-type="fig" rid="fig3">Figure 3</xref>, with a total of 21 agents. Varied learnability of different 21 sites was observed in <xref ref-type="fig" rid="fig3">Figure 3</xref> and the 21 reward curves were not clustered together at the end.</p><table-wrap id="table1" ><label><xref ref-type="table" rid="table1">Table 1</xref></label><caption><title> Reward function based on character similarity and molecule similarity. The Tanimoto similarity of ECFP4 fingerprints between two molecules (we called it molecule similarity in this paper) is calculated using software RDKit with fingerprint length = 2048, and fingerprint radius = 3</title></caption><table><tbody><thead><tr><th align="center" valign="middle" >reward(learning_representation, target_representation, j): score = TanimotoSimilarity(learning_representation[left,j,right], target_representation[left,j,right],) score = 2.0*score if learning_representation[j] == target_representation[j] score += 1 return score</th></tr></thead></tbody></table></table-wrap></sec><sec id="s3_1_2"><title>3.1.2. Character Similarity Based Multi-Agents</title><p>To understand the contributions of character similarity and molecule similarity respectively to the overall collaborative learning of multi-agents, we trained 21 agents with a reward function that was solely based on character similarity (<xref ref-type="table" rid="table2">Table 2</xref>). Three experiments were conducted and the average of the collaborative learning outcomes from the multi-agents were collected (<xref ref-type="fig" rid="fig4">Figure 4</xref>). The correlation between the increase of character similarity and molecule similarity in <xref ref-type="fig" rid="fig4">Figure 4</xref> was not as strong as that in <xref ref-type="fig" rid="fig2">Figure 2</xref>, which implied adding the molecule similarity in the reward function boosted the increase of molecule similarity. In other words, the multi-agents trained with character similarity and molecule similarity using the reward function in <xref ref-type="table" rid="table1">Table 1</xref> could find molecules that were closer to the target molecule in chemical structures, a desired property of machine learning for molecules. The learnability of different 21 sites was similar as displayed in <xref ref-type="fig" rid="fig5">Figure 5</xref> and the 21 reward curves were clustered together at the end, a clear contrast to the behaviors of the curves in <xref ref-type="fig" rid="fig3">Figure 3</xref>.</p><p>To visualize the learnability of different sites by multi-agents, we scaled the last cumulative rewards in <xref ref-type="fig" rid="fig3">Figure 3</xref> and <xref ref-type="fig" rid="fig5">Figure 5</xref> to the range of (0, 10) (<xref ref-type="fig" rid="fig6">Figure 6</xref>). In general, different sites were similarly learnable by character based agents in this section but they showed much varied learnability by character and molecule based agents in Section 3.1.1. This could be interpreted as characters in SELFIES string representation</p><table-wrap id="table2" ><label><xref ref-type="table" rid="table2">Table 2</xref></label><caption><title> Reward function based on character similarity</title></caption><table><tbody><thead><tr><th align="center" valign="middle" >reward(learning_representation, target_representation, j): score = −1 if learning_representation[j] == target_representation[j] score = 1 if learning_representation[left] == target_representation[left] score += 1 if learning_representation[right] == target_representation[right] score += 1 return score</th></tr></thead></tbody></table></table-wrap><p>play similar roles in the learning of character based agents whereas character and molecule based agents need to learn extra chemical information so the learning is harder and more varied.</p><p>From the documents for SELFIES https://selfies.readthedocs.io/en/latest/tutorial.html, we could analyze the varied learnability of different sites by the multi-agents in Section 3.1.1 as shown in <xref ref-type="fig" rid="fig6">Figure 6</xref>. Let’s recall that the Favipiravir SELFIES string is [C][=C][Branch1_1][P][N][=C][Branch1_1][Branch2_1][C] [Branch1_2][C][=O][N][Ring1][Branch1_3][C][Branch1_2][C][=O][N][F]. The character [P] at site 4 is after [Branch1_1] so it is used as a symbol (not as an atom) to specify the length of the branch. [P] has a Q value of 15, and therefore the length of the branch = Q + 1 = 16 that covers the 16 SELFIES characters from first [N] to the second [N] in Favipiravir, which correspond to the SMILES branch (N=C(C(=O)N1)C(=O)N) in the SMILES representation of Favipiravir C1=C(N=C(C(=O)N1)C(=O)N)F. Because [P] can have dual meanings so it is harder to learn. The same reason could be given to [Branch2_1] at site 8 in [Branch1_1][Branch2_1], because [Branch2_1] has a Q value of 6, so the length of the branch [Branch1_1] = Q + 1 = 7, which covers [C][Branch1_2][C][=O][N][Ring1][Branch1_3] that corresponds to the SMILES branch C(C(=O)N1). [Branch1_3] at site 15 has a Q value of 5, so [Ring1] has a length of Q + 1 that connects the current atom [N] with the 6th preceding atom through a single bond (Q + 1 = 6), which is C1 in the SMILES ring C1=C(N=C(C(=O)N1. In summary, these results suggested that the symbols after branch in SELFIES were harder to learn because of their dual meanings when the 21 agents were trained with a reward function that included both character similarity and molecule similarity.</p></sec><sec id="s3_1_3"><title>3.1.3. Chemical Structures of the Molecules Discovered by Multi-Agents</title><p>This section illustrates the molecules found by our multi-agents near the target Favipiravir in chemical space. We first display the chemical structure of the target molecule Favipiravir (<xref ref-type="fig" rid="fig7">Figure 7</xref>) and then several molecules discovered by our multi-agents (Figures 8-13). As expected, the molecule similarity of the learned molecule and the target decreased as the character similarity decreased. In this context, the similarity of the two molecules serves as a distance, which measures the closeness between the learned molecule and the target in chemical space.</p></sec></sec><sec id="s3_2"><title>3.2. Single-Agent</title><p>In this single-agent setting, at the start of training the agent was given a random SELFIES string of length 21, and the state at site j (j from 1 to 21) was the character at this site. The agent learned to take one action from the SELFIES alphabet of size 61 and used this action (character) to update the current character at j. One training episode consisted of 21 updates that went through 21 sites (from 1 to 21) of the string sequentially by the same agent.</p><sec id="s3_2_1"><title>3.2.1. Character and Molecule Similarity Based Single-Agent with Joint State Reward Function</title><p>To make a fair comparison between multi-agent and single-agent methods, the same reward function (<xref ref-type="table" rid="table3">Table 3</xref>) is used for the single-agent here as in Section 3.1 for the multi-agent experiments (<xref ref-type="table" rid="table1">Table 1</xref>). In this case, the input state of the single agent was a joint state made of three states, left state, state, right state. The SELFIES string was used as a circular linear list when defining left state and right state of the current state. Three experiments were conducted and the average of the learning outcomes from a single agent was collected. The learning curves based on the training with the joint state reward function in <xref ref-type="table" rid="table3">Table 3</xref> are in <xref ref-type="fig" rid="fig1">Figure 1</xref>4. As in the sections on multi-agents, a moving average of window size = 50 was used in <xref ref-type="fig" rid="fig1">Figure 1</xref>4, and we scaled the molecule similarity from (0, 1) to (0, 21).</p></sec><sec id="s3_2_2"><title>3.2.2. Character and Molecule Similarity Based Single-Agent with Single State Reward Function</title><p>In this section, we chose a reward function that is only meaningful in the single-agent setting (<xref ref-type="table" rid="table4">Table 4</xref>). In this case, the character at site j (j from 1 to 21) was a state, which was used as input state to the single agent. This input state was a single state as compared to the joint state in <xref ref-type="table" rid="table3">Table 3</xref>. However, the molecule similarity was measured from the first character to the current character j (<xref ref-type="table" rid="table4">Table 4</xref>). Three experiments were conducted and the average of the learning outcomes from a single agent was collected. The learning curves of the agents trained with the single state reward function in <xref ref-type="table" rid="table4">Table 4</xref> were in <xref ref-type="fig" rid="fig1">Figure 1</xref>5, and again a moving average of window size = 50 was used in <xref ref-type="fig" rid="fig1">Figure 1</xref>5, and we scaled the molecule similarity from (0, 1) to (0, 21). Compared to the results of multi-agents in Section 3.1, the performance of single-agent was not as good as that of multi-agents. Furthermore, in this single-agent setting, the joint state reward function (<xref ref-type="table" rid="table3">Table 3</xref>) seemed to be better than the single state reward function (<xref ref-type="table" rid="table4">Table 4</xref>), in the learning of character similarity and molecule similarity (<xref ref-type="fig" rid="fig1">Figure 1</xref>4 and <xref ref-type="fig" rid="fig1">Figure 1</xref>5).</p><table-wrap id="table3" ><label><xref ref-type="table" rid="table3">Table 3</xref></label><caption><title> Reward function based on character similarity and molecule similarity (joint state)</title></caption><table><tbody><thead><tr><th align="center" valign="middle" >reward(learning_representation, target_representation, j): score = TanimotoSimilarity(learning_representation[left,j,right], target_representation[left,j,right],) score = 2.0*score if learning_representation[j] == target_representation[j] score += 1 return score</th></tr></thead></tbody></table></table-wrap><table-wrap id="table4" ><label><xref ref-type="table" rid="table4">Table 4</xref></label><caption><title> Reward function based on character similarity and molecule similarity (single state)</title></caption><table><tbody><thead><tr><th align="center" valign="middle" >reward(learning_representation, target_representation, j): score = TanimotoSimilarity(learning_representation[0 to j], target_representation[0 to j],) score = 2.0*score if learning_representation[j] == target_representation[j] score += 1 return score</th></tr></thead></tbody></table></table-wrap><p>Our findings reported in this section suggested that multi-agent RL approach increased learning significantly compared to single-agent RL, which could be attributed to the synergy of team work of a group of agents and a more focused individual responsibility of each agent in the group.</p></sec></sec></sec><sec id="s4"><title>4. CONCLUSION</title><p>In drug discovery, study of the structural neighborhood of known molecules is critical for local optimization, and machine learning can be employed for this task to enhance structure-based molecule and drug design. By viewing molecule design as a game, RL can be applied to de nova drug design of molecules with desired properties without using a large training dataset. The RL agents can also search chemical space without any prior knowledge, which may lead to discovery of molecules unknown before. We proposed to use multi-agent RL to study the local chemical space of Favipiravir in this work. To assess the validity of our idea, we ran our algorithm on Favipiravir molecule with a group of RL agents. Our experiments confirmed multi-agents outperform single-agent in exploring local chemical space, and showed the collaborative learning of this team of agents as well as the individual learning of each agent therein. Multi-agents exhibit the advantage of concurrent learning of all sites of a molecule, compared to the single-agent approach that can only work on one site at a time and cannot change or improve any previous sites while working on the current site. Essentially, a single-agent approach can only learn to move forward, but cannot backward. But in a real learning task, we typically need to use our current learning outcome to better inform us of how to change or improve our previous learning, in addition to using history to help us move forward.</p></sec><sec id="s5"><title>CONFLICTS OF INTEREST</title><p>The author declares no conflicts of interest regarding the publication of this paper.</p></sec><sec id="s6"><title>REFERENCES</title></sec></body><back><ref-list><title>References</title><ref id="scirp.111957-ref1"><label>1</label><mixed-citation publication-type="other" xlink:type="simple">Thiede, L.A., Krenn, M., Nigam, A.K. and Aspuru-Guzik, A. (2020) Curiosity in Exploring Chemical Space: Intrinsic Rewards for Deep Molecular Reinforcement Learning. arXiv:2012.11293 [cs.LG]</mixed-citation></ref><ref id="scirp.111957-ref2"><label>2</label><mixed-citation publication-type="other" xlink:type="simple">Neil, D., Segler, M., Guasch, L., Ahmed, M., Plumbley, D., Sellwood, M. and Brown, N. (2018) Exploring Deep Recurrent Models with Reinforcement Learning for Molecule Design. ICLR 2018 Conference, Vancouver, BC, 30 April-3 May 2018.</mixed-citation></ref><ref id="scirp.111957-ref3"><label>3</label><mixed-citation publication-type="other" xlink:type="simple">Jeon, W. and Kim, D. (2020) Autonomous Molecule Generation Using Reinforcement Learning and Docking to Develop Potential Novel Inhibitors. Scientific Reports, 10, 22104. https://doi.org/10.1038/s41598-020-78537-2</mixed-citation></ref><ref id="scirp.111957-ref4"><label>4</label><mixed-citation publication-type="other" xlink:type="simple">Popova, M., Isayev, O. and Tropsha, A. (2018) Deep Reinforcement Learning for De Novo Drug Design. Science Advances, 4, eaap7885. https://doi.org/10.1126/sciadv.aap7885</mixed-citation></ref><ref id="scirp.111957-ref5"><label>5</label><mixed-citation publication-type="other" xlink:type="simple">Pereira, T., Abbasi, M., Ribeiro, B. and Arrais, J.P. (2021) Diversity Oriented Deep Reinforcement Learning for Targeted Molecule Generation. Journal of Cheminformatics, 13, Article Number: 21. https://doi.org/10.1186/s13321-021-00498-z</mixed-citation></ref><ref id="scirp.111957-ref6"><label>6</label><mixed-citation publication-type="other" xlink:type="simple">Agrawal, U., Raju, R. and Udwadiac, Z.F. (2020) Favipiravir: A New and Emerging Antiviral Option in COVID-19. Medical Journal Armed Forces India, 76, 370-376. https://doi.org/10.1016/j.mjafi.2020.08.004</mixed-citation></ref><ref id="scirp.111957-ref7"><label>7</label><mixed-citation publication-type="other" xlink:type="simple">Nigam, A.K., Pollice, R., Krenn, M., Gomes, G. dos P. and Aspuru-Guzik, A. (2021) Beyond Generative Models: Superfast Traversal, Optimization, Novelty, Exploration and Discovery (STONED) Algorithm for Molecules using SELFIES. Chemical Science, 12, 7079-7090. https://doi.org/10.1039/D1SC00231G</mixed-citation></ref><ref id="scirp.111957-ref8"><label>8</label><mixed-citation publication-type="other" xlink:type="simple">Li, X., Xu, Y., Yao, H. and Lin, K. (2020) Chemical Space Exploration Based on Recurrent Neural Networks: Applications in Discovering Kinase Inhibitors. Journal of Cheminformatics, 12, 42. https://doi.org/10.1186/s13321-020-00446-3</mixed-citation></ref><ref id="scirp.111957-ref9"><label>9</label><mixed-citation publication-type="other" xlink:type="simple">Sutton, R.S. and Barto, A.G. (2018) Reinforcement Learning: An Introduction (Adaptive Computation and Machine Learning) (Adaptive Computation and Machine Learning Series). 2nd Edition, MIT Press, Cambridge, MA.</mixed-citation></ref><ref id="scirp.111957-ref10"><label>10</label><mixed-citation publication-type="other" xlink:type="simple">Duryea, E., Ganger, M. and Hu, W. (2016) Exploring Deep Reinforcement Learning with Multi Q-Learning. Intelligent Control and Automation, 7, 129-144. https://doi.org/10.4236/ica.2016.74012</mixed-citation></ref><ref id="scirp.111957-ref11"><label>11</label><mixed-citation publication-type="other" xlink:type="simple">Hu, W. and Hu, J. (2020) Distributional Reinforcement Learning with Quantum Neural Networks. Intelligent Control and Automation, 10, Article ID: 91668, 16 p. https://doi.org/10.4236/ica.2019.102004</mixed-citation></ref><ref id="scirp.111957-ref12"><label>12</label><mixed-citation publication-type="other" xlink:type="simple">David, L., Thakkar, A., Mercado, R. and Engkvist, O. (2020) Molecular Representations in AI-Driven Drug Discovery: A Review and Practical Guide. Journal of Cheminformatics, 12, 56. https://doi.org/10.1186/s13321-020-00460-5</mixed-citation></ref><ref id="scirp.111957-ref13"><label>13</label><mixed-citation publication-type="other" xlink:type="simple">Krenn, M., Haese, F., Nigam, A.K., Friederich, P. and Aspuru-Guzik, A. (2020) Self-Referencing Embedded Strings (SELFIES): A 100% Robust Molecular String Representation. Machine Learning: Science and Technology, 1, Article ID: 045024. https://doi.org/10.1088/2632-2153/aba947</mixed-citation></ref></ref-list></back></article>