<?xml version="1.0" encoding="UTF-8"?><!DOCTYPE article  PUBLIC "-//NLM//DTD Journal Publishing DTD v3.0 20080202//EN" "http://dtd.nlm.nih.gov/publishing/3.0/journalpublishing3.dtd"><article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" dtd-version="3.0" xml:lang="en" article-type="research article"><front><journal-meta><journal-id journal-id-type="publisher-id">CMB</journal-id><journal-title-group><journal-title>Computational Molecular Bioscience</journal-title></journal-title-group><issn pub-type="epub">2165-3445</issn><publisher><publisher-name>Scientific Research Publishing</publisher-name></publisher></journal-meta><article-meta><article-id pub-id-type="doi">10.4236/cmb.2019.94009</article-id><article-id pub-id-type="publisher-id">CMB-97091</article-id><article-categories><subj-group subj-group-type="heading"><subject>Articles</subject></subj-group><subj-group subj-group-type="Discipline-v2"><subject>Biomedical&amp;Life Sciences</subject></subj-group></article-categories><title-group><article-title>
 
 
  Inferring Multi-Type Birth-Death Parameters for a Structured Host Population with Application to HIV Epidemic in Africa
 
</article-title></title-group><contrib-group><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Hassan</surname><given-names>W. Kayondo</given-names></name><xref ref-type="aff" rid="aff1"><sup>1</sup></xref><xref ref-type="corresp" rid="cor1"><sup>*</sup></xref></contrib><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Samuel</surname><given-names>Mwalili</given-names></name><xref ref-type="aff" rid="aff2"><sup>2</sup></xref></contrib><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>John</surname><given-names>M. Mango</given-names></name><xref ref-type="aff" rid="aff3"><sup>3</sup></xref></contrib></contrib-group><aff id="aff1"><addr-line>Pan African University, Institute of Basic Sciences, Technology and Innovation, Nairobi, Kenya</addr-line></aff><aff id="aff3"><addr-line>Department of Mathematics, Makerere University, Kampala, Uganda</addr-line></aff><aff id="aff2"><addr-line>Department of Statistics and Actuarial Sciences, Jomo Kenyatta University of Agriculture and Technology, Nairobi, Kenya</addr-line></aff><pub-date pub-type="epub"><day>25</day><month>11</month><year>2019</year></pub-date><volume>09</volume><issue>04</issue><fpage>108</fpage><lpage>131</lpage><history><date date-type="received"><day>2,</day>	<month>May</month>	<year>2019</year></date><date date-type="rev-recd"><day>10,</day>	<month>September</month>	<year>2019</year>	</date><date date-type="accepted"><day>13,</day>	<month>December</month>	<year>2019</year></date></history><permissions><copyright-statement>&#169; Copyright  2014 by authors and Scientific Research Publishing Inc. </copyright-statement><copyright-year>2014</copyright-year><license><license-p>This work is licensed under the Creative Commons Attribution International License (CC BY). http://creativecommons.org/licenses/by/4.0/</license-p></license></permissions><abstract><p>
 
 
  Human Immunodeficiency Virus (HIV) dynamics in Africa are purely characterised by sparse sampling of DNA sequences for individuals who are infected. There are some sub-groups that are more at risk than the general population. These sub-groups have higher infectivity rates. We came up with a likelihood inference model of multi-type birth-death process that can be used to make inference for HIV epidemic in an African setting. We employ a likelihood inference that incorporates a probability of removal from infectious pool in the model. We have simulated trees and made parameter inference on the simulated trees as well as investigating whether the model distinguishes between heterogeneous and homogeneous dynamics. The model makes fairly good parameter inference. It distinguishes between heterogeneous and homogeneous dynamics well. Parameter estimation was also performed under sparse sampling scenario. We investigated whether trees obtained from a structured population are more balanced than those from a non-structured host population using tree statistics that measure tree balance and imbalance. Trees from non-structured population were more balanced basing on Colless and Sackin indices.
 
</p></abstract><kwd-group><kwd>HIV</kwd><kwd> Likelihood Inference</kwd><kwd> Multi-Type Birth-Death Process</kwd><kwd> Probability of Removal</kwd><kwd> Structured Population</kwd></kwd-group></article-meta></front><body><sec id="s1"><title>1. Introduction</title><p>Human Immunodeficiency Virus (HIV) is among the most infectious diseases globally. According to [<xref ref-type="bibr" rid="scirp.97091-ref1">1</xref>], an estimated 36.7 million people were living with HIV globally in 2016. Sub-Saharan Africa is the most affected by the epidemic as of 2016, 19.4 million people were living with HIV in Eastern and Southern Africa. Of the total HIV infections in 2016, 6.1 million infections were registered in Western and Central Africa [<xref ref-type="bibr" rid="scirp.97091-ref1">1</xref>]. This implies that about 70% of all people living with HIV resided in Africa in 2016. The world population nearly hit 7.6 billion as of mid-2017 [<xref ref-type="bibr" rid="scirp.97091-ref2">2</xref>]. African continent had approximately 1.3 billion, translating into only 17% of the global population [<xref ref-type="bibr" rid="scirp.97091-ref2">2</xref>]. African continent therefore has less than twenty percent of the global population, yet having almost three quarters of the total number of people living with HIV Worldwide. With Africa having the most of HIV infections, the epidemic poses a threat to both health and socio-economic development of people in Sub-Saharan Africa region. African HIV epidemic is generalised in most parts of the continent. According to [<xref ref-type="bibr" rid="scirp.97091-ref3">3</xref>], a generalised HIV epidemic has a prevalence greater than 1% in the general population and the epidemic self-sustains through heterosexual transmission. For developed countries, HIV epidemic is concentrated. A concentrated HIV epidemic is defined in [<xref ref-type="bibr" rid="scirp.97091-ref3">3</xref>] as where HIV has spread in one or more sub-populations while the general population has a prevalence of less than 1%. The generalised HIV epidemic in Africa requires slightly different methods compared to a concentrated epidemic.</p><p>Many African countries are battling with HIV epidemic. For example, the total burden of HIV in Uganda is increasing according to spectrum estimates from the Ministry of Health as documented in [<xref ref-type="bibr" rid="scirp.97091-ref4">4</xref>]. The number of individuals living with HIV in Uganda rose from 1.4 million in 2013 to 1.5 million in 2015. This was due to continuous spread of the disease and increased life expectancy for people living with HIV [<xref ref-type="bibr" rid="scirp.97091-ref4">4</xref>]. From [<xref ref-type="bibr" rid="scirp.97091-ref4">4</xref>], the number of AIDS related deaths was 28,000, while new HIV infections were 83,000 in 2015. In Uganda, HIV infects individuals disproportionally according to both their geographical and risk behaviours [<xref ref-type="bibr" rid="scirp.97091-ref5">5</xref>]. Some groups are more at risk compared to the general population in Uganda. Such groups have high incidence for HIV. These groups include sexual workers, fish folk community and truck drivers [<xref ref-type="bibr" rid="scirp.97091-ref6">6</xref>] [<xref ref-type="bibr" rid="scirp.97091-ref7">7</xref>] [<xref ref-type="bibr" rid="scirp.97091-ref8">8</xref>] [<xref ref-type="bibr" rid="scirp.97091-ref9">9</xref>]. It was observed that sex workers, their clients and regular partners of clients, contributed between 7% and 11% of new infections in Uganda, Swaziland and Zambia [<xref ref-type="bibr" rid="scirp.97091-ref8">8</xref>]. Dis-proportionality is also visualised among females and males. Although the general prevalence of HIV among adults in Uganda was 6.2% in 2016, it was 7.6% among females and 4.7% among males [<xref ref-type="bibr" rid="scirp.97091-ref10">10</xref>]. HIV prevalence differs in various regions in Uganda as well. It was stated in [<xref ref-type="bibr" rid="scirp.97091-ref10">10</xref>] that the prevalence of HIV in Central, Eastern, West Nile, Northern and Western Uganda was 8.0%, 5.1%, 3.1%, 5.7% and 7.9%, respectively in 2016. Therefore, HIV epidemic in Uganda is structured as well as generalised.</p><p>In order to study such a host population so as to make inference about HIV epidemic, viewing it as a structured population gives more reliable parameter inference due to heterogeneity manifested by individuals in terms of infectivity. In this paper, we are interested in a model that depicts HIV dynamics in an African setting. Dis-proportionality in terms of infectivity is one of the characteristics significant for HIV dynamics in Africa. In order to account for this, we employ models that make inference for structured host population. Incorporating such models for HIV dynamics in Africa is still a challenge as the models should entail characteristics like sparse sampling and heterogeneity in infectivity among infected individuals. We incorporate such characteristics in the model for a structured host population. This aids in making parameter inference for HIV epidemic in an African setting. Molecular sequence data uncovers underlying disease dynamics in structured host populations much better than prevalence and incidence data.</p><p>Molecular sequence data is increasingly becoming available even in Africa, for example Southern African Treatment and Resistance Network has a growing database of greater than 70,000 HIV sequences [<xref ref-type="bibr" rid="scirp.97091-ref11">11</xref>]. The availability of this genomic data has resulted in analysis of DNA HIV sequences, for example, The Phylogenetics and Networks for Generalized Epidemics in Africa consortium (PANGEA-HIV) utilises viral sequence analyses to establish transmission patterns of HIV in Africa [<xref ref-type="bibr" rid="scirp.97091-ref11">11</xref>]. Analysis of sequence data is explored by many researchers for making inference about epidemiological parameters of epidemics like in computation of basic reproductive number. The basic reproductive number of the pathogen for HIV-1 epidemic in Switzerland was estimated using DNA sequences [<xref ref-type="bibr" rid="scirp.97091-ref12">12</xref>]. Some researchers have integrated both sequence data and compartment models to make inference about disease dynamics. The use of sequence data with a compartmental susceptible-infected-removed (SIR) enabled the joint estimation of epidemiological parameters and establishing disease origin [<xref ref-type="bibr" rid="scirp.97091-ref13">13</xref>]. The method was applied to HIV-1 sampled data from the UK to estimate the basic reproductive ratios within clusters and on hepatitis C virus data set from Argentina to establish the origin of the disease [<xref ref-type="bibr" rid="scirp.97091-ref13">13</xref>]. Analysis of sequence data involves using phylogenetic trees.</p><p>We extend the multi-type birth death model of likelihood inference implemented by [<xref ref-type="bibr" rid="scirp.97091-ref14">14</xref>]. We incorporate an additional parameter which is the probability of removal from an infectious pool. In this paper, a sampled individual becomes non-infectious immediately after sampling with a probability r and remains infectious with probability 1 − r . The added parameter accounts for scenarios where some infected individuals remain infectious even upon diagnosis like in most of Africa’s HIV epidemic. Most of individuals who are diagnosed with HIV are not necessarily removed from the infectious pool owing to risky behaviour practised by some individuals. We extend the likelihood inference due to its flexibility in varying other parameters like sampling proportions of different types or sub-groups when making parameter inference. We estimate parameters by varying the sampling proportions as well. The effect of sparse sampling on parameter inference is investigated as well. This is because, analysing parameter inference in such scenarios of sparse sampling is important since HIV epidemic in Africa is still characterized by limited DNA sequences for HIV positive individuals. The model employed in this paper was used to infer parameter estimation in a structured population with two types. The model distinguished between heterogeneous and homogeneous dynamics very well. We also succeeded in making inference about tree balance for both structured and non-structured populations.</p><p>The model used in this paper differs from the Bayesian inference that was implemented by [<xref ref-type="bibr" rid="scirp.97091-ref15">15</xref>] as it extends the likelihood techniques of a birth-death multi-type model developed by [<xref ref-type="bibr" rid="scirp.97091-ref14">14</xref>]. In the Bayesian framework, both epidemiology and fossil calibration of parameters were estimated. We extended the likelihood inference of [<xref ref-type="bibr" rid="scirp.97091-ref14">14</xref>] by adding an extra parameter that accounts for individuals that remain infectious even after being diagnosed with HIV. The extended model which is used for parameter inference in this paper is helpful in studying HIV epidemic in an African setting. Section 2 gives a detailed explanation of the method we are employing as well as explaining extension of the model in [<xref ref-type="bibr" rid="scirp.97091-ref14">14</xref>]. The results of our study are presented in Section 3, while discussion of the results and conclusions are presented in Section 4.</p></sec><sec id="s2"><title>2. Methods</title><sec id="s2_1"><title>2.1. Modelling Framework</title><p>Using ideas developed in [<xref ref-type="bibr" rid="scirp.97091-ref14">14</xref>] on multi-type birth-death process (MTBD), we develop a multi-type HIV model that captures the dynamics of the spread of HIV within the context of a generalised HIV epidemic in Africa. In this model, the host population comprises of different demes. A deme is a subdivision of a population that consists of individuals that possess closely related characteristics. The demes are represented by various types. The disease dynamics are modelled using birth-death processes, hence the name multi-type birth-death infection model. Within each deme, HIV dynamics are homogeneous. However, at the population level, the resultant dynamics are heterogeneous. <xref ref-type="fig" rid="fig1">Figure 1</xref> shows a structured host population with two demes and the parameters for the model.</p><p>For example, deme 1 can be a general population and deme 2 a sub-population at risk.</p><p>We describe the parameters in the first deme. Those for the second deme are described analogously. The following are the parameters for the model in deme 1;</p><p>1) Parameter λ 11 is the rate at which an individual in deme 1 transmits or infects an individual of deme 1. Similarly, λ 12 is the rate at which an individual in deme 1 transmits to an individual of deme 2. These are the speciation (transmission) rate parameters in deme 1. There is both within ( λ 11 ) and across ( λ 12 ) deme transmission in this multi-type birth-death model.</p><p>2) Parameter μ 1 is the rate at which an individual in deme 1 dies. This is the extinction rate parameter.</p><p>3) Parameter ψ 1 is the rate at which an individual of deme 1 is sampled. Sampling in this context is the way HIV sequences are included into the study.</p><p>4) Parameter r is the probability that an individual in either of the demes becomes non-infectious upon sampling (being removed from infectious pool). An individual therefore remains infectious with probability 1 − r . Parameter r is the added parameter in our model which is not in the model presented in [<xref ref-type="bibr" rid="scirp.97091-ref14">14</xref>]. In our model framework, an individual either remains infectious or is removed from the infectious pool upon sampling. Therefore, r μ 1 ψ 1 is the rate at which an individual in deme 1 is removed from the infectious pool upon sampling. The rate at which an individual in deme 1 who remains infectious, transmits to an individual in deme 1 upon sampling is ( 1 − r ) λ 11 ψ 1 . The rate, ( 1 − r ) λ 12 ψ 1 is analogously defined.</p><p>For the derivation of the likelihood inference for parameter estimation, we still build on ideas used in [<xref ref-type="bibr" rid="scirp.97091-ref14">14</xref>]. <xref ref-type="fig" rid="fig2">Figure 2</xref> illustrates the multi-type birth-death tree that we use in the derivation for likelihood of multi-type birth-death tree.</p><p>In <xref ref-type="fig" rid="fig2">Figure 2</xref>, black dots are sampled individuals, individuals either belong to deme 1 or 2. E is an edge which represents an individual. For parameters on the</p><p>dotted vertical line, 0 is the present, τ is the time at which a certain individual that belongs to deme 2 was sampled, t is the time taken to give raise to the sub-tree (dashed oval) that is subtended by an individual represented by an edge E. T is the time taken for the process or epidemic. This is the time from the origin of the epidemic to present. The model illustrated in <xref ref-type="fig" rid="fig1">Figure 1</xref> results in a multi-type tree as visualized in <xref ref-type="fig" rid="fig2">Figure 2</xref>. In our model, the multi-type tree can have a sampled descendant, while a sampled individual in the model presented in [<xref ref-type="bibr" rid="scirp.97091-ref14">14</xref>] had no sampled descendants.</p><p>In this derivation, D E 1 ( t ) is the probability density for an individual represented by an edge E in deme 1 giving raise to the sub-tree of age t (dotted oval). For D E 1 ( t ) , t &gt; τ as shown in <xref ref-type="fig" rid="fig2">Figure 2</xref>, where τ is the time at sampling of an individual. This implies that time increases backwards from present to the origin. We define another term Z 1 ( t ) , which is the probability that an individual in deme 1 is not sampled and does not have a descendant between t and 0. In order to obtain D E 1 ( t + Δ t ) , we assume that D E 1 ( t ) is known and computed where Δ t is a very small time step. Along an edge E for time step Δ t , either no event happens or birth events happen. This gives:</p><p>D E 1 ( t + Δ t ) = ( 1 − ( ∑ j = 1 2     λ 1 , j + μ 1 + ψ 1 ) Δ t ) D E 1 ( t ) + ∑ j = 1 2     λ 1 , j Δ t Z j ( t ) D E 1 ( t )     + ∑ j = 1 2     λ 1 , j Δ t Z 1 ( t ) D E j ( t ) + O ( Δ t 2 ) (1)</p><p>From Equation (1), the first term on the right is for no birth, death and sampling events, second term represents birth of an individual of deme j while lineage j produces no samples in time t, third term signifies birth of an individual of deme j while lineage 1 produces no samples in time t. The last term on right is the probability that more than one event occurs in the time interval Δ t .</p><p>After re-arranging the terms in Equation (1) and letting Δ t → 0 , we obtain the differential equation:</p><p>d d t D E 1 ( t ) = − ( ∑ j = 1 2     λ 1 , j + μ 1 + ψ 1 ) D E 1 ( t ) + ∑ j = 1 2     λ 1 , j Z j ( t ) D E 1 ( t )     + ∑ j = 1 2     λ 1 , j Z 1 ( t ) D E j ( t ) (2)</p><p>The initial condition at τ is given as:</p><p>D E 1 ( τ ) = r μ 1 ψ 1 + ( 1 − r ) ∑ j = 1 2     λ 1, j ψ 1 ,   for     j = 1 , and D E 1 ( τ ) = 0, for     j ≠ 1, (3)</p><p>The first term on the right of Equation (3) depicts individuals who become non-infectious upon sampling while the last term signifies those who are not removed from the infectious pool, i.e., they remain infectious even after being sampled. In typical African communities, many individuals infected with HIV remain infecting others either knowingly or unknowingly. When r is 1, the initial condition given in Equation (3) reduces to that used in [<xref ref-type="bibr" rid="scirp.97091-ref14">14</xref>].</p><p>Equation (2) is used to obtain the quantity D E 1 ( T ) which is the desired probability density for the multi-type birth-death tree. The derivation for Z 1 ( t ) is required to obtain the density of the tree. Z 1 ( t ) is derived using similar steps as those used in the derivation for Equation (2). The differential equation for Z 1 ( t ) is given as:</p><p>d d t Z 1 ( t ) = ( 1 − ψ 1 ) μ 1 − ( ∑ j = 1 2     λ 1 , j + μ 1 + ψ 1 ) Z 1 ( t ) + ∑ j = 1 2     λ 1 , j Z 1 ( t ) Z j ( t ) , (4)</p><p>with the initial condition,</p><p>Z 1 ( 0 ) = Z 2 ( 0 ) = 1. (5)</p><p>The first term on the right of Equation (4) represents death without sampling, the second signifies no birth, death and sampling as well as an individual in deme 1 producing no samples. The last term depicts a birth of an individual j but both individuals 1 and j producing no samples in time t. The initial condition given in Equation (5) implies that an individual at time t = 0 is not sampled. This is the case because the model used accounts for sampling through time only. Contemporaneous sampling, which is sampling at only one time point especially the present is not modelled in this inference.</p><p>The differential equations (2) and (4) with initial condition given by (3) are for deme 1. The corresponding differential equations for deme 2 were analogously derived. The system of differential equations with their initial conditions that is integrated along a given edge for 2 demes is given as:</p><p>d d t Z 1 ( t ) = ( 1 − ψ 1 ) μ 1 − ( λ 1 , 1 + λ 1 , 2 + μ 1 + ψ 1 ) Z 1 ( t ) + λ 1 , 1 Z 1 2 ( t ) + λ 1 , 2 Z 1 ( t ) Z 2 ( t ) d d t Z 2 ( t ) = ( 1 − ψ 2 ) μ 2 − ( λ 2 , 1 + λ 2 , 2 + μ 2 + ψ 2 ) Z 2 ( t ) + λ 2 , 1 Z 2 ( t ) Z 1 ( t ) + λ 2 , 2 Z 2 2 ( t ) d d t D E 1 ( t ) = − ( λ 1 , 1 + λ 1 , 2 + μ 1 + ψ 1 ) D E 1 ( t ) + ( λ 1 , 1 Z 1 ( t ) + λ 1 , 2 Z 2 ( t ) ) D E 1 ( t )                                 + λ 1 , 1 Z 1 ( t ) D E 1 ( t ) + λ 1 , 2 Z 1 ( t ) D E 2 (t)</p><p>d d t D E 2 ( t ) = − ( λ 2 , 1 + λ 2 , 2 + μ 2 + ψ 2 ) D E 2 ( t ) + ( λ 2 , 1 Z 1 ( t ) + λ 2 , 2 Z 2 ( t ) ) D E 2 ( t )                                 + λ 2 , 1 Z 2 ( t ) D E 1 ( t ) + λ 2 , 2 Z 2 ( t ) D E 2 ( t ) Z 1 ( 0 ) = Z 2 ( 0 ) = 1 D E 1 ( τ ) = r μ 1 ψ 1 + ( 1 − r ) ( λ 1 , 1 + λ 1 , 2 ) ψ 1 D E 2 ( τ ) = r μ 2 ψ 2 + ( 1 − r ) ( λ 2 , 1 + λ 2 , 2 ) ψ 2 (6)</p><p>Solving the system of equations in (6) numerically gives the calculations along an edge of a tree. This is then followed by pruning at the nodes and finally reaching the root of the tree. This results in obtaining either D E 1 ( T ) or D E 2 ( T ) depending on whether the root was in deme 1 or 2.</p></sec><sec id="s2_2"><title>2.2. Probability Density of the Tree</title><p>The probability density of the tree with the first individual at time T conditioned on observing at least an individual who is sampled for a structured population with 2 demes is given as:</p><p>p ( T | λ , μ , ψ , T ) = f 1 D E 1 ( T ) 1 − Z 1 ( T ) + f 2 D E 2 ( T ) 1 − Z 2 ( T ) (7)</p><p>The probability that an individual at time T is of deme 1 is given by f 1 and its distribution must be specified. Equation (7) is the likelihood of the parameters given the data. This is used for parameter inference. It should be noted that neither Z 1 ( T ) nor Z 2 ( T ) can be 1. This is because if any takes on 1, then the process will not give raise to the inferred tree. The conditioning is necessary since studies have shown that this conditioning gives more accurate parameter values [<xref ref-type="bibr" rid="scirp.97091-ref16">16</xref>]. The parameters that we estimated were:</p><p>λ = ( λ 1 , 1 , λ 1 , 2 , λ 2 , 1 , λ 2 , 2 ) μ = ( μ 1 , μ 2 ) , ψ = ( ψ 1 , ψ 2 ) , r and T .</p></sec><sec id="s2_3"><title>2.3. Parameter Inference</title><p>The necessary changes suggested in Equation (3) to reflect HIV dynamics in an African setting were effected in R using an R package, Treesim of [<xref ref-type="bibr" rid="scirp.97091-ref17">17</xref>]. The parameters for simulated trees were then estimated using an R package, TreePar of [<xref ref-type="bibr" rid="scirp.97091-ref18">18</xref>] with changes to suit our model. In the first set of simulations, 50 trees with 100 tips were simulated under multi-type birth-death model with 2 demes (MTBD-2), with sampling probability of 0.2 for each of the two sub-populations ( ψ 1 = ψ 2 = 0.2 ). We first set r = 0.2 and r = 0.8 to investigate the effect of the probability of removal on parameter inference. We then increased the number of simulated trees to 100 but kept r = 0.8. This was to investigate the effect of increasing the number of simulated trees on parameter estimates.</p><p>Another set of simulations was done for 100 simulated trees, with 200 sampled tips. We varied the probability of removal, taking on values, r = 0.8 and r = 0.5. The sampling probability remained at 0.2 for both sub-populations. In another set of simulations, the sampling probabilities were set to ψ 1 = ψ 2 = 0.05 for both the sub-populations. This depicted sparse sampling which is highly the case in several HIV epidemics for Africa. Again the probability of removal was varied as it was set at r = 0.2 and r = 0.8. To investigate the effect of varying the sampling proportions, we simulated trees when the sampling proportions were different. We set ψ 1 = 0.2 and ψ 2 = 0.01 for this kind of parameter inference.</p></sec><sec id="s2_4"><title>2.4. Heterogeneous and Homogeneous Dynamics</title><p>For comparison between two models, we used likelihood ratios. From [<xref ref-type="bibr" rid="scirp.97091-ref19">19</xref>], suppose that H 0 is derived from a full alternative ( H 1 ). Suppose further Θ 0 and Θ 1 are defined such that H 0 corresponds to a situation that a true parameter θ is in Θ 0 ⊆ θ , and H 1 corresponds to a case θ ∈ Θ 1 − Θ 0 . The log likelihood ratio statistic is given as:</p><p>L R = − 2 [ max θ ∈ Θ 0 log ( L ( θ ) ) − max θ ∈ Θ 1 log ( L ( θ ) ) ] (8)</p><p>In [<xref ref-type="bibr" rid="scirp.97091-ref20">20</xref>], they defined the likelihood ratio statistic as 2 ( ln ( X | H 1 ) ( X | H 0 ) ) , where</p><p>( X | H ) is the likelihood of the data (X) under hypothesis H. This likelihood statistic simplifies to the one given in Equation (8). Likelihood ratio statistic is asymptotically chi-squared distributed under the null hypothesis according to [<xref ref-type="bibr" rid="scirp.97091-ref19">19</xref>].</p><p>For heterogeneous dynamics, we assume that the population is structured into demes, say 2 as in our model. Each of the deme has different birth and death parameters. For homogeneous case, the population is not structured and individuals have the same birth and death rates. We investigated how well our model distinguished between heterogeneous and homogeneous dynamics. For simplicity, we had two sets of simulations, named A and B.</p><p>In A, both trees simulated and parameter inference were made under MTBD-2. For B, parameter inference was made for multi-type birth-death model with 1 deme (MTBD-1), yet trees were obtained under MTBD-2. We then performed likelihood tests to find out the number of trees that were rejected in support of the MTBD-2 (heterogeneous) dynamics. For homogeneous dynamics, we simulated trees under MTBD-1 and made parameter inference under both MTBD-1 and MTBD-2. Likelihood ratio tests were used to find out how many trees were rejected whose inference was made under MTBD-2.</p><p>Since the use of likelihood ratios requires that H 0 is derived from a full alternative H 1 , in A, we tested H 0 : Tree simulation under MTBD-2 and parameter inference under MTBD-1 against H 1 : Tree simulation under MTBD-2 and parameter inference under MTBD-2. In B, we had H 0 : Tree simulation under MTBD-1 and parameter inference under MTBD-1 against H 1 : Tree simulation under MTBD-1 and parameter inference under MTBD-2. H 0 and H 1 are defined as null and alternative models, respectively.</p></sec><sec id="s2_5"><title>2.5. Simulated DNA Sequences</title><p>In the absence of real DNA sequence data for a structured host population, DNA sequences were simulated using a stochastic agent based model called, the Discrete Spatial Phylo Simulator (DSPS). DSPS was also used in [<xref ref-type="bibr" rid="scirp.97091-ref21">21</xref>]. The simulated DNA sequences were for a structured population which was composed of homosexual men (MSM), heterosexuals and bisexual men. The homosexual men were further split into LowMSM, MedMSM and HighMSM that were of increasing risk behaviour. Heterosexuals consisted of Low Females, High Females, Low Males and High Males. Bisexual men contacted both heterosexuals and homosexual men. During the simulations, half of bisexual men were diagnosed as if they were homosexual men (this was more likely early in infection) and the other half like heterosexual men (this was more likely late in infection). It was assumed that all bisexual men contacted men and women at the same rate on average. This group was not sub-divided further. The sampling was at 100% for 10 years (25 - 35, equivalent to 1995-2005). The simulated DNA sequences were from pol gene.</p><p>In order to make analysis faster, we down sampled the set of sequences which we used in the analysis. In this set, there were 21532 sequences and we sampled 200 sequences uniformly at random. We further reduced the sub-groups to homosexuals, heterosexuals and bisexuals. Out of 200 sequences, 157 were for homosexuals, 38 for bisexuals and only 5 for heterosexuals. Since we were interested in MTBD-2 model, we discarded 5 sequences for heterosexuals. This resulted in analysis of 195 sequences in our set of simulated HIV sequences.</p></sec><sec id="s2_6"><title>2.6. BEAST Study of Simulated DNA Sequences</title><p>We made a Bayesian study for a sampled set of sequences using Bayesian Evolution Analysis by Sampling Trees (BEAST) developed by [<xref ref-type="bibr" rid="scirp.97091-ref22">22</xref>]. This was done using the multi-type birth-death model package version 0.2.0 of BEAST. The inferred parameters included reproductive number for each deme, sampling proportions, becoming non-infectious rate, among others. For this study, we used the HKY substitution model with 4 gamma categories and estimated base frequencies. We set 0.867 as a proportion of invariant sites. Since our alignment contained sequences sampled at different times and the times were measured in years, we used a strict clock model with starting value set to 0.005. We set sampling change time to 10.0 which was slightly greater than time for the oldest sample. The sampling proportions were set to 0 before the time of the first sample.</p><p>We used the mean estimates for the parameters obtained from this Bayesian study to simulate trees with 200 sampled tips. The death rate parameters used in the simulation were equivalent to becoming non-infectious parameters from the Bayesian study. The birth-rate parameters were not inferred directly from this study. The birth-rate parameters were however computed from the reproductive numbers for each deme. This was possible since the BEAST study inferred the reproductive numbers of each deme. Since the reproductive number of each was</p><p>obtained using R 0 i = ∑ j = 1 2     λ i , j μ i , and yet we obtained R 0 i and μ i from the BEAST study, we therefore set ∑ j = 1 2     λ i , j = μ i R 0 i during the simulations with the</p><p>mean parameter estimates from the BEAST study of the simulated sequences. From the simulated trees, we estimated the parameters to evaluate the performance of our model with parameters that were obtained from the simulated DNA sequences.</p></sec><sec id="s2_7"><title>2.7. Tree Balance versus Structured Host Population</title><p>We investigated whether there was a relationship between some tree statistics and the structuring of the host population. This was done using tree statistics that measure balance and imbalance of a phylogeny. In [<xref ref-type="bibr" rid="scirp.97091-ref23">23</xref>], they state that measuring the degree of imbalance or asymmetry of a tree topology may provide support for the hypothesis that species have different potential for speciation. It also serves as an indication of the patterns of speciation events for organisms under study. We used three tree statistics that measure tree balance and imbalance. We employed gamma-statistic, Sackin and Colles indices.</p><p>Gamma ( γ ) statistic was defined in [<xref ref-type="bibr" rid="scirp.97091-ref24">24</xref>]. Let g 2 , g 3 , ⋯ , g n be the inter-node distances of the reconstructed phylogeny with n taxa, the γ -statistic is defined as:</p><p>γ = ( 1 n − 2 ∑ i = 2 n − 1 ( ∑ k = 2 i     k g k ) ) − ( T 2 ) T 1 12 ( n − 2 ) (9)</p><p>where</p><p>T = ∑ j = 2 n     j g j</p><p>We analysed whether the values for γ -statistic are affected by the structuring of a host population. We computed gamma-statistic for both structured and non-structured host population. In [<xref ref-type="bibr" rid="scirp.97091-ref25">25</xref>], it is stated that under pure birth process, γ -statistic values of complete reconstructed phylogenies follow a standard normal distribution. If gamma-statistic values are greater than zero, then the internal nodes are closer to the leaves than expected under a pure birth process. On the other hand, if gamma-statistic values are less than zero, then internal nodes are closer to the origin (root) than expected under a pure birth model.</p><p>Another tree statistic which we used is called Sackin index which adds the number of internal nodes between each leaf and the root. It is defined as:</p><p>I s n = ∑ i = 1 n     N i (10)</p><p>where n is the number of leaves of the tree and N i is the number of internal nodes crossed in the path from leaf i to the root. It is well known in systematic biology that the expectation of I s n under Yule model is of order 2 n ln ( n ) . The normalized Sackin index converges in distribution as the number of leaves, n grows</p><p>to infinity. The normalized Sackin index is defined in [<xref ref-type="bibr" rid="scirp.97091-ref25">25</xref>] as I ^ s n = I s n − E ( I s n ) n .</p><p>Colless index is another tree statistic that assesses tree imbalance. It is defined as:</p><p>I c = 2 ( n − 1 ) ( n − 2 ) ∑ i = 1 n − 1 | T R − T L | (11)</p><p>where i is an internal node, n is the total number of leaves. For each internal node, T R is the number of terminal taxa subtended by the right hand branch and T L is the number of terminal taxa subtended by left hand branch. For a normalised I c , it ranges from 0 (perfect balance) to 1 (incomplete) balance.</p><p>Tree balance is an important consideration for phylogenies because the balance of the true phylogeny affects the accuracy of its estimates as stated in [<xref ref-type="bibr" rid="scirp.97091-ref26">26</xref>]. For Sackin and Colless indices, the higher the values for these two indices, the more unbalanced the trees are [<xref ref-type="bibr" rid="scirp.97091-ref26">26</xref>].</p><p>We simulated phylogenetic trees in R using TreeSim package of [<xref ref-type="bibr" rid="scirp.97091-ref17">17</xref>]. For a structured host population, we set our parameters as: λ 11 = 17 , λ 12 = 4 , λ 21 = 6 , λ 22 = 8 , <inline-formula><inline-graphic xlink:href="/html.scirp.org/file/2-2220090x108.png" xlink:type="simple"/></inline-formula>, <inline-formula><inline-graphic xlink:href="/html.scirp.org/file/2-2220090x109.png" xlink:type="simple"/></inline-formula>and <inline-formula><inline-graphic xlink:href="/html.scirp.org/file/2-2220090x110.png" xlink:type="simple"/></inline-formula> For a non-structured host population, the parameters were set to<inline-formula><inline-graphic xlink:href="/html.scirp.org/file/2-2220090x111.png" xlink:type="simple"/></inline-formula>, <inline-formula><inline-graphic xlink:href="/html.scirp.org/file/2-2220090x112.png" xlink:type="simple"/></inline-formula>, <inline-formula><inline-graphic xlink:href="/html.scirp.org/file/2-2220090x113.png" xlink:type="simple"/></inline-formula>, <inline-formula><inline-graphic xlink:href="/html.scirp.org/file/2-2220090x114.png" xlink:type="simple"/></inline-formula>, <inline-formula><inline-graphic xlink:href="/html.scirp.org/file/2-2220090x115.png" xlink:type="simple"/></inline-formula>, <inline-formula><inline-graphic xlink:href="/html.scirp.org/file/2-2220090x116.png" xlink:type="simple"/></inline-formula>and <inline-formula><inline-graphic xlink:href="/html.scirp.org/file/2-2220090x117.png" xlink:type="simple"/></inline-formula> Parameter <inline-formula><inline-graphic xlink:href="/html.scirp.org/file/2-2220090x118.png" xlink:type="simple"/></inline-formula> is the rate at which an individual in sub-population i infects an individual in sub-population j, <inline-formula><inline-graphic xlink:href="/html.scirp.org/file/2-2220090x119.png" xlink:type="simple"/></inline-formula>is the rate at which an individual in sub-population i dies and <inline-formula><inline-graphic xlink:href="/html.scirp.org/file/2-2220090x120.png" xlink:type="simple"/></inline-formula> is the rate at which an individual is sub-population i is sampled, for <inline-formula><inline-graphic xlink:href="/html.scirp.org/file/2-2220090x121.png" xlink:type="simple"/></inline-formula> and 2.</p><p>In the first set of simulations, we simulated 5000 trees with 200 sampled tips under both structured and non-structured host populations (A1 and A2). In the second set of simulations, we increased the number of sampled tips to 500 while other parameters remained the same as those for the first set of simulations under both structured and non-structured host populations (B1 and B2). In the third set of simulations, parameters were the same as those in first set of simulations but increased the number of sampled trees to 15,000 for both structured and non-structured populations (C1 and C2). Under each of the simulations, we computed gamma-statistic values for the simulated trees. Using an R package apTreeshape of [<xref ref-type="bibr" rid="scirp.97091-ref27">27</xref>], we computed Colless and Sackin indices for the simulated trees. We additionally computed their mean, median and 95% confidence intervals for the three tree statistics in each of the simulation sets.</p></sec></sec><sec id="s3"><title>3. Results</title><sec id="s3_1"><title>3.1. Parameter Inference</title><p><xref ref-type="fig" rid="fig3">Figure 3</xref> and <xref ref-type="fig" rid="fig4">Figure 4</xref> display the estimates for some of the selected parameters graphically. Parameter estimates for <inline-formula><inline-graphic xlink:href="/html.scirp.org/file/2-2220090x122.png" xlink:type="simple"/></inline-formula> and <inline-formula><inline-graphic xlink:href="/html.scirp.org/file/2-2220090x123.png" xlink:type="simple"/></inline-formula> when <inline-formula><inline-graphic xlink:href="/html.scirp.org/file/2-2220090x123.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x124.png" xlink:type="simple"/></inline-formula> and <inline-formula><inline-graphic xlink:href="/html.scirp.org/file/2-2220090x123.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x124.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x125.png" xlink:type="simple"/></inline-formula> are depicted in <xref ref-type="fig" rid="fig3">Figure 3</xref>. <xref ref-type="fig" rid="fig4">Figure 4</xref> shows parameter estimates for <inline-formula><inline-graphic xlink:href="/html.scirp.org/file/2-2220090x123.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x124.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x125.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x126.png" xlink:type="simple"/></inline-formula> and <inline-formula><inline-graphic xlink:href="/html.scirp.org/file/2-2220090x123.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x124.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x125.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x126.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x127.png" xlink:type="simple"/></inline-formula> when <inline-formula><inline-graphic xlink:href="/html.scirp.org/file/2-2220090x123.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x124.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x125.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x126.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x127.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x128.png" xlink:type="simple"/></inline-formula> and <inline-formula><inline-graphic xlink:href="/html.scirp.org/file/2-2220090x123.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x124.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x125.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x126.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x127.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x128.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x129.png" xlink:type="simple"/></inline-formula> Apart from <inline-formula><inline-graphic xlink:href="/html.scirp.org/file/2-2220090x123.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x124.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x125.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x126.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x127.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x128.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x129.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x130.png" xlink:type="simple"/></inline-formula> which is shown in lower part of <xref ref-type="fig" rid="fig4">Figure 4</xref>, other parameter estimates are uniformly scattered across their true values (solid red line).</p><p>For fixed sampling probabilities (<inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x131.png" xlink:type="simple"/></inline-formula>), parameter estimates for<inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x131.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x132.png" xlink:type="simple"/></inline-formula>, yielded slightly better estimates compared to estimates obtained for <inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x131.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x132.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x133.png" xlink:type="simple"/></inline-formula> as shown in <xref ref-type="table" rid="table1">Table 1</xref>, upper and middle parts. There was an overall poor estimation for the second death parameter (<inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x131.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x132.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x133.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x134.png" xlink:type="simple"/></inline-formula>). The lower part of <xref ref-type="table" rid="table1">Table 1</xref> shows parameter estimates when the number of simulated trees was increased from 50 to 100. Better estimates were realized when the number of simulated trees was increased. Parameter estimation for <inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x131.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x132.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x133.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x134.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x135.png" xlink:type="simple"/></inline-formula> greatly improved when the number of sampled trees was increased to 100. A close look at the middle and lower parts of <xref ref-type="table" rid="table1">Table 1</xref>, reveals that 95% confidence interval for <inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x131.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x132.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x133.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x134.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x135.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x136.png" xlink:type="simple"/></inline-formula> narrows by close to half of the interval for the same parameter when the number of simulated trees was increased from 50 to 100.</p><p>Having observed that increasing the number of simulated trees to 100 improved parameter estimates, we increased the number of sampled tips from 100 to 200 in the next set of simulations. This is shown in <xref ref-type="table" rid="table2">Table 2</xref>. On making comparisons between the lower part of <xref ref-type="table" rid="table1">Table 1</xref> and the upper part of <xref ref-type="table" rid="table2">Table 2</xref>, we note that increasing the number of sampled tips to 200 and keeping the number of simulated trees constant (100), registers a slight improvement in the parameter</p><table-wrap id="table1" ><label><xref ref-type="table" rid="table1">Table 1</xref></label><caption><title> Maximum likelihood parameter estimates from 50 and 100 simulated trees with 100 sampled tips (sampling probability<inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x143.png" xlink:type="simple"/></inline-formula>) under MTBD-2</title></caption><table><tbody><thead><tr><th align="center" valign="middle" >For <inline-formula><inline-graphic xlink:href="/html.scirp.org/file/2-2220090x144.png" xlink:type="simple"/></inline-formula> &amp; 50 simulated trees</th><th align="center" valign="middle" >True value</th><th align="center" valign="middle" >Mean</th><th align="center" valign="middle" >Median</th><th align="center" valign="middle" >95% CI</th></tr></thead><tr><td align="center" valign="middle" ><inline-formula><inline-graphic xlink:href="/html.scirp.org/file/2-2220090x145.png" xlink:type="simple"/></inline-formula></td><td align="center" valign="middle" >17.000</td><td align="center" valign="middle" >16.542</td><td align="center" valign="middle" >16.489</td><td align="center" valign="middle" >[11.007, 22.704]</td></tr><tr><td align="center" valign="middle" ><inline-formula><inline-graphic xlink:href="/html.scirp.org/file/2-2220090x146.png" xlink:type="simple"/></inline-formula></td><td align="center" valign="middle" >4.000</td><td align="center" valign="middle" >4.724</td><td align="center" valign="middle" >2.895</td><td align="center" valign="middle" >[0.000, 25.838]</td></tr><tr><td align="center" valign="middle" ><inline-formula><inline-graphic xlink:href="/html.scirp.org/file/2-2220090x147.png" xlink:type="simple"/></inline-formula></td><td align="center" valign="middle" >6.000</td><td align="center" valign="middle" >10.782</td><td align="center" valign="middle" >6.489</td><td align="center" valign="middle" >[0.000, 39.772]</td></tr><tr><td align="center" valign="middle" ><inline-formula><inline-graphic xlink:href="/html.scirp.org/file/2-2220090x148.png" xlink:type="simple"/></inline-formula></td><td align="center" valign="middle" >8.000</td><td align="center" valign="middle" >7.273</td><td align="center" valign="middle" >6.827</td><td align="center" valign="middle" >[0.000, 16.052]</td></tr><tr><td align="center" valign="middle" ><inline-formula><inline-graphic xlink:href="/html.scirp.org/file/2-2220090x149.png" xlink:type="simple"/></inline-formula></td><td align="center" valign="middle" >9.000</td><td align="center" valign="middle" >11.593</td><td align="center" valign="middle" >9.781</td><td align="center" valign="middle" >[0.000, 23.597]</td></tr><tr><td align="center" valign="middle" ><inline-formula><inline-graphic xlink:href="/html.scirp.org/file/2-2220090x150.png" xlink:type="simple"/></inline-formula></td><td align="center" valign="middle" >3.000</td><td align="center" valign="middle" >0.935</td><td align="center" valign="middle" >1.243</td><td align="center" valign="middle" >[−25.061, 20.470]</td></tr><tr><td align="center" valign="middle" >For <inline-formula><inline-graphic xlink:href="/html.scirp.org/file/2-2220090x151.png" xlink:type="simple"/></inline-formula> &amp; 50 simulated trees</td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td></tr><tr><td align="center" valign="middle" ><inline-formula><inline-graphic xlink:href="/html.scirp.org/file/2-2220090x152.png" xlink:type="simple"/></inline-formula></td><td align="center" valign="middle" >17.000</td><td align="center" valign="middle" >16.464</td><td align="center" valign="middle" >16.520</td><td align="center" valign="middle" >[11.207, 20.851]</td></tr><tr><td align="center" valign="middle" ><inline-formula><inline-graphic xlink:href="/html.scirp.org/file/2-2220090x153.png" xlink:type="simple"/></inline-formula></td><td align="center" valign="middle" >4.000</td><td align="center" valign="middle" >3.983</td><td align="center" valign="middle" >3.765</td><td align="center" valign="middle" >[0.000, 10.515]</td></tr><tr><td align="center" valign="middle" ><inline-formula><inline-graphic xlink:href="/html.scirp.org/file/2-2220090x154.png" xlink:type="simple"/></inline-formula></td><td align="center" valign="middle" >6.000</td><td align="center" valign="middle" >8.092</td><td align="center" valign="middle" >6.922</td><td align="center" valign="middle" >[0.000, 32.281]</td></tr><tr><td align="center" valign="middle" ><inline-formula><inline-graphic xlink:href="/html.scirp.org/file/2-2220090x155.png" xlink:type="simple"/></inline-formula></td><td align="center" valign="middle" >8.000</td><td align="center" valign="middle" >8.153</td><td align="center" valign="middle" >7.651</td><td align="center" valign="middle" >[0.000, 17.024]</td></tr><tr><td align="center" valign="middle" ><inline-formula><inline-graphic xlink:href="/html.scirp.org/file/2-2220090x156.png" xlink:type="simple"/></inline-formula></td><td align="center" valign="middle" >9.000</td><td align="center" valign="middle" >10.015</td><td align="center" valign="middle" >9.252</td><td align="center" valign="middle" >[0.000, 23.330]</td></tr><tr><td align="center" valign="middle" ><inline-formula><inline-graphic xlink:href="/html.scirp.org/file/2-2220090x157.png" xlink:type="simple"/></inline-formula></td><td align="center" valign="middle" >3.000</td><td align="center" valign="middle" >3.014</td><td align="center" valign="middle" >2.104</td><td align="center" valign="middle" >[−14.388, 33.727]</td></tr><tr><td align="center" valign="middle" >For <inline-formula><inline-graphic xlink:href="/html.scirp.org/file/2-2220090x158.png" xlink:type="simple"/></inline-formula> &amp; 100 simulated trees</td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td></tr><tr><td align="center" valign="middle" ><inline-formula><inline-graphic xlink:href="/html.scirp.org/file/2-2220090x159.png" xlink:type="simple"/></inline-formula></td><td align="center" valign="middle" >17.000</td><td align="center" valign="middle" >16.137</td><td align="center" valign="middle" >16.920</td><td align="center" valign="middle" >[5.484, 20.558]</td></tr><tr><td align="center" valign="middle" ><inline-formula><inline-graphic xlink:href="/html.scirp.org/file/2-2220090x160.png" xlink:type="simple"/></inline-formula></td><td align="center" valign="middle" >4.000</td><td align="center" valign="middle" >5.007</td><td align="center" valign="middle" >4.483</td><td align="center" valign="middle" >[0.000, 15.587]</td></tr><tr><td align="center" valign="middle" ><inline-formula><inline-graphic xlink:href="/html.scirp.org/file/2-2220090x161.png" xlink:type="simple"/></inline-formula></td><td align="center" valign="middle" >6.000</td><td align="center" valign="middle" >6.611</td><td align="center" valign="middle" >5.016</td><td align="center" valign="middle" >[0.000, 25.069]</td></tr><tr><td align="center" valign="middle" ><inline-formula><inline-graphic xlink:href="/html.scirp.org/file/2-2220090x162.png" xlink:type="simple"/></inline-formula></td><td align="center" valign="middle" >8.000</td><td align="center" valign="middle" >7.959</td><td align="center" valign="middle" >8.684</td><td align="center" valign="middle" >[0.000, 12.711]</td></tr><tr><td align="center" valign="middle" ><inline-formula><inline-graphic xlink:href="/html.scirp.org/file/2-2220090x163.png" xlink:type="simple"/></inline-formula></td><td align="center" valign="middle" >9.000</td><td align="center" valign="middle" >11.302</td><td align="center" valign="middle" >9.492</td><td align="center" valign="middle" >[3.754, 30.200]</td></tr><tr><td align="center" valign="middle" ><inline-formula><inline-graphic xlink:href="/html.scirp.org/file/2-2220090x164.png" xlink:type="simple"/></inline-formula></td><td align="center" valign="middle" >3.000</td><td align="center" valign="middle" >3.237</td><td align="center" valign="middle" >1.899</td><td align="center" valign="middle" >[−2.926, 16.476]</td></tr></tbody></table></table-wrap><table-wrap id="table2" ><label><xref ref-type="table" rid="table2">Table 2</xref></label><caption><title> Maximum likelihood parameter estimates from 100 simulated trees with 200 sampled tips (sampling probability<inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x165.png" xlink:type="simple"/></inline-formula>) under MTBD-2</title></caption><table><tbody><thead><tr><th align="center" valign="middle" >Parameters for <inline-formula><inline-graphic xlink:href="/html.scirp.org/file/2-2220090x166.png" xlink:type="simple"/></inline-formula></th><th align="center" valign="middle" >True value</th><th align="center" valign="middle" >Mean</th><th align="center" valign="middle" >Median</th><th align="center" valign="middle" >95% CI</th></tr></thead><tr><td align="center" valign="middle" ><inline-formula><inline-graphic xlink:href="/html.scirp.org/file/2-2220090x167.png" xlink:type="simple"/></inline-formula></td><td align="center" valign="middle" >17.000</td><td align="center" valign="middle" >16.438</td><td align="center" valign="middle" >16.898</td><td align="center" valign="middle" >[7.990, 19.241]</td></tr><tr><td align="center" valign="middle" ><inline-formula><inline-graphic xlink:href="/html.scirp.org/file/2-2220090x168.png" xlink:type="simple"/></inline-formula></td><td align="center" valign="middle" >4.000</td><td align="center" valign="middle" >4.407</td><td align="center" valign="middle" >3.598</td><td align="center" valign="middle" >[0.000, 15.162]</td></tr><tr><td align="center" valign="middle" ><inline-formula><inline-graphic xlink:href="/html.scirp.org/file/2-2220090x169.png" xlink:type="simple"/></inline-formula></td><td align="center" valign="middle" >6.000</td><td align="center" valign="middle" >6.508</td><td align="center" valign="middle" >6.108</td><td align="center" valign="middle" >[0.000, 18.591]</td></tr><tr><td align="center" valign="middle" ><inline-formula><inline-graphic xlink:href="/html.scirp.org/file/2-2220090x170.png" xlink:type="simple"/></inline-formula></td><td align="center" valign="middle" >8.000</td><td align="center" valign="middle" >7.693</td><td align="center" valign="middle" >8.484</td><td align="center" valign="middle" >[0.000, 11.829]</td></tr><tr><td align="center" valign="middle" ><inline-formula><inline-graphic xlink:href="/html.scirp.org/file/2-2220090x171.png" xlink:type="simple"/></inline-formula></td><td align="center" valign="middle" >9.000</td><td align="center" valign="middle" >10.983</td><td align="center" valign="middle" >9.038</td><td align="center" valign="middle" >[4.330, 26.605]</td></tr><tr><td align="center" valign="middle" ><inline-formula><inline-graphic xlink:href="/html.scirp.org/file/2-2220090x172.png" xlink:type="simple"/></inline-formula></td><td align="center" valign="middle" >3.000</td><td align="center" valign="middle" >2.985</td><td align="center" valign="middle" >2.058</td><td align="center" valign="middle" >[−2.162, 13.363]</td></tr><tr><td align="center" valign="middle" >For <inline-formula><inline-graphic xlink:href="/html.scirp.org/file/2-2220090x173.png" xlink:type="simple"/></inline-formula></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td></tr><tr><td align="center" valign="middle" ><inline-formula><inline-graphic xlink:href="/html.scirp.org/file/2-2220090x174.png" xlink:type="simple"/></inline-formula></td><td align="center" valign="middle" >17.000</td><td align="center" valign="middle" >16.869</td><td align="center" valign="middle" >16.905</td><td align="center" valign="middle" >[12.755, 20.473]</td></tr><tr><td align="center" valign="middle" ><inline-formula><inline-graphic xlink:href="/html.scirp.org/file/2-2220090x175.png" xlink:type="simple"/></inline-formula></td><td align="center" valign="middle" >4.000</td><td align="center" valign="middle" >3.766</td><td align="center" valign="middle" >3.510</td><td align="center" valign="middle" >[0.000, 10.938]</td></tr><tr><td align="center" valign="middle" ><inline-formula><inline-graphic xlink:href="/html.scirp.org/file/2-2220090x176.png" xlink:type="simple"/></inline-formula></td><td align="center" valign="middle" >6.000</td><td align="center" valign="middle" >7.258</td><td align="center" valign="middle" >6.429</td><td align="center" valign="middle" >[0.000, 21.580]</td></tr><tr><td align="center" valign="middle" ><inline-formula><inline-graphic xlink:href="/html.scirp.org/file/2-2220090x177.png" xlink:type="simple"/></inline-formula></td><td align="center" valign="middle" >8.000</td><td align="center" valign="middle" >7.959</td><td align="center" valign="middle" >8.107</td><td align="center" valign="middle" >[2.528, 11.630]</td></tr><tr><td align="center" valign="middle" ><inline-formula><inline-graphic xlink:href="/html.scirp.org/file/2-2220090x178.png" xlink:type="simple"/></inline-formula></td><td align="center" valign="middle" >9.000</td><td align="center" valign="middle" >10.094</td><td align="center" valign="middle" >9.433</td><td align="center" valign="middle" >[2.249, 23.875]</td></tr><tr><td align="center" valign="middle" ><inline-formula><inline-graphic xlink:href="/html.scirp.org/file/2-2220090x179.png" xlink:type="simple"/></inline-formula></td><td align="center" valign="middle" >3.000</td><td align="center" valign="middle" >2.335</td><td align="center" valign="middle" >2.395</td><td align="center" valign="middle" >[−11.925, 13.089]</td></tr></tbody></table></table-wrap><p>estimation. The 95% confidence intervals for parameters in the upper part of <xref ref-type="table" rid="table2">Table 2</xref> are narrower than those in the lower part of <xref ref-type="table" rid="table1">Table 1</xref>. When we varied the probability of removal (r) from 0.8 to 0.5, there was no significant difference in the parameter estimates as seen in the lower part of <xref ref-type="table" rid="table2">Table 2</xref>.</p><p>To investigate whether sampling intensity had an effect on parameter estimates, we varied the sampling proportion from 0.2 to 0.05. Parameter estimates when the sampling intensity was fixed at <inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x180.png" xlink:type="simple"/></inline-formula> are shown in <xref ref-type="table" rid="table3">Table 3</xref>. Slightly wider intervals were obtained for the parameter estimates for this setting. Parameter <inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x180.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x181.png" xlink:type="simple"/></inline-formula> was poorly estimated.</p><p>When parameter inference was made with different sampling proportions, i.e, for <inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x182.png" xlink:type="simple"/></inline-formula> and<inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x182.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x183.png" xlink:type="simple"/></inline-formula>, there was no difference noted in parameter inference for the simulated trees, results not shown here. Parameter <inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x182.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x183.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x184.png" xlink:type="simple"/></inline-formula> was poorly estimated again.</p><p>For the error bars shown in <xref ref-type="fig" rid="fig5">Figure 5</xref> and <xref ref-type="fig" rid="fig6">Figure 6</xref>, the intervals for parameter estimates obtained when each simulated tree had 200 sampled tips were narrower than those when the simulated trees had 100 sampled tips. The error bars for two parameters (<inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x185.png" xlink:type="simple"/></inline-formula>and<inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x185.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x186.png" xlink:type="simple"/></inline-formula>) are shown. The two selected parameters gave a representation for the rest of the parameters that were estimated.</p></sec><sec id="s3_2"><title>3.2. Heterogeneous and Homogeneous Dynamics</title><p>To establish whether our model distinguished between heterogeneous population (MTBD-2) from homogeneous population (MTBD-1), we performed likelihood ratio tests. In the likelihood ratio inference, we used a decision rule of rejecting the null-hypothesis if the likelihood ratio (test statistic) is greater than</p><table-wrap id="table3" ><label><xref ref-type="table" rid="table3">Table 3</xref></label><caption><title> Maximum likelihood parameter estimates from 50 simulated trees with 100 sampled tips (sampling probability<inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x187.png" xlink:type="simple"/></inline-formula>) under MTBD-2</title></caption><table><tbody><thead><tr><th align="center" valign="middle" >Parameters for <inline-formula><inline-graphic xlink:href="/html.scirp.org/file/2-2220090x188.png" xlink:type="simple"/></inline-formula></th><th align="center" valign="middle" >True value</th><th align="center" valign="middle" >Mean</th><th align="center" valign="middle" >Median</th><th align="center" valign="middle" >95% CI</th></tr></thead><tr><td align="center" valign="middle" ><inline-formula><inline-graphic xlink:href="/html.scirp.org/file/2-2220090x189.png" xlink:type="simple"/></inline-formula></td><td align="center" valign="middle" >17.000</td><td align="center" valign="middle" >16.347</td><td align="center" valign="middle" >15.770</td><td align="center" valign="middle" >[0.000, 29.305]</td></tr><tr><td align="center" valign="middle" ><inline-formula><inline-graphic xlink:href="/html.scirp.org/file/2-2220090x190.png" xlink:type="simple"/></inline-formula></td><td align="center" valign="middle" >4.000</td><td align="center" valign="middle" >4.351</td><td align="center" valign="middle" >2.190</td><td align="center" valign="middle" >[0.000, 14.730]</td></tr><tr><td align="center" valign="middle" ><inline-formula><inline-graphic xlink:href="/html.scirp.org/file/2-2220090x191.png" xlink:type="simple"/></inline-formula></td><td align="center" valign="middle" >6.000</td><td align="center" valign="middle" >12.800</td><td align="center" valign="middle" >6.120</td><td align="center" valign="middle" >[0.000, 41.502]</td></tr><tr><td align="center" valign="middle" ><inline-formula><inline-graphic xlink:href="/html.scirp.org/file/2-2220090x192.png" xlink:type="simple"/></inline-formula></td><td align="center" valign="middle" >8.000</td><td align="center" valign="middle" >7.482</td><td align="center" valign="middle" >5.842</td><td align="center" valign="middle" >[0.000, 25.795]</td></tr><tr><td align="center" valign="middle" ><inline-formula><inline-graphic xlink:href="/html.scirp.org/file/2-2220090x193.png" xlink:type="simple"/></inline-formula></td><td align="center" valign="middle" >9.000</td><td align="center" valign="middle" >11.718</td><td align="center" valign="middle" >9.402</td><td align="center" valign="middle" >[2.312, 34.556]</td></tr><tr><td align="center" valign="middle" ><inline-formula><inline-graphic xlink:href="/html.scirp.org/file/2-2220090x194.png" xlink:type="simple"/></inline-formula></td><td align="center" valign="middle" >3.000</td><td align="center" valign="middle" >3.865</td><td align="center" valign="middle" >0.886</td><td align="center" valign="middle" >[−18.917, 36.919]</td></tr><tr><td align="center" valign="middle" >For <inline-formula><inline-graphic xlink:href="/html.scirp.org/file/2-2220090x195.png" xlink:type="simple"/></inline-formula></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td></tr><tr><td align="center" valign="middle" ><inline-formula><inline-graphic xlink:href="/html.scirp.org/file/2-2220090x196.png" xlink:type="simple"/></inline-formula></td><td align="center" valign="middle" >17.000</td><td align="center" valign="middle" >16.788</td><td align="center" valign="middle" >16.200</td><td align="center" valign="middle" >[7.715, 22.555]</td></tr><tr><td align="center" valign="middle" ><inline-formula><inline-graphic xlink:href="/html.scirp.org/file/2-2220090x197.png" xlink:type="simple"/></inline-formula></td><td align="center" valign="middle" >4.000</td><td align="center" valign="middle" >4.499</td><td align="center" valign="middle" >3.623</td><td align="center" valign="middle" >[0.000, 16.305]</td></tr><tr><td align="center" valign="middle" ><inline-formula><inline-graphic xlink:href="/html.scirp.org/file/2-2220090x198.png" xlink:type="simple"/></inline-formula></td><td align="center" valign="middle" >6.000</td><td align="center" valign="middle" >12.477</td><td align="center" valign="middle" >8.373</td><td align="center" valign="middle" >[0.000, 36.119]</td></tr><tr><td align="center" valign="middle" ><inline-formula><inline-graphic xlink:href="/html.scirp.org/file/2-2220090x199.png" xlink:type="simple"/></inline-formula></td><td align="center" valign="middle" >8.000</td><td align="center" valign="middle" >5.452</td><td align="center" valign="middle" >5.410</td><td align="center" valign="middle" >[0.000, 12.472]</td></tr><tr><td align="center" valign="middle" ><inline-formula><inline-graphic xlink:href="/html.scirp.org/file/2-2220090x200.png" xlink:type="simple"/></inline-formula></td><td align="center" valign="middle" >9.000</td><td align="center" valign="middle" >13.710</td><td align="center" valign="middle" >11.637</td><td align="center" valign="middle" >[0.000, 43.433]</td></tr><tr><td align="center" valign="middle" ><inline-formula><inline-graphic xlink:href="/html.scirp.org/file/2-2220090x201.png" xlink:type="simple"/></inline-formula></td><td align="center" valign="middle" >3.000</td><td align="center" valign="middle" >−1.834</td><td align="center" valign="middle" >−1.238</td><td align="center" valign="middle" >[−18.265, 11.133]</td></tr></tbody></table></table-wrap><p><inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x208.png" xlink:type="simple"/></inline-formula>(9.236) for a significance level (<inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x208.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x209.png" xlink:type="simple"/></inline-formula>) of 0.1. We used <inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x208.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x209.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x210.png" xlink:type="simple"/></inline-formula> since in MTBD-2 inference, 7 parameters were estimated including the likelihood value of the tree, while 4 parameters were estimated for MTBD-1. This gave an average of 6 parameters during the inference and since we needed a theoretical value of<inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x208.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x209.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x210.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x211.png" xlink:type="simple"/></inline-formula>, we used <inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x208.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x209.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x210.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x211.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x212.png" xlink:type="simple"/></inline-formula> for the inference. In the first scenario, 41 out of 50 trees rejected MTBD-1 model in favour of MTBD-2 using likelihood ratio tests at 0.9 level (<inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x208.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x209.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x210.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x211.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x212.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x213.png" xlink:type="simple"/></inline-formula>). All the 50 trees supported MTBD-1 while rejecting MTBD-2. Some of the results are shown in <xref ref-type="table" rid="table4">Table 4</xref> and <xref ref-type="table" rid="table5">Table 5</xref>.</p><table-wrap id="table4" ><label><xref ref-type="table" rid="table4">Table 4</xref></label><caption><title> Maximum likelihood parameter estimates from 50 simulated trees with 100 sampled tips (<inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x214.png" xlink:type="simple"/></inline-formula>and<inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x214.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x215.png" xlink:type="simple"/></inline-formula>). In A, both tree simulation and inference is performed under MTBD-2 while in B, tree simulation is under MTBD-2 and inference under MTBD-1</title></caption><table><tbody><thead><tr><th align="center" valign="middle" >(A)</th><th align="center" valign="middle" >True value</th><th align="center" valign="middle" >Mean</th><th align="center" valign="middle" >Median</th><th align="center" valign="middle" >95% CI</th></tr></thead><tr><td align="center" valign="middle" ><inline-formula><inline-graphic xlink:href="/html.scirp.org/file/2-2220090x216.png" xlink:type="simple"/></inline-formula></td><td align="center" valign="middle" >17.000</td><td align="center" valign="middle" >16.586</td><td align="center" valign="middle" >17.067</td><td align="center" valign="middle" >[8.634, 23.338]</td></tr><tr><td align="center" valign="middle" ><inline-formula><inline-graphic xlink:href="/html.scirp.org/file/2-2220090x217.png" xlink:type="simple"/></inline-formula></td><td align="center" valign="middle" >4.000</td><td align="center" valign="middle" >4.285</td><td align="center" valign="middle" >3.076</td><td align="center" valign="middle" >[0.000, 12.518]</td></tr><tr><td align="center" valign="middle" ><inline-formula><inline-graphic xlink:href="/html.scirp.org/file/2-2220090x218.png" xlink:type="simple"/></inline-formula></td><td align="center" valign="middle" >6.000</td><td align="center" valign="middle" >10.645</td><td align="center" valign="middle" >7.853</td><td align="center" valign="middle" >[0.000, 33.991]</td></tr><tr><td align="center" valign="middle" ><inline-formula><inline-graphic xlink:href="/html.scirp.org/file/2-2220090x219.png" xlink:type="simple"/></inline-formula></td><td align="center" valign="middle" >8.000</td><td align="center" valign="middle" >6.939</td><td align="center" valign="middle" >6.774</td><td align="center" valign="middle" >[0.000, 15.088]</td></tr><tr><td align="center" valign="middle" ><inline-formula><inline-graphic xlink:href="/html.scirp.org/file/2-2220090x220.png" xlink:type="simple"/></inline-formula></td><td align="center" valign="middle" >9.000</td><td align="center" valign="middle" >12.285</td><td align="center" valign="middle" >9.990</td><td align="center" valign="middle" >[0.019, 28.269]</td></tr><tr><td align="center" valign="middle" ><inline-formula><inline-graphic xlink:href="/html.scirp.org/file/2-2220090x221.png" xlink:type="simple"/></inline-formula></td><td align="center" valign="middle" >3.000</td><td align="center" valign="middle" >0.519</td><td align="center" valign="middle" >−0.396</td><td align="center" valign="middle" >[−21.688, 18.662]</td></tr><tr><td align="center" valign="middle" >(B)</td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td></tr><tr><td align="center" valign="middle" ><inline-formula><inline-graphic xlink:href="/html.scirp.org/file/2-2220090x222.png" xlink:type="simple"/></inline-formula></td><td align="center" valign="middle" ></td><td align="center" valign="middle" >35.774</td><td align="center" valign="middle" >33.637</td><td align="center" valign="middle" >[2.879, 61.085]</td></tr><tr><td align="center" valign="middle" ><inline-formula><inline-graphic xlink:href="/html.scirp.org/file/2-2220090x223.png" xlink:type="simple"/></inline-formula></td><td align="center" valign="middle" ></td><td align="center" valign="middle" >4.139</td><td align="center" valign="middle" >3.489</td><td align="center" valign="middle" >[1.741, 8.112]</td></tr><tr><td align="center" valign="middle" ><inline-formula><inline-graphic xlink:href="/html.scirp.org/file/2-2220090x224.png" xlink:type="simple"/></inline-formula></td><td align="center" valign="middle" ></td><td align="center" valign="middle" >−0.969</td><td align="center" valign="middle" >−1.999</td><td align="center" valign="middle" >[−7.964, 9.120]</td></tr></tbody></table></table-wrap><table-wrap id="table5" ><label><xref ref-type="table" rid="table5">Table 5</xref></label><caption><title> Maximum likelihood parameter estimates from 50 simulated trees with 100 sampled tips (<inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x225.png" xlink:type="simple"/></inline-formula>and<inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x225.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x226.png" xlink:type="simple"/></inline-formula>). In A, both tree simulation and inference is performed under MTBD-1 while in B, parameter estimation is made still under MTBD-1 but with some constraints</title></caption><table><tbody><thead><tr><th align="center" valign="middle" >(A)</th><th align="center" valign="middle" >True value</th><th align="center" valign="middle" >Mean</th><th align="center" valign="middle" >Median</th><th align="center" valign="middle" >95% CI</th></tr></thead><tr><td align="center" valign="middle" ><inline-formula><inline-graphic xlink:href="/html.scirp.org/file/2-2220090x227.png" xlink:type="simple"/></inline-formula></td><td align="center" valign="middle" >15.000</td><td align="center" valign="middle" >14.707</td><td align="center" valign="middle" >14.514</td><td align="center" valign="middle" >[9.530, 20.433]</td></tr><tr><td align="center" valign="middle" ><inline-formula><inline-graphic xlink:href="/html.scirp.org/file/2-2220090x228.png" xlink:type="simple"/></inline-formula></td><td align="center" valign="middle" >7.500</td><td align="center" valign="middle" >5.625</td><td align="center" valign="middle" >5.871</td><td align="center" valign="middle" >[0.000, 15.468]</td></tr><tr><td align="center" valign="middle" ><inline-formula><inline-graphic xlink:href="/html.scirp.org/file/2-2220090x229.png" xlink:type="simple"/></inline-formula></td><td align="center" valign="middle" >15.000</td><td align="center" valign="middle" >24.068</td><td align="center" valign="middle" >18.018</td><td align="center" valign="middle" >[2.885, 81.330]</td></tr><tr><td align="center" valign="middle" ><inline-formula><inline-graphic xlink:href="/html.scirp.org/file/2-2220090x230.png" xlink:type="simple"/></inline-formula></td><td align="center" valign="middle" >7.500</td><td align="center" valign="middle" >6.973</td><td align="center" valign="middle" >6.460</td><td align="center" valign="middle" >[0.000, 20.342]</td></tr><tr><td align="center" valign="middle" ><inline-formula><inline-graphic xlink:href="/html.scirp.org/file/2-2220090x231.png" xlink:type="simple"/></inline-formula></td><td align="center" valign="middle" >6.000</td><td align="center" valign="middle" >6.515</td><td align="center" valign="middle" >3.488</td><td align="center" valign="middle" >[0.000, 30.117]</td></tr><tr><td align="center" valign="middle" ><inline-formula><inline-graphic xlink:href="/html.scirp.org/file/2-2220090x232.png" xlink:type="simple"/></inline-formula></td><td align="center" valign="middle" >6.000</td><td align="center" valign="middle" >9.578</td><td align="center" valign="middle" >4.414</td><td align="center" valign="middle" >[−35.129, 96.456]</td></tr><tr><td align="center" valign="middle" >(B)</td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td></tr><tr><td align="center" valign="middle" ><inline-formula><inline-graphic xlink:href="/html.scirp.org/file/2-2220090x233.png" xlink:type="simple"/></inline-formula></td><td align="center" valign="middle" >15.000</td><td align="center" valign="middle" >8.079</td><td align="center" valign="middle" >6.877</td><td align="center" valign="middle" >[0.873, 19.679]</td></tr><tr><td align="center" valign="middle" ><inline-formula><inline-graphic xlink:href="/html.scirp.org/file/2-2220090x234.png" xlink:type="simple"/></inline-formula></td><td align="center" valign="middle" >7.500</td><td align="center" valign="middle" >17.461</td><td align="center" valign="middle" >16.330</td><td align="center" valign="middle" >[3.898, 35.268]</td></tr><tr><td align="center" valign="middle" ><inline-formula><inline-graphic xlink:href="/html.scirp.org/file/2-2220090x235.png" xlink:type="simple"/></inline-formula></td><td align="center" valign="middle" >6.000</td><td align="center" valign="middle" >5.303</td><td align="center" valign="middle" >2.467</td><td align="center" valign="middle" >[−12.899, 45.491]</td></tr></tbody></table></table-wrap><p>However, according to [<xref ref-type="bibr" rid="scirp.97091-ref28">28</xref>], he suggests using a <inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x236.png" xlink:type="simple"/></inline-formula> with 1 degree of freedom. So at <inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x236.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x237.png" xlink:type="simple"/></inline-formula> of 0.1, we reject the null hypothesis if the likelihood ratio is greater than 2.706. With this rule, instead of 41 out of 50, it is 49 trees that rejected MTBD-1 model in favour of MTBD-2 in first scenario and results for the second scenario remain the same when we use the idea of [<xref ref-type="bibr" rid="scirp.97091-ref28">28</xref>].</p></sec><sec id="s3_3"><title>3.3. Results from BEAST Study of Simulated DNA Sequences</title><p>We analysed 195 sequences from the sample. These sequences were structured into two groups, namely MSM and Bisexuals. For this sample, we run MCMC in BEAST which was of 12,000,000 length. The mixing was not very good as some parameters had effective sample sizes (ESS) which were low although some had high ESS of above 400. Our interest was mainly on three parameters, namely<inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x238.png" xlink:type="simple"/></inline-formula>, becoming non-infectious rate and sampling proportion.</p><p><xref ref-type="fig" rid="fig7">Figure 7</xref> shows one of the trees sampled in the MCMC runs from BEAST. We also simulated 100 trees with 200 tips using the mean values of the parameters obtained from the analysis of the simulated sequences. These results are shown in <xref ref-type="table" rid="table6">Table 6</xref>.</p></sec><sec id="s3_4"><title>3.4. Tree Balance versus Structured Host Population</title><p>From results shown in <xref ref-type="table" rid="table7">Table 7</xref>, generally the gamma-statistic values for simulated trees between structured and non-structured populations were all negative values apart from upper limits of the confidence intervals. The mean values for gamma-statistic values were generally more negative under non-structured population when we considered simulation cases A1, A2, B1 and B2. This was not true for simulation sets C1 and C2. When the number of tips was increased from 200 to 500, the mean and median values for gamma-statistic became more</p><table-wrap id="table6" ><label><xref ref-type="table" rid="table6">Table 6</xref></label><caption><title> Maximum likelihood parameter estimates from 100 simulated trees with 200 sampled tips using median estimates from for the simulated sequences (sampling probability<inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x240.png" xlink:type="simple"/></inline-formula>)</title></caption><table><tbody><thead><tr><th align="center" valign="middle" >Parameters for <inline-formula><inline-graphic xlink:href="/html.scirp.org/file/2-2220090x241.png" xlink:type="simple"/></inline-formula></th><th align="center" valign="middle" >True value</th><th align="center" valign="middle" >Mean</th><th align="center" valign="middle" >Median</th><th align="center" valign="middle" >95% CI</th></tr></thead><tr><td align="center" valign="middle" ><inline-formula><inline-graphic xlink:href="/html.scirp.org/file/2-2220090x242.png" xlink:type="simple"/></inline-formula></td><td align="center" valign="middle" >0.100</td><td align="center" valign="middle" >0.374</td><td align="center" valign="middle" >0.134</td><td align="center" valign="middle" >[0.000, 1.813]</td></tr><tr><td align="center" valign="middle" ><inline-formula><inline-graphic xlink:href="/html.scirp.org/file/2-2220090x243.png" xlink:type="simple"/></inline-formula></td><td align="center" valign="middle" >0.140</td><td align="center" valign="middle" >0.178</td><td align="center" valign="middle" >0.147</td><td align="center" valign="middle" >[0.000, 0.759]</td></tr><tr><td align="center" valign="middle" ><inline-formula><inline-graphic xlink:href="/html.scirp.org/file/2-2220090x244.png" xlink:type="simple"/></inline-formula></td><td align="center" valign="middle" >1.280</td><td align="center" valign="middle" >1.306</td><td align="center" valign="middle" >1.315</td><td align="center" valign="middle" >[0.991, 1.613]</td></tr><tr><td align="center" valign="middle" ><inline-formula><inline-graphic xlink:href="/html.scirp.org/file/2-2220090x245.png" xlink:type="simple"/></inline-formula></td><td align="center" valign="middle" >0.800</td><td align="center" valign="middle" >0.746</td><td align="center" valign="middle" >0.627</td><td align="center" valign="middle" >[0.035, 2.274]</td></tr><tr><td align="center" valign="middle" ><inline-formula><inline-graphic xlink:href="/html.scirp.org/file/2-2220090x246.png" xlink:type="simple"/></inline-formula></td><td align="center" valign="middle" >0.215</td><td align="center" valign="middle" >0.325</td><td align="center" valign="middle" >0.264</td><td align="center" valign="middle" >[0.000, 1.009]</td></tr><tr><td align="center" valign="middle" ><inline-formula><inline-graphic xlink:href="/html.scirp.org/file/2-2220090x247.png" xlink:type="simple"/></inline-formula></td><td align="center" valign="middle" >0.215</td><td align="center" valign="middle" >0.175</td><td align="center" valign="middle" >0.110</td><td align="center" valign="middle" >[−0.950, 1.354]</td></tr></tbody></table></table-wrap><table-wrap id="table7" ><label><xref ref-type="table" rid="table7">Table 7</xref></label><caption><title> Tree statistics under structured and non-structured populations</title></caption><table><tbody><thead><tr><th align="center" valign="middle"  colspan="4"  >A1: 5000 simulated trees with 200 sampled tips in a structured population</th></tr></thead><tr><td align="center" valign="middle" >Tree statistic</td><td align="center" valign="middle" >Mean</td><td align="center" valign="middle" >Median</td><td align="center" valign="middle" >95% CI</td></tr><tr><td align="center" valign="middle" >Gamma-statistic</td><td align="center" valign="middle" >−8.0765</td><td align="center" valign="middle" >−35.0077</td><td align="center" valign="middle" >[−95.1967, 36.576]</td></tr><tr><td align="center" valign="middle" >Colless’ index</td><td align="center" valign="middle" >1848.324</td><td align="center" valign="middle" >1763</td><td align="center" valign="middle" >[1111, 3090]</td></tr><tr><td align="center" valign="middle" >Sackin’s index</td><td align="center" valign="middle" >2873.392</td><td align="center" valign="middle" >2785</td><td align="center" valign="middle" >[2191, 4062]</td></tr><tr><td align="center" valign="middle"  colspan="4"  >A2: 5000 simulated trees with 200 sampled tips in a non-structured population</td></tr><tr><td align="center" valign="middle" >Tree statistic</td><td align="center" valign="middle" >Mean</td><td align="center" valign="middle" >Median</td><td align="center" valign="middle" >95% CI</td></tr><tr><td align="center" valign="middle" >Gamma-statistic</td><td align="center" valign="middle" >−25.4391</td><td align="center" valign="middle" >−34.6674</td><td align="center" valign="middle" >[−125.8019, 55.931]</td></tr><tr><td align="center" valign="middle" >Colless’ index</td><td align="center" valign="middle" >1558.561</td><td align="center" valign="middle" >1490</td><td align="center" valign="middle" >[947, 2532]</td></tr><tr><td align="center" valign="middle" >Sackin’s index</td><td align="center" valign="middle" >2609.295</td><td align="center" valign="middle" >2539</td><td align="center" valign="middle" >[2050, 3533]</td></tr><tr><td align="center" valign="middle"  colspan="4"  >B1: 5000 simulated trees with 500 sampled tips in a structured population</td></tr><tr><td align="center" valign="middle" >Tree statistic</td><td align="center" valign="middle" >Mean</td><td align="center" valign="middle" >Median</td><td align="center" valign="middle" >95% CI</td></tr><tr><td align="center" valign="middle" >Gamma-statistic</td><td align="center" valign="middle" >−52.1761</td><td align="center" valign="middle" >−51.7531</td><td align="center" valign="middle" >[−109.566, 4.93]</td></tr><tr><td align="center" valign="middle" >Colless’ index</td><td align="center" valign="middle" >5964.97</td><td align="center" valign="middle" >5714</td><td align="center" valign="middle" >[3821, 9482]</td></tr><tr><td align="center" valign="middle" >Sackin’s index</td><td align="center" valign="middle" >8974.741</td><td align="center" valign="middle" >8734</td><td align="center" valign="middle" >[6955, 12375]</td></tr><tr><td align="center" valign="middle"  colspan="4"  >B2: 5000 simulated trees with 500 sampled tips in a non-structured population</td></tr><tr><td align="center" valign="middle" >Tree statistic</td><td align="center" valign="middle" >Mean</td><td align="center" valign="middle" >Median</td><td align="center" valign="middle" >95% CI</td></tr><tr><td align="center" valign="middle" >Gamma-statistic</td><td align="center" valign="middle" >−103.5037</td><td align="center" valign="middle" >−52.006</td><td align="center" valign="middle" >[−154.277, 36.237]</td></tr><tr><td align="center" valign="middle" >Colless’ index</td><td align="center" valign="middle" >4934.575</td><td align="center" valign="middle" >4747</td><td align="center" valign="middle" >[3247, 7631]</td></tr><tr><td align="center" valign="middle" >Sackin’s index</td><td align="center" valign="middle" >8016.174</td><td align="center" valign="middle" >7830</td><td align="center" valign="middle" >[6441, 10616]</td></tr><tr><td align="center" valign="middle"  colspan="4"  >C1: 15,000 simulated trees with 200 sampled tips in a structured population</td></tr><tr><td align="center" valign="middle" >Tree statistic</td><td align="center" valign="middle" >Mean</td><td align="center" valign="middle" >Median</td><td align="center" valign="middle" >95% CI</td></tr><tr><td align="center" valign="middle" >Gamma-statistic</td><td align="center" valign="middle" >−44.139</td><td align="center" valign="middle" >−35.071</td><td align="center" valign="middle" >[−105.172, 40.067]</td></tr><tr><td align="center" valign="middle" >Colless’ index</td><td align="center" valign="middle" >1854</td><td align="center" valign="middle" >1764</td><td align="center" valign="middle" >[1092, 3121]</td></tr><tr><td align="center" valign="middle" >Sackin’s index</td><td align="center" valign="middle" >2878.715</td><td align="center" valign="middle" >2788</td><td align="center" valign="middle" >[2173, 4090]</td></tr><tr><td align="center" valign="middle"  colspan="4"  >C2: 15,000 simulated trees with 200 sampled tips in a non-structured population</td></tr><tr><td align="center" valign="middle" >Tree statistic</td><td align="center" valign="middle" >Mean</td><td align="center" valign="middle" >Median</td><td align="center" valign="middle" >95% CI</td></tr><tr><td align="center" valign="middle" >Gamma-statistic</td><td align="center" valign="middle" >−25.312</td><td align="center" valign="middle" >−34.754</td><td align="center" valign="middle" >[−124.382, 71.212]</td></tr><tr><td align="center" valign="middle" >Colless’ index</td><td align="center" valign="middle" >1553.154</td><td align="center" valign="middle" >1484</td><td align="center" valign="middle" >[948, 2539]</td></tr><tr><td align="center" valign="middle" >Sackin’s index</td><td align="center" valign="middle" >2603.978</td><td align="center" valign="middle" >2534</td><td align="center" valign="middle" >[2052, 3549]</td></tr></tbody></table></table-wrap><p>negative. However, the confidence intervals became narrower. When the sampled trees were increased from 5000 to 15,000 trees, while kept number of sampled tips to 200, the mean values for gamma-statistic became more negative under structured population though remained almost constant under non structured population. The median values remained almost constant.</p><p>For Colless index, the mean and median values were higher under structured population compared to non-structured population. Increasing the sampled tips from 200 to 500 resulted in higher mean and median Colless index values. Increasing simulated trees from 5000 to 15,000 resulted in no big difference in the mean and median values. Sackin index values have the same pattern as the Colless index values.</p><p>The densities for the three tree statistics are shown in <xref ref-type="fig" rid="fig8">Figure 8</xref>. From these densities, there was no significant difference between structured and non-structured populations. The only difference is in the densities for Colless index as the median value for a structured population is higher than that for a non-structured one.</p></sec></sec><sec id="s4"><title>4. Discussion and Conclusion</title><p>From the results shown in Tables 1-6 and Figures 3-6, parameter estimation was better when sampling probability was fixed at 0.2 compared to when it was set at 0.05 to depict sparse sampling. It was also observed that although parameter estimation for sparse sampling (<inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x249.png" xlink:type="simple"/></inline-formula>) resulted in wider parameter intervals, parameters were fairly estimated. The results for sparse sampling being less accurate may suggest that we need to be mindful when making inference in situations with sparse sampling of DNA sequences. It was also observed that the second death parameter (<inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x249.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x250.png" xlink:type="simple"/></inline-formula>) was poorly estimated. It is only the second death parameter (<inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x249.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x250.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x251.png" xlink:type="simple"/></inline-formula>) that is worst estimated with the lower limit for the intervals being negative. The first death rate parameter (<inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x249.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x250.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x251.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x252.png" xlink:type="simple"/></inline-formula>) is fairly estimated with all estimates being positive. The mean and median estimates for <inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x249.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x250.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x251.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x252.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x253.png" xlink:type="simple"/></inline-formula> were however fairly estimated. Poor extinction rate parameter estimates may be due to the fact that generally the extinction (death) rate parameters are poorly estimated by general constant birth-death models [<xref ref-type="bibr" rid="scirp.97091-ref29">29</xref>] [<xref ref-type="bibr" rid="scirp.97091-ref30">30</xref>] [<xref ref-type="bibr" rid="scirp.97091-ref31">31</xref>]. When the probability of removal (r) was varied from 0.2 to 0.8, narrower 95% confidence intervals for parameter estimates were realized. We defined r as the probability that an individual becomes non-infectious upon sampling (removed from infectious pool). For<inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x249.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x250.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x251.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x252.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x253.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x254.png" xlink:type="simple"/></inline-formula>, our model reduces to that of [<xref ref-type="bibr" rid="scirp.97091-ref14">14</xref>] which the authors employed for their parameter inference. When<inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x249.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x250.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x251.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x252.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x253.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x254.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x255.png" xlink:type="simple"/></inline-formula>, it implies that an individual i remains infectious upon sampling with the same rate (<inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x249.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x250.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x251.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x252.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x253.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x254.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x255.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x256.png" xlink:type="simple"/></inline-formula>). Our model resulted in less accurate results when <inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x249.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x250.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x251.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x252.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x253.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x254.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x255.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x256.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x257.png" xlink:type="simple"/></inline-formula> simply because it had an extra parameter that affected the identifiability of the parameters. The estimates were however better when <inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x249.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x250.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x251.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x252.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x253.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x254.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x255.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x256.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x257.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x258.png" xlink:type="simple"/></inline-formula> because as r approached 1, the model approached that used in [<xref ref-type="bibr" rid="scirp.97091-ref14">14</xref>] and it was a parameter less compared to the model used in our study. Our model is handy in making inference that depicts an African setting which is characterized by both sparse sampling of DNA sequences and a large proportion of individuals that always remain infectious even after diagnosis with HIV.</p><p>Our proposed model distinguished between heterogeneous and homogeneous populations. This is because 41 trees out of 50 supported the heterogeneous model dynamics, rejecting the homogeneous type of model inference. For homogeneous dynamics, all the 50 trees accepted the homogeneous dynamics and rejected the heterogeneous type dynamics. We employed the likelihood test statistic for this model selection using a level of significance (<inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x259.png" xlink:type="simple"/></inline-formula>) of 0.1. This was close to that used in [<xref ref-type="bibr" rid="scirp.97091-ref14">14</xref>] which was 0.2. In their work, a <inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x259.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="//html.scirp.org/file/2-2220090x260.png" xlink:type="simple"/></inline-formula> distribution at 0.8 level was used for the inference. It should be noted that even after incorporating a new parameter (probability of removal), the model still accounted for a structured host population. The probability of removal was included to reflect a situation where individuals fail to change behaviour upon diagnosis with HIV. Whereas studies from Europe have reported that individuals become non-infectious upon diagnosis with HIV as reported in [<xref ref-type="bibr" rid="scirp.97091-ref14">14</xref>], it is still a challenge in many African communities. It is therefore imperative to emphasize that structured populations characterize HIV epidemics in many African countries due to the fact that there are some groups who are more at risk compared to the general population in many African countries.</p><p>For the investigation between structuring of a population and tree balance, it was concluded from the results shown in <xref ref-type="table" rid="table7">Table 7</xref> that simulated trees from a non-structured host population are more balanced compared to those from a structured host population. This conclusion was based on Colless and Sackin indices. It is not surprising that when we increased the number of simulated trees from 5000 to 15,000, we did not realise a significant change. This is because the definitions of Colless and Sackin indices use internal nodes of a tree. This is the reason we observed a change in the value of indices when we increased the number of sampled tips from 200 to 500. When tips of a tree are increased, the internal nodes are increased as well. For a structured population, we had only 2 demes, it might be interesting to analyse tree balance in a structured population with more than 2 demes. The relationship between timing of bifurcation events and tree structure was assessed using the gamma-statistic values. From densities shown in <xref ref-type="fig" rid="fig8">Figure 8</xref>, those for gamma-statistic are identical regarding to whether the underlying population is structured or not. A general conclusion from <xref ref-type="table" rid="table7">Table 7</xref> is that the gamma-statistic values could not be used to uncover the underlying population structure. This could be because the timing of the bifurcation events may be having very less influence on the overall tree shape of a phylogeny.</p></sec><sec id="s5"><title>Acknowledgements</title><p>We thank the Editor and reviewers for their comments which greatly improved this paper. The authors thank Pan-African University Institute of Basic Sciences, Technology and Innovation (PAUSTI) for funding this research.</p></sec><sec id="s6"><title>Conflicts of Interest</title><p>The authors declare no conflicts of interest regarding the publication of this paper.</p></sec><sec id="s7"><title>Cite this paper</title><p>Kayondo, H.W., Mwalili, S. and Mango, J.M. (2019) Inferring Multi-Type Birth-Death Parameters for a Structured Host Population with Application to HIV Epidemic in Africa. Computational Molecular Bioscience, 9, 108-131. https://doi.org/10.4236/cmb.2019.94009</p></sec></body><back><ref-list><title>References</title><ref id="scirp.97091-ref1"><label>1</label><mixed-citation publication-type="other" xlink:type="simple">UNAIDS (2017) Unaids Data. Technical Report, Joint United Nations Programme on HIV/AIDS.</mixed-citation></ref><ref id="scirp.97091-ref2"><label>2</label><mixed-citation publication-type="other" xlink:type="simple">UNDESA (2017) United Nations Department of Economic and Social Affairs, Population Division. World Population Prospects: The 2017 Revision, Key Findings and Advance Tables. Technical Report, United Nations, New York.</mixed-citation></ref><ref id="scirp.97091-ref3"><label>3</label><mixed-citation publication-type="other" xlink:type="simple">UNAIDS (2011) Unaids Terminology Guidelines. Technical Report, Joint United Nations Programme on HIV/AIDS.</mixed-citation></ref><ref id="scirp.97091-ref4"><label>4</label><mixed-citation publication-type="other" xlink:type="simple">UAC (2016) Uganda HIV and Aids Country Progress Report. Technical Report, Uganda AIDS Commission.</mixed-citation></ref><ref id="scirp.97091-ref5"><label>5</label><mixed-citation publication-type="other" xlink:type="simple">UAC (2014) Uganda HIV and Aids Country Progress Report. Technical Report, Uganda AIDS Commission.</mixed-citation></ref><ref id="scirp.97091-ref6"><label>6</label><mixed-citation publication-type="other" xlink:type="simple">Kiwanuka, N., Ssetaala, A., Nalutaaya, A., Mpendo, J., Wambuzi, M., Nanvubya, A., Sigirenda, S., Kitandwe, P.K., Nielsen, L.E. and Balyegisawa, A. (2014) High Incidence of HIV-1 Infection in a General Population of Fishing Communities around Lake Victoria, Uganda. PLoS ONE, 9, e94932.https://doi.org/10.1371/journal.pone.0094932</mixed-citation></ref><ref id="scirp.97091-ref7"><label>7</label><mixed-citation publication-type="other" xlink:type="simple">Kribs-Zaleta, C.M., Lee, M., Rom an, C., Wiley, S. and Hernandez-Suarez, C.M. (2005) The Effect of the HIV/AIDS Epidemic on Africa’s Truck Drivers. Mathematical Biosciences and Engineering: MBE, 2, 771-788.https://doi.org/10.3934/mbe.2005.2.771</mixed-citation></ref><ref id="scirp.97091-ref8"><label>8</label><mixed-citation publication-type="other" xlink:type="simple">Gouws, E. and Cuchi, P. (2012) Focusing the HIV Response through Estimating the Major Modes of HIV Transmission: A Multi-Country Analysis. Sexually Transmitted Infections, 88, i76-i85. https://doi.org/10.1136/sextrans-2012-050719</mixed-citation></ref><ref id="scirp.97091-ref9"><label>9</label><mixed-citation publication-type="other" xlink:type="simple">Bbosa, N., Ssemwanga, D., Nsubuga, R.N., Salazar-Gonzalez, J.F., Salazar, M.G., Nanyonjo, M., Kuteesa, M., Seeley, J., Kiwanuka, N., Bagaya, B.S., et al. (2019) Phylogeography of HIV-1 Suggests that Ugandan Fishing Communities Are a Sink for, Not a Source of, Virus from General Populations. Scientific Reports, 9, 1051.https://doi.org/10.1038/s41598-018-37458-x</mixed-citation></ref><ref id="scirp.97091-ref10"><label>10</label><mixed-citation publication-type="other" xlink:type="simple">UPHIA (2017) Uganda Population Based HIV Impact Assessment. Technical Report, Government of Uganda.</mixed-citation></ref><ref id="scirp.97091-ref11"><label>11</label><mixed-citation publication-type="other" xlink:type="simple">Pillay, D., Herbeck, J., Cohen, M.S., Tulio de Oliveira, Fraser, C., Ratmann, O., Leigh Brown, A. and Kellam, P. (2015) The Pangea-HIV Consortium: Phylogenetics and Networks for Generalised HIV Epidemics in Africa. The Lancet Infectious Diseases, 15, 259. https://doi.org/10.1016/S1473-3099(15)70036-8</mixed-citation></ref><ref id="scirp.97091-ref12"><label>12</label><mixed-citation publication-type="other" xlink:type="simple">Stadler, T., Kouyos, R., von Viktor, W., Yerly, S., Boni, J., Burgisser, P., Klimkait, T., Joos, B., Rieder, P., Xie, D., et al. (2011) Estimating the Basic Reproductive Number from Viral Sequence Data. Molecular Biology and Evolution, 29, 347-357.https://doi.org/10.1093/molbev/msr217</mixed-citation></ref><ref id="scirp.97091-ref13"><label>13</label><mixed-citation publication-type="other" xlink:type="simple">Kuhnert, D., Stadler, T., Vaughan, T.G. and Drummond, A.J. (2014) Simultaneous Reconstruction of Evolutionary History and Epidemiological Dynamics from Viral Sequences with the Birth-Death Sir Model. Journal of the Royal Society Interface, 11, Article ID: 20131106. https://doi.org/10.1098/rsif.2013.1106</mixed-citation></ref><ref id="scirp.97091-ref14"><label>14</label><mixed-citation publication-type="other" xlink:type="simple">Stadler, T. and Bonhoeffer, S. (2013) Uncovering Epidemiological Dynamics in Heterogeneous Host Populations Using Phylogenetic Methods. Philosophical Transactions of the Royal Society B: Biological Sciences, 368, Article ID: 20120198.https://doi.org/10.1098/rstb.2012.0198</mixed-citation></ref><ref id="scirp.97091-ref15"><label>15</label><mixed-citation publication-type="other" xlink:type="simple">Gavryushkina, A., Welch, D., Stadler, T. and Drummond, A.J. (2014) Bayesian Inference of Sampled Ancestor Trees for Epidemiology and Fossil Calibration. PLoS Computational Biology, 10, e1003919. https://doi.org/10.1371/journal.pcbi.1003919</mixed-citation></ref><ref id="scirp.97091-ref16"><label>16</label><mixed-citation publication-type="other" xlink:type="simple">Stadler, T. (2012) How Can We Improve Accuracy of Macroevolutionary Rate Estimates? Systematic Biology, 62, 321-329. https://doi.org/10.1093/sysbio/sys073</mixed-citation></ref><ref id="scirp.97091-ref17"><label>17</label><mixed-citation publication-type="other" xlink:type="simple">Stadler, T. (2014) Treesim: Simulating Trees under the Birth-Death Model. R Package Version 2.0.</mixed-citation></ref><ref id="scirp.97091-ref18"><label>18</label><mixed-citation publication-type="other" xlink:type="simple">Stadler, T. and Maintainer Stadler, T. (2015) Estimating Birth and Death Rates Based on Phylogenies. R Package Version 3.3.</mixed-citation></ref><ref id="scirp.97091-ref19"><label>19</label><mixed-citation publication-type="other" xlink:type="simple">Gascuel, O. (2005) Mathematics of Evolution and Phylogeny. OUP, Oxford.</mixed-citation></ref><ref id="scirp.97091-ref20"><label>20</label><mixed-citation publication-type="other" xlink:type="simple">Ota, R., Waddell, P.J., Hasegawa, M., Shimodaira, H. and Kishino, H. (2000) Appropriate Likelihood Ratio Tests and Marginal Distributions for Evolutionary Tree Models with Constraints on Parameters. Molecular Biology and Evolution, 17, 798-803. https://doi.org/10.1093/oxfordjournals.molbev.a026358</mixed-citation></ref><ref id="scirp.97091-ref21"><label>21</label><mixed-citation publication-type="other" xlink:type="simple">Yebra, G., Hodcroft, E.B., Ragonnet-Cronin, M.L., Pillay, D., Leigh Brown, A.J., Fraser, C., Kellam, P., De Oliveira, T., Dennis, A., Hoppe, A., et al. (2016) Using Nearly Full-Genome HIV Sequence Data Improves Phylogeny Reconstruction in a Simulated Epidemic. Scientific Reports, 6, 39489.https://doi.org/10.1038/srep39489</mixed-citation></ref><ref id="scirp.97091-ref22"><label>22</label><mixed-citation publication-type="other" xlink:type="simple">Drummond, A.J. and Rambaut, A. (2007) Beast: Bayesian Evolutionary Analysis by Sampling Trees. BMC Evolutionary Biology, 7, 214.https://doi.org/10.1186/1471-2148-7-214</mixed-citation></ref><ref id="scirp.97091-ref23"><label>23</label><mixed-citation publication-type="other" xlink:type="simple">Blum, M.G.B. and Francois, O. (2005) On Statistical Tests of Phylogenetic Tree Imbalance: The Sackin and Other Indices Revisited. Mathematical Biosciences, 195, 141-153. https://doi.org/10.1016/j.mbs.2005.03.003</mixed-citation></ref><ref id="scirp.97091-ref24"><label>24</label><mixed-citation publication-type="other" xlink:type="simple">Pybus, O.G. and Harvey, P.H. (2000) Testing Macro-Evolutionary Models Using Incomplete Molecular Phylogenies. Proceedings of the Royal Society of London B: Biological Sciences, 267, 2267-2272. https://doi.org/10.1098/rspb.2000.1278</mixed-citation></ref><ref id="scirp.97091-ref25"><label>25</label><mixed-citation publication-type="other" xlink:type="simple">Blum, M.G.B., Francois, O. and Janson, S. (2006) The Mean, Variance and Limiting Distribution of Two Statistics Sensitive to Phylogenetic Tree Balance. The Annals of Applied Probability, 16, 2195-2214. https://doi.org/10.1214/105051606000000547</mixed-citation></ref><ref id="scirp.97091-ref26"><label>26</label><mixed-citation publication-type="other" xlink:type="simple">Shao, K.T. (1990) Tree Balance. Systematic Zoology, 39, 266-276.https://doi.org/10.2307/2992186</mixed-citation></ref><ref id="scirp.97091-ref27"><label>27</label><mixed-citation publication-type="other" xlink:type="simple">Bortolussi, N., Durand, E., Blum, M. and Francois, O. (2005) Aptreeshape: Statistical Analysis of Phylogenetic Tree Shape. Bioinformatics, 22, 363-364.https://doi.org/10.1093/bioinformatics/bti798</mixed-citation></ref><ref id="scirp.97091-ref28"><label>28</label><mixed-citation publication-type="other" xlink:type="simple">Maddison, W.P., Midford, P.E. and Otto, S.P. (2007) Estimating a Binary Character’s Effect on Speciation and Extinction. Systematic Biology, 56, 701-710.https://doi.org/10.1080/10635150701607033</mixed-citation></ref><ref id="scirp.97091-ref29"><label>29</label><mixed-citation publication-type="other" xlink:type="simple">Rabosky, D.L (2010). Extinction Rates Should Not Be Estimated from Molecular Phylogenies. Evolution: International Journal of Organic Evolution, 64, 1816-1824.https://doi.org/10.1111/j.1558-5646.2009.00926.x</mixed-citation></ref><ref id="scirp.97091-ref30"><label>30</label><mixed-citation publication-type="other" xlink:type="simple">Beaulieu, J.M. and O’Meara, B.C. (2015) Extinction Can Be Estimated from Moderately Sized Molecular Phylogenies. Evolution, 69, 1036-1043.https://doi.org/10.1111/evo.12614</mixed-citation></ref><ref id="scirp.97091-ref31"><label>31</label><mixed-citation publication-type="other" xlink:type="simple">Rabosky, D.L. (2016) Challenges in the Estimation of Extinction from Molecular Phylogenies: A Response to Beaulieu and O’Meara. Evolution, 70, 218-228.https://doi.org/10.1111/evo.12820</mixed-citation></ref></ref-list></back></article>