<?xml version="1.0" encoding="UTF-8"?><!DOCTYPE article  PUBLIC "-//NLM//DTD Journal Publishing DTD v3.0 20080202//EN" "http://dtd.nlm.nih.gov/publishing/3.0/journalpublishing3.dtd"><article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" dtd-version="3.0" xml:lang="en" article-type="research article"><front><journal-meta><journal-id journal-id-type="publisher-id">OALibJ</journal-id><journal-title-group><journal-title>Open Access Library Journal</journal-title></journal-title-group><issn pub-type="epub">2333-9705</issn><publisher><publisher-name>Scientific Research Publishing</publisher-name></publisher></journal-meta><article-meta><article-id pub-id-type="doi">10.4236/oalibj.2023.1012960</article-id><article-id pub-id-type="publisher-id">OALibJ-130125</article-id><article-categories><subj-group subj-group-type="heading"><subject>Articles</subject></subj-group><subj-group subj-group-type="Discipline-v2"><subject>Biomedical&amp;Life Sciences</subject><subject> Business&amp;Economics</subject><subject> Chemistry&amp;Materials Science</subject><subject> Computer Science&amp;Communications</subject><subject> Earth&amp;Environmental Sciences</subject><subject> Engineering</subject><subject> Medicine&amp;Healthcare</subject><subject> Physics&amp;Mathematics</subject><subject> Social Sciences&amp;Humanities</subject></subj-group></article-categories><title-group><article-title>
 
 
  Analysis of a Two-Stage Adaptive Negative Binomial Group Testing Model for Estimating Prevalence of a Rare Trait
 
</article-title></title-group><contrib-group><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Jackline</surname><given-names>Akomboh</given-names></name><xref ref-type="aff" rid="aff1"><sup>1</sup></xref></contrib><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Wanyonyi</surname><given-names>Ronald Waliaula</given-names></name><xref ref-type="aff" rid="aff1"><sup>1</sup></xref></contrib><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Cox</surname><given-names>Tamba</given-names></name><xref ref-type="aff" rid="aff1"><sup>1</sup></xref></contrib><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Justine</surname><given-names>Obwoge Okenye</given-names></name><xref ref-type="aff" rid="aff1"><sup>1</sup></xref></contrib></contrib-group><aff id="aff1"><addr-line>Department of Mathematics, Egerton University, Egerton, Kenya</addr-line></aff><pub-date pub-type="epub"><day>04</day><month>12</month><year>2023</year></pub-date><volume>10</volume><issue>12</issue><fpage>1</fpage><lpage>14</lpage><history><date date-type="received"><day>3,</day>	<month>November</month>	<year>2023</year></date><date date-type="rev-recd"><day>24,</day>	<month>December</month>	<year>2023</year>	</date><date date-type="accepted"><day>27,</day>	<month>December</month>	<year>2023</year></date></history><permissions><copyright-statement>&#169; Copyright  2014 by authors and Scientific Research Publishing Inc. </copyright-statement><copyright-year>2014</copyright-year><license><license-p>This work is licensed under the Creative Commons Attribution International License (CC BY). http://creativecommons.org/licenses/by/4.0/</license-p></license></permissions><abstract><p>
 
 
  Group testing is an efficient method for classifying observations and estimating trait prevalence in a population. However, using appropriate group sizes is crucial for maximizing its benefits. Adaptive schemes have been developed to address improper group size selection issues. Existing adaptive schemes are based on a Binomial sampling model, requiring testing of all groups before recording successes. In certain scenarios, like infectious diseases, immediate reporting of estimates upon detection is necessary. A two-stage adaptive Negative Binomial group testing model for such cases was constructed. This adaptive model adjusted group sizes based on estimates from previous stages thus using optimal sizes to minimize the mean squared error and variance of the prevalence rate estimate. The maximum likelihood estimation method was employed to find the model’s parameter estimate, and its properties were also investigated. The comparative analysis highlighted the superiority of the adaptive model over the non-adaptive model especially under low prevalence emphasizing the importance of incorporating adaptivity in group testing procedures, particularly in disease screening and surveillance, such as for COVID-19. 
 
</p></abstract><kwd-group><kwd>Negative Binomial</kwd><kwd> Group Testing</kwd><kwd> Maximum Likelihood Estimation</kwd><kwd> Prevalence</kwd></kwd-group></article-meta></front><body><sec id="s1"><title>1. Introduction</title><p>Group testing, also known as pool testing or batch testing, is a method that involves combining individuals into various pools and conducting tests on these pools simultaneously. The tests are used to detect the presence of infections or defects, as seen in epidemiological studies. The idea was first introduced by Dorfman in 1943 to improve cost-saving techniques in detecting soldiers with syphilis [<xref ref-type="bibr" rid="scirp.130125-ref1">1</xref>] . Since Dorfman’s pioneering work, group testing has been applied in various fields, including epidemiology, quality control, and genetics, with two main objectives including classification [<xref ref-type="bibr" rid="scirp.130125-ref2">2</xref>] [<xref ref-type="bibr" rid="scirp.130125-ref3">3</xref>] and estimation of prevalence [<xref ref-type="bibr" rid="scirp.130125-ref4">4</xref>] . The classification or identification of individuals as either positive or negative of a trait serves as the first objective of Group testing championed by [<xref ref-type="bibr" rid="scirp.130125-ref1">1</xref>] . To enhance cost-effectiveness and reduce the required tests, an adaptation of the Dorfman testing scheme has been investigated and extended to incorporate multi-stage testing [<xref ref-type="bibr" rid="scirp.130125-ref5">5</xref>] . Apart from classification problems, group testing has also been used extensively in the estimation problem of the prevalence rate of a trait. This is the second objective of group testing pioneered by [<xref ref-type="bibr" rid="scirp.130125-ref6">6</xref>] and served as the main focus of this study. In his study, he used the maximum likelihood estimation method, which was found to be reliable when the population was small [<xref ref-type="bibr" rid="scirp.130125-ref7">7</xref>] . Extended the estimation work using the MLE method and incorporating testing errors to account for real practical situations where errors are bound to happen during testing procedure.</p><p>Subsequent studies on estimation focused on design matters especially based on selection of group sizes used in group testing procedures [<xref ref-type="bibr" rid="scirp.130125-ref8">8</xref>] . The main aim of putting more emphasis on selection of group sizes was to reduce the chances of obtaining all negative or all positive groups during testing. Another aim was to minimize MSE by incorporating prior information in choosing k [<xref ref-type="bibr" rid="scirp.130125-ref9">9</xref>] . All these studies were carried out using the Binomial sampling model [<xref ref-type="bibr" rid="scirp.130125-ref10">10</xref>] . Later suggested the use of Negative Binomial sampling for estimating efficient p when the prevalence rate is small. Point and interval estimation has been explored under the negative binomial model using equal group sizes and has been found to be efficient in surveillance cases where quick response is desired [<xref ref-type="bibr" rid="scirp.130125-ref11">11</xref>] . Recent works that have considered inverse binomial models include estimation using the Bayesian approach as well as confidence interval estimation [<xref ref-type="bibr" rid="scirp.130125-ref12">12</xref>] .</p><p>Group testing can be categorized into two forms: non-adaptive and adaptive group testing schemes. Non-adaptive group testing involves testing groups of a fixed size to obtain dichotomous results. On the other hand, adaptive group testing adjusts the group sizes from one stage to the next, allowing for more flexibility and efficiency [<xref ref-type="bibr" rid="scirp.130125-ref13">13</xref>] . Adaptive estimators have been developed to improve efficiency by reducing the Mean Squared Error under the binomial model [<xref ref-type="bibr" rid="scirp.130125-ref9">9</xref>] . Other probabilistic models, such as the beta-binomial, geometric, and hyper-geometric models, have also been considered [<xref ref-type="bibr" rid="scirp.130125-ref14">14</xref>] . For urgent situations where the prevalence rate of a trait needs to be estimated quickly, the Negative Binomial sampling model has been found to be preferable over the Binomial model [<xref ref-type="bibr" rid="scirp.130125-ref11">11</xref>] . The Negative Binomial model has been applied in emergency situations like disease outbreaks and natural disasters to measure risk promptly. Despite the benefits of group testing, its success depends on choosing appropriate group sizes. Inadequate group size selection can lead to deficiencies in the procedure. To address these deficiencies, this study aimed to investigate a two-stage adaptive Negative Binomial group testing procedure for estimating the prevalence rate of a rare trait.</p></sec><sec id="s2"><title>2. Methodology</title><sec id="s2_1"><title>2.1. The Model</title><p>The technique known as Negative Binomial sampling holds significant importance in the context of biological sample collection. If the proportion of individuals possessing a specific character trait is denoted as p, and sampling continues until a predetermined number, such as x individuals, is observed, the distribution of the sampled individuals follows a negative binomial distribution. Combining Negative Binomial sampling with group testing provides an attractive approach for delivering early estimates during the screening process.</p><p>This model operates under the assumption that the number of pools having a trait of interest is pre-established, and the testing process continues until the desired number of positive pools is identified.</p><p>The adaptive scheme involved testing groups in stages and adjusting the group sizes from one stage to the next. The group size used at a stage depended entirely on the outcome of the preceding stages. This implies that k’s were determined sequentially as the experiment progressed. The value of k<sub>1</sub> is determined by optimizing the variance of the estimator obtained from the non-adaptive Negative Binomial group testing scheme, which serves as prior information. In Stage One, the number of positive pools to be observed X<sub>1</sub> is fixed, and T<sub>1</sub> the number of groups to be tested before observing X<sub>1</sub> positive pools follows a Negative Binomial distribution.</p><p>In Stage Two, the group size k<sub>2</sub> is constructed by minimizing the variance of the estimator obtained in Stage One.</p><p>k 2 = arg min k [ var ( p ^ 1 ) ] p = p ^ 1 (1)</p><p>where k<sub>2</sub> was the value which minimizes the variance of p ^ 1 which is the estimator obtained in stage 1. The goal is to select the group size that minimizes the variance of the estimator from Stage One. This approach is crucial for enhancing the precision and accuracy of the overall estimation procedure. By minimizing the variance, the estimation process aims to achieve a more reliable and stable outcome, contributing to the effectiveness of the adaptive estimation model.</p><p>The number of positive pools to be observed in Stage Two X<sub>2</sub> is fixed, and the number of groups to be tested T<sub>2</sub> to achieve X<sub>2</sub> positive pools follows a Negative Binomial distribution, which depends on the output of Stage One. This adaptive model ensures an efficient allocation of resources and provides a robust strategy for large-scale screening and identification of positive groups in group testing scenarios.</p><p>The derivation presented focused on two models; the usual Negative Binomial model and the proposed two-stage adaptive Negative Binomial group testing model for estimating the prevalence of a rare trait.</p><p>Using the usual non-adaptive Negative Binomial to get p ^ N as in Katholi (2006). T follows a Negative Binomial with parameter x and π .</p><p>f ( t / p ) = ( t − 1 x − 1 ) [ π ( p ) ] x [ 1 − π ( p ) ] t − x (2)</p><p>The Likelihood function of Equation (2) is expressed as;</p><p>L ( p / t , x ) ∝ [ π ( p ) ] x [ 1 − π ( p ) ] t − x (3)</p><p>The log Likelihood function to base 10 is given as;</p><p>log L ( p / t , x ) ∝ x log [ π ( p ) ] + ( t − x ) log [ 1 − π ( p ) ] (4)</p><p>The maximum likelihood estimator of the non-adaptive model is obtained as the solution to</p><p>∂ log L ( . ) ∂ ( p ) = 0 (5)</p><p>Which is equivalent to;</p><p>x π ( p ) − t − x 1 − π ( p ) = 0 (6)</p><p>where π = 1 − ( 1 − p ) k which is the probability of obtaining a positive group. Equation (6) yields the results obtained by Katholi (2006) as;</p><p>p ^ N = 1 − ( 1 − x t ) 1 k (7)</p><p>Finding the variance of p ^ N the Cramer Rao Lower Bound was used where the Fisher’s information of the likelihood function was utilized;</p><p>Varianceof   p ^ N = 1 I ( p ^ N ) (8)</p><p>where the Fisher’s information is given as;</p><p>I ( p ^ N ) = − E [ ∂ 2 ∂ p 2 log L ( . ) ] − 1 (9)</p><p>The second derivative of the log likelihood function is given as;</p><p>= − t ( π ′ ( p ) ) 2 π ( p ) ( 1 − π ( p ) ) (10)</p><p>It is worth noting that π ( p ) = 1 − ( 1 − p ) k and π ′ ( p ) = k ( 1 − p ) k − 1 . Substituting in Equation (10) and taking expectation gives;</p><p>= t k 2 ( 1 − p ) 2 k − 2 1 − ( 1 − p ) k ( 1 − p ) k (11)</p><p>The E ( T ) = x π . Thus, taking the expectation and the inverse of Equation (11) will yield</p><p>Var ( p ^ 0 ) = 1 − ( 1 − p ) k t k 2 ( 1 − p ) k − 2 (12)</p><p>The proposed two-stage adaptive model aimed at optimizing resource allocation and estimation efficiency. The study considered two sets of desired positive groups, X<sub>1</sub> and X<sub>2</sub>. The first stage estimator was based on the Negative Binomial distribution with a prior derived from the non-adaptive estimator. In the second stage, the study introduces the two-stage adaptive estimator.</p><p>Stage two proceeds by testing groups of size k<sub>2</sub> and T<sub>2</sub> is the number of the groups to be tested to obtain X<sub>2</sub> positive groups. Thus, T<sub>2</sub> is conditioned on T<sub>1</sub>. This follows that T<sub>2</sub> has a negative binomial distribution. Specifically,</p><p>T 2 / T 1 ~ NegativeBinomial ( X 2 , π 2 / 1 = 1 − ( 1 − p ) k 2 ( t 1 ) )</p><p>Equation (6) gives the joint distribution of T<sub>1</sub> and T<sub>2</sub> as;</p><p>f ( T 2 , T 1 ) = f ( T 2 / T 1 ) &#215; f ( T 1 ) = NegativeBinomial ( X 1 , π = 1 − ( 1 − p ) k 1 )     &#215; NegativeBinomial ( X 2 , π 2 / 1 = 1 − ( 1 − p ) k 2 ( t 1 ) ) (13)</p><p>The joint distribution of T<sub>1</sub> and T<sub>2</sub> was used to derive the final two-stage adaptive estimator p ^ 2 and is given as;</p><p>f ( t 1 , t 2 ) = ( t 1 − 1 x 1 − 1 ) [ 1 − ( 1 − p ) k 1 ] x 1 ( 1 − p ) k 1 ( t 1 − x 1 )   &#215; ( t 2 − 1 x 2 − 1 ) [ 1 − ( 1 − p ) k 2 ] x 2 ( 1 − p ) k 2 ( t 2 − x 2 ) (14)</p><p>( t 1 − 1 x 1 − 1 ) and ( t 2 − 1 x 2 − 1 ) are constants of proportionality thus we replace with ∝ giving;</p><p>f ( t 1 , t 2 ) ∝ [ 1 − ( 1 − p ) k 1 ] x 1 ( 1 − p ) k 1 ( t 1 − x 1 ) &#215; [ 1 − ( 1 − p ) k 2 ] x 2 ( 1 − p ) k 2 ( t 2 − x 2 ) (15)</p><p>The log likelihood function to base 10 for the joint distribution of T<sub>1</sub> and T<sub>2</sub> was obtained as.</p><p>ln L ∝ x 1 ln [ 1 − ( 1 − p ) k 1 ] + k 1 ( t 1 − x 1 ) ln ( 1 − p )     + x 2 [ 1 − ( 1 − p ) k 2 ] + k 2 ( t 2 − x 2 ) ln ( 1 − p ) (16)</p><p>The adaptive estimator is obtained by solving Equation (17) iteratively since the solution equated to zero was not in a closed form thus was not tractable.</p><p>∂ log L ∂ p ∝ x 1 π 1 π ′ 1 − t 1 − x 1 1 − π 1 π ′ 1 + x 2 π 2 π ′ 2 − t 2 − x 2 1 − π 2 π ′ 2 (17)</p><p>where π 1 = 1 − ( 1 − p ) k 1 and π 2 = 1 − ( 1 − p ) k 2 . Also π ′ 1 = k 1 ( 1 − p ) k 1 − 1 and π ′ 2 = k 2 ( 1 − p ) k 2 − 1 .</p><p>The variance of the adaptive estimator was derived using Fisher’s information. Where the Fisher’s information is given as;</p><p>I = − E [ ∂ 2 ∂ p 2 log L ( . ) ] (18)</p><p>Thus, the variance of the adaptive estimator was derived and obtained as;</p><p>Var ( p ^ A ) = 1 ∑ i = 1 2 x i k i 2 ( 1 − p ) 2 k i − 2 π i 2 ( 1 − π i ) (19)</p><p>This variance was used to construct the Wald confidence interval as;</p><p>p ^ A &#177; Z α 2 Var ( p ^ A ) (20)</p></sec><sec id="s2_2"><title>2.2. Simulation</title><p>Data was simulated using an algorithm which mimics the Negative Binomial process as illustrated in <xref ref-type="fig" rid="fig1">Figure 1</xref>.</p><p>Steps for Simulation</p><p>Step 1: Specify p, k and X then set x = 0.</p><p>Step 2: Generate k Bernoulli random variables.</p><p>Set Y = (Y<sub>1</sub>, …, Y<sub>k</sub>).</p><p>Step 3: If the sum of y<sub>is</sub> is greater than 0. A success is considered. If the sum of y<sub>is</sub> is less than the group is considered negative.</p><p>Step 4: If the success was recorded, repeat the loop if x is not equal to X. If x = X, the procedure stops.</p><p>Step 5: Report T and calculate p estimate.</p></sec></sec><sec id="s3"><title>3. Results and Discussion</title><sec id="s3_1"><title>3.1. Relationship between t, p, and k</title><p>The Negative Binomial sampling method used in group testing experiments involves a random number of trials, while the required positive pools and group size are fixed. The number of tests needed to obtain the required positive groups depends on various variables, so it’s important to understand how changing the values of p and k affects the number of testing trials required.</p><p>As the probability of success increases in the negative binomial group testing model, the number of trials required to obtain the desired number of positive groups generally decreases as illustrated in <xref ref-type="fig" rid="fig2">Figure 2</xref>. Examining the plots, for a fixed value of k, it can be observed that as p increases, the number of trials decreases. This trend holds true across different values of k. This behavior is expected because a higher probability of success implies a greater likelihood of encountering positive groups during testing. Therefore, fewer trials are needed to reach the desired number of positive groups when the probability of success is higher. It is worth noting that increasing the group size from 5 to 100 in the negative binomial group testing model typically leads to a reduction in the number of trials required to achieve the desired number of positive groups as well.</p></sec><sec id="s3_2"><title>3.2. Adaptive Estimator and Its Properties</title><p>The results of the maximum likelihood estimator and its properties including the variance, bias, and mean squared error for the adaptive group testing model are presented in <xref ref-type="table" rid="table1">Table 1</xref>. The results are organized based on different group sizes and true probabilities while the number of predetermined desired positive groups set at X = 30 as set by [<xref ref-type="bibr" rid="scirp.130125-ref15">15</xref>] .</p><table-wrap id="table1" ><label><xref ref-type="table" rid="table1">Table 1</xref></label><caption><title> Adaptive estimator with its properties for k = 5, 10, 20, 50, 100 when X = 30</title></caption><table><tbody><thead><tr><th align="center" valign="middle" >p</th><th align="center" valign="middle" >Mle</th><th align="center" valign="middle" >Var</th><th align="center" valign="middle" >Bias</th><th align="center" valign="middle" >MSE</th></tr></thead><tr><td align="center" valign="middle"  colspan="5"  >k = 5</td></tr><tr><td align="center" valign="middle" >0.001</td><td align="center" valign="middle" >0.000488</td><td align="center" valign="middle" >7.94E−09</td><td align="center" valign="middle" >−0.0004</td><td align="center" valign="middle" >4.0938E−08</td></tr><tr><td align="center" valign="middle" >0.005</td><td align="center" valign="middle" >0.002677</td><td align="center" valign="middle" >2.38E−07</td><td align="center" valign="middle" >−0.0025</td><td align="center" valign="middle" >1.02E−07</td></tr><tr><td align="center" valign="middle" >0.01</td><td align="center" valign="middle" >0.004862</td><td align="center" valign="middle" >7.82E−07</td><td align="center" valign="middle" >−0.0040</td><td align="center" valign="middle" >4.05E−06</td></tr><tr><td align="center" valign="middle" >0.05</td><td align="center" valign="middle" >0.035433</td><td align="center" valign="middle" >3.94E−05</td><td align="center" valign="middle" >−0.0209</td><td align="center" valign="middle" >7.98E−05</td></tr><tr><td align="center" valign="middle" >0.1</td><td align="center" valign="middle" >0.075796</td><td align="center" valign="middle" >0.000167</td><td align="center" valign="middle" >−0.0275</td><td align="center" valign="middle" >0.000322241</td></tr><tr><td align="center" valign="middle"  colspan="5"  >k = 10</td></tr><tr><td align="center" valign="middle" >0.001</td><td align="center" valign="middle" >0.000551</td><td align="center" valign="middle" >1.01E−08</td><td align="center" valign="middle" >−0.0005</td><td align="center" valign="middle" >3.88E−08</td></tr><tr><td align="center" valign="middle" >0.005</td><td align="center" valign="middle" >0.002264</td><td align="center" valign="middle" >1.70E−07</td><td align="center" valign="middle" >−0.0022</td><td align="center" valign="middle" >8.89E−07</td></tr><tr><td align="center" valign="middle" >0.01</td><td align="center" valign="middle" >0.005933</td><td align="center" valign="middle" >1.16E−06</td><td align="center" valign="middle" >−0.0042</td><td align="center" valign="middle" >4.18E−06</td></tr><tr><td align="center" valign="middle" >0.05</td><td align="center" valign="middle" >0.027808</td><td align="center" valign="middle" >2.31E−05</td><td align="center" valign="middle" >−0.0177</td><td align="center" valign="middle" >8.29E−05</td></tr><tr><td align="center" valign="middle" >0.1</td><td align="center" valign="middle" >0.070342</td><td align="center" valign="middle" >0.000142</td><td align="center" valign="middle" >−0.0258</td><td align="center" valign="middle" >0.00034779</td></tr><tr><td align="center" valign="middle"  colspan="5"  >k = 20</td></tr><tr><td align="center" valign="middle" >0.001</td><td align="center" valign="middle" >0.000634</td><td align="center" valign="middle" >1.34E−08</td><td align="center" valign="middle" >−0.0006</td><td align="center" valign="middle" >3.59E−08</td></tr><tr><td align="center" valign="middle" >0.005</td><td align="center" valign="middle" >0.003144</td><td align="center" valign="middle" >3.28E−07</td><td align="center" valign="middle" >−0.0019</td><td align="center" valign="middle" >1.05E−06</td></tr><tr><td align="center" valign="middle" >0.01</td><td align="center" valign="middle" >0.005263</td><td align="center" valign="middle" >9.15E−07</td><td align="center" valign="middle" >−0.0057</td><td align="center" valign="middle" >3.63E−06</td></tr><tr><td align="center" valign="middle" >0.05</td><td align="center" valign="middle" >0.029922</td><td align="center" valign="middle" >2.82E−05</td><td align="center" valign="middle" >−0.0198</td><td align="center" valign="middle" >9.04E−05</td></tr><tr><td align="center" valign="middle" >0.1</td><td align="center" valign="middle" >0.074906</td><td align="center" valign="middle" >0.000158</td><td align="center" valign="middle" >−0.0268</td><td align="center" valign="middle" >0.00036145</td></tr><tr><td align="center" valign="middle"  colspan="5"  >k = 50</td></tr><tr><td align="center" valign="middle" >0.001</td><td align="center" valign="middle" >0.000588</td><td align="center" valign="middle" >1.15E−08</td><td align="center" valign="middle" >−0.0004</td><td align="center" valign="middle" >3.84E−08</td></tr><tr><td align="center" valign="middle" >0.005</td><td align="center" valign="middle" >0.002204</td><td align="center" valign="middle" >1.61E−07</td><td align="center" valign="middle" >−0.0026</td><td align="center" valign="middle" >8.88E−07</td></tr><tr><td align="center" valign="middle" >0.01</td><td align="center" valign="middle" >0.004198</td><td align="center" valign="middle" >5.84E−07</td><td align="center" valign="middle" >−0.0042</td><td align="center" valign="middle" >3.28E−06</td></tr><tr><td align="center" valign="middle" >0.05</td><td align="center" valign="middle" >0.026812</td><td align="center" valign="middle" >2.27E−05</td><td align="center" valign="middle" >−0.0223</td><td align="center" valign="middle" >9.28E−05</td></tr><tr><td align="center" valign="middle" >0.1</td><td align="center" valign="middle" >0.075437</td><td align="center" valign="middle" >0.000165</td><td align="center" valign="middle" >−0.0506</td><td align="center" valign="middle" >0.001282513</td></tr><tr><td align="center" valign="middle"  colspan="5"  >k = 100</td></tr><tr><td align="center" valign="middle" >0.001</td><td align="center" valign="middle" >0.000488</td><td align="center" valign="middle" >7.94E−09</td><td align="center" valign="middle" >−0.0005</td><td align="center" valign="middle" >3.93E−08</td></tr><tr><td align="center" valign="middle" >0.005</td><td align="center" valign="middle" >0.002503</td><td align="center" valign="middle" >2.08E−07</td><td align="center" valign="middle" >−0.0018</td><td align="center" valign="middle" >8.49E−07</td></tr><tr><td align="center" valign="middle" >0.01</td><td align="center" valign="middle" >0.00432</td><td align="center" valign="middle" >6.18E−07</td><td align="center" valign="middle" >−0.0048</td><td align="center" valign="middle" >3.70E−06</td></tr><tr><td align="center" valign="middle" >0.05</td><td align="center" valign="middle" >0.033403</td><td align="center" valign="middle" >3.49E−05</td><td align="center" valign="middle" >−0.0164</td><td align="center" valign="middle" >0.00030666</td></tr><tr><td align="center" valign="middle" >0.1</td><td align="center" valign="middle" >0.068592</td><td align="center" valign="middle" >0.000137</td><td align="center" valign="middle" >−0.0175</td><td align="center" valign="middle" >0.00438814</td></tr></tbody></table></table-wrap><p>A scrutiny of <xref ref-type="table" rid="table1">Table 1</xref> shows that the estimated probabilities tend to increase as the probability increases, although the increase is generally small. It is important to note that the MLE values of the adaptive model exhibits monotonic behavior as the model dynamically adjusts the group size based on stage one’s outcomes which lead to more consistent and accurate estimations. The MLE generally increases as p increases for all values of k as found by [<xref ref-type="bibr" rid="scirp.130125-ref16">16</xref>] . The variance of the estimated probabilities remains relatively small across different values of p. The bias of the estimation shows negative values indicating a slight underestimation of the true population probability. However, the bias remains relatively small across all values of p. The MSE combines the variance and bias to provide an overall measure of estimation accuracy. The bias remains relatively small and consistent across different values of p indicating the effectiveness of the adaptive group testing model in reducing bias.</p></sec><sec id="s3_3"><title>3.3. Relationship between p ^ and p</title><p>The results presented in this section examine the relationship between the adaptive maximum likelihood estimates and the true probability for different values of group size. The results are represented in four graphs for k = 5, 20, 50, 100. These findings highlighted that the adaptive nature of the model in adjusting the estimations was based on observed outcomes and the varying performance of the adaptive approach across different group sizes.</p><p>The relationship was further investigated by plotting the values p ^ against p while varying the waiting parameter X for different values of group size k.</p><p><xref ref-type="fig" rid="fig3">Figure 3</xref> illustrates the relationship between the Maximum Likelihood Estimation values and the true probability for different combinations of X in the adaptive approach. The MLE values generally increase as p increases, as well as with larger values of X and k. However, the rate and pattern of this increase vary depending on the specific combination of X and k. For X = 30, the MLE graph</p><p>shows a relatively steep curve, indicating that small changes in p lead to noticeable changes in the estimated probabilities of success. For X = 20, the MLE values also increase as p increases, but the curve is relatively flat, suggesting a less pronounced change in the MLE values as p varies. The adaptive MLE in this scenario exhibits a slower rate of increase compared to when X = 30, implying a less sensitive response to changes in p.</p><p>Similar patterns are observed for the other combinations of X and k. The adaptive approach tends to provide more conservative estimates with low MLE values. The MLE values increase with increasing p, but the specific patterns and sensitivities depend on the combination of X and k. The adaptive approach consistently provides more conservative estimates with lower MLE values, indicating a cautious approach in estimating the prevalence of the rare trait.</p></sec><sec id="s3_4"><title>3.4. Relationship between Variance of p ^ and p</title><p>We examined the relationship between the variance of the estimated proportion p ^ and the true proportion p in the context of the two-stage adaptive negative binomial model. <xref ref-type="fig" rid="fig4">Figure 4</xref> provides insights into this relationship for different combinations of X and k.</p><p><xref ref-type="fig" rid="fig4">Figure 4</xref> presents the interplay between X, k, and the variance of p ^ in the two-stage adaptive negative binomial model. They highlight the impact of the true proportion p, the desired number of positive groups X, and the group size k on the variability of the estimates. As the true proportion p increases, the variance of the estimated proportion p ^ also tends to increase, although the magnitude of increase varies across different X and k values. This suggests that higher probabilities of success lead to greater variability in the estimates, highlighting the increased uncertainty associated with higher p values. Comparing different X values, we find that as x increases, the variance of p ^ generally tends to increase. This implies that aiming for a higher number of positive groups introduces more variability into the estimates. Analyzing the effect of k, we notice</p><p>that for a fixed x value, as k increases, the variance of p ^ tends to decrease. This indicates that larger group sizes result in more precise estimates and lower variability. Larger group sizes provide more information, reducing the sampling error and enhancing the precision of the estimates. The insights obtained from this analysis can inform the selection of appropriate values for X and k to optimize the accuracy and reliability of the model in estimating the proportion of successes in a population.</p></sec></sec><sec id="s4"><title>4. Comparison of the Model</title><p>We conducted a model comparison to evaluate the performance of the two-stage adaptive negative binomial group testing model for estimating the prevalence of a rare trait over the non-adaptive model. In this study, we used two statistical measures, Asymptotic Relative Efficiency and the Relative Mean Squared Error, to compare the efficiency and accuracy of the proposed two-stage adaptive negative binomial group testing model with an existing non-adaptive model by [<xref ref-type="bibr" rid="scirp.130125-ref6">6</xref>] .</p><p>The estimator of non-adaptive model was denoted by p ^ N since it is developed under the usual Negative Binomial model while the computed estimator was denoted as p ^ A since is developed under the adaptive Negative Binomial model. Then, ARE was obtained as</p><p>ARE = Var ( p ^ N ) Var ( p ^ A ) (21)</p><p>ARE values of greater than one implied that our model is more efficient than the non-adaptive model. ARE measures how much more efficient the adaptive model is compared to the non-adaptive model as the sample size approaches infinity. Higher ARE values indicate that the adaptive model provides better estimates and inferences. The comparison was done for different combinations of group size k and the desired number of positive groups X.</p><p>The study found that as X increased, there was an overall increasing trend in the ARE values, indicating improved performance in detecting the desired outcome (<xref ref-type="fig" rid="fig5">Figure 5</xref>). On the other hand, increasing k for a fixed X value led to a decreasing trend in the ARE values, implying that larger group sizes result in better performance in detecting the desired outcome.</p><p>The RMSE was used to compare the mean squared errors of the estimators obtained from the constructed adaptive model with the one by [<xref ref-type="bibr" rid="scirp.130125-ref6">6</xref>] . This is a convenient way of comparing the MSE of the estimates obtained using different procedures. It is expected that a good model to produce an estimator with a small MSE. For this study, RMSE was computed as;</p><p>RMSE = MSE ( p ^ N ) MSE ( p ^ A ) (22)</p><p>The study found that the computed estimator was more efficient as compared to the [<xref ref-type="bibr" rid="scirp.130125-ref10">10</xref>] since the values of RMSE were greater than one (<xref ref-type="fig" rid="fig6">Figure 6</xref>). It was worth noting also that as the true probability of success p increases, the RMSE</p><p>decreases, indicating better fit and improved estimation accuracy at low prevalence. In addition, aiming for a greater number of positive groups and increasing the group size also contributed to higher values of RMSE values, suggesting more accurate estimation and improved model accuracy especially when the prevalence is low.</p></sec><sec id="s5"><title>5. Conclusion and Recommendations</title><p>In conclusion, this study successfully achieved its objectives by developing and analyzing a two-stage adaptive negative binomial model in group testing for estimating the prevalence of a rare trait using MLE. The adaptive estimator demonstrated superior performance compared to the non-adaptive estimator, providing more accurate and smaller estimates, while maintaining low variance and bias. The comprehensive simulations further confirmed the superiority of the adaptive model, showing better efficiency, lower mean squared error MSE, and improved fit to the data. The study recommends future research to incorporate imperfect tests in the model to reflect real-world scenarios and evaluate their impact on estimation accuracy. In addition, exploring the extension of the adaptive estimator to multi-stage group testing procedures could enhance the model’s applicability to larger populations and improve logistical considerations for estimation.</p></sec><sec id="s6"><title>Conflicts of Interest</title><p>The authors declare no conflicts of interest.</p></sec><sec id="s7"><title>Cite this paper</title><p>Akomboh, J., Waliaula, W.R., Tamba, C. and Okenye, J.O. (2023) Analysis of a Two-Stage Adaptive Negative Binomial Group Testing Model for Estimating Prevalence of a Rare Trait. Open Access Library Journal, 10: e10960. https://doi.org/10.4236/oalib.1110960</p></sec></body><back><ref-list><title>References</title><ref id="scirp.130125-ref1"><label>1</label><mixed-citation publication-type="other" xlink:type="simple">Dorfman, R. (1943) The Detection of Defective Members of Large Populations. The Annals of Mathematical Statistics, 14, 436-440.  
https://doi.org/10.1214/aoms/1177731363</mixed-citation></ref><ref id="scirp.130125-ref2"><label>2</label><mixed-citation publication-type="other" xlink:type="simple">Sobel, M. and Elashoff, R.M. (1975) Group Testing with a New Goal, Estimation. Biometrika, 62, 181-193. https://doi.org/10.1093/biomet/62.1.181</mixed-citation></ref><ref id="scirp.130125-ref3"><label>3</label><mixed-citation publication-type="other" xlink:type="simple">Tamba, C.L. (2012) Computational Statistical Model for Group Testing with Retesting. Ph.D. Thesis, Egerton University, Kenya.</mixed-citation></ref><ref id="scirp.130125-ref4"><label>4</label><mixed-citation publication-type="other" xlink:type="simple">Wanyonyi, R.W. (2015) Estimation of Proportion of a Trait by Batch Testing Model in a Quality Control Process. American Journal of Theoretical and Applied Statistics, 4, 610-613. https://doi.org/10.11648/j.ajtas.20150406.34</mixed-citation></ref><ref id="scirp.130125-ref5"><label>5</label><mixed-citation publication-type="other" xlink:type="simple">Bilder, C.R., Tebbs, J.M. and Chen, P. (2010) Informative Retesting. Journal of the American Statistical Association, 105, 942-955.  
https://doi.org/10.1198/jasa.2010.ap09231</mixed-citation></ref><ref id="scirp.130125-ref6"><label>6</label><mixed-citation publication-type="other" xlink:type="simple">Thompson, K.H. (1962) Estimation of the Proportion of Vectors in a Natural Population of Insects. Biometrics, 18, 568-576. https://doi.org/10.2307/2527902</mixed-citation></ref><ref id="scirp.130125-ref7"><label>7</label><mixed-citation publication-type="other" xlink:type="simple">Nyongesa, L.K. (2004) Testing for the Presence of Disease by Pooling Samples. New Zealand Journal of Statistics, 46, 383-390.  
https://doi.org/10.1111/j.1467-842X.2004.00337.x</mixed-citation></ref><ref id="scirp.130125-ref8"><label>8</label><mixed-citation publication-type="other" xlink:type="simple">Chiang, C.L. and Reeves, W.C. (1962) Statistical Estimation of Virus Infection Rates in Mosquito Vector Populations. American Journal of Epidemiology, 75, 377-391. https://doi.org/10.1093/oxfordjournals.aje.a120259</mixed-citation></ref><ref id="scirp.130125-ref9"><label>9</label><mixed-citation publication-type="other" xlink:type="simple">Hughes-Oliver, J.M. and Swallow, W.H. (1994) A Two-Stage Adaptive Group-Testing Procedure for Estimating Small Proportions. Journal of the American Statistical Association, 89, 982-993. https://doi.org/10.1080/01621459.1994.10476832</mixed-citation></ref><ref id="scirp.130125-ref10"><label>10</label><mixed-citation publication-type="other" xlink:type="simple">Katholi, C.R. and Unnasch, T.R. (2006) Important Experimental Parameters for Determining Infection Rates in Arthropod Vectors Using Pool Screening Approaches. The American Journal of Tropical Medicine and Hygiene, 74, 779-785.  
https://doi.org/10.4269/ajtmh.2006.74.779</mixed-citation></ref><ref id="scirp.130125-ref11"><label>11</label><mixed-citation publication-type="other" xlink:type="simple">Pritchard, N.A. and Tebbs, J.M. (2010) Estimating Disease Prevalence Using Inverse Binomial Pooled Testing. Journal of Agricultural, Biological, and Environmental Statistics, 16, 70-87. https://doi.org/10.1007/s13253-010-0036-4</mixed-citation></ref><ref id="scirp.130125-ref12"><label>12</label><mixed-citation publication-type="other" xlink:type="simple">Pritchard, N.A. and Tebbs, J.M. (2011) Bayesian Inference for Disease Prevalence Using Negative Binomial Group Testing. Biometrical Journal, 53, 40-56.  
https://doi.org/10.1002/bimj.201000148</mixed-citation></ref><ref id="scirp.130125-ref13"><label>13</label><mixed-citation publication-type="other" xlink:type="simple">Okoth, A.W., Nyongesa, L.K. and Kwatch, B.O. (2017) Multi-Stage Adaptive Pool Testing Model with Test Errors; Improved Efficiency. IOSR Journal of Mathematics, 13, 43-55. https://doi.org/10.9790/5728-1301024355</mixed-citation></ref><ref id="scirp.130125-ref14"><label>14</label><mixed-citation publication-type="other" xlink:type="simple">Turechek, W.W. and Madden, L.V. (2003) A Generalized Linear Modeling Approach for Characterizing Disease Incidence in a Spatial Hierarchy. Phytopathology, 93, 458-466. https://doi.org/10.1094/PHYTO.2003.93.4.458</mixed-citation></ref><ref id="scirp.130125-ref15"><label>15</label><mixed-citation publication-type="other" xlink:type="simple">Xiong, W. (2015) The Optimal Group Size Using Inverse Binomial Group Testing Considering Misclassification. Communications in Statistics—Theory and Methods, 45, 4600-4610. https://doi.org/10.1080/03610926.2014.923461</mixed-citation></ref><ref id="scirp.130125-ref16"><label>16</label><mixed-citation publication-type="other" xlink:type="simple">Swallow, W.H. (1985) Group Testing for Estimating Infection Rates and Probabilities of Disease Transmission. Phytopathology, 75, 882.  
https://doi.org/10.1094/Phyto-75-882</mixed-citation></ref></ref-list></back></article>