<?xml version="1.0" encoding="UTF-8"?><!DOCTYPE article  PUBLIC "-//NLM//DTD Journal Publishing DTD v3.0 20080202//EN" "http://dtd.nlm.nih.gov/publishing/3.0/journalpublishing3.dtd"><article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" dtd-version="3.0" xml:lang="en" article-type="research article"><front><journal-meta><journal-id journal-id-type="publisher-id">AM</journal-id><journal-title-group><journal-title>Applied Mathematics</journal-title></journal-title-group><issn pub-type="epub">2152-7385</issn><publisher><publisher-name>Scientific Research Publishing</publisher-name></publisher></journal-meta><article-meta><article-id pub-id-type="doi">10.4236/am.2024.153012</article-id><article-id pub-id-type="publisher-id">AM-132175</article-id><article-categories><subj-group subj-group-type="heading"><subject>Articles</subject></subj-group><subj-group subj-group-type="Discipline-v2"><subject>Physics&amp;Mathematics</subject></subj-group></article-categories><title-group><article-title>
 
 
  A Bayesian Mixture Model Approach to Disparity Testing
 
</article-title></title-group><contrib-group><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Gary</surname><given-names>C. McDonald</given-names></name><xref ref-type="aff" rid="aff1"><sub>1</sub></xref><xref ref-type="corresp" rid="cor1"><sup>*</sup></xref></contrib></contrib-group><aff id="aff1"><label>1</label><addr-line>Department of Mathematics and Statistics, Oakland University, Rochester, USA</addr-line></aff><pub-date pub-type="epub"><day>29</day><month>03</month><year>2024</year></pub-date><volume>15</volume><issue>03</issue><fpage>214</fpage><lpage>234</lpage><history><date date-type="received"><day>21,</day>	<month>February</month>	<year>2024</year></date><date date-type="rev-recd"><day>26,</day>	<month>March</month>	<year>2024</year>	</date><date date-type="accepted"><day>29,</day>	<month>March</month>	<year>2024</year></date></history><permissions><copyright-statement>&#169; Copyright  2014 by authors and Scientific Research Publishing Inc. </copyright-statement><copyright-year>2014</copyright-year><license><license-p>This work is licensed under the Creative Commons Attribution International License (CC BY). http://creativecommons.org/licenses/by/4.0/</license-p></license></permissions><abstract><p>
 
 
  The topic of this article is one-sided hypothesis testing for disparity, 
  <em>i.e.</em>, the mean of one group is larger than that of another when there is uncertainty as to which group a datum is drawn. For each datum, the uncertainty is captured with a given discrete probability distribution over the groups. Such situations arise, for example, in the use of Bayesian imputation methods to assess race and ethnicity disparities with certain insurance, health, and financial data. A widely used method to implement this assessment is the Bayesian Improved Surname Geocoding (BISG) method which assigns a discrete probability over six race/ethnicity groups to an individual given the individual’s surname and address location. Using a Bayesian framework and Markov Chain Monte Carlo sampling from the joint posterior distribution of the group means, the probability of a disparity hypothesis is estimated. Four methods are developed and compared with an illustrative data set. Three of these methods are implemented in an R-code and one method in WinBUGS. These methods are programed for any number of groups between two and six inclusive. All the codes are provided in the appendices.
 
</p></abstract><kwd-group><kwd>Bayesian Improved Surname and Geocoding (BISG)</kwd><kwd> Mixture Likelihood Function</kwd><kwd> Posterior Distribution</kwd><kwd> Metropolis-Hastings Algorithms</kwd><kwd>  Random Walk Chain</kwd><kwd> Independence Chain</kwd><kwd> Gibbs Sampling</kwd><kwd> WinBUGS</kwd></kwd-group></article-meta></front><body><sec id="s1"><title>1. Introduction</title><p>This article deals with the issue of disparity testing, i.e., testing the hypothesis that the mean of one group is larger than that of a second group when the groups from which samples are drawn are uncertain and given with a probability distribution. McDonald [<xref ref-type="bibr" rid="scirp.132175-ref1">1</xref>] presented relevant approaches and calculations for such testing, for both frequentist and Bayesian formulations, with two groups. In both formulations, use is made of all possible data configurations along with their corresponding probabilities for small sample sizes (≤6). Elkadry and McDonald [<xref ref-type="bibr" rid="scirp.132175-ref2">2</xref>] provide extensive discussion of the background giving rise to such disparity testing problems, and give an R-code to analyze sample sizes up to 22 using all possible data configurations. McDonald and Oakley [<xref ref-type="bibr" rid="scirp.132175-ref3">3</xref>] use a Bayesian framework and Markov Chain Monte Carlo sampling from the joint posterior distribution of the group means to estimate the probability of a disparity hypothesis. They provide the applicable R-codes and a WinBUGS code, and greatly extend sample size limitations of previous methods given in the literature. McDonald and Willard [<xref ref-type="bibr" rid="scirp.132175-ref4">4</xref>] , using a frequentist approach, employ a bootstrap methodology to generate summary statistics for the population of p-values arising from all possible configurations of the data to assess the credibility of the disparity hypothesis. Using their provided R-code, previous limitations on sample sizes are substantially eliminated. The p-value calculations for statistical hypotheses testing are presented in numerous texts (e.g., see Navidi [<xref ref-type="bibr" rid="scirp.132175-ref5">5</xref>] ).</p><p>The primary motivation for this work is the use of imputation of race/ethnicity of an individual based on available information such as surname and address. For example, in testing for disparity among race/ethnicity loan applicants where such information is not directly available, imputation methods are being used to assign one of six race/ethnicity groups to an application. There has been considerable research in developing such imputation methodology following the work of Fiscella and Fremont [<xref ref-type="bibr" rid="scirp.132175-ref6">6</xref>] . See, for example, the articles by Adjaye-Gbewonyo, et al. [<xref ref-type="bibr" rid="scirp.132175-ref7">7</xref>] , Brown, et al. [<xref ref-type="bibr" rid="scirp.132175-ref8">8</xref>] , Consumer Financial Protection Bureau [<xref ref-type="bibr" rid="scirp.132175-ref9">9</xref>] [<xref ref-type="bibr" rid="scirp.132175-ref10">10</xref>] , Elliott, et al. [<xref ref-type="bibr" rid="scirp.132175-ref11">11</xref>] , McDonald and Rojc [<xref ref-type="bibr" rid="scirp.132175-ref12">12</xref>] [<xref ref-type="bibr" rid="scirp.132175-ref13">13</xref>] [<xref ref-type="bibr" rid="scirp.132175-ref14">14</xref>] , Zhang [<xref ref-type="bibr" rid="scirp.132175-ref15">15</xref>] , and Zavez, et al. [<xref ref-type="bibr" rid="scirp.132175-ref16">16</xref>] for a wide variety of applications of such imputation methods in healthcare and finance using geocoding and surname analysis. Voicu, et al. [<xref ref-type="bibr" rid="scirp.132175-ref17">17</xref>] include the first name information, along with surname and geocoding, to improve the accuracy of race/ethnicity classification. To assess disparity of, say, auto loan rates extended to two race/ethnicity groups, a one-sided statistical hypothesis test could be used to assess the plausibility of the mean rate of one group being larger than that of another. This article addresses the issue of such statistical testing with imputed data as noted above.</p><p>This article expands the models used in references given above from two groups based on a binomial model to one using c groups (2 ≤ c ≤ 6) based on a mixture model within a Bayesian context. The methodologies developed in McDonald and Oakley [<xref ref-type="bibr" rid="scirp.132175-ref3">3</xref>] are applied with the more encompassing mixture likelihood function. Thus, this work is directly applicable to the output of proxy methods such as that from the Bayesian Improved Surname Geocoding (BISG) (see, e.g., Elliott, et al. [<xref ref-type="bibr" rid="scirp.132175-ref11">11</xref>] ) used extensively by the Consumer Financial Protection Bureau (CFPB) and others. Using BISG, the CFPB [<xref ref-type="bibr" rid="scirp.132175-ref9">9</xref>] ordered Ally Financial Inc. and Ally Bank to pay $80 million in damages to African-American, Hispanic, and Asian and Pacific Islander consumers harmed by Ally’s alleged discriminatory auto loan pricing, and $18 million in civil money penalties. The Wall Street Journal provides a website (published in 2015), http://graphics.wsj.com/ally-settlement-race-calculator/, implementing BISG for six race/ethnicity groups.</p><p>A Bayesian approach to the statistical hypothesis testing problem is utilized here. With this approach, the group means of the variable of interest (e.g., auto loan interest rates) are modeled with a probability distribution from which the probability of the disparity hypothesis can be calculated. Prior knowledge on the group means is combined with the likelihood function of the sample data to yield a so-called posterior probability distribution of the group means. Several approaches to these calculations are herein described and illustrated using computational codes given in the appendices. The illustrative calculations utilize one set of assumptions on the model and data (“noninformative” prior knowledge on the group means, and normal likelihood function for the sample data). However, the approach and computational tools given in the appendices are easily adapted to other assumptions. Hence the robustness of conclusions can be assessed with several other specifications of model assumptions. Some robustness considerations are explicitly considered in Sections 4 and 5.</p><p>Section 2 of this article describes the data set to be used subsequently to illustrate the computations of the various methodologies herein presented. It also applies a method labeled Laplace (Albert [<xref ref-type="bibr" rid="scirp.132175-ref18">18</xref>] ) to develop a multivariate normal distribution approximation to the posterior distribution of the group means. Section 3 details the use of “BayesMix”, an R-code, integrating several Metropolis-Hastings algorithms so as to generate draws from the posterior distribution of the group means. Using these draws, the mean values of the group means are calculated to assess the probabilities of linear contrasts of these mean values. The probability of a disparity hypothesis is calculated. Section 4 uses the publicly available software package WinBUGS to calculate quantities similar to those of Section 3. Gill [<xref ref-type="bibr" rid="scirp.132175-ref19">19</xref>] and Christensen, et al. [<xref ref-type="bibr" rid="scirp.132175-ref20">20</xref>] provide excellent introductions to Bayesian modeling emphasizing the use of both R and WinBUGS to analyze real data. Section 5 provides a summary and the concluding remarks.</p></sec><sec id="s2"><title>2. Normal Approximation to the Posterior Distribution</title><sec id="s2_1"><title>2.1. The Data</title><p>To illustrate the following methodologies a simulated data set consisting of n = 18 observations, given in Appendix A, will be used. The data can be entered into the programs given in the Appendices B-D in several different manners, e.g., excel spreadsheet, list format, etc. The excel spreadsheet format will be used here with the R-code BayesMix given in Appendix B. The excel spreadsheet with the example data for this article is labeled “TestData.xlsx”. Two R packages are required for using this data format with the BayesMix: “readxl” and “LearnBayes”. The data of Appendix A were randomly drawn from a normal distribution in c = 6 groups of three—with means of 2, 4, 6, 8, 10, and 12, and with standard deviations of 1. The probabilities for the six categories of each datum are given as q = (q<sub>1</sub>, &#183;&#183;&#183;, q<sub>6</sub>). For each datum the category from which it was drawn is assigned the max probability. The objective of the analysis is to assess stochastic properties of the group means based on the data and group probabilities. These analyses will utilize algorithms given in Albert [<xref ref-type="bibr" rid="scirp.132175-ref18">18</xref>] and utilized in McDonald and Oakley [<xref ref-type="bibr" rid="scirp.132175-ref3">3</xref>] for c = 2. The user input required to execute BayesMix is given in <xref ref-type="table" rid="table1">Table 1</xref>. The specific values of the variables given in Appendix B used in this article are noted. The vector A, as given here, is used to formulate the disparity hypothesis “mu[<xref ref-type="bibr" rid="scirp.132175-ref2">2</xref>] is greater than mu[<xref ref-type="bibr" rid="scirp.132175-ref1">1</xref>]” for which its probability is estimated. Similarly, the contrast vector A1 is used for the hypothesis “mu[<xref ref-type="bibr" rid="scirp.132175-ref3">3</xref>] is greater than the average of mu[<xref ref-type="bibr" rid="scirp.132175-ref4">4</xref>] and mu[<xref ref-type="bibr" rid="scirp.132175-ref5">5</xref>]”. The quantities mu[i] refer to the group[i] means.</p><p>The output of BayesMix contains many data summaries and diagnostics useful in checking the analyses structures and calculations. The user can suppress one or more of these if so wished. However, the primary output is <xref ref-type="table" rid="table2">Table 2</xref> (less the column WinBUGS).</p></sec><sec id="s2_2"><title>2.2. Laplace</title><p>As described in the above references, the R-code Laplace computes the posterior mode of the vector mu = (mu[<xref ref-type="bibr" rid="scirp.132175-ref1">1</xref>], &#183;&#183;&#183;, mu[c]), the associated variance-covariance matrix, and an estimate of the logarithm of the normalizing constant for a general posterior density, and an indication of algorithm convergence (true or false). Three inputs are required: logpost, a function that defines the logarithm of the posterior density (denoted by loglike in Appendix B); mode, an initial guess at the posterior mode; par, a list of parameters associated with the function logpost.</p><p>The probability (posterior) density for the data is a mixture distribution given by</p><p>f ( x / q , μ , σ ) = ∑ i = 1 c q i ⋅ dnorm ( x , μ i , σ ) (1)</p><p>where 0 ≤ q<sub>i</sub> ≤ 1, ∑q<sub>i</sub> = 1, |x| &lt; ∞, 0 &lt; σ &lt; ∞, and dnorm ( x , μ i , σ ) is the normal probability density with mean and sigma equal to ( μ i , σ ) evaluated at datum x.</p><table-wrap id="table1" ><label><xref ref-type="table" rid="table1">Table 1</xref></label><caption><title> User required input for BayesMix (Appendix B)</title></caption><table><tbody><thead><tr><th align="center" valign="middle" >Variable</th><th align="center" valign="middle" >Description</th><th align="center" valign="middle" >Values used in this article</th></tr></thead><tr><td align="center" valign="middle" >sig</td><td align="center" valign="middle" >Std dev of mixture normal densities</td><td align="center" valign="middle" >1</td></tr><tr><td align="center" valign="middle" >mcn</td><td align="center" valign="middle" >Sample size of MCMC chains</td><td align="center" valign="middle" >52,000</td></tr><tr><td align="center" valign="middle" >sca</td><td align="center" valign="middle" >Scale parameter for MCMC random walk</td><td align="center" valign="middle" >1.2</td></tr><tr><td align="center" valign="middle" >dis</td><td align="center" valign="middle" >Burn-in for MCMC draws</td><td align="center" valign="middle" >2000</td></tr><tr><td align="center" valign="middle" >A</td><td align="center" valign="middle" >Vector to compare two group means</td><td align="center" valign="middle" >(−1, 1, 0, 0, 0, 0)</td></tr><tr><td align="center" valign="middle" >A1</td><td align="center" valign="middle" >Vector specifying linear transform of group means</td><td align="center" valign="middle" >(0, 0, 1, −0.5, −0.5, 0)</td></tr></tbody></table></table-wrap><table-wrap id="table2" ><label><xref ref-type="table" rid="table2">Table 2</xref></label><caption><title> Output (to four dp) of BayesMix (Appendix B) and of WinBUGS (Appendix C) with Appendix A data</title></caption><table><tbody><thead><tr><th align="center" valign="middle" ></th><th align="center" valign="middle" >Lap</th><th align="center" valign="middle" >Ran</th><th align="center" valign="middle" >Ind</th><th align="center" valign="middle" >RanRed</th><th align="center" valign="middle" >IndRed</th><th align="center" valign="middle" >WinBUGS</th></tr></thead><tr><td align="center" valign="middle" >Draws</td><td align="center" valign="middle" >NA</td><td align="center" valign="middle" >52,000</td><td align="center" valign="middle" >52,000</td><td align="center" valign="middle" >50,000</td><td align="center" valign="middle" >50,000</td><td align="center" valign="middle" >50,000</td></tr><tr><td align="center" valign="middle" >est1</td><td align="center" valign="middle" >2.7351</td><td align="center" valign="middle" >2.7107</td><td align="center" valign="middle" >2.7327</td><td align="center" valign="middle" >2.7117</td><td align="center" valign="middle" >2.7320</td><td align="center" valign="middle" >2.728</td></tr><tr><td align="center" valign="middle" >sd1</td><td align="center" valign="middle" >0.5977</td><td align="center" valign="middle" >0.6289</td><td align="center" valign="middle" >0.6190</td><td align="center" valign="middle" >0.6312</td><td align="center" valign="middle" >0.6184</td><td align="center" valign="middle" >0.6154</td></tr><tr><td align="center" valign="middle" >est2</td><td align="center" valign="middle" >4.8257</td><td align="center" valign="middle" >4.8337</td><td align="center" valign="middle" >4.8220</td><td align="center" valign="middle" >4.8307</td><td align="center" valign="middle" >4.8211</td><td align="center" valign="middle" >4.833</td></tr><tr><td align="center" valign="middle" >sd2</td><td align="center" valign="middle" >0.6052</td><td align="center" valign="middle" >0.6328</td><td align="center" valign="middle" >0.6214</td><td align="center" valign="middle" >0.6320</td><td align="center" valign="middle" >0.6212</td><td align="center" valign="middle" >0.6515</td></tr><tr><td align="center" valign="middle" >est3</td><td align="center" valign="middle" >7.0174</td><td align="center" valign="middle" >7.1319</td><td align="center" valign="middle" >7.1520</td><td align="center" valign="middle" >7.1305</td><td align="center" valign="middle" >7.1504</td><td align="center" valign="middle" >7.146</td></tr><tr><td align="center" valign="middle" >sd3</td><td align="center" valign="middle" >0.7207</td><td align="center" valign="middle" >0.8304</td><td align="center" valign="middle" >0.8338</td><td align="center" valign="middle" >0.8294</td><td align="center" valign="middle" >0.8321</td><td align="center" valign="middle" >0.8727</td></tr><tr><td align="center" valign="middle" >est4</td><td align="center" valign="middle" >7.7113</td><td align="center" valign="middle" >7.6265</td><td align="center" valign="middle" >7.6050</td><td align="center" valign="middle" >7.6289</td><td align="center" valign="middle" >7.6036</td><td align="center" valign="middle" >7.605</td></tr><tr><td align="center" valign="middle" >sd4</td><td align="center" valign="middle" >0.6837</td><td align="center" valign="middle" >0.8464</td><td align="center" valign="middle" >0.7738</td><td align="center" valign="middle" >0.8466</td><td align="center" valign="middle" >0.7733</td><td align="center" valign="middle" >0.8089</td></tr><tr><td align="center" valign="middle" >est5</td><td align="center" valign="middle" >10.1977</td><td align="center" valign="middle" >10.3075</td><td align="center" valign="middle" >10.3140</td><td align="center" valign="middle" >10.3063</td><td align="center" valign="middle" >10.3149</td><td align="center" valign="middle" >10.27</td></tr><tr><td align="center" valign="middle" >sd5</td><td align="center" valign="middle" >0.7285</td><td align="center" valign="middle" >0.8609</td><td align="center" valign="middle" >0.8218</td><td align="center" valign="middle" >0.8634</td><td align="center" valign="middle" >0.8207</td><td align="center" valign="middle" >0.9251</td></tr><tr><td align="center" valign="middle" >est6</td><td align="center" valign="middle" >11.8836</td><td align="center" valign="middle" >11.8867</td><td align="center" valign="middle" >11.8753</td><td align="center" valign="middle" >11.8856</td><td align="center" valign="middle" >11.8753</td><td align="center" valign="middle" >11.87</td></tr><tr><td align="center" valign="middle" >sd6</td><td align="center" valign="middle" >0.5824</td><td align="center" valign="middle" >0.6040</td><td align="center" valign="middle" >0.5982</td><td align="center" valign="middle" >0.6016</td><td align="center" valign="middle" >0.5985</td><td align="center" valign="middle" >0.603</td></tr><tr><td align="center" valign="middle" >P(mu[<xref ref-type="bibr" rid="scirp.132175-ref2">2</xref>] &gt; mu[<xref ref-type="bibr" rid="scirp.132175-ref1">1</xref>])</td><td align="center" valign="middle" >0.9934</td><td align="center" valign="middle" >0.9904</td><td align="center" valign="middle" >0.9901</td><td align="center" valign="middle" >0.9904</td><td align="center" valign="middle" >0.9902</td><td align="center" valign="middle" >0.9905</td></tr><tr><td align="center" valign="middle" >P(Y &gt; 0)</td><td align="center" valign="middle" >0.0146</td><td align="center" valign="middle" >NA</td><td align="center" valign="middle" >NA</td><td align="center" valign="middle" >0.0491</td><td align="center" valign="middle" >0.0485</td><td align="center" valign="middle" >0.0533</td></tr><tr><td align="center" valign="middle" >Accept. Rate</td><td align="center" valign="middle" >NA</td><td align="center" valign="middle" >0.2191</td><td align="center" valign="middle" >0.8091</td><td align="center" valign="middle" >NA</td><td align="center" valign="middle" >NA</td><td align="center" valign="middle" >NA</td></tr></tbody></table></table-wrap><p>The initial guess mode, a vector of length c, is denoted by mu in BayesMix. This initial guess can affect the convergence of the Laplace algorithm indicated by “true” or “false” in the output. One method for specifying this guess is to put each of the mu values equal to the average of the data. This has worked well for many of the calculations made in preparing this manuscript, but not all. There have been instances when the output of Laplace indicated “false” for convergence. In these cases, a good strategy is to rerun the Laplace function with the initial guess restated with the mode values output for the false convergence output. Another strategy that has worked very well, and is included in BayesMix, is to form the initial guess of mu[i] by taking the average of the data for which the category[i] has the maximum value over the c groups.</p><p>Laplace estimates of the mode values of the mu variables along with their standard deviations are given in <xref ref-type="table" rid="table2">Table 2</xref> in the column denoted Lap. The estimates and standard deviations for mu[i] are denoted by esti and sdi, respectively. Since A is specified as (−1, 1, 0, 0, 0, 0), the vector (Ahi, Alo) = (2, 1) and the probability of mu[<xref ref-type="bibr" rid="scirp.132175-ref2">2</xref>] being greater than mu[<xref ref-type="bibr" rid="scirp.132175-ref1">1</xref>] is approximated with the normal distribution and the Laplace variance-covariance output as 0.9934. Also, A1 is specified by (0, 0, 1, −0.5, −0.5, 0), so Y = mu[<xref ref-type="bibr" rid="scirp.132175-ref3">3</xref>] − 0.5(mu[<xref ref-type="bibr" rid="scirp.132175-ref4">4</xref>] + mu[<xref ref-type="bibr" rid="scirp.132175-ref5">5</xref>]). Thus, Y &gt; 0 is equivalent to mu[<xref ref-type="bibr" rid="scirp.132175-ref3">3</xref>] greater than the average of mu[<xref ref-type="bibr" rid="scirp.132175-ref4">4</xref>] and mu[<xref ref-type="bibr" rid="scirp.132175-ref5">5</xref>], the estimated probability of which is given in <xref ref-type="table" rid="table2">Table 2</xref> as 0.0146. The entries in the other columns of <xref ref-type="table" rid="table2">Table 2</xref>, and the row “Accept Rate”, are explained in subsequent sections of this article.</p><p>While the Laplace algorithm with the normal approximation may not be most appropriate for the posterior distribution, the Laplace output provides very useful input to the algorithms to be utilized subsequently. The output of Laplace is explicitly used in the “proposal” and “start” inputs for the two Markov Chain Monte Carlo (MCMC) algorithms to be described in Section 3. BayesMix, which generated <xref ref-type="table" rid="table2">Table 2</xref>, ran in about 6.5 minutes on a laptop computer with Windows 10.</p></sec></sec><sec id="s3"><title>3. Exploring the Posterior Distribution with Metropolis-Hastings Algorithms</title><p>In this Section, random draws from the posterior distribution of the mean vector mu are used to estimate the probability of disparity as well as the cumulative probability distribution of any contrast of the category means. As was developed by McDonald and Oakley [<xref ref-type="bibr" rid="scirp.132175-ref3">3</xref>] , these draws are obtained using the popular MCMC methods in the Bayesian literature. As stated in Section 3 of the McDonald and Oakley reference, these methods consist of specifying an initial value for the parameter vector of interest and then generating a chain of values each of which depends only on the previous value in the chain. A rule for generating the sequence is a function of a proposal density and an acceptance probability. Not all generated values in the sequence are accepted as they must pass a probability criterion to be retained in the final sample. These components, the proposal density and acceptance probability criterion, are constructed so that the accepted simulated draws converge to draws from a random variable having the target posterior distribution. Different choices of the proposal density may result in slightly different distribution summaries.</p><p>Two such proposal density implementations described in detail by Albert [<xref ref-type="bibr" rid="scirp.132175-ref18">18</xref>] , and executed in the R package LearnBayes, are incorporated here in BayesMix given in Appendix B. The R functions “rwmetrop” and “indepmetrop” implement the so-called random walk and independence Metropolis-Hastings algorithms for special choices of the proposal densities. The methodology descriptions given in the following two subsections follow closely that of Section 3 of McDonald and Oakley [<xref ref-type="bibr" rid="scirp.132175-ref3">3</xref>] where the application is developed for c = 2 with a binomial likelihood. Here the applications are built for 2 ≤ c ≤ 6 with the likelihood function consisting of a mixture of normal densities given in Equation (1). The R-code BayesMix can be extended in a straightforward manner for c &gt; 6, i.e., for more than six groups. However, the use of WinBUGS, described in Section 4, can be easily applied to cases of c &gt; 6 by suitably stating the number of groups and appropriately entering the data in Appendix C and Appendix D.</p><p>Since the MCMC algorithms converge to sampling from the posterior distribution, initial draws may not be from the stationary distribution of the Markov chain, i.e., not drawn from the target posterior distribution. Thus, it is common practice to discard an initial portion of the draws and utilize the remainder. This practice is frequently referred to as “burn-in.” There are several diagnostics that can be viewed to help decide when the chain has progressed sufficiently far to stabilize on the posterior distribution (e.g., see Section 9.6 of Albert and Hu [<xref ref-type="bibr" rid="scirp.132175-ref21">21</xref>] ; Section 4.4.2 of Lunn et al. [<xref ref-type="bibr" rid="scirp.132175-ref22">22</xref>] ). <xref ref-type="table" rid="table2">Table 2</xref> provides results based on 52,000 random draws using MCMC algorithms described in the following two subsections. Results are also provided following a burn-in of 2000 draws, i.e., results based on the last 50,000 random draws. As noted, the burn-in deletions result in very minor changes in the posted estimates for this illustrative data set. The two chains utilized in subsections 3.1 and 3.2 are described fully in Albert [<xref ref-type="bibr" rid="scirp.132175-ref18">18</xref>] .</p><sec id="s3_1"><title>3.1. Random Walk Metropolis Chain</title><p>Within LearnBayes, the function rwmetrop (logpost, proposal, start, m, par) requires five inputs: logpost, function defining the log posterior density; proposal, a list containing var, an estimated variance-covariance matrix, and scale, the Metropolis scale factor; start, a vector giving the starting value of the parameter; m, the number of iterations of the chain; par, the data used in the function logpost. The output is par, a matrix of the simulated values where each row corresponds to a value of the vector parameter; accept, the acceptance rate of the algorithm. A summary of the outputs of rwmetrop run with the specification of the data described in Section 2.1 with 52,000 draws from the posterior distribution is given in <xref ref-type="table" rid="table2">Table 2</xref> in the column designated Ran.</p><p>As noted in the output, the acceptance rate here is 0.2191. The input “scale” should be chosen so that the acceptance rate is around 25%. Acceptance rates are discussed on page 121 of Albert [<xref ref-type="bibr" rid="scirp.132175-ref18">18</xref>] and on page 253 of Rizzo [<xref ref-type="bibr" rid="scirp.132175-ref23">23</xref>] . The acceptance rate is a decreasing function of scale.</p><p>The results following a burn-in of 2000 are given in <xref ref-type="table" rid="table2">Table 2</xref> column RanRed. The results in this column are thus based on 50,000 draws and differ very little than those in column Ran.</p></sec><sec id="s3_2"><title>3.2. Independence Metropolis Chain</title><p>The function indepmetrop (logpost, proposal, start, m, data) also requires five inputs: logpost, as above; proposal, a list containing mu, an estimated mean, and var, an estimated variance-covariance matrix for the normal proposal density; start, array with a single row that gives the starting value for the parameter vector; m, the number of iterations of the chain; data, data used in the function logpost. A summary of the outputs of indepmetrop run with the illustrative data set with 52,000 draws is given in <xref ref-type="table" rid="table2">Table 2</xref> in the column designated Ind. Column IndRed provides analogous results for the 50,000 draws following the burn-in of 2000. As is the case with rwmetrop algorithm, there is negligible difference in the results with and without burn-in deletion.</p></sec></sec><sec id="s4"><title>4. Applying WinBUGS for Bayesian Simulation</title><p>Another approach to generating draws from the posterior distribution is to use a MCMC software package, WinBUGS, designed specifically for Bayesian computation. WinBUGS is the MS Windows operating system version of BUGS, Bayesian Analysis Using Gibbs Sampling. There is a free download of this Bayesian software package available at the site (http://www.mrc-bsu.cam.ac.uk/bugs/overview/contents.shtml). Clear discussions of the applicability, limitations, and use of WinBUGS are given in Hahn [<xref ref-type="bibr" rid="scirp.132175-ref24">24</xref>] , Woodworth [<xref ref-type="bibr" rid="scirp.132175-ref25">25</xref>] , Lunn et al. [<xref ref-type="bibr" rid="scirp.132175-ref22">22</xref>] , and in many other references. The setup of this approach for the illustrative example considered here, where c = 6 and n = 18, is given in Appendix C. Note that in WinBUGS the normal distribution density is specified by dnorm(&#181;, τ), where &#181; is the mean and τ is the precision (= 1/variance). This differs from the specification in R, as given in Equation (1), where the normal distribution density at x is denoted by dnorm(x, &#181;, σ) with &#181; as the mean and σ as the standard deviation.</p><p>Output from WinBugs is given in <xref ref-type="table" rid="table2">Table 2</xref> based on a sample of 50,000 draws from the posterior distribution following a burn-in of an initial 2000 draws (similar to that done for the RanRed and IndRed entries). A “count1” node is included to estimate P(mu[<xref ref-type="bibr" rid="scirp.132175-ref2">2</xref>] &gt; mu[<xref ref-type="bibr" rid="scirp.132175-ref1">1</xref>]) as shown in <xref ref-type="table" rid="table2">Table 2</xref>. A similar “count2” node is included to estimate P(Y &gt; 0) for a user specified transformation of the group means, A1 = (0, 0, 1, −0.5, −0.5, 0), in this example. To execute WinBugs, initial values are required for mu, group, and sigma. For mu, a reasonable choice for mu inits are the mode values given by Laplace. For group, a reasonable choice is to assign that group which has the largest probability for the datum. In case several groups share the largest value, choose one of those at random. For sigma, simply use a reasonable guess (e.g., use sigma = 1 as done here). The robustness of the results can easily be checked by running WinBUGS with other choices. The run time for WinBUGS as given in Appendix C is about forty seconds.</p><p>Appendix D provides a modification of the Appendix C WinBUGS program by treating sigma as an unknown value, common to all six groups, and to be estimated along with the other nodes. With this program, the mean of the 50,000 sigma draws is 0.9886 with a standard deviation of 0.3024. The output for the other nodes is very close to those given in <xref ref-type="table" rid="table2">Table 2</xref> for WinBUGS with sigma fixed at one (i.e., the output using the Appendix C program). The P(m[<xref ref-type="bibr" rid="scirp.132175-ref2">2</xref>] &gt; mu[<xref ref-type="bibr" rid="scirp.132175-ref1">1</xref>]) is estimated to be 0.9811, and P(Y &gt; 0) to be 0.0689. A further extension of Appendix D can be made easily to accommodate the case where the group sigma values are not assumed known or to be equal.</p></sec><sec id="s5"><title>5. Summary and Concluding Remarks</title><p><xref ref-type="table" rid="table2">Table 2</xref> shows a great deal of consistency with the entries, especially among the last five columns of the table. The results for the Laplace column also are in close agreement with those of the other columns, with perhaps the row entries corresponding to the standard deviations (i.e., sdi’s) where the Laplace values are uniformly a bit smaller than the others. The last five columns are based on statistics calculated from the random draws from the posterior distribution, whereas the Laplace column is based on a numerical fit to the data. These observations are based on the one data set given in Appendix A and may not generalize to other data sets. All the results given in this article apply to a mixture likelihood function of normal densities. Other densities could be substituted with appropriate changes in Equation (1) and appropriate modifications to the computer codes given in the Appendices.</p><p>An important question to be addressed with the methodologies used in this article is simply one of sample size. What is the tradeoff between the sample size (n) and the computing time? To partially address this issue, a sample of size n = 72 was constructed by pooling together four copies of the data given in Appendix A. The R-code in Appendix B and the WinBUGS code in Appendix C were run with exactly the same inputs used in Sections 3 and 4, with the expanded data set, generating an analogous <xref ref-type="table" rid="table2">Table 2</xref>. The laptop computer time for BayesMix was approximately forty-five minutes and for WinBUGS approximately forty seconds. The corresponding computer times with n = 18 (Appendix A) were approximately nine minutes for BayesMix and forty seconds for WinBUGS. There was no meaningful difference in computing times for WinBUGS between the two data sets. The output for the larger sample size closely matched that for the smaller sample size with the exception of the standard deviations (sdi’s) whose values with the larger data set were approximately half of the values with those obtained with the smaller set. For BayesMix, the computer time was approximately five times longer for n = 72 vs. n = 18. For much larger sample sizes, the bootstrap approach developed by McDonald and Willard [<xref ref-type="bibr" rid="scirp.132175-ref4">4</xref>] could be adapted to limit the required computing time.</p><p>As with all analyses, especially Bayesian, the robustness of the conclusions with respect to choices of model and prior specifications should be explored. With WinBUGS, a feature called “chains” facilitates running multiple MCMCs with different prior initializations given in the inits list. The priors used in this article were chosen to be so-called “noninformative” (e.g., distributions with very large variances). Other graphical diagnostics provided by WinBUGS are given in Appendix E for mu[<xref ref-type="bibr" rid="scirp.132175-ref1">1</xref>] and are very helpful in addressing the issue of stability of MCMC draws from a stationary posterior distribution. These diagnostics, along with similar ones for the other nodes, support the assessment that the MCMC draws are from a stable stationary posterior distribution. Chapter 6 of Hahn [<xref ref-type="bibr" rid="scirp.132175-ref24">24</xref>] and Chapter 14 of Gill [<xref ref-type="bibr" rid="scirp.132175-ref19">19</xref>] provide extensive discussion of assessing MCMC performance in WinBUGS along with some interfaces to R packages.</p><p>In conclusion, the Bayesian methods herein presented with appropriate computer codes provide a statistically sound basis for drawing conclusions about group disparity using race and ethnicity uncertainty data such as that arising with BISG and similarly based methodologies.</p></sec><sec id="s6"><title>Conflicts of Interest</title><p>The author declares no conflicts of interest regarding the publication of this paper.</p></sec><sec id="s7"><title>Cite this paper</title><p>McDonald, G.C. (2024) A Bayesian Mixture Model Approach to Disparity Testing. Applied Mathematics, 15, 214-234. https://doi.org/10.4236/am.2024.153012</p></sec><sec id="s8"><title>Appendix A. Excel Spreadsheet Data, TestData.xlsx</title></sec><sec id="s9"><title>Appendix B. BayesMix Implementing Laplace, Metropolis-Hastings Algorithms</title><disp-formula id="scirp.132175-formula20"><graphic  xlink:href="//html.scirp.org/file/2-7405230x5.png?20240328165809735"  xlink:type="simple"/></disp-formula><disp-formula id="scirp.132175-formula21"><graphic  xlink:href="//html.scirp.org/file/2-7405230x6.png?20240328165809735"  xlink:type="simple"/></disp-formula><disp-formula id="scirp.132175-formula22"><graphic  xlink:href="//html.scirp.org/file/2-7405230x7.png?20240328165809735"  xlink:type="simple"/></disp-formula><disp-formula id="scirp.132175-formula23"><graphic  xlink:href="//html.scirp.org/file/2-7405230x8.png?20240328165809735"  xlink:type="simple"/></disp-formula><disp-formula id="scirp.132175-formula24"><graphic  xlink:href="//html.scirp.org/file/2-7405230x9.png?20240328165809735"  xlink:type="simple"/></disp-formula><disp-formula id="scirp.132175-formula25"><graphic  xlink:href="//html.scirp.org/file/2-7405230x10.png?20240328165809735"  xlink:type="simple"/></disp-formula><disp-formula id="scirp.132175-formula26"><graphic  xlink:href="//html.scirp.org/file/2-7405230x11.png?20240328165809735"  xlink:type="simple"/></disp-formula></sec><sec id="s10"><title>Appendix C. WinBUGS, Assuming σ = 1 for the c = 6 Groups</title><disp-formula id="scirp.132175-formula27"><graphic  xlink:href="//html.scirp.org/file/2-7405230x12.png?20240328165809735"  xlink:type="simple"/></disp-formula><disp-formula id="scirp.132175-formula28"><graphic  xlink:href="//html.scirp.org/file/2-7405230x13.png?20240328165809735"  xlink:type="simple"/></disp-formula></sec><sec id="s11"><title>Appendix D. WinBUGS, Assuming a Common Unknown σ for the c = 6 Groups</title><disp-formula id="scirp.132175-formula29"><graphic  xlink:href="//html.scirp.org/file/2-7405230x14.png?20240328165809735"  xlink:type="simple"/></disp-formula><disp-formula id="scirp.132175-formula30"><graphic  xlink:href="//html.scirp.org/file/2-7405230x15.png?20240328165809735"  xlink:type="simple"/></disp-formula></sec><sec id="s12"><title>Appendix E. WinBUGS Node Statistics Output and Graphical Diagnostics for mu[<xref ref-type="bibr" rid="scirp.132175-ref1">1</xref>] for Appendix 3 Run-From Top Left to Lower Right: Dynamic Trace, Running Quantiles, Time Series History, Kernel Density, and Autocorrelation Function</title></sec></body><back><ref-list><title>References</title><ref id="scirp.132175-ref1"><label>1</label><mixed-citation publication-type="other" xlink:type="simple">McDonald, G.C. (2018) Statistical Testing When the Populations from Which Samples Are Drawn Are Uncertain. Health Services and Outcomes Research Methodology, 18, 155-174. https://doi.org/10.1007/s10742-018-0182-7</mixed-citation></ref><ref id="scirp.132175-ref2"><label>2</label><mixed-citation publication-type="other" xlink:type="simple">Elkadry, A. and McDonald, G.C. (2020) Hypothesis Testing When Data Sources Are Uncertain. Journal of Statistical Theory and Practice, 14, Article No. 66. https://doi.org/10.1007/s42519-020-00132-5</mixed-citation></ref><ref id="scirp.132175-ref3"><label>3</label><mixed-citation publication-type="other" xlink:type="simple">McDonald, G.C. and Oakley, R.H. (2023) Extending Computations for Disparity Testing When Data Sources Are Uncertain. Health Services and Outcomes Research Methodology, 23, 207-226. https://doi.org/10.1007/s10742-022-00286-8</mixed-citation></ref><ref id="scirp.132175-ref4"><label>4</label><mixed-citation publication-type="other" xlink:type="simple">McDonald, G.C. and Willard, J.F. (2023) Bootstrap Approach to Disparity Testing with Source Uncertainty in the Data. Health Services and Outcomes Research Methodology. https://doi.org/10.1007/s10742-023-00318-x</mixed-citation></ref><ref id="scirp.132175-ref5"><label>5</label><mixed-citation publication-type="other" xlink:type="simple">Navidi, W. (2024) Statistics for Engineers and Scientists. 6th Edition, McGraw Hill, New York.</mixed-citation></ref><ref id="scirp.132175-ref6"><label>6</label><mixed-citation publication-type="other" xlink:type="simple">Fiscella, K. and Fremont, A.M. (2006) Use of Geocoding and Surname Analysis to Estimate Race and Ethnicity. Health Services Research, 41, 1482-1500. https://doi.org/10.1111/j.1475-6773.2006.00551.x</mixed-citation></ref><ref id="scirp.132175-ref7"><label>7</label><mixed-citation publication-type="other" xlink:type="simple">Adjaye-Gbewonyo, D., Bednarczyk, R.A., Davis, R.L. and Omer, S.B. (2014) Using the Bayesian Improved Surname Geocoding Method (BISG) to Create a Working Classification of Race and Ethnicity in a Diverse Managed Care Population: A Validation Study. Health Services Research, 49, 268-283. https://doi.org/10.1111/1475-6773.12089</mixed-citation></ref><ref id="scirp.132175-ref8"><label>8</label><mixed-citation publication-type="other" xlink:type="simple">Brown, D.P., Knapp, C., Baker, K. and Kaufmann, M. (2016) Using Bayesian Imputation to Assess Racial and Ethnic Disparities in Pediatric Performance Measures. Health Services Research, 51, 1095-1108. https://doi.org/10.1111/1475-6773.12405</mixed-citation></ref><ref id="scirp.132175-ref9"><label>9</label><mixed-citation publication-type="other" xlink:type="simple">Consumer Financial Protection Bureau (2013, December 20) CFPB and DOJ Order Ally to Pay $80 Million to Consumers Harmed by Discriminatory Auto Loan Pricing. https://www.consumerfinance.gov/enforcement/actions/ally-financial-ally-bank/</mixed-citation></ref><ref id="scirp.132175-ref10"><label>10</label><mixed-citation publication-type="other" xlink:type="simple">Consumer Financial Protection Bureau (2014) Using Publicly Available Information to Proxy for Unidentified Race and Ethnicity. https://www.consumerfinance.gov/data-research/research-reports/using-publicly-available-information-to-proxy-for-unidentified-race-and-ethnicity/</mixed-citation></ref><ref id="scirp.132175-ref11"><label>11</label><mixed-citation publication-type="other" xlink:type="simple">Elliott, M.N., Morrison, P.A., Fremont, A., McCaffrey, D.F., Pantoja, P. and Lurie, N. (2009) Using the Census Bureau’s Surname List to Improve Estimates of Race/Ethnicity and Associated Disparities. Health Services and Outcomes Research Methodology, 9, 69-83. https://doi.org/10.1007/s10742-009-0047-1</mixed-citation></ref><ref id="scirp.132175-ref12"><label>12</label><mixed-citation publication-type="other" xlink:type="simple">McDonald, K.M. and Rojc, K.J. (2015) Automotive Finance Regulation: Warning Lights Flashing. The Business Lawyer, 70, 617-624.</mixed-citation></ref><ref id="scirp.132175-ref13"><label>13</label><mixed-citation publication-type="other" xlink:type="simple">McDonald, K.M. and Rojc, K.J. (2017) Accelerating Regulation of Automotive Finance. The Business Lawyer, 72, 559-566.</mixed-citation></ref><ref id="scirp.132175-ref14"><label>14</label><mixed-citation publication-type="other" xlink:type="simple">McDonald, K.M. and Rojc, K.J. (2022) Ladies and Gentlemen: Rev Your Regulatory Engines! The Business Lawyer, 77, 581-590.</mixed-citation></ref><ref id="scirp.132175-ref15"><label>15</label><mixed-citation publication-type="other" xlink:type="simple">Zhang, Y. (2018) Assessing Fair Lending Risks Using Race/Ethnicity Proxies. Management Science, 64, 178-197. https://doi.org/10.1287/mnsc.2016.2579</mixed-citation></ref><ref id="scirp.132175-ref16"><label>16</label><mixed-citation publication-type="other" xlink:type="simple">Zavez, K., Harel, O. and Aseltine, R.H. (2022) Imputing Race and Ethnicity in Healthcare Claims Databases. Health Services and Outcomes Research Methodology, 22, 493-507. https://doi.org/10.1007/s10742-022-00273-z</mixed-citation></ref><ref id="scirp.132175-ref17"><label>17</label><mixed-citation publication-type="other" xlink:type="simple">Voicu, I. (2018) Using First Name Information to Improve Race and Ethnicity Classification. Statistics and Public Policy, 5, 1-13. https://doi.org/10.1080/2330443X.2018.1427012</mixed-citation></ref><ref id="scirp.132175-ref18"><label>18</label><mixed-citation publication-type="other" xlink:type="simple">Albert, J. (2009) Bayesian Computation with R. 2nd Edition, Springer, New York. https://doi.org/10.1007/978-0-387-92298-0</mixed-citation></ref><ref id="scirp.132175-ref19"><label>19</label><mixed-citation publication-type="other" xlink:type="simple">Gill, J. (2015) Bayesian Methods: A Social and Behavioral Sciences Approach. 3rd Edition, CRC Press, Boca Raton.</mixed-citation></ref><ref id="scirp.132175-ref20"><label>20</label><mixed-citation publication-type="other" xlink:type="simple">Christensen, R., Johnson, W., Branscum, A. and Hanson, T.E. (2011) Bayesian Ideas and Data Analysis. CRC Press, Boca Raton. https://doi.org/10.1201/9781439894798</mixed-citation></ref><ref id="scirp.132175-ref21"><label>21</label><mixed-citation publication-type="other" xlink:type="simple">Albert, J. and Hu, J. (2020) Probability and Bayesian Modeling. CRC Press, Boca Raton. https://doi.org/10.1201/9781351030144</mixed-citation></ref><ref id="scirp.132175-ref22"><label>22</label><mixed-citation publication-type="other" xlink:type="simple">Lunn, D., Jackson, C., Best, N., Thomas, A. and Spiegelhalter, D. (2013) The BUGS Book: A Practical Introduction to Bayesian Analysis. CRC Press, Boca Raton. https://doi.org/10.1201/b13613</mixed-citation></ref><ref id="scirp.132175-ref23"><label>23</label><mixed-citation publication-type="other" xlink:type="simple">Rizzo, M.L. (2008) Statistical Computing with R. Chapman &amp; Hall/CRC, Boca Raton.</mixed-citation></ref><ref id="scirp.132175-ref24"><label>24</label><mixed-citation publication-type="other" xlink:type="simple">Hahn, E.D. (2014) Bayesian Methods for Management and Business: Pragmatic Solutions for Real Problems. Wiley, Hoboken.</mixed-citation></ref><ref id="scirp.132175-ref25"><label>25</label><mixed-citation publication-type="other" xlink:type="simple">Woodworth, G.G. (2004): Biostatistics: A Bayesian Introduction. Wiley, Hoboken.</mixed-citation></ref></ref-list></back></article>