A Note Comparing Two Subset Selection Procedures for the Threshold Parameters of Two Exponential Populations ()
1. Introduction
McDonald and Hodaj [1] investigated two subset selection rules for
exponential populations that possibly differ in their threshold parameters. The scale parameters of the populations were assumed to be known and equal. Without loss of generality, the scale parameters can be taken to be one as dividing the data by the known common scale parameter leads to a data set possessing a unity scale parameter for all populations. Let
be the
independent sample value drawn from the
population, denoted by
,
,
. Let
denote the threshold parameter of the
population. The cumulative distribution function (cdf) and the probability density function (pdf) for the sample value
is given, respectively, by
(1.1)
and
(1.2)
Subset selection and ranking problems have a long history in the statistical decision literature; see for example Bechhofer [2], Gupta [3], Gupta and Panchapakesan [4], and Gibbons et al. [5], among others. These works develop general multiple decision procedures for selecting the best population under various distributional and sampling assumptions. In two recent papers, McDonald and Hodaj [1] [6] specialize these ideas to exponential populations with possibly different threshold parameters. They develop the operating characteristics of several subset selection rules under both exact and approximate computation requirements. The present note focuses on the special case, where exact analytic calculations are feasible, and compares in detail two natural selection rules based on the sample minimum and the sample mean.
The goal of the subset selection procedure is to choose a subset of the k populations such that the probability of the population with the largest threshold parameter is included in the subset with a probability no less than a user specified value
(
) no matter what the underlying configuration of the threshold parameters. That is, the probability of a Correct Selection (CS) is greater than or equal to
,
(1.3)
Two subset selection procedures are considered by McDonald and Hodaj [1]. The first, denoted by
, is based on the sample minimum statistic from each population. Let
denote the minimum order statistic from the sample values
. The expected value of
is
, and the variance is
. The selection rule is given by
(1.4)
The
-value is chosen to be as small as possible while still preserving the
condition (1.3).
A second selection procedure is based on the sample means. This is motivated by the expected value of the sample mean for the
population being
, and the variance
. The selection rule is similar to that based on the sample minimum values. Let
denote the sample mean of
,
, for
. Then
(1.5)
As with
, the
-value is chosen to be as small as possible while still maintaining the inequality (1.3).
The articles by McDonald and Hodaj [1] and [6] address computational methods to determine the values of
and
to satisfy the
condition and operating characteristics (OCs) of the selection rules
and
. The OCs include the Pr(CS), the probability of an incorrect selection (Pr(ICS)) and the expected size of the selected subset (ESS) under several parametric configurations (slippage and equi-spaced) of the threshold parameters. Computational methods include simulations for arbitrary
, calculations based on the Central Limit Theorem (CLT) for
, exact calculations for
with
. Section 2 provides exact computational methods for the selection rule
for the
-value and the OCs.
2. Special Case for k = 2 Populations for Computing the OCs for R1 and R2
Let
be the minimum values of the samples from
respectively; and
be the corresponding threshold parameters for the two populations. The ordered values of the
are denoted by
, and
denotes the minimum sample value that emanates from the population associated with
. For a sample of size
, the cumulative density function (cdf) and probability density function (pdf) of the minimum value are given respectively by (see, e.g., McDonald and Hodaj [1])
(2.1)
and
(2.2)
For the calculation of the OCs for
, set
and
. Denote the Pr(CS) for selection rule
by Pr(CS1) and for an incorrect selection by Pr(ICS1), and the expected subset size by ESS1. Then
(2.3)
The Pr(CS1) is minimized when
, so for a given value of
,
(2.4)
Following the same line of derivation for Pr(CS1) with
, gives
(2.5)
And with
, then
(2.6)
and
(2.7)
The OCs described in this Section for
and the selection rule
based on the minimum order statistics, and the corresponding OCs for
based on the sample means, are codified in the R-code of Appendix A. Table 1 provides the output of this code for
,
, and a slippage value for
denoted by
. Calculations are given to 5 dp. The constants for the selection rules
and
are
and
. The entries in Table 1 for
are given in Section 2, Table 1 of McDonald and Hodaj [6], and are repeated here for comparison to the corresponding values for
.
Table 1. OCs for rules
(red) and
(blue):
,
,
,
.
δ= |
0 |
0.05 |
0.15 |
0.20 |
0.25 |
0.30 |
0.35 |
0.40 |
0.45 |
0.50 |
0.55 |
Pr (CS) |
0.95 |
0.98567 |
0.99882 |
0.99966 |
0.9999 |
0.99997 |
0.99999 |
1 |
1 |
1 |
1 |
Pr (CS) |
0.95 |
0.96574 |
0.98519 |
0.99066 |
0.99428 |
0.99659 |
0.99803 |
0.99889 |
0.99939 |
0.99968 |
0.99983 |
Pr (ICS) |
0.95 |
0.82548 |
0.11759 |
0.03369 |
0.00965 |
0.00277 |
0.00079 |
0.00023 |
0.00007 |
0.00002 |
0.00001 |
Pr (ICS) |
0.95 |
0.92896 |
0.86747 |
0.82581 |
0.77666 |
0.72045 |
0.65815 |
0.5912 |
0.52148 |
0.45109 |
0.38221 |
ESS |
1.9 |
1.81116 |
1.11641 |
1.03335 |
1.00956 |
1.00274 |
1.00078 |
1.00022 |
1.00006 |
1.00002 |
1.00001 |
ESS |
1.9 |
1.8947 |
1.85266 |
1.81648 |
1.77094 |
1.71704 |
1.65618 |
1.59009 |
1.52087 |
1.45077 |
1.38204 |
The probabilities of correct and incorrect selection for
are denoted, respectively, by Pr(CS2) and Pr(ICS2), and the expected subset size by ESS2. Since the exponential distribution has finite variance, the CLT applies to the sample mean. As noted in the earlier publication [6], by the CLT (see Navidi [7], Section 4.11), the distribution of the sample mean,
, is approximately normal with mean and variance equal to
and
respectively. In McDonald and Hodaj [6], CLT-based probabilities were found to be numerically very close to those obtained from the exact gamma distribution for moderate sample sizes supporting the use of the normal approximation here. Then using the CLT, after noting that
, and that a linear combination of normal random variables follows a normal distribution, it follows that
(2.8)
Since the Pr(CS2) is minimized when
, the value of
is given by
(2.9)
The expressions for Pr(ICS2) and ESS2 follow in a similar manner.
(2.10)
and
(2.11)
Here the functions
and
are the cumulative distribution function (cdf) and the inverse cdf for a standardized normal random variable.
The OCs for
are substantially better than those for
: higher Pr(CS), lower Pr(ICS), and lower ESS.
Table 2 and Table 3 examine the OCs of the two selection rules, as done in Table 1, for different choices of sample sizes and the
values. Table 2 uses a sample size of 10 with
equal to 0.90 and selection constants
and
. Table 3 uses a sample size of 50 with
equal to 0.99 and selection constants
and
. The OC results in these two additional cases are qualitatively the same as those given in Table 1. The selection rule
, compared to
, displays consistently higher Pr(CS), lower Pr(ICS), and lower ESS.
Table 2. OCs for rules
(red) and
(blue):
,
,
,
.
δ= |
0 |
0.05 |
0.15 |
0.20 |
0.25 |
0.30 |
0.35 |
0.40 |
0.45 |
0.50 |
0.55 |
Pr (CS) |
0.90 |
0.93935 |
0.97769 |
0.98647 |
0.99179 |
0.99502 |
0.99698 |
0.99817 |
0.99889 |
0.99933 |
0.99959 |
Pr (CS) |
0.90 |
0.91824 |
0.94706 |
0.95807 |
0.96716 |
0.97455 |
0.9805 |
0.98522 |
0.98992 |
0.99179 |
0.99399 |
Pr (ICS) |
0.90 |
0.83513 |
0.55183 |
0.33834 |
0.20521 |
0.12447 |
0.07549 |
0.04579 |
0.02777 |
0.01684 |
0.01022 |
Pr (ICS) |
0.90 |
0.87895 |
0.82796 |
0.79795 |
0.76502 |
0.72931 |
0.69108 |
0.65067 |
0.60847 |
0.56494 |
0.52062 |
ESS |
1.8 |
1.77447 |
1.52952 |
1.3248 |
1.197 |
1.11949 |
1.07247 |
1.04396 |
1.02666 |
1.01617 |
1.00981 |
ESS |
1.8 |
1.79719 |
1.77502 |
1.75603 |
1.73217 |
1.70386 |
1.67158 |
1.63589 |
1.59739 |
1.55674 |
1.51461 |
Table 3. OCs for rules
(red) and
(blue):
,
,
,
.
δ= |
0 |
0.05 |
0.15 |
0.20 |
0.25 |
0.30 |
0.35 |
0.40 |
0.45 |
0.50 |
0.55 |
Pr (CS) |
0.99 |
0.99918 |
0.99999 |
1 |
1 |
1 |
1 |
1 |
1 |
1 |
1 |
Pr (CS) |
0.99 |
0.99501 |
0.99895 |
0.99956 |
0.99983 |
0.99993 |
0.99998 |
0.99999 |
1 |
1 |
1 |
Pr (ICS) |
0.99 |
0.87818 |
0.01383 |
0.00113 |
0.00009 |
0.00001 |
0 |
0 |
0 |
0 |
0 |
Pr (ICS) |
0.99 |
0.98107 |
0.94253 |
0.90764 |
0.85911 |
0.7957 |
0.71781 |
0.62792 |
0.53043 |
0.43107 |
0.33591 |
ESS |
1.98 |
1.87735 |
1.01382 |
1.00113 |
1.00009 |
1.00001 |
1 |
1 |
1 |
1 |
1 |
ESS |
1.98 |
1.97608 |
1.94148 |
1.9072 |
1.85894 |
1.79563 |
1.71779 |
1.62791 |
1.53043 |
1.43107 |
1.33591 |
3. Conclusions
The results herein obtained further reinforce the conclusion that the subset selection rule based on the sample minimum values outperforms that based on the sample means when the CLT is applicable. Both selection rules are based on unbiased estimators of the exponential threshold parameter. The variance of the sample minimum value is
, while that of the sample mean is
making the min value a more efficient estimator of the threshold parameter. When applying subset selection within the context herein described, the expected number of populations selected using the minimum values will be no greater than that using the mean values while still preserving the required
condition.
From a practical standpoint, such subset selection rules are directly relevant in reliability and life-testing studies where experimenters must select a design, component, or treatment with the largest threshold (or guaranteed minimum lifetime) among several candidates. Using the rule based on sample minima allows practitioners to achieve a desired probability of correct selection with an expected subset size no greater than that obtained with a selection rule based on sample means. This can translate into fewer units being carried forward for further testing or field deployment while still maintaining prespecified reliability guarantees; see Meeker et al. [8] for a discussion of applications of exponential and related lifetime models.
The sum of independent exponential random variables follows a gamma distribution which may be used with small samples when the use of the CLT might not be appropriate. This approach was investigated by McDonald and Hodaj [6], and results using the gamma distribution were shown to be comparable to those using the CLT.
This note focuses on analytic expressions for the operating characteristics of two subset selection rules with
permitting quantitative comparisons of the rules with computer evaluations. Further challenges and avenues for future work arise in several areas. Might there be a mathematical proof that Pr(CS1) is never less than that of the rule based on sample means? And can similar results be derived for greater than two populations? More broadly, the exponential distribution applies in applications where the hazard function is constant. This suggests that the population of units is not wearing out over time. The exponential distribution is a special case of the Weibull distribution which is another direction in which to expand these analyses.
Appendix
#OCs for R1 and R2, k=2 #R1 values end in 1; R2 end in 2 #Set n, delta, and P n<-50; delta<-c(0,0.05,0.15,0.20,0.25,0.30,0.35,0.40,0.45,0.50,0.55); P<-0.99 len<-length(delta) PCS1<-rep(0,len);PCS2<-rep(0,len);PICS1<-rep(0,len) PICS2<-rep(0,len) d<--log(2*(1-P))/n b<-sqrt(2/n)*qnorm(P) params<-c(d,b) params for (i in 1:len){ PCS1[i]<-1-0.5*exp(-n*delta[i]-n*d) PCS2[i]<-pnorm((b+delta[i])/sqrt(2/n)) if(delta[i]>=d){PICS1[i]<-(0.5)*exp(-n*(delta[i]-d))} if(d>delta[i]){PICS1[i]<-1-0.5*exp(-n*(d-delta[i]))} PICS2[i]<-pnorm((b-delta[i])/sqrt(2/n)) } ESS1<-PCS1+PICS1 ESS2<-PCS2+PICS2 df1<-data.frame(n,d,b,P,delta,PCS1,PCS2,PICS1,PICS2,ESS1,ESS2) df1<-round(df1,4) df1 length(delta) length(PCS1) |