Selection Rules for Exponential Population Threshold Parameters ()
1. Introduction
The Weibull distribution is one of the most widely utilized probability distributions in applied reliability statistics. Its development spanned from 1922 to 1943, during which three separate groups worked independently toward different objectives. Among these researchers was Waloddi Weibull, whose name has since become synonymous with the distribution (see Rinne [1]). The distribution form that was introduced by Weibull in 1939, depends on three parameters. The cumulative distribution function, the density function and the hazard function or the failure rate of a three-parameter Weibull distribution for
are given as:
(1.1)
(1.2)
(1.3)
where:
is the location parameter, also called the threshold or shift parameter.
is the scale parameter.
is the shape parameter.
The mean and the variance of a Weibull random variable are respectively,
(1.4)
(1.5)
where
is the gamma function defined as,
(1.6)
Weibull percentiles are important in many applications, as they represent the time at which a certain percentage of the population is expected to fail. When
= 0 the scale parameter
is the 63.2% percentile of the Weibull distribution, i.e., the time at which 63.2% of the population will fail. The percentile function is
(1.7)
The most used Weibull model is the two-parameter model, when the threshold parameter is not present, i.e.
. So the functions that characterize this form of the distribution are, for
,
(1.8)
(1.9)
(1.10)
The mean in this case is
, and the variance remains the same as in the three-parameter case, since it does not depend on the threshold parameter.
The practical significance of the Weibull distribution lies in its versatility to model failure patterns across various commonly observed shapes. When
, the Weibull distribution exhibits a decreasing hazard function. Such a model is applied in electronic testing for components that have a high chance of early failure. For example, testing microchips or circuit boards where early failure often occurs due to minor defects introduced in manufacturing (Meeker, et al. [2]). At
, the hazard function remains constant, representing a failure rate that does not change over time. For this value of the shape parameter, Weibull reduces to the exponential distribution, for
,
(1.11)
and when
, for
the more familiar form
(1.12)
This model is useful for items where the probability of failure does not depend on age (memoryless property). Some examples of such items are, light bulbs, electronic resistors, or other items where failure events are random and independent of how long they’ve been in use.
For
the hazard function is increasing. When
the failure rate is moderately increasing and such a model is applied to fatigue testing in mechanical engineering. Components subject to moderate stress, like springs, bearings, or light-duty machinery parts, may experience a gradually increasing failure rate as minor wear gradually contributes to eventual failure (Meeker et al. [2]).
When
, a special case of the Weibull distribution, known as the Rayleigh distribution, has a linearly increasing failure rate. This model is applied in meteorology, for wind speed modeling. An application might be calculating probabilities of extreme weather events like strong gusts, which are more likely with increasing wind speed.
For
the distribution has a rapidly increasing failure rate. It models aging or wear-out phase where failure rate accelerates as the item ages, indicating that the likelihood of failure increases significantly over time. This is common for high-stress mechanical systems that are subject to significant wear such as engines, turbines, or heavy equipment (Meeker et al. [2]). For the case when
is very large, the Weibull distribution models sudden wear out. It approximates a deterministic lifespan where failure occurs almost predictably after a certain period. This is useful in setting maintenance schedules for parts that must be replaced after a defined usage time to avoid sudden failure.
The extensive applicability of the Weibull distribution in reliability analysis underscores its importance, often making it the preferred choice over other distributions. A notable example is the study by McDonald, et al. [3], which demonstrates the advantages of using the Weibull distribution over the lognormal distribution in analyzing car emission data, highlighting its superior fit and interpretative value in that context.
The goal of this paper is to derive selection rules for the special case of the Weibull distribution when
. This is the exponential distribution with location and scale parameters
. In an early paper, An Hsu [4] gives some procedures for Weibull populations. The author constructs optimal selection procedures in order to select a subset of the
populations containing the best population. The procedures control the size of the selected subset and also maximize the minimum probability of a correct selection. The author considers the cases when the shape parameter
is common and known among the
populations and the sample sizes are not necessarily equal. The selection procedures are based on the scale parameter
, where a population is considered best if it has the largest scale parameter. Hence selecting the best population means selecting the population with the largest scale
. So, the main two problems of An Hsu’s paper are, to maximize the probability of correct selection and to minimize the subset size.
Two excellent and comprehensive books on ranking and selection procedures are authored by Gupta and Panchapkesan [5] and Gibbons, et al. [6].
2. Selection Rules for the Exponential Threshold Model (
,
)
Let
be
three-parameter Weibull populations with common, known, shape and scale parameters, (
,
) and threshold parameters
, for
. Let
be a random variable associated with population
, for
, then
follows the distribution
(2.1)
Two approaches to selection rules will be developed. The first approach considers subset selection, where the goal is to select a subset of the
populations that contains the best population with a specified probability
. The second, involves the indifference zone, which focuses on selecting the single best population among the
, again with a specified probability
. In both these scenarios, the best population is the one having the largest (the smallest) threshold parameter, depending on the context. In the exponential model considered here, the threshold parameter is additive in the expression for the mean and for any percentile. Thus, selection for the largest threshold is equivalent to selection for the largest mean or for any of the percentiles.
2.1. Subset Selection for Populations to Contain the One with the Largest Threshold
The subset selection approaches developed herein follow the seminal work by Gupta [7]. Let
be
independent identically distributed(iid) random variables from population
, for
and
. Let
be the ordered thresholds, where
corresponds to the best population. Since
, then the minimum order statistic is a typical estimator for
. So, let
be the minimum of the sample from the population with threshold parameter
,
(2.2)
The selection rule is then,
(2.3)
The distribution of the minimum order statistic for independent random variables
following (2.1) can easily be derived.
If
, then
(2.4)
From here, the pdf of
is
,
. The probability of a (CS) now follows:
Now, let
, and so
. Then since
,
(2.5)
The parameter configuration
is referred to as the Least Favorable Configuration (LFC) since it yields the minimum value of
. Setting the expression in (2.5) equal to a specified value
(
), the value of
can now be determined. Now the selection rule
will ensure that
no matter the configuration of
.
Common and Known Scale Parameter
We assume that the scale parameter is common for all the populations
for
and it is known, but it’s not 1. Does this change the selection rule for subset selection? Let
. The random variable
is still exponentially distributed with scale parameter
,
(2.6)
Or otherwise,
(2.7)
Hence, we get back to the previous case where scale parameter
and the same selection rule,
, can be applied.
2.2. Selection Based on Sample Means
If
, then
and
. Suppose
are
independent populations with
distributed as Weibull
, and
is a common value known. Let
be an independent random sample from
,
. Our goal is to select a subset of the populations such that the population having the largest
value is contained in the subset with a specified probability
(
). Denote the ordered
-values by
Since
is a threshold, as covered in Section 2 it is reasonable to consider a selection rule based on the minimum sample values from the population, such as
given in 0.15.
Since the population means are ordered as the population
-values, it is also reasonable to consider a selection rule based on the population sample means,
. That is,
(2.8)
where
and
is a nonnegative number chosen to satisfy the
condition,
(2.9)
where
.
Let
be the sample mean drawn from the population possessing the mean value
. Using the Central Limit Theorem, the distribution of
is approximately normal with mean
and variance
. So for large
,
(2.10)
where
,
, are independent standardized normal random variables and
. Thus,
is determined by setting Equation (2.10) equal to
and solving for
.
For small values of
, the selection procedure can be based on the sum of the random variables,
,
, and utilizing the property that a sum of
independent exponential random variables follows the gamma distribution with parameters
assuming the threshold parameter,
, is equal to 0 (Casella and Berger [8]).
To assess the sensitivity of the subsets selected using selection rules
and
to the assumption of a common known scale parameter,
, set a bound on possible values of the parameter. That is, say the assessment of the analyst is
. Now discretize the interval
into
discrete values, i.e.,
. Follow the procedures in Sections (2.1.1) and (2.2) with
assumed to be equal to
to obtain the selected subsets. Repeat the process using the remaining values of
. Compare the
selected subsets and assess the sensitivity to the selections based on the uncertainty of the scale parameter. To assess the sensitivity to the assumption that the shape parameter,
, is equal to one is not possible since the LFC has not been determined for values of
not equal to one.
2.3. Application of the Two Selection Rules
In this section, an example for each of the selection rules that were developed previously, rules (2.3) and (2.8) are given. The R-code for these rules is given in the Appendix section.
2.3.1. Application of the Minimum Order Statistics Selection Rule
Using the selection rule for the largest threshold based on minimum order statistics from the
populations, the
-values for different levels of the probability
are computed.
The example will be for
populations. From each population, a sample of size
is drawn, with an equi-spaced configuration for the threshold parameters, i.e.,
, by 1. Simulations are set for
. The
-values obtained are given in Table 1. The R-code for generating the
-values is found in Appendix A. The
values based on simulations can be checked using the R-code in Appendix D.
Table 1.
and
-values.
|
0.75 |
0.90 |
0.95 |
0.975 |
0.99 |
|
0.1083 |
0.1501 |
0.1785 |
0.2077 |
0.2449 |
Using the R-code in Appendix B random samples of size 25 are generated from 10 exponential populations having
-values equal to 1 up to 10 with step size 1. Table 2 gives the minimum and mean values of these samples, respectively. The populations selected using rule
are given in Table 4 for the five values of
given in Table 1.
Table 2. Minimums and means for each of the k populations.
Population |
Minimum |
Mean |
1 |
0.0100 |
0.8842 |
2 |
0.1626 |
2.3289 |
3 |
0.4438 |
3.2933 |
4 |
0.3080 |
4.1323 |
5 |
0.0774 |
4.5331 |
6 |
0.0683 |
5.5062 |
7 |
0.2490 |
6.2474 |
8 |
0.5265 |
9.1137 |
9 |
0.2727 |
6.7970 |
10 |
0.3481 |
9.6343 |
2.3.2. Application of the Sample Means Rule
For the selection rule based on the sample means, the
-value is calculated using the R-code in Appendix C, which for
,
is given in Table 3:
Table 3.
and
-values.
|
0.75 |
0.90 |
0.95 |
0.975 |
0.99 |
|
0.4528 |
0.5970 |
0.6836 |
0.7598 |
0.8500 |
The selected populations using
with the example simulated exponential dataset,
, step size 1, are given in Table 4. The sample means procedure yields subsets with smaller size in most of the
values. For this particular simulation
would be preferred over
since the selected subset size is favorable. Another similarly simulated data set could well result in
being preferred over
. Simulation studies, yet to be published, show that the expected subset size of the selected populations is smaller for
than for
with equi-spaced threshold parameters (or slippage parameter) greater than 0.
Table 4.
and the selected populations using rules
and
.
|
Min Procedure,
|
Means Procedure,
|
0.75 |
3, 8 |
10 |
0.90 |
3, 8 |
8, 10 |
0.95 |
3, 8, 10 |
8, 10 |
0.975 |
3, 8, 10 |
8, 10 |
0.99 |
3, 4, 8, 10 |
8, 10 |
2.4. Indifference Zone Approach
Following the Indifference Zone approach formulated by Bechhofer [9], the population yielding the largest minimum order statistics,
would be asserted to be the “best” population. The statistical goal is to have the probability of this assertion to be correct with a specified probability,
, over the parameter space
, where
is a user specified meaningful difference between the best population and all the remainder ones. The minimum of the sample size
needs to be determined so that the
for the parameter configuration
. Other parameter configurations comprise the “indifference zone” in the sense that the
is not necessarily applicable.
For this decision rule,
(2.11)
The LFC is
. The above integral expression is calculated with the R-code found in Appendix 5 and through iteration we can determine the minimum sample size required to meet the
requirement.
For example, for
populations and requirement
with
, Table 5 gives evaluations of (2.11). Thus, a sample size of 12 would be the minimum size requirement to meet the experimental goal.
3. Summary and Conclusions
The importance of the Weibull distribution sparked our interest to investigate
Table 5. Sample size
and the
requirement.
|
|
10 |
0.8488 |
11 |
0.8801 |
12 |
0.9053 |
13 |
0.9254 |
14 |
0.9414 |
procedures for selecting the best among several Weibull populations. In this paper, the case when the shape parameter is common, known and
is considered. Under this consideration, the distribution reduces to the exponential distribution with scale and threshold parameters
. Further, the scale parameter
is assumed common and known.
Two different approaches, subset selection and indifference zone, are developed. For the subset selection approach, two different procedures, one based on the minimum order statistic and the other based on the sample means are considered. The two selection rules are given by Equations (2.3) and (2.8). Using the R-codes provided in the appendices, there is little difference in the computational requirements for implementation of the selection rules
and
.
In the work presented here, the sample sizes for the
populations are equal. As is often the case in practice, samples are not equal. In these cases, Gibbons, et al. [6], Section (4), suggest using a generalized average sample size for each of the populations. In particular, they recommend using the square-mean-root,
, given by
(3.1)
Note that the square root of
is the arithmetic mean of the square root of the actual sample sizes. Gibbons, et al. [6] remark that
is always greater than the geometric mean of the actual sample sizes, and always smaller than the arithmetic mean. In practice,
is rarely very different from the arithmetic mean; however (3.1) gives more accurate results most of the time and is therefore preferable to the arithmetic mean.
These theoretical results are illustrated with simulated data to demonstrate the selection procedures. Using the R-codes in Appendices A-C, the values of
and
given the number of populations
, the sample size
and the probability of a correct selection
, can be determined. Once these values were determined, selection rules (2.3) and (2.8) determined the subset that contains the best population. The size of the selected subset is a nondecreasing function of
. For our example, the sample means procedure
, seems to be superior to the minimum order statistic procedure,
, because it selected a subset of smaller size, for almost all the values of
. However, this is just one illustrative example. The conditions under which one of the two procedures is guaranteed better, in the sense of yielding a subset size no larger than the other, have yet to be determined.
For the indifference zone approach formulated by Bechhofer [9], our procedure was based on the minimum order statistics. In this case, the probability of a correct selection is specified over the parameter space determined by
, where
is user specified. Here the minimum sample size
is determined so that the probability of a correct selection is at least
. This is given by Equation (2.11). An example is provided for this approach. The R-code for its implementation can be found in Appendix D. For a given number of populations
,
and
a minimum sample size of
is required in order to achieve the experimental goal.
This article naturally leads to further investigations about selection procedures. What are the conditions under which
is preferred to
? One can also be interested in developing procedures when the shape parameter is common and known, but not equal to 1, or when the scale parameter is common and unknown and the sample sizes are not necessarily equal.
A. Appendix
B. Appendix
C. Appendix
D. Appendix