Selection Rules for Exponential Population Threshold Parameters

Abstract

This article constructs statistical selection procedures for exponential populations that may differ in only the threshold parameters. The scale parameters of the populations are assumed common and known. The independent samples drawn from the populations are taken to be of the same size. The best population is defined as the one associated with the largest threshold parameter. In case more than one population share the largest threshold, one of these is tagged at random and denoted the best. Two procedures are developed for choosing a subset of the populations having the property that the chosen subset contains the best population with a prescribed probability. One procedure is based on the sample minimum values drawn from the populations, and another is based on the sample means from the populations. An “Indifference Zone” (IZ) selection procedure is also developed based on the sample minimum values. The IZ procedure asserts that the population with the largest test statistic (e.g., the sample minimum) is the best population. With this approach, the sample size is chosen so as to guarantee that the probability of a correct selection is no less than a prescribed probability in the parameter region where the largest threshold is at least a prescribed amount larger than the remaining thresholds. Numerical examples are given, and the computer R-codes for all calculations are given in the Appendices.

Share and Cite:

McDonald, G. and Hodaj, J. (2025) Selection Rules for Exponential Population Threshold Parameters. Applied Mathematics, 16, 1-14. doi: 10.4236/am.2025.161001.

1. Introduction

The Weibull distribution is one of the most widely utilized probability distributions in applied reliability statistics. Its development spanned from 1922 to 1943, during which three separate groups worked independently toward different objectives. Among these researchers was Waloddi Weibull, whose name has since become synonymous with the distribution (see Rinne [1]). The distribution form that was introduced by Weibull in 1939, depends on three parameters. The cumulative distribution function, the density function and the hazard function or the failure rate of a three-parameter Weibull distribution for tγ are given as:

F( t|γ,η,β )=1exp[ ( tγ η ) β ] (1.1)

f( t|γ,η,β )= β η ( tγ η ) β1 exp[ ( tγ η ) β ] (1.2)

H( t;γ,η,β )= β η ( tγ η ) β , (1.3)

where:

  • γ is the location parameter, also called the threshold or shift parameter.

  • η>0 is the scale parameter.

  • β>0 is the shape parameter.

The mean and the variance of a Weibull random variable are respectively,

E( T )=ηΓ( 1 β +1 )+γ (1.4)

V( T )= η 2 [ Γ( 2 β +1 ) Γ 2 ( 1 β +1 ) ], (1.5)

where Γ( z ) is the gamma function defined as,

Γ( z )= 0 t z1 e t dt . (1.6)

Weibull percentiles are important in many applications, as they represent the time at which a certain percentage of the population is expected to fail. When γ = 0 the scale parameter η is the 63.2% percentile of the Weibull distribution, i.e., the time at which 63.2% of the population will fail. The percentile function is

t p =γ+η [ log( 1p ) ] 1/β . (1.7)

The most used Weibull model is the two-parameter model, when the threshold parameter is not present, i.e. γ=0 . So the functions that characterize this form of the distribution are, for t0 ,

F( t|η,β )=1exp[ ( t η ) β ] (1.8)

f( t|γ,η,β )= β η ( t η ) β1 exp[ ( t η ) β ] (1.9)

H( t;γ,η,β )= β η ( t η ) β . (1.10)

The mean in this case is E( T )=ηΓ( 1 β +1 ) , and the variance remains the same as in the three-parameter case, since it does not depend on the threshold parameter.

The practical significance of the Weibull distribution lies in its versatility to model failure patterns across various commonly observed shapes. When 0<β<1 , the Weibull distribution exhibits a decreasing hazard function. Such a model is applied in electronic testing for components that have a high chance of early failure. For example, testing microchips or circuit boards where early failure often occurs due to minor defects introduced in manufacturing (Meeker, et al. [2]). At β=1 , the hazard function remains constant, representing a failure rate that does not change over time. For this value of the shape parameter, Weibull reduces to the exponential distribution, for tγ ,

f( t|γ,η )= 1 η exp[ ( tγ η ) ] (1.11)

and when γ=0 , for t0 the more familiar form

f( t|η )= 1 η exp( t η ). (1.12)

This model is useful for items where the probability of failure does not depend on age (memoryless property). Some examples of such items are, light bulbs, electronic resistors, or other items where failure events are random and independent of how long they’ve been in use.

For β>1 the hazard function is increasing. When 1<β<2 the failure rate is moderately increasing and such a model is applied to fatigue testing in mechanical engineering. Components subject to moderate stress, like springs, bearings, or light-duty machinery parts, may experience a gradually increasing failure rate as minor wear gradually contributes to eventual failure (Meeker et al. [2]).

When β=2 , a special case of the Weibull distribution, known as the Rayleigh distribution, has a linearly increasing failure rate. This model is applied in meteorology, for wind speed modeling. An application might be calculating probabilities of extreme weather events like strong gusts, which are more likely with increasing wind speed.

For β>2 the distribution has a rapidly increasing failure rate. It models aging or wear-out phase where failure rate accelerates as the item ages, indicating that the likelihood of failure increases significantly over time. This is common for high-stress mechanical systems that are subject to significant wear such as engines, turbines, or heavy equipment (Meeker et al. [2]). For the case when β is very large, the Weibull distribution models sudden wear out. It approximates a deterministic lifespan where failure occurs almost predictably after a certain period. This is useful in setting maintenance schedules for parts that must be replaced after a defined usage time to avoid sudden failure.

The extensive applicability of the Weibull distribution in reliability analysis underscores its importance, often making it the preferred choice over other distributions. A notable example is the study by McDonald, et al. [3], which demonstrates the advantages of using the Weibull distribution over the lognormal distribution in analyzing car emission data, highlighting its superior fit and interpretative value in that context.

The goal of this paper is to derive selection rules for the special case of the Weibull distribution when β=1 . This is the exponential distribution with location and scale parameters ( γ,η ) . In an early paper, An Hsu [4] gives some procedures for Weibull populations. The author constructs optimal selection procedures in order to select a subset of the k populations containing the best population. The procedures control the size of the selected subset and also maximize the minimum probability of a correct selection. The author considers the cases when the shape parameter β is common and known among the k populations and the sample sizes are not necessarily equal. The selection procedures are based on the scale parameter η , where a population is considered best if it has the largest scale parameter. Hence selecting the best population means selecting the population with the largest scale η . So, the main two problems of An Hsu’s paper are, to maximize the probability of correct selection and to minimize the subset size.

Two excellent and comprehensive books on ranking and selection procedures are authored by Gupta and Panchapkesan [5] and Gibbons, et al. [6].

2. Selection Rules for the Exponential Threshold Model ( β=1 , η=1 )

Let Π 1 , Π 2 ,, Π k be k three-parameter Weibull populations with common, known, shape and scale parameters, ( β=1 , η=1 ) and threshold parameters γ i , for i=1,2,,k . Let T i be a random variable associated with population Π i , for i=1,2,,k , then T i follows the distribution

F( t| γ i )=1exp[ ( t γ i ) ],t γ i . (2.1)

Two approaches to selection rules will be developed. The first approach considers subset selection, where the goal is to select a subset of the k populations that contains the best population with a specified probability P * . The second, involves the indifference zone, which focuses on selecting the single best population among the k , again with a specified probability P * . In both these scenarios, the best population is the one having the largest (the smallest) threshold parameter, depending on the context. In the exponential model considered here, the threshold parameter is additive in the expression for the mean and for any percentile. Thus, selection for the largest threshold is equivalent to selection for the largest mean or for any of the percentiles.

2.1. Subset Selection for Populations to Contain the One with the Largest Threshold

The subset selection approaches developed herein follow the seminal work by Gupta [7]. Let T ij be n independent identically distributed(iid) random variables from population Π i , for i=1,2,,k and j=1,2,,n . Let γ [ 1 ] γ [ 2 ] γ [ k ] be the ordered thresholds, where γ [ k ] corresponds to the best population. Since tγ , then the minimum order statistic is a typical estimator for γ . So, let Y ( i ) be the minimum of the sample from the population with threshold parameter γ [ i ] ,

Y ( i ) = min 1jn ( X ij ),i=1,2,,k. (2.2)

The selection rule is then,

R 1 :Select Π i iff Y i max 1jk ( Y j )d,d0. (2.3)

The distribution of the minimum order statistic for independent random variables X i ,i=1,2,,n following (2.1) can easily be derived.

If Y=min( X 1 , X 2 ,, X n ) , then

Pr( Yy )=1Pr( Y>y ) =1Pr( X i >y,i=1,2,,n ) =1 i=1 n Pr( X i >y )byindependence =1 i=1 n [ 1Pr( X i y ) ] =1 i=1 n { 1[ 1 e ( yγ ) ] } =1 e n( yγ ) ,yγ. (2.4)

From here, the pdf of Y is f Y ( y )=n e n( yγ ) , yγ . The probability of a (CS) now follows:

Pr( CS )=Pr( Y ( k ) max 1jk1 ( Y ( j ) )d ) =Pr( max 1jk1 ( Y ( j ) ) Y ( k ) +d ) = γ [ k ] j=i k1 Pr( Y ( j ) y+d )d F Y ( k ) ( y ) = γ [ k ] j=i k1 Pr( Y ( j ) y+d )n e n( y γ [ k ] ) dy .

Now, let x=y γ [ k ] , and so y=x+ γ [ k ] . Then since γ [ k ] γ [ j ] 0 , j=1,2,,k1

Pr( CS )=n 0 j=i k1 Pr( Y ( j ) γ [ j ] x+ γ [ k ] γ [ j ] +d ) e nx dx n 0 j=i k1 Pr( Y ( j ) γ [ j ] x+d ) e nx dx . (2.5)

The parameter configuration γ 1 = γ 2 == γ k is referred to as the Least Favorable Configuration (LFC) since it yields the minimum value of Pr( CS ) . Setting the expression in (2.5) equal to a specified value P * ( k 1 < P * <1 ), the value of d can now be determined. Now the selection rule R 1 will ensure that Pr( CS ) P * no matter the configuration of γ i s .

Common and Known Scale Parameter η1

We assume that the scale parameter is common for all the populations Π i for i=1,2,,k and it is known, but it’s not 1. Does this change the selection rule for subset selection? Let T~exp( η,γ ) . The random variable T η is still exponentially distributed with scale parameter η=1 ,

Pr( T η t )=Pr( Tηt ) =1exp[ ηtγ η ] =1exp[ ( t γ η ) ] =1exp[ ( t γ ) ]. (2.6)

Or otherwise,

F( t η |γ )=1exp[ ( t γ ) ],t γ , γ = γ η . (2.7)

Hence, we get back to the previous case where scale parameter η=1 and the same selection rule, R 1 , can be applied.

2.2. Selection Based on Sample Means

If β=1 , then E( T )=ηΓ( 2 )+γ=η+γ and V( T )= η 2 [ Γ( 3 ) Γ 2 ( 2 ) ]= η 2 . Suppose Π i ,i=1,2,,k are k independent populations with Π i distributed as Weibull ( γ i ,η,β=1 ) , and η is a common value known. Let X ij be an independent random sample from Π i , j=1,2,,n . Our goal is to select a subset of the populations such that the population having the largest γ value is contained in the subset with a specified probability P * ( k 1 < P * <1 ). Denote the ordered γ -values by

γ [ 1 ] γ [ 2 ] γ [ k ] .

Since γ is a threshold, as covered in Section 2 it is reasonable to consider a selection rule based on the minimum sample values from the population, such as R 1 given in 0.15.

Since the population means are ordered as the population γ -values, it is also reasonable to consider a selection rule based on the population sample means, X ¯ i . That is,

R 2 :Choose Π i iff X ¯ i max 1jk ( X ¯ j )b, (2.8)

where X ¯ i = ( j=1 n X ij )/n and b=b( k,n, P * ) is a nonnegative number chosen to satisfy the P * condition,

min Ω Pr( CS ) P * , (2.9)

where Ω=( γ 1 , γ 2 ,, γ k ) .

Let X ¯ ( i ) be the sample mean drawn from the population possessing the mean value η+ γ [ k ] . Using the Central Limit Theorem, the distribution of X ¯ ( i ) is approximately normal with mean η+ γ [ k ] and variance V( T )/n = η 2 /n . So for large n ,

Pr( CS )=Pr( X ¯ ( k ) max 1jk1 X ¯ ( j ) b ) =Pr( Z k Z j +( γ [ j ] γ [ k ] )( n /η )b,j=1,2,,k1 ) Φ k1 ( x+cb )ϕ( x ) dx , (2.10)

where Z i , i=1,2,,k , are independent standardized normal random variables and c= n /η . Thus, b=b( k,n, P * ) is determined by setting Equation (2.10) equal to P * and solving for b .

For small values of n , the selection procedure can be based on the sum of the random variables, n X ¯ i , i=1,2,,k , and utilizing the property that a sum of n independent exponential random variables follows the gamma distribution with parameters ( n,η ) assuming the threshold parameter, γ , is equal to 0 (Casella and Berger [8]).

To assess the sensitivity of the subsets selected using selection rules R 1 and R 2 to the assumption of a common known scale parameter, η , set a bound on possible values of the parameter. That is, say the assessment of the analyst is L<η<U . Now discretize the interval ( L,U ) into m discrete values, i.e., L= v 1 < v 2 << v m =U . Follow the procedures in Sections (2.1.1) and (2.2) with η assumed to be equal to v i to obtain the selected subsets. Repeat the process using the remaining values of v=( v 1 ,, v m ) . Compare the m selected subsets and assess the sensitivity to the selections based on the uncertainty of the scale parameter. To assess the sensitivity to the assumption that the shape parameter, β , is equal to one is not possible since the LFC has not been determined for values of β not equal to one.

2.3. Application of the Two Selection Rules

In this section, an example for each of the selection rules that were developed previously, rules (2.3) and (2.8) are given. The R-code for these rules is given in the Appendix section.

2.3.1. Application of the Minimum Order Statistics Selection Rule

Using the selection rule for the largest threshold based on minimum order statistics from the k populations, the d -values for different levels of the probability P * are computed.

The example will be for k=10 populations. From each population, a sample of size n=25 is drawn, with an equi-spaced configuration for the threshold parameters, i.e., γ i s=110 , by 1. Simulations are set for N=10000 . The d -values obtained are given in Table 1. The R-code for generating the d -values is found in Appendix A. The P * values based on simulations can be checked using the R-code in Appendix D.

Table 1. P * and d -values.

P *

0.75

0.90

0.95

0.975

0.99

d

0.1083

0.1501

0.1785

0.2077

0.2449

Using the R-code in Appendix B random samples of size 25 are generated from 10 exponential populations having γ -values equal to 1 up to 10 with step size 1. Table 2 gives the minimum and mean values of these samples, respectively. The populations selected using rule R 1 are given in Table 4 for the five values of P * given in Table 1.

Table 2. Minimums and means for each of the k populations.

Population

Minimum

Mean

1

0.0100

0.8842

2

0.1626

2.3289

3

0.4438

3.2933

4

0.3080

4.1323

5

0.0774

4.5331

6

0.0683

5.5062

7

0.2490

6.2474

8

0.5265

9.1137

9

0.2727

6.7970

10

0.3481

9.6343

2.3.2. Application of the Sample Means Rule

For the selection rule based on the sample means, the b -value is calculated using the R-code in Appendix C, which for k=10 , n=25 is given in Table 3:

Table 3. P * and b -values.

P *

0.75

0.90

0.95

0.975

0.99

b

0.4528

0.5970

0.6836

0.7598

0.8500

The selected populations using R 2 with the example simulated exponential dataset, γ=110 , step size 1, are given in Table 4. The sample means procedure yields subsets with smaller size in most of the P * values. For this particular simulation R 2 would be preferred over R 1 since the selected subset size is favorable. Another similarly simulated data set could well result in R 1 being preferred over R 2 . Simulation studies, yet to be published, show that the expected subset size of the selected populations is smaller for R 1 than for R 2 with equi-spaced threshold parameters (or slippage parameter) greater than 0.

Table 4. P * and the selected populations using rules R 1 and R 2 .

P *

Min Procedure, R 1

Means Procedure, R 2

0.75

3, 8

10

0.90

3, 8

8, 10

0.95

3, 8, 10

8, 10

0.975

3, 8, 10

8, 10

0.99

3, 4, 8, 10

8, 10

2.4. Indifference Zone Approach

Following the Indifference Zone approach formulated by Bechhofer [9], the population yielding the largest minimum order statistics, Y [ k ] would be asserted to be the “best” population. The statistical goal is to have the probability of this assertion to be correct with a specified probability, P * , over the parameter space γ [ k ] γ [ k1 ] d * , where d * is a user specified meaningful difference between the best population and all the remainder ones. The minimum of the sample size n=n( k, P * , d * ) needs to be determined so that the Pr( CS ) P * for the parameter configuration γ [ k ] γ [ k1 ] d * . Other parameter configurations comprise the “indifference zone” in the sense that the Pr( CS ) is not necessarily applicable.

For this decision rule,

Pr( CS )=Pr( Y ( k ) max Y ( j ) ,j=1,2,,k1 ) =Pr( Y ( j ) Y ( k ) ,j=1,2,,k1 ) =Pr( Y ( j ) γ [ j ] Y ( k ) γ [ k ] + γ [ k ] γ [ j ] ,j=1,2,,k1 ) Pr( Y ( j ) γ [ j ] Y ( k ) γ [ k ] + d * ,j=1,2,,k1 ) =n 0 [ 1 e n( x+ d * ) ] k1 e nx dx . (2.11)

The LFC is γ [ 1 ] == γ [ k1 ] = γ [ k ] d * . The above integral expression is calculated with the R-code found in Appendix 5 and through iteration we can determine the minimum sample size required to meet the Pr( CS ) requirement.

For example, for k=5 populations and requirement P * =0.90 with d * =0.25 , Table 5 gives evaluations of (2.11). Thus, a sample size of 12 would be the minimum size requirement to meet the experimental goal.

3. Summary and Conclusions

The importance of the Weibull distribution sparked our interest to investigate

Table 5. Sample size n and the P * requirement.

n

P *

10

0.8488

11

0.8801

12

0.9053

13

0.9254

14

0.9414

procedures for selecting the best among several Weibull populations. In this paper, the case when the shape parameter is common, known and β=1 is considered. Under this consideration, the distribution reduces to the exponential distribution with scale and threshold parameters ( η,γ ) . Further, the scale parameter η is assumed common and known.

Two different approaches, subset selection and indifference zone, are developed. For the subset selection approach, two different procedures, one based on the minimum order statistic and the other based on the sample means are considered. The two selection rules are given by Equations (2.3) and (2.8). Using the R-codes provided in the appendices, there is little difference in the computational requirements for implementation of the selection rules R 1 and R 2 .

In the work presented here, the sample sizes for the k populations are equal. As is often the case in practice, samples are not equal. In these cases, Gibbons, et al. [6], Section (4), suggest using a generalized average sample size for each of the populations. In particular, they recommend using the square-mean-root, N 0 , given by

N 0 = ( ( i=1 k n i )/k ) 2 . (3.1)

Note that the square root of N 0 is the arithmetic mean of the square root of the actual sample sizes. Gibbons, et al. [6] remark that N 0 is always greater than the geometric mean of the actual sample sizes, and always smaller than the arithmetic mean. In practice, N 0 is rarely very different from the arithmetic mean; however (3.1) gives more accurate results most of the time and is therefore preferable to the arithmetic mean.

These theoretical results are illustrated with simulated data to demonstrate the selection procedures. Using the R-codes in Appendices A-C, the values of d and b given the number of populations k , the sample size n and the probability of a correct selection P * , can be determined. Once these values were determined, selection rules (2.3) and (2.8) determined the subset that contains the best population. The size of the selected subset is a nondecreasing function of Pr( CS ) . For our example, the sample means procedure R 2 , seems to be superior to the minimum order statistic procedure, R 1 , because it selected a subset of smaller size, for almost all the values of P * . However, this is just one illustrative example. The conditions under which one of the two procedures is guaranteed better, in the sense of yielding a subset size no larger than the other, have yet to be determined.

For the indifference zone approach formulated by Bechhofer [9], our procedure was based on the minimum order statistics. In this case, the probability of a correct selection is specified over the parameter space determined by γ [ k ] γ [ k1 ] d * , where d * is user specified. Here the minimum sample size n is determined so that the probability of a correct selection is at least P * . This is given by Equation (2.11). An example is provided for this approach. The R-code for its implementation can be found in Appendix D. For a given number of populations k=5 , d * =0.25 and P * =0.90 a minimum sample size of n=12 is required in order to achieve the experimental goal.

This article naturally leads to further investigations about selection procedures. What are the conditions under which R 1 is preferred to R 2 ? One can also be interested in developing procedures when the shape parameter is common and known, but not equal to 1, or when the scale parameter is common and unknown and the sample sizes are not necessarily equal.

A. Appendix

B. Appendix

C. Appendix

D. Appendix

Conflicts of Interest

The authors declare no conflicts of interest regarding the publication of this paper.

References

[1] Rinne, H. (2009) The Weibull Distribution, a Handbook. Chapman & Hall/CRC.
[2] Meeker, W.Q., Escobar, L.A. and Pascual, F.G. (2022) Statistical Methods for Reliability Data. 2nd Edition, John Wiley & Sons.
[3] McDonald, G.C. and Vance, L.C. and Gibbons, D.I. (1995) Some Tests for Dis-criminating Between Lognormal and Weibull Distributions: An Application to Emissions Data. In: Balakrishnan, N., Ed., Recent Advances in Life-Testing and Reliability, CRC Press, 475-490.[CrossRef]
[4] An Hsu, T. (1982) On Some Optimal Selection Procedures for Weibull Populations. Communications in StatisticsTheory and Methods, 11, 2657-2668.[CrossRef]
[5] Gupta, S.S. and Panchapakesan, S. (1979) Multiple Decision Procedures. John Wiley & Sons.
[6] Gibbons, J.D., Olkin, I. and Sobel, M. (1999) Selecting and Ordering Populations: A New Statistical Methodology. Society for Industrial and Applied Mathematics.[CrossRef]
[7] Gupta, S.S. (1965) On Some Multiple Decision (Selection and Ranking) Rules. Technometrics, 7, 225-245.[CrossRef]
[8] Casella, G. and Berger, R.L. (2002) Statistical Inference. Duxbury.
[9] Bechhofer, R.E. (1954) A Single-Sample Multiple Decision Procedure for Ranking Means of Normal Populations with Known Variances. The Annals of Mathematical Statistics, 25, 16-39.[CrossRef]

Copyright © 2026 by authors and Scientific Research Publishing Inc.

Creative Commons License

This work and the related PDF file are licensed under a Creative Commons Attribution 4.0 International License.