New Probability Distributions in Astrophysics: XIII. Truncation for the Benini Distribution ()
1. Introduction
The Benini distribution with three parameters was introduced in 1905 [1] in order to generalize the Pareto distribution with two parameters introduced in 1896 [2]. The Benini distribution has not been well studied, and only recently, 2013, was the sequence of its moments analysed [3] when the number of parameters is one. Another study in 2021 derived, for the Benini distribution two parameters, the following quantities: its random generation, median, and how to determine its parameters through the maximum likelihood estimator [4]. The above two references outline that the Beninini distribution was poorly analysed in the fields of economics and actuarial science and absent from the fields of physics and astrophysics. The above studies allow posing some questions:
Is possible to derive the main statistical properties of the truncated Benini distribution?
Can the Benini distribution model the initial mass function for stars?
Is using the truncated Benini distribution better than using the untruncated one?
In order to answer these questions, Section 2 treats the untruncated Benini distribution and Section 3 introduces its truncation. Section 4 applies the obtained results to the initial mass function for stars.
2. The Benini Distribution
Let X be a random variable taking values x in the interval
. The Benini probability density function (PDF), after [1] [5], is
(1)
where
,
and
are
. Its distribution function (DF) is
(2)
The genesis of this variate can be found in a generalization of the Pareto distribution, derived in 1896 [2], which has a DF
(3)
The survival function, SF, is defined as
(4)
where
is the distribution function and the natural logarithm for the Pareto’s survival function is
(5)
which is a polynomial of first degree in
. The natural logarithm for the Benini’s survival function is
(6)
which is a polynomial of second degree in
. In other words, the degree for the natural logarithm of the survival function is increased by one in the Benini distribution. The Benini PDF is presented in Figure 1 for different parameters.
The average value or mean of the Benini distribution,
, is
(7)
The variance is derived through the formula (1) and its value is
(8)
where erfc is the complementary error function
Figure 1. Benini PDF. The parameters are
,
,
for the red line,
,
,
for the green line and
,
,
for the blue line.
(9)
and erf is the error function [6]. The standard deviation, std, is
(10)
and the kth moment about the origin,
, is
(11)
The skewness can be derived through the implicit definition as in formula (2) and its explicit value is
(12)
The kurtosis can be derived through the implicit definition as in formula (3) and its explicit value is
(13)
A 3D display of the skewness is presented in Figure 2.
Figure 2. Benini skewness as function of
and
when
.
The random generation of the Benini variate X is given by
(14)
where R is the unit rectangular variate. The median,
, is at
(15)
and the mode is at
(16)
The three parameters
and
are obtained in the following way. Consider a sample
and let
denote their order statistics, so that
,
. Then
(17)
The two remaining parameters
and
are found by solving the two following equations which arise from the MLE:
(18)
(19)
3. The Right Truncated Benini Distribution
Let X be a random variable taking values in
, where
. The DF,
, of the right truncated Bernini distribution is
(20)
and its PDF,
, is
(21)
The survival function,
is
(22)
Its average value or mean,
, is
(23)
The increasing value of the right truncated mean as a function of the upper value in
,
, is shown in Figure 3.
Figure 3. Mean of the right truncated Benini PDF as function of
when
,
,
.
The kth moment about the origin,
, is
(24)
The above formula allows deriving the variance, skewness, and kurtosis, through the implicit formulas (A.1), (A.2) and (A.3) but they have complicated expressions which we do not present. The random generation of the right truncated Benini variate X is given by
(25)
where R is the unit rectangular variate. The median,
, is at
(26)
and the mode is at the same position as for the standard Benini distribution, see Equation (16). We now outline how to determine the four parameters. The parameter
is
(27)
and the parameter
is
(28)
The two remaining parameters
and
are found by solving numerically the two following equations which arise from the MLE:
(29)
(30)
where
(31)
(32)
(33)
(34)
(35)
4. Application to the Stars
This section reviews the lognormal distribution, Salpeter’s exponent, the Pareto distribution, the truncated Pareto distribution, the adopted statistics, applies the obtained results to the initial mass function for stars (IMF) and explores the survival function of massive stars.
4.1. Lognormal Distribution
Let X be a random variable taking values x in the interval
; the first lognormal PDF, following [7] or formula (14.2) in [8], is
(36)
its DF is
(37)
and its SF,
is
(38)
The second definition has PDF
(39)
where
and
. The DF of the second definition is
(40)
and its SF is
(41)
4.2. Salpeter’s Exponent
The distribution in mass of the stars has been fitted with a power law starting with [9]. Salpeter suggested
where
denotes the probability of having a mass between m and
. He found
in the range
; this value has changed little with time and a recent evaluation quotes 2.35, for stars with mass greater than few
, see [10].
4.3. Pareto Distribution
Let X be a random variable taking values x in the interval
,
. The Pareto PDF is defined by
(42)
with
, see formula (20.3) in [8]. The traditional Salpeter slope is therefore
. Its DF is
(43)
and its SF, S, is
(44)
4.4. Truncated Pareto distribution
An upper truncated Pareto random variable is defined in the interval
, and its corresponding PDF, following [11]-[14], is
(45)
Its DF is
(46)
and its SF is
(47)
4.5. Statistics
The merit function
is given by
(48)
where n is the number of bins,
is the theoretical value, and
is the experimental value as given by the frequencies. The theoretical frequency distribution is given by
(49)
where N is the number of elements of the sample,
is the magnitude of the size interval, and
is the PDF under examination. A reduced merit function
is given by
(50)
where
is the number of degrees of freedom, n is the number of bins, and k is the number of parameters. The goodness of the fit can be expressed by the probability Q, see equation 15.2.12 in [15], which involves the number of degrees of freedom and
. According to [15] p. 658, the fit “may be acceptable” if
. The Akaike information criterion (AIC), see [16], is defined by
(51)
where L is the likelihood function and k the number of free parameters in the model. We assume a Gaussian distribution for the errors. The likelihood function
can then be derived from the
statistic:
where
is given by Equation (48), see [17] and [18]. Now the AIC becomes
(52)
The Kolmogorov-Smirnov test (K-S), see [19]-[21], does not require the data to be binned. The K-S test, as implemented by the FORTRAN subroutine KSONE in [15], finds the maximum distance, D, between the theoretical and the astronomical DFs, as well as the significance level
; see formulas 14.3.5 and 14.3.9 in [15]. If
, then the goodness of the fit is believable.
4.6. The IMF for Stars
The first test is performed on NGC 2362, where the 271 stars have a range of
, see [22] and CDS catalog J/MNRAS/384/675/table1. According to [23], the distance of NGC 2362 is 1480 pc. The second test is performed on the low-mass IMF in the young cluster NGC 6611, see [24] and CDS catalog J/MNRAS/392/1034. This massive cluster has an age of 2 - 3 Myr and contains masses from
. Therefore, the brown dwarfs (BD) region,
, is covered. The third test is performed on the
Velorum cluster, where the 237 stars have a range of
, see [25] and CDS catalog J/A+A/589/A70/table5. The fourth test is performed on the young cluster Berkeley 59, where the 420 stars have a range of
, see [26] and CDS catalog J/AJ/155/44/table3. The fifth test is performed on the Hyades, where the 602 stars have a range of
, see [27] and CDS catalog J/AJ/165/108/table1.
The results are presented in Table 1 for the Benini distribution and Table 2 for the right truncated Benini distribution. In Table 1 and Table 2, the last column shows whether the results of the K-S test are better when compared to the lognormal distribution (Y) or worse (N).
As an example, the empirical DF visualized through histograms and the theoretical Benini DF for the
Velorum cluster are presented in Figure 4.
Another example is given by the PDF of the truncated Benini distribution, see Figure 5.
Table 1. Numerical values of
, AIC, probability Q, D, the maximum distance between theoretical and observed DF, and
, the significance level, in the K-S test of the Beninini distribution, see Equation (1), for different astrophysical environments. The last column (F) indicates a
higher (Y) or lower (N) than that for the lognormal distribution. The number of linear bins, n, is 10.
| Cluster |
parameters |
AIC |
|
Q |
D |
|
F |
| NGC 2362 |
,
,
|
134 |
18.38 |
1.16 × 10−24 |
0.196 |
1.11 × 10−9 |
N |
| NGC 6611 |
,
,
|
92.61 |
12.37 |
6.11 × 10−16 |
0.198 |
1.19 × 10−7 |
N |
| γ Velorum |
,
,
|
20.64 |
2.09 |
4 × 10−2 |
0.0372 |
0.89 |
Y |
| Berkeley 59 |
,
,
|
27.99 |
3.14 |
2.54 × 10−3 |
0.07 |
0.038 |
Y |
| Hyades |
,
,
|
19.89 |
1.98 |
5.3 × 10−2 |
0.043 |
0.2 |
Y |
Table 2. Numerical values of
, AIC, probability Q, D, the maximum distance between theoretical and observed DF, and
, the significance level, in the K-S test of the right truncated Beninini distribution, see Equation (21), for different astrophysical environments. The last column (F) indicates a
higher (Y) or lower (N) than that for the lognormal distribution. The number of linear bins, n, is 10.
| Cluster |
parameters |
AIC |
|
Q |
D |
|
F |
| NGC 2362 |
,
,
|
122.88 |
19.14 |
1.92 × 10−22 |
0.252 |
1 × 10−15 |
N |
| NGC 6611 |
,
,
|
83.177 |
12.52 |
3.52 × 10−14 |
0.261 |
6.09 × 10−13 |
N |
| γ Velorum |
,
,
|
22.48 |
2.41 |
2.46 × 10−2 |
0.042 |
0.779 |
Y |
| Berkeley 59 |
,
,
|
30.07 |
3.67 |
1.17 × 10−3 |
0.069 |
0.034 |
Y |
| Hyades |
,
,
|
22.47 |
2.41 |
2.47 × 10−2 |
0.056 |
0.387 |
Y |
Figure 4. Empirical DF of the mass distribution for γ Velorum cluster (blue histogram) with a superposition of the Benini DF (red dashed line). Theoretical parameters as in Table 1.
Figure 5. Empirical PDF of the mass distribution for Berkeley 59 (blue histogram) with a superposition of the truncated Benini PDF (red dashed line). Theoretical parameters as in Table 2.
4.7. Massive Stars
We analyse the more massive stars,
in the framework of the survival function (SF). We therefore analyse four different SFs:
1) the power law SF, see Equation (44),
2) the truncated power law SF, see Equation (47),
3) the lognormal SF, see Equations (38) or (41),
4) the right truncated Benini SF, see Equation (22).
The behaviour of the large masses,
, for Hyades is presented in Figure 6 and that for Hyades in Figure 7.
Figure 6. Survival function of NGC 6611 cluster data as
plot when
: data (empty circles), survival function of the truncated Pareto pdf (red full line) (
,
,
) and survival function of the Pareto pdf (green dashed line) (
, Salpeter slope −2.3). The lognormal (blue dot-dash-dot-dash line) and the SF of the truncated Benini distribution with parameters as in Table 2.
Figure 7. Survival function of Hyades data as
plot when
: data (empty circles), survival function of the truncated Pareto pdf (red full line) (
,
,
) and survival function of the Pareto pdf (green dashed line) (
, Salpeter slope −2.09). The lognormal (blue dot-dash-dot-dash line) and the SF of the truncated Benini distribution with parameters as in Table 2.
5. Conclusions
The truncated distribution
We derived the PDF, the DF, the average value, the kth moment about the origin, the median, a random number generator, and the MLE for the Benini distribution truncated on the right.
Application to the IMF
The application of the Benini distribution to the IMF for stars gives better results than the lognormal distribution for three out of five samples, see Table 1. The truncated Benini distribution does not improve the results over those of the regular Benini distribution for the five samples here considered, see Table 1 and Table 2.
The results for the mass distribution of the
Velorum cluster compared with other distributions are shown in Table 3, in which the truncated Benini distribution occupies the last position.
Table 3. Numerical values of D, the maximum distance between theoretical and observed DF, and
, the significance level, in the K–S test for different distributions in the case of
Velorum cluster.
| Distribution |
Reference |
D |
|
| Benini |
here |
0.0372 |
0.89 |
| Benini rigth truncated |
here |
0.042 |
0.779 |
| truncated Gompertz |
[28] |
0.173 |
9.27 × 10−7 |
| truncated Topp-Leone |
[29] |
6.09 × 10−2 |
0.25 |
| Frècet |
[30] |
0.125 |
3.13 × 10−4 |
| truncated Frècet |
[30] |
0.077 |
0.07 |
| truncated Weibull |
[31] |
0.046 |
0.576 |
| truncated Sujatha |
[32] |
0.0485 |
0.534 |
| truncated Lindley |
[33] |
0.11 |
0.48 |
| generalized gamma |
[34] |
0.11 |
1.24 × 10−3 |
| truncated generalized gamma |
[34] |
0.062 |
0.24 |
| lognormal |
[35] |
0.0729 |
0.11 |
| truncated lognormal |
[35] |
0.047 |
0.55 |
| gamma |
[36] |
0.059 |
0.28 |
| truncated gamma |
[36] |
0.0754 |
0.08 |
| beta |
[37] |
0.059 |
0.28 |
The most massive stars, see the SF reported in Figure 6 and Figure 7, are better modeled by the truncated distributions, right truncated Benini and truncated Pareto, when compared to the regular distributions, lognormal and Pareto.
Appendix
Implicit Formulas
The implicit formulae for the variance, skewness and kurtosis are
(A.1)
(A.2)
(A.3)
where
is the kth moment about the origin.