New Probability Distributions in Astrophysics: XVI. Truncation for the Kiang Distribution and the VAST Catalog for Cosmic Voids ()
1. Introduction
The VAST void catalog for SDSS DR7, see [1], computes the effective radii and centres of cosmic voids for two cosmologies with three different methods. This catalog has been used to explain the gamma-ray dark matter emission from halos of galaxies, see [2], and the calibration of the background universe, see [3] [4]. A useful quantity to model the cosmic voids is the spherically averaged stacked density profile
(1)
where
is the radius of the void,
is the spherical density of galaxies at
and
is the average density of galaxies of the catalog, see [1]. The ratio
is zero at
and reaches the maximum value of 1 at
. The availability of reliable data for the cosmic voids requires a precise analysis of the data. There are, roughly speaking, two methods of fitting data. One method adopts a generic probability distribution ad hoc, e.g. the lognormal distribution. The other one explores the probability distributions which arise from the particular problem analysed. In the case of cosmic voids, a useful tool is the distribution connected with the Voronoi diagrams, i.e. the Kiang distribution [5], which dates back to 1961. Two targets of the research on cosmic voids are the determination of the radius of the maximum sphere, which is ≈ 50 Mpc, and the evaluation of the radius of the minimum sphere which can be detected, which is ≈ 8 Mpc. The maximum radius is connected with the physical process which produces the voids, but the minimum radius is connected with the observational techniques. The presence of the two boundaries in the distribution of radii requires the introduction of the truncated PDFs.
In order to explore the two above distributions, Section 2 reviews the lognormal distribution and the other three modified lognormal distributions. Section 3 applies the lognormal family to the VAST catalog of voids. Section 4 reviews the Kiang function, its particularization to the radius of the Voronoi diagrams in 3D and introduces the truncation for both the distributions.
2. The Lognormal Family
In the following, PDF means probability density function and DF distribution function. We now review the lognormal distribution, the truncated lognormal distribution, the exponentiated lognormal distribution and the exponentiated lognormal distribution with truncation.
2.1. The Lognormal Distribution
Let
be a random variable taking values
in the interval
; the first definition for the lognormal PDF, following [6] or formula (14.2) in [7], is
(2)
The average value,
, is
(3)
and the distribution function,
, is given by
(4)
The second definition is
(5)
where
and
.
2.2. The Truncated Lognormal Distribution
Let
be a random variable defined in
; the truncated lognormal PDF (
) is based on the first definition of the lognormal as given by Equation (2)
(6)
where
is the scale parameter,
is the shape parameter,
denotes the minimal value, and
denotes the maximal value. The introduction of the following coefficients allows a compact notation:
In this compact notation, the PDF is
(7)
the DF is
(8)
and the mean,
, is
(9)
More details can be found in [8].
2.3. The Exponentiated Lognormal Distribution
The modified lognormal distribution with a power-law (MLP) defined in the interval
has the PDF
(10)
where
is the complementary error function
(11)
see the handbook [9]. Its DF is
(12)
see formula (16) in [10]. The first moment, or mean,
, is defined for
:
(13)
see formula (19) in [10].
The variance,
, is defined for
:
(14)
see formula (21) in [10]. The mode should be evaluated numerically. The random generation of the variate
of the truncated MLP is accomplished by solving the following nonlinear equation in
:
(15)
where
is the unit rectangular variate. More details can be found in [11].
2.4. Truncated MLP
The truncated version of the MLP defined in the interval
has the following PDF:
(16)
where
(17)
The DF is
(18)
The first moment, or average value, is
(19)
More details can be found in [11].
3. Astrophysical Applications
This section reviews the adopted statistics and the various catalogs for cosmic voids. The application of four types of lognormal to the last and more complete catalog of voids is reported.
3.1. Adopted Statistics
The Kolmogorov-Smirnov test (K-S), see [12]-[14], does not require the data to be binned. The K-S test, as implemented by the FORTRAN subroutine KSONE in [15], given ordered statistics
,
, where
defines the number of elements in the sample, finds the maximum distance,
, between the theoretical and the astronomical DFs. The significance is
(20)
and the significance level to have an observed
is
(21)
see formulas 14.3.5 and 14.3.9 in [15]. If
, then the goodness of the fit is believable.
3.2. Catalogs of Voids
The first catalog of cosmic voids can be found in [16], where the effective radius of the voids,
, has been derived to be
Pam et al. 2012.(22)
The second catalog is that with radii up to redshift
Mpc in (SDSS-DR7), see [17],
Varela et al. 2012.(23)
The third catalog is that of the Baryon Oscillation Spectroscopic Survey, see [18],
Mao et al. 2017.(24)
The fourth catalog is that of the VAST void catalog for SDSS DR7 which uses three algorithms, VoidFinder, V2/VIDE and V2/REVOLVER in the framework of two cosmologies, Planck2018 and WMAP5, see [1]. The data with Cartesian coordinates and effective radius
in Mpc are available at the following address https://zenodo.org/records/11043278. A graphical display of the voids is reported in Figure 1.
3.3. Statistics of the Voids
The statistics for the lognormal distribution of the VAST catalog are given in Table 1, for the truncated lognormal distribution in Table 2, for the MLP distribution in Table 3 and for the truncated MLP distribution in Table 4.
Figure 1. 3D display through spheres of the 1184 voids derived with the VoidFinder method in WMAP5 cosmology.
Table 1. Numerical values of
, the maximum distance between theoretical and observed DFs, and
, significance level, in the K-S test for the lognormal distribution, see Equation (5), for different cosmologies and methods.
Cosmology |
Method |
Parameters |
|
|
WMAP5 |
VoidFinder,
|
,
|
0.0652 |
7.84 × 10−5 |
Planck 2018 |
V2/VIDE,
|
,
|
0.0669 |
0.0163 |
Planck 2018 |
V2/REVOLVER,
|
,
|
0.0677 |
0.0162 |
Table 2. Numerical values of
, the maximum distance between theoretical and observed DFs, and
, significance level, in the K-S test for the truncated lognormal distribution, see Equation (7), for different cosmologies and methods.
Cosmology |
Method |
Parameters |
|
|
WMAP5 |
VoidFinder,
|
,
|
0.0551 |
1.05 × 10−21 |
Planck 2018 |
V2/VIDE,
|
,
|
0.0193 |
0.988 |
Planck 2018 |
V2/REVOLVER,
|
,
|
0.021 |
0.972 |
Table 3. Numerical values of
, the maximum distance between theoretical and observed DFs, and
, significance level, in the K-S test for the exponentiated lognormal distribution, see Equation (10), for different cosmologies and methods.
Cosmology |
Method |
Parameters |
|
|
WMAP5 |
VoidFinder,
|
,
,
|
0.0416 |
3.16 × 10−2 |
Planck 2018 |
V2/VIDE,
|
,
,
|
0.063 |
2.78 × 10−2 |
Planck 2018 |
V2/REVOLVER,
|
,
,
|
0.0665 |
1.94 × 10−2 |
Table 4. Numerical values of
, the maximum distance between theoretical and observed DFs, and
, significance level, in the K-S test for the exponentiated and truncated lognormal distribution or truncated MLP see Equation (16), for different cosmologies and methods.
Cosmology |
Method |
Parameters |
|
|
WMAP5 |
VoidFinder,
|
,
,
|
0.0363 |
8.57 × 10−2 |
Planck 2018 |
V2/VIDE,
|
,
,
|
0.0211 |
0.969 |
Planck 2018 |
V2/REVOLVER,
|
,
,
|
0.022 |
0.961 |
We now report the best results for the three cases here analysed as empirical and theoretical PDFs: Figure 2 reports the WMAP5 and VoidFinder data in the case of the truncated MLP distribution, Figure 3 reports the Planck 2018 and V2/VIDE data in the case of the truncated lognormal distribution and Figure 4 reports the Planck 2018 and V2/REVOLVER data in the case of the truncated lognormal distribution.
Figure 2. Logarithmic histogram of distribution of voids as given by WMAP5 and VoidFinder data (red) with a superposition of the truncated MLP distribution when the number of bins, m, is 30 (green line). Parameters as in Table 4. Vertical and horizontal axes have logarithmic scales.
Figure 3. Logarithmic histogram of distribution of voids as given by Planck 2018 and V2/VIDE data (red) with a superposition of the truncated lognormal distribution when the number of bins, m, is 30 (green line). Parameters as in Table 2. Vertical and horizontal axes have logarithmic scales.
Figure 4. Logarithmic histogram of distribution of voids as given by Planck 2018 and V2/REVOLVER data (red) with a superposition of the truncated lognormal distribution when the number of bins, m, is 30 (green line). Parameters as in Table 2. Vertical and horizontal axes have logarithmic scales.
4. Voronoi Diagrams
This section first deals with the generic Kiang distribution and then particularizes it to the distribution in radius. The effect of truncation on both of these two distributions is reported.
4.1. Generic Distribution
Let
be a random variable taking values
in the interval
; the gamma PDF is
(25)
where
(26)
is the gamma function,
is the scale parameter and
is the shape parameter, see formula (17.23) in [7]. Its expected value is
(27)
We now re-scale the above PDF in such a way that the expected value is 1 and the gamma variate,
, ([5]) is obtained
(28)
where
,
. The DF of the Kiang function is
(29)
where
is the incomplete Gamma function defined as
(30)
see [9]. The Kiang PDF has a mean,
, of
(31)
the variance is
(32)
the skewness is
(33)
the kurtosis is
(34)
and the mode is at
(35)
An approximate expression for the median can be obtained by an order 3 Taylor series for the DF about the mean
(36)
The percent error of the above median is 0.013% at
.
In the case of a 1D Poissonian Voronoi tessellation (PVT),
is an exact analytical result, but
is supposed to be 4 or 6 for 2D or 3D PVTs, respectively, the so called Kiang conjecture [5]. The value
in 3D was successively refined to 5.5 due to a change in the generator of random numbers [19], or to 5.78 as deduced by [20]. A first method to derive the numerical value of
is to equalize the variance of the sample, var, with the theoretical variance
(37)
A second method to derive the numerical value of
is the maximum likelihood (MLE) method, which maximizes
(38)
where
is the number of elements in the sample
. The value of
is found solving the following non-linear equation
(39)
where
is the digamma or Psi function defined as
(40)
where
, see [9].
Figure 5 reports a typical evaluation of the volume distribution of the Voronoi diagrams in the presence of Poissonian seeds.
Figure 5. Histogram of volume distribution for the Voronoi diagrams generated by
. Poissonian seeds (blue line) with a superposition of a red dashed line representing the Kiang PDF as given by Equation (28) when
as deduced with the first method. The two parameters of the K-S test are
and
.
4.2. Radius Distribution
We approximate the volume of each cell of the Voronoi diagrams,
, with
(41)
The Kiang function in volumes is
(42)
and that in radius,
,
(43)
where
is the variable to be found from the data of the volumes. We now introduce the scale
: the scaled PDF is
(44)
and the scaled DF is
(45)
The above distribution has a mean
(46)
variance
(47)
skewness
(48)
and kurtosis
(49)
where
(50)
and
(51)
An approximate expression for the median can be obtained by an order two Taylor series for the DF about the mean:
(52)
One way to deduce the parameters associates
with the maximum of the sample;
is found by solving the non-linear Equation (37) with the theoretical variance as given by Equation (47). Another way, the MLE method, maximizes
(53)
The values of
and
are found by solving the two following non-linear equations:
(54)
and
(55)
The statistics for the Kiang function in radius of the VAST catalog are given in Table 5.
Table 5. Numerical values of
, the maximum distance between theoretical and observed DFs,
, significance level, in the K-S test for the Kiang distribution in radius, see Equation (44), and percent error for the approximated median, see Equation (52), for different cosmologies and methods.
Cosmology |
Method |
Parameters |
|
|
% error median |
WMAP5 |
VoidFinder,
|
Mpc
|
0.103 |
1.73 × 10−11 |
4.37% |
Planck 2018 |
V2/VIDE,
|
Mpc
|
0.129 |
3.41 × 10−8 |
11.57% |
Planck 2018 |
V2/REVOLVER,
|
Mpc
|
0.019 |
6.3 × 10−5 |
5.87% |
As a visual example, Figure 6 reports the Planck 2018 and V2/REVOLVER data in the case of the Kiang distribution in radius.
Figure 6. Logarithmic histogram of distribution of voids as given by Planck 2018 and V2/REVOLVER data (red) with a superposition of the Kiang distribution in radius when the number of bins, m, is 30 (green line). Parameters as in Table 5. Vertical and horizontal axes have logarithmic scales.
4.3. The Truncated Generic Distribution
The truncated gamma variate
is
(56)
where
,
.
The truncated distribution function is
(57)
The truncated Kiang PDF has a mean,
, given by
(58)
and the
th moment about the origin is,
, is
(59)
The variance is
(60)
One way to find
uses Equation (37) in which
is given by the previous equation. Another way to find
is the MLE method, which maximizes
(61)
In this case,
has a complicated expression which is not reported. As an application, Figure 7 reports a typical evaluation of the volume distribution of the Voronoi diagrams in the presence of Poissonian seeds.
Figure 7. Histogram of volume distribution for the Voronoi diagrams generated by
. Poissonian seeds (blue line) with a superposition of a red dashed line representing the truncated Kiang PDF as given by Equation (56) when c = 5.385 as deduced with the second method. The two parameters of the K–S test are
and
.
4.4. Radius Distribution with Truncation
We now introduce the truncation in the radius distribution with scale, see Equation (44). The truncated PDF is
(62)
where
,
,
and
(63)
The distribution function is
(64)
and the average value
(65)
The median can be found by numerically solving the equation
(66)
The statistics for the truncated Kiang function in radius of the VAST catalog are given in Table 6, the high values of the probability do not mean overfitting or limitations in the data-model fit. Figure 8 reports the Planck 2018 and V2/REVOLVER data.
5. Conclusions
Lognormal family
We tested the progressive increase in the number of parameters for the lognormal family here represented by the lognormal, the truncated lognormal, the modified lognormal and the modified lognormal with truncation on the radius distribution for cosmic voids as given by the VAST catalog. This increase produces progressive best fits, see the last column in Tables 1-4.
Table 6. Numerical values of
, the maximum distance between theoretical and observed DFs,
, significance level, in the K-S test for the truncated Kiang distribution in radius, see Equation (62), and percent error for the numerical median, see Equation (66), for different cosmologies and methods.
Cosmology |
Method |
Parameters |
|
|
% error median |
WMAP5 |
VoidFinder,
|
Mpc
|
0.083 |
1.08 × 10−7 |
3.025% |
Planck 2018 |
V2/VIDE,
|
Mpc
|
0.051 |
0.121 |
4.27% |
Planck 2018 |
V2/REVOLVER,
|
Mpc
|
0.0192 |
0.99 |
0.623% |
Figure 8. Histogram of distribution of voids as given by Planck 2018 and V2/REVOLVER data (blue) with a superposition of the truncated Kiang distribution in radius when the number of bins, m, is 30 (red dashed line). Parameters as in Table 6. Vertical and horizontal axes have linear scales.
Kiang distribution
The introduction of a truncation in the Kiang distribution in radius increases the reliability of the fits, see the last column in Table 5 and Table 6. We also derived an approximation for the median of the Kiang distribution, see Equation (36), and for the Kiang distribution in radius, see Equation (52). The results for the VAST catalog of the median in terms of percent error are reported in the last column of Table 6. The positive results for the Kiang distribution require a refinement of the model, i.e. the stereological properties of the Voronoi diagrams.
Comparison of the two distributions
The modified lognormal distribution with truncation for the radius distribution in cosmic voids produces better results for two of the three cases analysed. The Kiang distribution in radius with truncation reports a surprising
which is the highest level of significance for the VAST catalog, see Table 4 and Table 6.
Maximum void
The Boötes void, sometimes called ‘the great nothing’ or great void, was discovered in 1987. Its radius is 62 Mpc [21], which is the maximum at the moment of writing. In our framework, this can be identified with the upper radius,
, of the truncated Kiang distribution in radius, see formula 63.