Kumaraswamy-Adaptive Normal Kernel Densities for Robust Smoothing of Skewed, Outlier-Contaminated Data ()
1. Introduction
Kernel Density Estimation (KDE) is a fundamental non-parametric method for smoothing and estimating unknown probability density functions. KDE constructs a continuous density approximation by averaging localized kernel contributions scaled by a bandwidth parameter ℎ. The statistical foundations of KDE originate in Rosenblatt [1] and Parzen [2].
The Epanechnikov kernel, introduced by Epanechnikov [3], is known to be optimal in a mean squared error (MSE) sense among compact-support kernels. Subsequent work by Silverman [4] and Scott [5] established KDE as a central tool in applied statistics. Further refinements in bandwidth selection were developed by Jones, Marron, and Sheather [6], and later expanded by Sheather [7].
Despite these advances, classical kernels remain fixed-shape smoothers and may perform poorly in the presence of skewness and outlier contamination. This limitation motivates the development of flexible kernel constructions capable of adapting to non-standard data features.
The present work builds upon earlier contributions by Adepoju et al. [8], where transformation-based approaches and robustness under contamination were explored. In particular, the development of exponentiated test statistics in the presence of outliers [8] and the introduction of the Kumaraswamy Fisher-Snedecor distribution [9] provide a conceptual foundation for incorporating Kumaraswamy-based transformations into kernel density estimation. These prior studies motivate the current proposal of a Kumaraswamy-Adaptive Normal kernel designed to enhance robustness and flexibility in density smoothing.
2. Literature Review of Existing Kernel Densities
Before introducing the proposed adaptive kernel, it is important to situate the work within the broader framework of classical kernel density estimation. Over the years, several kernel functions have been developed and studied extensively, each possessing distinct smoothness, support, and efficiency characteristics. Although many kernels share similar asymptotic properties, their finite-sample performance may differ, particularly under skewness and contamination.
This section briefly reviews the most commonly used kernel functions in the literature. These standard kernels serve as benchmarks for comparison with the proposed Kumaraswamy-Normal kernel and provide a foundation for understanding its relative advantages.
Let
denote the standardized distance between the evaluation point
and the observation
, where
is the bandwidth parameter. The general kernel density estimator is given by
(1)
2.1. Gaussian Kernel
(2)
2.2. Epanechnikov Kernel
(3)
2.3. Uniform Kernel
(4)
2.4. Triangular Kernel
(5)
2.5. Biweight Kernel
(6)
2.6. Triweight Kernel
(7)
2.7. Cosine Kernel
(8)
2.8. Logistic Kernel
(9)
2.9. Sigmoid Kernel
(10)
3. Kumaraswamy-Normal Kernel
3.1. Definition
Let
(11)
Apply the Kumaraswamy transformation:
(12)
The induced kernel density is
(13)
when
, the Gaussian kernel is recovered.
3.2. Motivation for the Kumaraswamy Link
The Kumaraswamy link function is an established asymmetric transformation in the statistical literature, well known for its flexibility through its two shape parameters
and
. By varying these parameters, the function can generate left-skewed, right-skewed, and heavy-tailed distributions.
When the underlying density deviates from normality, purely symmetric kernels may introduce bias in the estimation process. Embedding the Gaussian kernel within the Kumaraswamy link provides a mechanism for introducing controlled asymmetry while preserving the smoothness and analytical tractability of the Normal kernel.
Thus, the Kumaraswamy-Normal kernel retains the stability of the Gaussian kernel while adapting to skewness and complex tail behavior, making it particularly suitable for density estimation in the presence of non-standard data structures.
4. Kumaraswamy-Normal Kernel
4.1. Definition
Let
(14)
denote respectively the standard normal density and cumulative distribution function.
Applying the Kumaraswamy transformation to the standard normal distribution gives
(15)
Differentiating with respect to
, the corresponding Kumaraswamy-Normal density is
(16)
when
, we recover the classical Gaussian kernel:
(17)
Thus, the Kumaraswamy-Normal kernel extends the Gaussian kernel through two shape parameters
and
, allowing greater flexibility in handling skewness and tail behavior.
4.2. Likelihood Function of the Kumaraswamy-Normal Distribution
Suppose
is a random sample from the Kumaraswamy-Normal distribution with parameters
and
. Then the joint likelihood function is
(18)
Hence,
(19)
Taking logarithms yields the log-likelihood function
(20)
Since
does not depend on
or
, it is constant for optimization purposes.
4.3. Score Equations
The score vector is
(21)
Derivative with respect to
Differentiating
with respect to
gives
(22)
Now,
(23)
Therefore,
(24)
Derivative with respect to
Differentiating
with respect to
gives
(25)
Hence, the likelihood equations are
(26)
and
(27)
From the second equation,
(28)
Thus,
and
are obtained by solving the nonlinear score equations numerically.
4.4. Hessian Matrix
The Hessian matrix is defined by
(29)
Second derivative with respect to
We have
(30)
Let
(31)
Then
(32)
Therefore,
(33)
Second derivative with respect to
(34)
Mixed derivative
Differentiating
with respect to
, we obtain
(35)
Hence,
(36)
Thus, the Hessian matrix becomes
(37)
4.5. Maximum Likelihood Estimation
The maximum likelihood estimators
are the values satisfying
(38)
Because the score equations are nonlinear, closed-form solutions do not generally exist for
and
. Therefore, iterative procedures such as the Newton-Raphson algorithm are employed:
(39)
Iteration continues until convergence.
4.6. Asymptotic Normality of the MLE
Under standard regularity conditions for maximum likelihood estimation, the MLE
(40)
is consistent and asymptotically normal. Specifically,
(41)
where
(42)
and
is the Fisher information matrix defined by
(43)
Equivalently,
(44)
In practice, the covariance matrix of
can be estimated using the observed information matrix
(45)
4.7. Estimation of the Kumaraswamy-Transformed Normal Kernel
Let
be a random sample from an unknown density
. The Kumaraswamy–transformed Normal kernel density estimator is defined as
(46)
where
is the bandwidth.
Substituting the kernel form,
(47)
Hence, the final estimator depends on three unknown quantities:
A practical estimation procedure is as follows:
1. Select an initial bandwidth
using a classical method such as Silverman’s rule of thumb, plug-in estimation, or cross-validation.
2. Estimate the shape parameters
and
by maximum likelihood.
3. Substitute
,
, and
into the estimator
(48)
Thus, the proposed kernel estimator generalizes the Gaussian kernel estimator by incorporating adaptive shape parameters that respond to skewness and contamination.
4.8. Bandwidth Estimation
A simple bandwidth choice is Silverman’s rule of thumb:
(49)
where
is the sample standard deviation.
Alternatively,
may be chosen by least-squares cross-validation:
(50)
where
denotes the leave-one-out estimator.
4.9. Bias and Variance of the Proposed Kernel Estimator
Let
(51)
be the second moment of the Kumaraswamy-Normal kernel, and let
(52)
Under the usual smoothness assumptions on
, the bias of the estimator is
(53)
Its variance is approximately
(54)
Therefore, the mean squared error is
(55)
Hence, the asymptotic mean integrated squared error is
(56)
Minimizing the AMISE with respect to
yields the asymptotically optimal bandwidth
(57)
Thus, the proposed Kumaraswamy-Normal kernel estimator retains the classical
bandwidth rate while allowing enhanced adaptability through the parameters
and
.
4.10. Special Case: Reduction to the Gaussian Kernel
When
,
(58)
and therefore the estimator reduces to the standard Gaussian kernel estimator:
(59)
Hence, the proposed estimator is a genuine extension of the Gaussian kernel density estimator.
5. Simulation Study
5.1. Mathematical Description of the Simulation Design
For each Monte Carlo replication, data were generated from a finite mixture of
lognormal components:
(60)
where
(61)
The mixture weights satisfy
(62)
5.2. Parameter Generation for Each Monte Carlo Run
For each replication
:
(63)
(64)
This guarantees strictly positive, highly skewed, multimodal densities.
5.3. Outlier Contamination Mechanism
Observed data were generated from
(65)
with heavy-tailed contaminant
(66)
(67)
5.4. Estimation of Shape Parameters
The shape parameters
and
were estimated separately for each experiment rather than fixed globally.
For each replication and sample size
,
(68)
This allows full adaptation to each dataset.
5.5. Monte Carlo Procedure
For each
(69)
500 datasets were generated from
.
The integrated squared error (ISE) was computed as
(70)
The estimator with the smallest ISE was recorded as the winner.
6. Results and Discussion
Kernel comparison at
Table 1. Kernel smoothing performance at
.
Kernel |
Win-Rate |
Mean ISE |
KwNormal |
0.490 |
0.03643 |
Gaussian |
0.180 |
0.03701 |
Triangular |
0.105 |
0.03763 |
KwEpan |
0.100 |
0.05164 |
Rectangular |
0.055 |
0.04204 |
Epanechnikov |
0.070 |
0.03900 |
Biweight |
0.000 |
0.03824 |
Cosine |
0.000 |
0.03800 |
Kernel comparison at
Table 2. Kernel smoothing performance at
.
Kernel |
Win-Rate |
Mean ISE |
KwNormal |
0.480 |
0.02537 |
Gaussian |
0.250 |
0.02595 |
Triangular |
0.090 |
0.02762 |
KwEpan |
0.095 |
0.03532 |
Epanechnikov |
0.050 |
0.02650 |
Biweight |
0.005 |
0.02701 |
Cosine |
0.005 |
0.02682 |
Rectangular |
0.025 |
0.02993 |
Kernel comparison at
Table 3. Kernel smoothing performance at
.
Kernel |
Win-Rate |
Mean ISE |
KwNormal |
0.615 |
0.01587 |
Gaussian |
0.195 |
0.01808 |
KwEpan |
0.115 |
0.02086 |
Triangular |
0.035 |
0.01863 |
Epanechnikov |
0.020 |
0.01957 |
Biweight |
0.005 |
0.01906 |
Cosine |
0.010 |
0.01889 |
Rectangular |
0.005 |
0.02161 |
Kernel comparison at
Table 4. Kernel smoothing performance at
.
Kernel |
Win-Rate |
Mean ISE |
KwNormal |
0.630 |
0.01076 |
Gaussian |
0.190 |
0.01241 |
KwEpan |
0.120 |
0.01371 |
Triangular |
0.035 |
0.01276 |
Epanechnikov |
0.010 |
0.01342 |
Biweight |
0.000 |
0.01307 |
Rectangular |
0.015 |
0.01504 |
Cosine |
0.000 |
0.01296 |
Kernel comparison at
Table 5. Kernel smoothing performance at
.
Kernel |
Win-Rate |
Mean ISE |
KwNormal |
0.605 |
0.00697 |
Gaussian |
0.220 |
0.00808 |
KwEpan |
0.140 |
0.00892 |
Triangular |
0.025 |
0.00833 |
Epanechnikov |
0.005 |
0.00879 |
Biweight |
0.000 |
0.00855 |
Cosine |
0.005 |
0.00847 |
Rectangular |
0.000 |
0.01035 |
Kernel comparison at
Table 6. Kernel smoothing performance at
.
Kernel |
Win-Rate |
Mean ISE |
KwNormal |
0.680 |
0.00477 |
Gaussian |
0.155 |
0.00592 |
KwEpan |
0.150 |
0.00602 |
Triangular |
0.015 |
0.00607 |
Epanechnikov |
0.000 |
0.00645 |
Biweight |
0.000 |
0.00627 |
Cosine |
0.000 |
0.00622 |
Rectangular |
0.000 |
0.00837 |
7. Kernel Comparison Plot
Tables 1-6 present the comparison of mean integrated squared error (ISE) across different sample sizes for all competing kernel estimators. As the sample size increases, the ISE decreases for all kernels, reflecting the expected improvement in estimation accuracy. Figure 1 is the corresponding plot.
However, the Kumaraswamy-Normal (KwNormal) kernel consistently achieves the lowest ISE across all sample sizes. This indicates its superior ability to adapt to skewness and heavy-tailed contamination in the data. Unlike classical symmetric kernels, the KwNormal kernel incorporates flexible shape parameters that allow it to adjust to the underlying structure of the distribution.
Furthermore, the gap between the KwNormal kernel and traditional kernels such as the Gaussian and Epanechnikov kernels becomes more pronounced as the sample size increases. This highlights not only its robustness but also its scalability for larger datasets.
These results visually confirm the findings from the simulation tables, demonstrating that the proposed Kumaraswamy-Normal kernel provides a stable and consistently superior performance in density estimation.
Figure 1. Comparison of mean integrated squared error (ISE) across sample sizes for different kernel estimators.
8. Conclusions
This study introduced the Kumaraswamy-Normal kernel as an adaptive extension of the classical Gaussian kernel for density estimation. By incorporating two shape parameters, the proposed kernel is able to effectively capture skewness and heavy-tailed behavior in data.
Theoretical analysis showed that the estimator retains desirable asymptotic properties, including consistency, asymptotic normality, and the optimal bandwidth rate. Simulation results further demonstrated that the KwNormal kernel consistently outperforms traditional kernels in terms of win-rate and mean ISE across all sample sizes.
Overall, the proposed kernel provides a flexible and robust alternative for non-parametric density estimation, particularly in the presence of skewed and contaminated data. This makes it a valuable tool for practical applications where classical kernel methods may be inadequate.