General Depth-Based Trimmed Means and Trimmed Scatter Matrices ()
1. Introduction and Preliminaries
Multivariate descriptive measures for location and scatter are the foundation of multivariate statistics and underpin many methods in the field. The classical notions of multivariate location and scatter are moment-based and are estimated by the sample mean vector and the sample covariance matrix. However, these estimators are not robust and are extremely sensitive to outlier(s). Many efforts have been made to solve this problem, and various estimators for multivariate location and scatter have been proposed. The widely used ones include the M-estimators [1], the minimum volume ellipsoid (MVE) and the minimum covariance determinant (MCD) estimators [2], the S-estimators [3] [4], the
-estimators [5], the constrained M-estimators [6], and so on.
In addition, various nonparametric estimators for multivariate location and scatter have been developed via depth functions. Generally, a depth function
is a nonnegative real-valued mapping which provides a distribution-based center-outward ordering of points
in
. Desirable properties for a depth function are: 1) affine invariance (i.e.,
for any nonsingular
matrix
and
-vector
); 2) maximality at “center”; 3) monotonicity relative to the deepest point; 4) vanishing at infinity. Widely used depth functions include the halfspace depth [7], the projection depth [8] [9], the simplicial depth [10], the
depth [11], and the Mahalanobis depth [8] [12]. The general discussions on depth functions and their applications can be found in [11] [13] [14].
Given a depth function
, the deepest point plays the role of “center” of the distribution
and thus is considered as a multivariate median, denoted by
. The
depth inner region is defined as
,
, and the pth central region as
with
, i.e., the smallest
-depth inner region having a probability weight of at least
,
. The boundary
of
is called the
depth contour and can be characterized by the radius function,
,
when
is the origin. If the underlying distribution
is continuous, the random depth
is continuous for all the depths mentioned above, and then
. Therefore,
and
.
Utilizing a depth function
, one can define a depth weighted mean as
and a depth weighted scatter matrix as
where
and
are suitable weight functions and can be different. Conceptionally, the general depth weighted means and scatter matrices include depth trimmed means and trimmed scatter matrices as a special case. However, when studying the properties of the depth weighted means and scatter matrices, some conditions on the weight functions were invoked such that the usual trimmed means and scatter matrices were excluded. For example, to establish the asymptotic distributions of the sample version of
, Dümbgen [15], Massé [16], and Zuo, Cui, and He [17] all made the assumption that the weight function
is continuously differentiable, which excluded the usual trimmed mean for which the weight function is an indicator function. For the finite sample breakdown point of the sample version of
, Zuo and Cui [18] assumed that both weight functions are positive, which ruled out any trimmed scatter matrices.
It is well-known that the univariate trimmed means are not only robust but also highly efficient (much more efficient than the univariate median). It is conjectured that multivariate trimmed means and scatter matrices have the same properties, which motivates this work. In fact, the
depth trimmed means based on the projection depth have been studied by Zuo [19], and the ones based on the halfspace depth have been studied by Donoho and Gasko [20], Massé [21], and Wang [22], with the
depth trimmed means defined as
However, since the depth value of a point depends on both the depth function and the underlying distribution, it is very hard to choose an appropriate
value for
in practice, and thus their applications are thwarted. Table 1 lists the theoretical trimming proportions
for the
based on the halfspace depth with
when the underlying distributions are the
-dimensional standard normal distribution
, the
-dimensional
distribution with 5 degrees of freedom
, and the
-dimensional spherically symmetric uniform distribution
, respectively. The formula for the calculation is:
, where random vector
follows a spherically symmetric distribution centered at
,
denotes the Euclidean norm, and
is its univariate marginal quantile function. The details for the formula can be seen from the proof of Corollary 1 and the following remark.
![]()
Table 1. The theoretical trimming proportions
for the
based on the halfspace depth with
when the underlying distributions are
,
, and
, respectively.
From the table, we see that the theoretical trimming proportion
varies enormously from distribution to distribution. It is even harder to choose the
value for many practical problems, in which the underlying distribution is unknown. Therefore, it is necessary to restudy multivariate trimmed means. As for multivariate trimmed scatter matrices, we have not seen any investigation in the literature.
In this paper, we propose and study some new general depth-based trimmed means and trimmed scatter matrices. Their definitions, basic properties, and algorithms are given in Section 2. The asymptotic distributions of their sample versions are established in Section 3. Using the asymptotic distributions, we study the asymptotic efficiencies of the sample trimmed means and the sample trimmed scatter matrices in Section 4. The results show that both the trimmed means and trimmed scatter matrices are highly efficient. Section 5 is devoted to the robustness of the sample trimmed means and the sample trimmed scatter matrices, which is explored through finite-sample breakdown point. The proofs of main results are reserved for the Appendix.
Throughout this paper, matrices are denoted by bold letters and vectors by bold italicized letters. We use uppercase letters to denote distribution functions and their lowercase counterparts to denote density functions if they exist. For example, we denote by
and
the cdf and density of a random vector
in
. When
is a random variable, the quantile function of
is denoted by
. Without confusion, we will omit the subscript. The indicator function of a set
is denoted by
and the transpose of a matrix or a vector
is denoted by
. For an
matrix
,
denotes the
-dimensional vector formed by stacking the columns of
, and
denotes the Kronecker product of matrices
and
.
2. Definitions, Basic Properties, and Algorithms
Let
and
be suitable weight functions on
, where
. To reduce or eliminate the impacts of outliers, we assume that
and
are nondecreasing. Given
and
in
, we define the general depth-based trimmed mean of trimming proportion
as
and define the general depth-based trimmed scatter matrix of trimming proportion
using the location measure
as
As a special case,
and
. Furthermore,
if
on
, and
if
on
. In general, if
and
,
and
are much more broadly defined than
and
, respectively. They do not require existence of any moments.
Remark. Unlike
and
, we separate weighting and trimming in
and
with trimming controlled by
and
, and thus weight functions are not affected by trimming. For example, the weight function for the usual trimmed mean can be
on
, which is continuously differentiable.
For simplicity, we will confine attention to the case
and
, and call
and
the
proportion or
general depth-based trimmed mean and trimmed scatter matrix, respectively. Given a sample
, the sample versions of
,
,
,
,
, and
are obtained by replacing
with
, the empirical distribution function of
, and are denoted by
,
,
,
,
, and
, respectively. For convenience, we will suppress
in
,
,
, and
in the following discussions unless it is necessary for clarity.
,
, and their sample versions have the following basic properties. The proofs are relatively straightforward and thus omitted.
Theorem 1. Suppose that the depth function
is affine invariant. Then we have the following results:
1)
and
are affine equivariant, that is,
and
for any nonsingular
matrix
and
-vector
, and so are their sample versions.
2) If
is centrally symmetric about
and has the first moment, then for any
,
(i.e.,
is Fisher consistent for
) and
.
3) If
follows an elliptically symmetric distribution
and has the second moment, then
and
, where
and
are some positive constants.
Computation of
and
is quite easy. The following are the algorithms by R.
1) Compute the depths of all points in the sample. Several R packages are available for depth-based statistical analysis. For the halfspace depth, the projection depth, the
depth, and the Mahalanobis depth, one can use the command “depth” in the R package “DepthProc” of Zawadzki, Kosiorowski, Slomczynski, Bocian, and Wegrzynkiewicz [23].
2) Order the data points according to their depths decreasingly, denoted by
, and get the points in
,
.
3) Compute
as
and compute
as
.
3. Asymptotic Distributions
To use an estimator or a statistic for inferences, its distribution or asymptotic distribution is necessary. To establish the asymptotic distributions of
and
, we invoke the following conditions.
is affine invariant, and
uniformly in
, where
;
Then, by Theorem 1 1), both
and
are invariant for any translation. Thus, we can assume without loss of generality that the origin is the deepest point.
has a continuous partial derivative with respect to
, and
with
being continuous in
;
There exists a Donsker class
of sets such that
and
almost surely for any
and
;
has a positive continuous density
in some neighborhood of
;
is continuous and
for
with
.
Remark. Uniform weak convergence of
has been established for the simplicial depth by Dümbgen [15] and Arcones and Giné [24], for the halfspace depth by Massé [16], for the projection depth and some generalized halfspace depth by Arcones, Cui, and Zuo [25], and thus
is satisfied for these depth functions. For a fixed
, the asymptotic representation of
was derived for the halfspace depth by Nolan [26] and for the projection depth by Zuo [19]. Considering
there as a variable, we see that
is satisfied by these depth functions. As for Condition
, we know that
is an ellipsoid for the Mahalanobis depth, a convex set formed by halfspaces for the halfspace depth and projection depth, and a union of some convex sets formed by halfspaces for the simplicial depth. Since {all ellipsoids in
} is a Vapnik-Červonenkis (VC) class [27] and so is {all sets in
formed by halfspaces} [24] [28] [29]. Condition
is satisfied for these depths.
Remark. It can be shown that every Donsker class is a Glivenko-Cantelli class almost surely [27].
Define
. Then we have the following results.
Theorem 2. Under the conditions
-
, for any
1)
, and thus
where
with
,
,
, where
,
where
and
is the Jacobian of
the transformation
with
and
, and
where
.
2)
, and thus
where
with
,
,
where
.
Both
and
are quite complicated in general. However, the results are relatively simple when
on
and
is an elliptically symmetric distribution. As an example, we concretize
and
for the halfspace depth-based trimmed means and scatter matrices for the very important special case. A continuous random vector
in
is said to have an elliptically symmetric distribution, denoted by
, if it has a density of the form
for a nonnegative function
with
and a positive definite matrix
. If the first moment of
exists,
. If the second moment of
exists,
, where
denotes the Euclidean norm. When
, the
identity matrix,
is said to have a spherically symmetric distribution with center
. If
,
. Let
. Then
has a uniform distribution on
, and thus
and
. In addition,
and
are independent. See Fang, Kotz, and Ng [30] for broad discussion about elliptically symmetric distributions.
Because the trimming based on the halfspace depth is a natural extension of the usual univariate trimming, we choose the halfspace depth here. The halfspace depth is first introduced by Tukey [7] and is defined as
where
= {all closed halfspaces}. As a leading depth function, the halfspace depth has been widely used. The finite-sample breakdown point of the halfspace depth was obtained by Donoho and Gasko [20]. Romanazzi [31] derived the influence function of
. The asymptotic distribution of its sample version
was established by Massé [16]. Nolan [26] studied the asymptotics of the sample
halfspace depth inner regions. An important feature of the trimming based on the halfspace depth is that it trims off the same proportion of data points in a closed halfspace at each direction normal to the boundary of
and outward with respect to
.
Corollary 1. Suppose that
with a positive continuous density
in a small neighborhood of
and
on
. Let
and
be the univariate marginal density and the marginal distribution function of
. If the trimming is based on the halfspace depth, then for any
and
.
1)
, where
with
, and
.
2)
, where
with
,
,
and
being a
-block matrix with
-block being
, which is a
matrix with 1 at entry
and 0 elsewhere.
4. Asymptotic Relative Efficiencies
When we evaluate the performance of an estimator, two important criteria are efficiency and robustness. Are
and
efficient? We will answer this question in this section. It is well-known that the sample mean vector
and the sample covariance matrix
are the most efficient unbiased estimators for the population mean vector
and covariance matrix
, respectively, when the population distribution is
. It is natural to compare
with
and
with
for efficiency.
4.1. Asymptotic Relative Efficiencies of the Sample Trimmed Means
Given the asymptotic distribution of
, it is straightforward to compute its asymptotic efficiency. For comparison, we compute the asymptotic relative efficiency (ARE) of the
based on the halfspace depth with
on
with respect to
for some elliptically symmetric distributions
,
where
and
are the asymptotic covariance matrices of
and
, respectively. Noting that
and
by Corollary 1, where
, we have
When the population distribution
is
, the results are reported in Table 2. It shows that
decreases as the dimension
or the trimming proportion
increases. Overall,
is efficient. Zuo [19] computed the asymptotic relative efficiencies of the halfspace depth-induced median (HM) and the projection depth-induced median (PM) with respect to
for
, which are 0.76 and 0.77, respectively. Even if the trimming proportion
, the trimmed mean is much more efficient than HM and PM. Table 3 lists the results for the
-dimensional
distribution with 5 degrees of freedom
. Again,
decreases as the dimension
increases for each trimming proportion
. The striking finding is that
increases first and then decreases as the trimming proportion
increases for each
. In addition,
is more efficient than
when
.
![]()
Table 2.
, where
is based on the halfspace depth with
on
, when the underlying distribution
is
.
![]()
Table 3.
, where
is based on the halfspace depth with
on
, when the underlying distribution
is
.
4.2. Asymptotic Relative Efficiencies of the Sample Trimmed Scatter Matrices
For the asymptotic efficiency of
, we will focus on how well it estimates some shape component of
. Generally speaking, any function
of a scatter matrix
can be considered as a shape component of
if
is invariant for a positive scalar multiple, i.e.,
for any
. See Tyler [32] and Kent and Tyler [6] for detailed discussions. For efficiency studies, the shape component of most interest is the nonsphericity of
and the measure most used for nonsphericity is the function
originating from the likelihood ratio test statistic for nonsphericity,
which was first derived for multivariate normal populations with
by Mauchly [33]. Muirhead [34] showed that if
is an elliptically symmetric distribution
with marginal kurtosis
and the null hypothesis
is true,
For
, we have the following result.
Theorem 3. Suppose that the conditions of Corollary 1 hold and the trimming is based on the halfspace depth. Then, if
and
,
where
is given in Corollary 1 and
.
Based on this result and the result of Muirhead [34], the asymptotic relative efficiency of the
based on the halfspace depth with
on
with respect to
is given as
We compute
for
and
, which are listed in Table 4 and Table 5, respectively. When the population distribution
is
,
decreases as the trimming proportion
increases for each dimension
. However, it increases first and then decreases as dimension
increases for each trimming proportion
. Overall,
is efficient.
![]()
Table 4.
, where
is based on the halfspace depth with
on
, when the underlying distribution
is
.
![]()
Table 5.
, where
is based on the halfspace depth with
on
, when the underlying distribution
is
.
When the population distribution
is
,
increases first and then decreases as trimming proportion
increases for each dimension
. Meanwhile, it increases first and then decreases as dimension
increases for each trimming proportion
.
is much more efficient than
.
5. Finite-Sample Breakdown Points
Now we explore the robustness of
and
based on the projection depth by finite-sample breakdown point. Donoho and Huber [35] discussed two types of finite-sample breakdown points: replacement breakdown point (RBP) and addition breakdown point (ABP). We employ the finite-sample replacement breakdown point in this section. Generally speaking, the finite-sample replacement breakdown point of an estimator is the minimum fraction of contaminated points (outliers or inliers) in a data set such that the estimator breaks down. Specifically, suppose that any
points in a sample
are replaced (contaminated) by
arbitrary points with the contaminated data set denoted by
. Then, the finite-sample replacement breakdown point of a location estimator
at
is defined as
and the finite-sample replacement breakdown point of a scatter estimator
at
is defined as
Essentially, the breakdown of
means that the largest eigenvalue
of
can be made arbitrarily large (explosion breakdown) or the smallest eigenvalue
of
can be made arbitrarily close to 0 (implosion breakdown). Usually, explosion breakdown is caused by outliers and implosion breakdown by inliers. However, the breakdown of a location estimator is only caused by outliers unless it is related to some scatter estimator. Before we give our main result, we briefly recall the projection depth.
The projection depth was first introduced by Liu [8] and was thoroughly studied by Zuo [9]. Given a univariate location estimator
and a univariate scale estimator
, the outlyingness of a point
in
with respect to a sample
is defined as
where
and
. Then the projection depth of
with respect to
is defined as
It is clear that
depends on
and
, and so does any estimator based on
. For the best robustness of an estimator based on
, the pair
is often chosen, where Med is the univariate median and MADk is a modified median absolute deviation (MAD):
with
for
and
being the ordered values of
in
from the smallest to the largest, where
is the ceiling function and
the floor function. When
,
the usual median absolute deviation. It is well-known that
. As for
, Gather and Hilker [36] showed that

when
, where
denotes the maximal multiplicity of points in
.
Unlike the trimming based on the halfspace depth, the trimming based on the projection depth is done based on only the depth values of data points regardless of their directions. This feature can significantly increase the finite-sample breakdown points of
and
. That is why we choose the projection depth here. To make things clear,
,
, and
at
are denoted by
,
, and
, respectively, in the following discussion. If a data set
in
is in general position, that is, no more than
points in
are contained in any
-dimensional hyperplane, we have the following results.
Theorem 4. Let
and
,
, be the
proportion trimmed mean and trimmed scatter matrix based on the projection depth with
. Suppose that
is in general position and
as
for any
in
. Then
1)
for
, and
2)
for
such that
.
From this theorem, we see that RBP
for all the trimmed means with
, i.e., with at least
points trimmed off. However, RBP
for the only trimmed scatter matrix with
, i.e., with
points trimmed off. It is well-known that
is the upper bound of RBP of any affine equivariant scatter estimators [3]. Also,
is the highest RBP of almost all existing affine equivariant location estimators.