1. Introduction
Consider a random experiment with
different outcomes, with the corresponding probabilities
with
. For such an experiment, the discrete Shannon entropy, also known as the information theoretic entropy
is defined by [1]-[3]
(1)
Here we restrict ourselves to the natural logarithm, but in general, Shannon entropy can have the logarithm of any base. If all possible outcomes are equally likely, Shannon entropy is also related to thermodynamic configurational entropy, or Boltzmann entropy, by [4] [5]
(2)
where
is the Boltzmann constant and Ω is the number of possible outcomes or microstates.
For a continuous random variable
, the common practice is to replace the discrete probabilities
in Equation (1) by the probability density function of the variable
, and change the summation to integral. Therefore, the continuous Shannon entropy, also known as differential entropy, for the continuous random variable, is written as [6]-[8]
(3)
where the limits
and
define the support set or the interval of the random variable
. This extension of Shannon entropy, however, suffers from certain shortcomings. For example, it is dimensionally inconsistent. Thus, if the random variable has a physical dimension such as length L, then the dimension of
will be L−1. However, the argument of a logarithmic function must be dimensionless. More specifically, if the variable
is length, depending on whether we use centimeters or meters for
, the value of the entropy changes. Another shortcoming is that for certain functions
, the entropy becomes negative whereas the discrete Shannon entropy is always non-negative. For example, for the probability density function
(4)
the continuous Shannon entropy of Equation (3) becomes negative. Nevertheless, continuous Shannon entropy has been applied in many areas, including statistical mechanics, thermodynamics, as well as signal processing and information theory.
Some of the inconsistencies described above have been pointed out in the literature [9]. In this article, we re-examine the transition from discrete to continuous Shannon entropy in a rigorous way. We clearly show that the continuous Shannon entropy described by Equation (3) is not the direct limit of the discrete entropy.
2. Transition from Discrete to Continuous Shannon Entropy
Let
be the probability density function of a random variable
defined on the interval
. Then the probability of the outcome being between
and
is approximately
, which becomes exact when
.
Let us divide the interval
into
equal slices, each of width
(5)
Then the discrete Shannon entropy of the outcomes would be
(6)
which can be written as
(7)
where
indicates approximation of
when
is discretized by
terms. Now, if we let
such that
, then each term in the second sum vanishes since
. The first sum, on the other hand, by definition becomes a Riemann integral,
(8)
Therefore, it seems plausible that the discrete Shannon entropy reduces to the continuous entropy given by Equation (3).
The problem, however, is that if we compare the corresponding terms in the two sums of Equation (7), we see that as
, the terms of the second sum are much larger than those in the first sum. This is because the ratio of a term in the second sum to the corresponding term in the first sum is
(9)
Then as
, the denominator of this fraction remains finite whereas the numerator approaches
. Consequently, in the above transition from discrete to continuous Shannon entropy, the main term is discarded. This accounts for the shortcomings of the continuous Shannoon entropy given by Equation (3).
We now take a closer look at a more fundamental question: Can the definition of Shannon entropy be extended to continuous random variables as described above? To do so, we start with Equation (6), which is the correct Shannon entropy for a discretized continuous probability distribution. Using Equation (5), this equation becomes
(10)
which reduces to
(11)
or
(12)
where the over-bar represents average value. Now, as
, every term on the
right hand side of this expression remains finite except
which becomes
. Therefore,
(13)
Consequently, as stated earlier, differential entropy is not the direct limit of discrete entropy [9]. We also verified this result computationally.
3. Conclusions
Although continuous Shannon entropy is applied in many areas, it lacks some of the properties of entropy. Extending the expression of discrete Shannon entropy to a continuous random variable results in an infinite entropy; therefore, such a simple extension is not possible. Furthermore, Equation (3), which is commonly used to express differential or continuous entropy is dimensionally inconsistent, scale dependent, and for certain probability density functions, becomes negative. Despite the fact that its shortcomings have been pointed out in some references, the expression continues to be applied in many fields.
It is also worth pointing out that some authors have suggested an alternative definition for continuous entropy by introducing an invariant factor
into Equation (3) [7],
(14)
However, because of this added function, this equation should not be called continuous “Shannon” entropy.