Theory of Regression Lines and Associated Linear Systems and Channels ()
1. Introduction: Regression Lines Can Model Linear Systems and Channels
Scientific research frequently employs scatter plots in which the relationship between the dependent variable
and the independent variable
can be modeled linearly—that is, by means of a regression line characterized by a slope
and a regression coefficient
of the sample considered.
This article aims to summarize the results of a theory regarding linear regression, whose relationships can always be interpreted as describing the input-output characteristic of a linear system. In the following discussion, I will refer to “linear channels” instead of “linear systems” because of the terminology I use in the article; however, the reader will appreciate that these are simply two different ways of describing the same functional relationship.
I have developed the theory in a series of articles concerning the mathematical properties of literary texts—both ancient and modern—particularly in relation to translations [1]-[17]. Another important application was reported in Reference [18].
The conviction that the theory is applicable to any scientific discipline—provided a linear input-output relationship exists—has led me to summarize the main results in a general form that the readers can easily apply to their specific problem. In any case, in the articles mentioned, there are numerous experimental applications demonstrating how the theory can be applied to real-world data.
After this introduction, Section 2 defines the single channel and its noise-to-signal power ratio; Section 3 presents and discusses the geometrical representation of the noise-to-signal power ratio; Section 4 defines the cross channel; Section 5 compares a stochastic channel versus its deterministic version; Section 6 studies the parallel connection of channels; Section 7 studies the series connection of channels; Section 8 reports a summary and a conclusion. Appendix A lists the mathematical symbols used in the article and their meaning (Table A1); Appendix B proves a fundamental relationship; Appendix C details the calculation of a formula of the main text.
2. Single Channel and Noise-to-Signal Power Ratio
In this section, I recall the general theory on stochastic variables linearly connected [19]. Let
(dependent variable) and
(independent variable) linked by the line:
(1)
Equation (1) models a deterministic relationship through the slope
. The intercept
; however, if
, the theory can be fully applied by defining a new dependent variable:
(2)
In the real world, the relationship between
and
is never deterministic, i.e., given by Equation (1), but stochastic (random). Equation (1) models, in fact, two variables perfectly correlated, with correlation coefficient
and slope
. As these conditions are never found in any application, Equation (1) transforms into:
. (3)
In Equation (3),
is a multiplicative distortion (bias);
is an additive stochastic variable, with zero mean value, which renders the linear channel partially stochastic, namely “noisy”, using a communication system terminology [20]. In other words, Equation (3) models a noisy channel with multiplicative bias.
Figure 1 shows the flow chart describing Equation (3) with a system/channel representation. The black box indicated with
represents a distorted deterministic channel; the black box indicated with
represents the parallel channel due to the scattering of
around the regression line.
Figure 1. Flow chart of linear systems. Upper panel: Deterministic channel with multiplicative distortion (bias)
. Lower panel: Noisy channel with multiplicative bias and additive noise source.
Now, let us consider:
a) The variance of the difference
, referred to as the “regression noise” power
.
b) The variance of the difference between the values not lying on Equation (3) (
), and those lying on it (
), referred to as the “correlation noise” power
. This “noise” is due to the spread of
around the line given by Equation (3), due to the stochastic noise
.
c) The variances
and
of
and
involved in Equation (3).
In case (a), we get the difference
, therefore the variance (or power) of the values of
lying on the regression line is given by:
. (4)
Now, we define the “regression noise-to-signal power ratio”
(NSRm) between the output and the input:
(5)
In case (b), the fraction of the variance
due to the values of
not lying on the regression line is given by:
. (6)
The parameter
is called the coefficient of determination and it gives the fraction of the variance of
explained by the regression line
[19]. However, this variance is correlated with the slope
because
is related to
according to:
. (7)
Figure 2 shows the flow chart of variances.
Figure 2. Flow chart of variances:
is the output variance of the values lying on the regression line;
is the output variance due to the values of
not lying on the regression line.
Therefore, inserting Equation (7) in Equation (6), we get the “correlation noise-to-signal power ratio”,
(NSRr):
(8)
Since the two noise variances (powers) are additive, the total noise-to-signal (NSR)
of the channel shown in Figure 2 is given by:
(9)
Therefore,
depends only on the two parameters
and
of the regression line, independently of their sign, explicitly given by:
(10)
Finally, the signal-to-noise ratio (SNR)
, is given by:
(11)
(12)
Of course, we expect that no real channel can yield
and
, in which case
, a channel purely deterministic. In experimental scatterplots, very likely,
,
. It should be noted that the SNR can be considered an indicator of the degree of deterministic nature of a channel: the higher the SNR, the greater the “degree” of the channel’s deterministic nature.
In conclusion, the slope
measures the multiplicative “distortion” of the dependent variable
with respect to the independent variable
in the deterministic channel; the correlation coefficient
measures how “precise” the best linear fit is.
Finally, notice the more direct and insightful analysis that can be achieved by using the NSR instead of the more common SNR because, in Equations (9) and (10), the single channel NSRs simply add together. This makes easier to study, for example, which addend determines
, and thus
, while this is by far less easy with Equation (11).
Figure 3 shows the signal-to-noise ratio drawn with varying multiplicative bias
and
, by assuming
(blue lines),
(red lines);
(green lines). Notice that at
,
and
. At these points, R is determined only by
,
dashed lines. The lower continuous lines refer the full signal-to-noise ratio Γ.
Figure 3. Signal-to-noise ratio drawn as a function of
, by assuming
for all, and
(blue lines),
(red lines);
(green lines). Notice that at
,
and
, see line labelled SNRm. At these points,
, see line labelled SNRr, dashed lines. The lower continuous lines refer to the full signal-to-noise ratio Γ.
Figure 4 shows the full signal-to-noise ratio Γ as a function of
and
. We can notice that as
,
, at
. In this region, Γ changes vary rapidly as small variations in
give very large variations in Γ, for any
.
The choice of using the NSR instead of the more common SNR also leads to a useful graphical representation of Equation (9) that can guide analysis and design, first used in my deep-space communication studies [20], as shown in the next section.
Figure 4. Γ versus
and
; as
,
, at
. In this region, Γ changes vary rapidly as small variations in
give very large variations in Γ, for any
.
3. Geometrical Representation of the Noise-to-Signal Power Ratio
As discussed in [20] for deep-space radio link design, among other features not of interest here, adopting the noise-to-signal power ratio instead of the more common signal-to-noise ratio allows a geometrical representation, which immediately shows how
and
, through their square roots, contribute to the total R.
In fact, we can represent Equation (10) graphically by considering the two orthogonal variables:
(13)
(14)
By setting the noise-to-signal ratio
, being
a constant,
and
trace an arch of the circle with radius
in the first Cartesian quadrant. All points inside the circle correspond to
(
); the origin of the axes corresponds to
of the purely deterministic channel,
and
. The signal-to-noise power ratio
becomes infinite at the origin and decreases as the radius of the circle increases. Figure 5 shows a universal chart with lines of equal signal-to-noise ratio Γ.
Channels can be combined and compared in various ways; therefore, in the following sections, I will examine the most common combinations and make some comparisons.
Figure 5. Lines of equal signal-to-noise ratio drawn in the universal chart of coordinates
and
.
4. Cross Channel
Let us study how the output variable
of channel
relates to the output variable
of another similar channel
for the same input
. This new channel is termed “cross channel”.
Let us consider the scatterplot
, referred to the couple
and the scatterplot
, referred to the couple
, in which the independent variable
and the dependent variable
are linked by the following regression lines in what we can define a “deterministic channel”:
(15)
(16)
As discussed in Section 2, Equations (15) and (16) do not give the full relationship between the two variables because they link only conditional average values, measured by the slopes
and
. According to Equation (3), we can write more general linear relationships by considering the scattering of the data, always present in experiments, modelled by zero-mean noise sources
and
.
(17)
(18)
Now, we can develop a series of interesting investigations on these equations. By eliminating
, we can compare the dependent variable
of Equation (18) to the dependent variable
of Equation (17) for
. In doing so, we can find the regression line and the correlation coefficient of the new scatterplot linking
to
without the availability of the scatterplot itself, assuming the same value of the independent variable
.
By eliminating
between Equation (17) and Equation (18), we get:
. (19)
Compared to the new independent variable
, the slope
of the regression line is given by:
(20)
Because the two noise sources are additive, the total noise is given by:
(21)
Figure 6 shows the flow chart describing the cross-channel.
Figure 6. Chart describing the cross-channel signal flow.
Now, defining the expected (deterministic) ratio between the two slopes
, we get
of the cross channel:
(22)
To proceed to find
, it is necessary to determine the correlation coefficient
between
and
because in Equation (21), the two noise sources might be correlated, with unknown correlation coefficient
. Appendix B shows that:
(23)
Therefore,
of the cross channel is:
(24)
In conclusion, we can determine the slope and the correlation coefficient of the scatterplot between
and
for the same value of the independent variable
. Now, the availability of this scatterplot is experimentally very rare because it is unlikely to find values of
and
for exactly the same value of
, therefore cross channels can reveal relationships very difficult to discover experimentally.
In the next section, we study how the output of a deterministic channel relates to the output of its stochastic version.
5. Stochastic Channel versus Deterministic Channel
We compare a deterministic channel
(slope
) with a stochastic channel
derived from channel
by adding noise (slope
, correlation coefficient
). This is a particular cross channel, therefore, from the theory of Section 4, we get:
(25)
(26)
(27)
, (28)
(29)
(30)
In conclusion, in transforming the deterministic channel into its stochastic version, only the correlation noise is present, therefore the SNR is given by:
. (31)
Equation (31) coincides with the ratio between the variance explained by the regression line (proportional to
) and the variance due to the scattering (correlation noise), proportional to
.
The development of this section confirms that the SNR can be considered—in any case—an indicator of the degree of deterministic nature of a channel: the higher the SNR, the greater the “degree” of the channel’s deterministic nature.
So far, we have considered single channels. In the following sections, we examine the series and parallel connection of channels and determine the signal-to-noise ratio of the overall channel.
6. Parallel Connection of Channels
Let us consider the parallel connection of the channels, Figure 7. The parallel connection gives a single channel with the following input-output relationship:
, (32)
Now, since
, Equation (32) is equivalent to the single channel
(33)
Therefore, Equation (33) coincides with Equation (3) in which:
(34)
(35)
Figure 7. Parallel connection of channels.
Let us calculate the correlation coefficient of this channel. From Equation (7)
. (36)
Therefore, after a few mathematical steps (Appendix C), we get:
(37)
Notice that if
then
(38)
Figure 8 shows how
varies as a function of
and
according to Equation (38). Notice that if
, then
.
In conclusion, the noise-to-signal ratios
and
are given, respectively, by Equation (5) and Equation (8), with
given by Equation (34) and
given by Equation (37). Of course, the theory of parallel channels can be extended to the parallel of
single channels by applying Equations (34) and (37) recursively.
7. Series Connection of Channels
In this section, we consider a channel made by a series of single channels, Figure 9. The “experimental” slope of the overall channel is given by:
(39)
Therefore, in a series of
channels, the expected overall slope is given by;
(40)
Therefore, the noise-to-signal regression noise
is given by:
(41)
From Figure 9, it is evident that the output noise of a preceding channel produces additive noise at the output of the next channels. Let us calculate the correlation noise-to-signal ratio
at the output of the overall channel.
Figure 8. Curves showing how
varies as a function of
versus
, by assuming
, according to Equation (38). Cyan line
; black line
; magenta line
; green
; red line
; blue line
.
Figure 9. Series connection of channels.
Theorem. The correlation noise-to-signal ratio
of
linear channels in series, each characterized by the correlation noise-to-signal ratio
, is given by:
(42)
Proof. Let the three linear input-output relationships of the isolated channels of Figure 9 (i.e., before connecting them in series) given by:
, (43)
, (44)
. (45)
Let
,
,
the variances (power) of the input variables to the single channels, and let
,
,
the variances (power) of the zero-mean noise
,
,
, then the NSRs of the isolated channels are given by:
(46)
(47)
(48)
When the first two channels are connected in series, the input to the second channel must also include the output noise of the first channel, therefore we get the modified output variable
:
(49)
In Equation (49),
is the output “signal” and
is the output noise, therefore, the NSR at the output of the second channel is:
(50)
Now, for three channels in series, it is sufficient to consider
given by Equation (50) as the input NSR to the third single channel to obtain the final NSR and prove Equation (42):
(51)
□
Finally, notice that
of Equation (51) is proportional to the mean
:
(52)
In other words, the series connection of channels averages the single correlation noise-to-signal ratios
.
8. Summary and Conclusions
I have summarized the results of a theory regarding linear regression developed in series of papers [1]-[17]. The linear relationship between the dependent variable
and the independent variable
is interpretable as describing the input-output characteristic of a linear system, or a linear channel.
The conviction that the theory is applicable to any scientific discipline—provided a linear input-output relationship exists—has led me to summarize the main results in a general form that the readers can easily apply to their specific problem.
Defining the regression noise-to-signal power ratio (NSR)
and the correlation noise-to-signal power ratio,
, I have shown that the overall noise-to-signal ratio
is the sum of the two. I have highlighted the more direct and insightful analysis achievable by using the NSR instead of the more common signal-to-noise ratio (SNR) because it makes easier to study which addend determines
and allows a useful graphical representation.
I have underlined that the SNR can be considered an indicator of the degree of deterministic nature of a channel, because the higher the SNR, the greater the “degree” of the channel’s deterministic nature.
In the paper, I have considered various connections of single channels (single, cross, deterministic, parallel, series), whose main results are summarized in Table 1.
Table 1. Channels configuration parameters. For the meaning of the mathematical symbols, see the main text and Appendix A.
Channels configuration |
|
|
Correlation coefficient |
Slope |
Single |
|
|
- |
- |
Cross |
|
|
|
|
Parallel |
|
|
|
|
Series |
|
|
- |
|
In conclusion, the theory and its mathematical results concerning the input-output characteristics of linear channels, described by experimental regression lines, should be useful in any scientific discipline in which the experimental problem is modelled linearly by regression lines.
Acknowledgements
I wish to thank Lucia Matricciani for drawing Figure 1, Figure 2, Figure 6, Figure 7, Figure 9, Figure A1, Figure A2.
Appendix A
Table A1. lists of the mathematical symbols used in the article with their meaning.
Symbol |
Definition |
|
Slope of regression line |
|
Expected slope |
|
Slope in combination of single channels |
|
Correlation coefficient in combination of single channels |
|
Correlation coefficient of linear variables |
|
Coefficient of determination |
|
Standard deviation |
|
Variance |
|
Regression noise power |
|
Correlation noise power |
|
Noise-to-signal overall power ratio |
|
Regression noise-to-signal power ratio |
|
Correlation noise-to-signal power ratio |
|
Dependent variable with intercept zero |
|
Signal-to-noise ratio (linear) |
|
Signal-to-noise ratio (dB) |
|
Mean value |
Appendix B
Consider the stochastic variables
and
with variances
and
respectively. Consider two vectors in the plane whose lengths are proportional to
and
and separated by an angle
whose cosine is given by the correlation coefficient
(Figure A1). The length of the diagonal of the parallelogram determined by them can be computed using the law of cosines ([21], p. 127):
(A1)
Now Equation (A1) is just the variance of the stochastic variable
, therefore the length of the diagonal
is the standard deviation of
. Thus, the addition of the stochastic variables
and
corresponds to vector addition of the vectors representing their standard deviations.
Figure A1. Geometrical representation of standard deviations ([21], p. 127). The standard deviations can be represented by two vectors whose lengths are proportional to
and
respectively and separated by an angle
whose cosine is given by the correlation coefficient
, according to the law of cosines.
Let us now apply this representation to the relationships of the problem in this article:
(A2)
(A3)
We get
(A4)
(A5)
We wish to calculate the correlation coefficient
between the two noise sources, here rewritten by changing the negative sign into positive without any impact on the calculations when dealing with noise:
(A6)
Figure A2 shows the vector representations of the standard deviations of Equations (A4) and (A5).
Figure A2. Vector representations of the standard deviations of Equations (A4) and (A5) to find the correlation coefficient
between
and
.
Now
is the cosine of the angle
between the vectors
and
, therefore:
(A7)
Now, since
(A8)
(A9)
We get the final relationship, Equation (23):
(A10)
Appendix C
In Equation (36), the variance,
, of the two correlated stochastic variables in Equation (32), is given by Equation (A1), here written for our problem:
(A11)
In Equation (A11),
is given by Equation (23). Therefore:
(A12)
(A13)
(A14)
From Equation (A14), we get Equation (37).