<?xml version="1.0" encoding="UTF-8"?><!DOCTYPE article  PUBLIC "-//NLM//DTD Journal Publishing DTD v3.0 20080202//EN" "http://dtd.nlm.nih.gov/publishing/3.0/journalpublishing3.dtd"><article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" dtd-version="3.0" xml:lang="en" article-type="research article"><front><journal-meta><journal-id journal-id-type="publisher-id">AM</journal-id><journal-title-group><journal-title>Applied Mathematics</journal-title></journal-title-group><issn pub-type="epub">2152-7385</issn><publisher><publisher-name>Scientific Research Publishing</publisher-name></publisher></journal-meta><article-meta><article-id pub-id-type="doi">10.4236/am.2022.136036</article-id><article-id pub-id-type="publisher-id">AM-118128</article-id><article-categories><subj-group subj-group-type="heading"><subject>Articles</subject></subj-group><subj-group subj-group-type="Discipline-v2"><subject>Physics&amp;Mathematics</subject></subj-group></article-categories><title-group><article-title>
 
 
  Links between Kalman Filtering and Data Assimilation with Generalized Least Squares
 
</article-title></title-group><contrib-group><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>William</surname><given-names>Menke</given-names></name><xref ref-type="aff" rid="aff1"><sub>1</sub></xref><xref ref-type="corresp" rid="cor1"><sup>*</sup></xref></contrib></contrib-group><aff id="aff1"><label>1</label><addr-line>Lamont-Doherty Earth Observatory of Columbia University, Palisades, USA</addr-line></aff><pub-date pub-type="epub"><day>16</day><month>06</month><year>2022</year></pub-date><volume>13</volume><issue>06</issue><fpage>566</fpage><lpage>584</lpage><history><date date-type="received"><day>18,</day>	<month>May</month>	<year>2022</year></date><date date-type="rev-recd"><day>26,</day>	<month>June</month>	<year>2022</year>	</date><date date-type="accepted"><day>29,</day>	<month>June</month>	<year>2022</year></date></history><permissions><copyright-statement>&#169; Copyright  2014 by authors and Scientific Research Publishing Inc. </copyright-statement><copyright-year>2014</copyright-year><license><license-p>This work is licensed under the Creative Commons Attribution International License (CC BY). http://creativecommons.org/licenses/by/4.0/</license-p></license></permissions><abstract><p>
 
 
  Kalman filtering (KF) is a popular form of data assimilation, especially in real-time applications. It combines observations with an equation that describes the dynamic evolution of a system to produce an estimate of its present-time state. Although KF does not use future information in producing an estimate of the state vector, later reanalysis of the archival data set can produce an improved estimate, in which all data, past, present and future, contribute. We examine the case in which the reanalysis is performed using generalized least squares (GLS), and establish the relationship between the real-time Kalman estimate and the GLS reanalysis. We show that the KF solution at a given time is equal to the GLS solution that one would obtain if data excluded future times. Furthermore, we show that the recursive procedure in KF is exactly equivalent to the solution of the GLS problem via Thomas’ algorithm for solving the block-tridiagonal matrix that arises in the reanalysis problem. This connection suggests that GLS reanalysis is better considered the final step of a single process, rather than a “different method” arbitrarily being applied, post factor. The connection also allows the concept of resolution, so important in other areas of inverse theory, to be applied to KF formulations. In an exemplary thermal diffusion problem, model resolution is found to be somewhat localized in both time and space, but with an extremely rough averaging kernel.
 
</p></abstract><kwd-group><kwd>Kalman Filter</kwd><kwd> Generalized Least Squares</kwd><kwd> Bayesian Inference</kwd><kwd> Data Assimilation</kwd><kwd> Real-Time</kwd><kwd> Resolution</kwd></kwd-group></article-meta></front><body><sec id="s1"><title>1. Introduction</title><p>In this paper, we compare two data assimilation methods that are routinely applied to monitor time-dependent of linear systems, one based on Generalized Least Squares (GLS) [<xref ref-type="bibr" rid="scirp.118128-ref1">1</xref>] and the other on Kalman Filtering (KF) [<xref ref-type="bibr" rid="scirp.118128-ref2">2</xref>] [<xref ref-type="bibr" rid="scirp.118128-ref3">3</xref>]. We are motivated by our anecdotal observation that, while both are widely used tools that do broadly similar things, GSL and KF tend to be used by different communities who often think of their method as the “best”. Our goal is to enumerate and study the similarities and differences between the GLS and KF. Especially, we wish to determine whether or not the same prior information is used in each. Establishing a link between KF and GLS provides a clear pathway for applying GLS concepts, and especially resolution analysis [<xref ref-type="bibr" rid="scirp.118128-ref4">4</xref>] [<xref ref-type="bibr" rid="scirp.118128-ref5">5</xref>], to KF, and also a pathway for extending GLS analysis to the real-time scenarios for which KF is already well-suited.</p><p>A synopsis of variables used in this paper is provided in <xref ref-type="table" rid="table1">Table 1</xref>. At any time, t i , 1 ≤ i ≤ K , a linear system is described by a state vector (model parameter vector), m ( i ) ∈ ℝ M . This state vector evolves away from the initial condition</p><p>m ( 1 ) = m A ( 1 ) (1)</p><p>according to the dynamical equation:</p><p>m ( i ) = D m ( i − 1 ) + s ( i − 1 ) (2)</p><p>Here, D is the dynamics matrix and s ∈ ℝ M is the source vector. Neither</p><table-wrap id="table1" ><label><xref ref-type="table" rid="table1">Table 1</xref></label><caption><title> List of variables</title></caption><table><tbody><thead><tr><th align="center" valign="middle" >Variable</th><th align="center" valign="middle" >Explanation</th><th align="center" valign="middle" >Variable</th><th align="center" valign="middle" >Explanation</th></tr></thead><tr><td align="center" valign="middle" >a</td><td align="center" valign="middle" >Gram right-hand-side vector</td><td align="center" valign="middle" >G</td><td align="center" valign="middle" >union of data kernels</td></tr><tr><td align="center" valign="middle" >a ^ i</td><td align="center" valign="middle" >Thomas right-hand-side vector at time, i</td><td align="center" valign="middle" >G ( i )</td><td align="center" valign="middle" >data kernel at time, i</td></tr><tr><td align="center" valign="middle" >a ˜ i</td><td align="center" valign="middle" >Kalman right-hand-side vector at time, i</td><td align="center" valign="middle" >G − g</td><td align="center" valign="middle" >generalized inverse</td></tr><tr><td align="center" valign="middle" >A</td><td align="center" valign="middle" >Gram matrix</td><td align="center" valign="middle" >h</td><td align="center" valign="middle" >prior right-hand-side vector</td></tr><tr><td align="center" valign="middle" >A i</td><td align="center" valign="middle" >th diagonal element of Gram matrix</td><td align="center" valign="middle" >H</td><td align="center" valign="middle" >union of prior information</td></tr><tr><td align="center" valign="middle" >A ^ i</td><td align="center" valign="middle" >Thomas matrix at time, i</td><td align="center" valign="middle" >f</td><td align="center" valign="middle" >GLS right-hand-side vector</td></tr><tr><td align="center" valign="middle" >A ˜ i</td><td align="center" valign="middle" >Kalman matrix at time, i</td><td align="center" valign="middle" >F</td><td align="center" valign="middle" >GLS matrix</td></tr><tr><td align="center" valign="middle" >B</td><td align="center" valign="middle" >Off-main-diagonal of Gram matrix</td><td align="center" valign="middle" >K</td><td align="center" valign="middle" >number of time steps</td></tr><tr><td align="center" valign="middle" >C A 1</td><td align="center" valign="middle" >prior covariance of initial condition</td><td align="center" valign="middle" >m</td><td align="center" valign="middle" >union of states</td></tr><tr><td align="center" valign="middle" >C d</td><td align="center" valign="middle" >covariance of data at time, i</td><td align="center" valign="middle" >m ( i )</td><td align="center" valign="middle" >state at time, i</td></tr><tr><td align="center" valign="middle" >C h</td><td align="center" valign="middle" >union of prior covariances</td><td align="center" valign="middle" >m A ( 1 )</td><td align="center" valign="middle" >initial condition</td></tr><tr><td align="center" valign="middle" >C m</td><td align="center" valign="middle" >posterior covariance of state</td><td align="center" valign="middle" >m G</td><td align="center" valign="middle" >GLS estimate of state</td></tr><tr><td align="center" valign="middle" >C o</td><td align="center" valign="middle" >union of data covariances</td><td align="center" valign="middle" >m K</td><td align="center" valign="middle" >Kalman estimate of state</td></tr><tr><td align="center" valign="middle" >C s</td><td align="center" valign="middle" >prior covariance of source</td><td align="center" valign="middle" >M</td><td align="center" valign="middle" >length of state vector</td></tr><tr><td align="center" valign="middle" >d</td><td align="center" valign="middle" >union of data</td><td align="center" valign="middle" >N</td><td align="center" valign="middle" >number of data at time, i</td></tr><tr><td align="center" valign="middle" >d ( i )</td><td align="center" valign="middle" >data at time, i</td><td align="center" valign="middle" >N</td><td align="center" valign="middle" >data resolution matrix</td></tr><tr><td align="center" valign="middle" >D</td><td align="center" valign="middle" >dynamics matrix</td><td align="center" valign="middle" >R</td><td align="center" valign="middle" >model resolution matrix</td></tr><tr><td align="center" valign="middle" >Δ 2</td><td align="center" valign="middle" >second difference matrix</td><td align="center" valign="middle" >s ( i )</td><td align="center" valign="middle" >source at time, i</td></tr></tbody></table></table-wrap><p>the initial conditions nor the source is known exactly, but rather have uncertainty described by their respective covariance matrices, C A 1 and C s .</p><p>This formulation well approximates the behavior of systems described by linear partial differential equations that are first order in time. For example, let m n ( i ) = m ( x n , t i ) be temperature at time, t i = i Δ t , and position, x n = n Δ x , where Δ t and Δ x are small increments, and suppose that m ( x , t ) satisfies the thermal diffusion equation, ∂ m / ∂ t = c ∂ 2 m / ∂ x 2 + q (with zero boundary conditions). This partial differential equation can be approximated by:</p><p>m ( i ) − m ( i − 1 ) Δ t = c ( Δ x ) 2 Δ 2 m ( i − 1 ) + q ( i − 1 )     or     m ( i ) = D m ( i − 1 ) + s ( i − 1 )</p><p>with     D = ( c Δ t ( Δ x ) 2 Δ 2 + I ) m ( i − 1 )     and     s ( i − 1 ) = Δ t q ( i − 1 ) (3)</p><p>which has the form of dynamical Equation (2). Here, the choices:</p><p>Δ 2 ≡ [ 1 1 − 2 1 1 − 2 1 ⋱ 1 − 2 1 1 ]     and     q ( i − 1 ) = [ 0 q 2 ( i − 1 ) ⋮ q M − 1 ( i − 1 ) 0 ] (4)</p><p>encode both the differential equation and the boundary conditions. The matrix, D , is sparse in this example, as well as in many other cases in which it approximates a differential operator.</p><p>The data equation expresses the relationship between the state vector and the observables:</p><p>G ( i ) m ( i ) = d ( i ) (5)</p><p>Here, d ( i ) ∈ ℝ N is the data vector, determined to covariance, C d , and G ( i ) is the data kernel. In the simplest case, the observations may be of selected elements of the state vector, itself, in which case, each row of G is zero, except for a single element, say in column, k, which is unity. Here, k ( i , n ) is a function that associates d n ( i ) with m k ( i ) . Other more complicated relationships are possible. For instance, in tomography, the data is a line integral through m ( x , y , t ) (with y another spatial dimension). The data kernel, G , is sparse in these two cases. In other cases, it may not be sparse.</p><p>The data assimilation problem is to estimate the set of state vectors, { m ( i ) : 1 &lt; i ≤ K } , using the dynamical equation, the data equation and the initial condition. One possible approach is based on Generalized Least Squares (GLS); another upon Kalman Filtering (KF). In this paper, we demonstrate that these two apparently very different methods are, in fact, exactly equivalent. In order to simplify notation, we concatenate the state vectors for times into an overall vector, m = [ m ( i ) , ⋯ , m ( K ) ] , and data vectors into an overall vector, d = [ d ( 2 ) , ⋯ , d ( N ) ] . By assumption, no observations are made at time, i = 1 .</p></sec><sec id="s2"><title>2. Generalized Least Squares Applied to the Data Assimilation Problem</title><p>Generalized Least Squares (GLS) [<xref ref-type="bibr" rid="scirp.118128-ref1">1</xref>] [<xref ref-type="bibr" rid="scirp.118128-ref6">6</xref>] [<xref ref-type="bibr" rid="scirp.118128-ref7">7</xref>] [<xref ref-type="bibr" rid="scirp.118128-ref8">8</xref>] is a technique used to estimate m when two types of information are available: prior information and data. By prior information, we mean expectations about the behavior of m that are based on past experience or general physical considerations. The dynamical equation and initial condition discussed in the previous section are examples of prior information. By data we mean direct observations, as typified by the data equation discussed in the previous section.</p><p>Prior information can be represented by the linear equation, H m = h (with h ∈ ℝ L and where the equation is accurate to covariance, C h ) and observations can be represented by the linear equation, G m = d (with covariance C o ). The Bayesian principle leads to the optimal solution, which we denote the Generalize Least Squares (GLS) solution, m G [<xref ref-type="bibr" rid="scirp.118128-ref1">1</xref>] [<xref ref-type="bibr" rid="scirp.118128-ref6">6</xref>] [<xref ref-type="bibr" rid="scirp.118128-ref7">7</xref>] [<xref ref-type="bibr" rid="scirp.118128-ref8">8</xref>]. It minimizes a combination of the weighted L 2 error in prior information and the weighted L 2 error in the data (where the weighting depends upon the covariances).</p><p>Several equivalent forms of the GLS solution, m G , and its posterior variance, C m , are common in the literature. We enumerate a few of the more commonly-used forms here:</p><p>Form 1 [<xref ref-type="bibr" rid="scirp.118128-ref8">8</xref>] groups the prior information and data equations into a single equation, F m = h :</p><p>m G = [ F T F ] − 1 F T f     with     F ≡ [ C h − 1 / 2 H C o − 1 / 2 G ]     and     f ≡ [ C h − 1 / 2 h C o − 1 / 2 d ]</p><p>C m = [ F T F ] − 1 (6)</p><p>Note that the factors of C h − 1 / 2 amd C o − 1 / 2 are weights proportional to “certainty”.</p><p>Form 2 [<xref ref-type="bibr" rid="scirp.118128-ref1">1</xref>] introduces generalized inverses, G − g and H − g :</p><p>m G = G − g d + H − g h</p><p>with     G − g ≡ A − 1 G T C o − 1     and     H − g ≡ A − 1 H T C h − 1     and     A ≡ G T C o − 1 G + H T C h − 1 H</p><p>C m = A − 1 (7)</p><p>Here, m G is the estimated solution and C m is its posterior covariance.</p><p>Form 3 [<xref ref-type="bibr" rid="scirp.118128-ref7">7</xref>] organizes the solution in terms of the prior state vector, m A ; that is, the state vector implied by the prior information, acting alone:</p><p>m G = G − g d + P G m A     with   P G ≡ ( I − G − g G )</p><p>and   with   m A ≡ { H − 1 h ∃ H − 1 [ H T C h − 1 H ] − 1 H T C h − 1 h ∄ H − 1     and     C A = [ H T C h − 1 H ] − 1</p><p>C m = P G C A (8)</p><p>The matrix, P G , plays the role of a projection matrix. See Appendix A.1 for a deviation of the covariance equation.</p><p>Form 4 [<xref ref-type="bibr" rid="scirp.118128-ref1">1</xref>] introduces the deviation, Δ m , of the solution from the prior state vector, and the corresponding deviation, Δ d , of the data from that predicted by the prior state vector:</p><p>Δ m = G − g Δ d     and     Δ m ≡ m G − m A     and     Δ d ≡ d − G m A (9)</p><p>Finally, Form 5 [<xref ref-type="bibr" rid="scirp.118128-ref6">6</xref>] [<xref ref-type="bibr" rid="scirp.118128-ref9">9</xref>] uses as alternate form of the generalized inverse:</p><p>Δ m = G ′ − g Δ d     with     G ′ − g ≡ C A G T A ∗ − 1     and     A ∗ ≡ [ C o + G C A G T ] (10)</p><p>The equality of the two forms was proven by [<xref ref-type="bibr" rid="scirp.118128-ref6">6</xref>] using a matrix identity that we denote TV82-A (Appendix A.2). Because A is M &#215; M and A ∗ is N &#215; N , the first form is most useful when M &lt; N ; the second when M &gt; N . However, a decision to use one or the other must also take in consideration the sparsity of the various matrix products.</p><p>Form 4 is derived from Form 3 by subtracting A m A from both sides of the Gram equation,</p><p>A ( m G − m A ) = a − A m A</p><p>[ G T C o − 1 G + H T C h − 1 H ] ( m G − m A ) = G T C o − 1 ( d − G m A ) + H T C h − 1 ( h − H m A ) (11)</p><p>and then by requiring that the second term on the right-hand side vanish, which leads to:</p><p>A Δ m = G T C o − 1 Δ d     with     Δ d = d − G m A     and     m A = [ H T C h − 1 H ] − 1 H T C h − 1 h (12)</p><p>That is, m A is due to the prior information acting along. The deviatoric manipulation is completely general; alternately the first term could have been made to vanish, in which leads to:</p><p>A Δ m = H T C h − 1 Δ h     with     Δ h = h − H m G       and       m G = [ G T C o − 1 G ] − 1 G T C o − 1 d (13)</p><p>Here, m G is due to the data acting alone. Note that deviatoric manipulations of this type never change the form of the matrix, A . We will apply this principle later in the paper.</p><p>In the subsequent analysis, we will focus on the Gram equations:</p><p>A m = G T C o − 1 d ≡ a (14)</p><p>The initial condition and the dynamical equation can be into a single prior information equation of the form, H m = h :</p><p>[ I − D I − D I ⋱ − D I ] [ m ( 1 ) m ( 2 ) m ( 3 ) ⋮ m ( K ) ] = [ m A ( 1 ) s ( 1 ) s ( 2 ) ⋮ s ( K − 1 ) ]     withcovariance     C h (15)</p><p>Here, C h ≡ diag ( C A 1 , C s , C s , ⋯ , C s ) . Several quantities derived from H , and which we will use later, are:</p><p>H T = [ I − D T I − D T I − D T ⋱ I − D T I ]     and     H − 1 = [ I D I D 2 D I D 3 D 2 D I ⋱ ]</p><p>H T C h − 1 H = [ [ C A 1 − 1 + D T C s − 1 D ] − D T C s − 1 0 − C s − 1 D [ C s − 1 + D T C s − 1 D ] − D T C s − 1 − C s − 1 D [ C s − 1 + D T C s − 1 D ] ⋱ − D T C s − 1 − D T C s − 1 C s − 1 ]</p><p>H T C h − 1 h = [ C A 1 − 1 m A ( 1 ) − D T C s − 1 s ( 1 ) C s − 1 s ( 1 ) − D T C s − 1 s ( 2 ) ⋮ C s − 1 s ( N − 1 ) − D T C s − 1 s ( N ) C s − 1 s ( N ) ] (16)</p><p>The existence of H − 1 implies that m A can be uniquely specified. The data equation expresses the relationship between the state vector and the observables, and presuming that no data are available for time, i = 1 , has the form:</p><p>[ 0 G ( 2 ) ⋱ G ( N ) ] [ m ( 1 ) m ( 2 ) ⋮ m ( M ) ] = [ d ( 2 ) d ( 3 ) ⋮ d ( N ) ] (17)</p><p>This equation is taken to have an accuracy described by the summary covariance matrix, C o ≡ diag ( C d , C d , C d , ⋯ , C d ) . We note that:</p><p>G T C d − 1 G = [ 0 G ( 2 ) T C d − 1 G ( 2 ) T ⋱ G ( K ) T C d − 1 G ( K ) T ] and     G T C d − 1 d = [ 0 G ( 2 ) T C d − 1 d ( 2 ) G ( 3 ) T C d − 1 d ( 3 ) ⋮ G ( K ) T C d − 1 d ( N ) ] (18)</p><p>The matrix, A , in the Gram equation is symmetric and block-triagonal:</p><p>[ A 1 B T B A 2 B T ⋱ B T B A N ] [ m ( 1 ) m ( 2 ) m ( 3 ) ⋮ m ( K ) ] = [ a 1 a 2 a 3 ⋮ a K ] (19)</p><p>with elements:</p><p>A i = { [ D T C s − 1 D + C A 1 − 1 ]       ( i = 1 ) [ D T C s − 1 D + C s − 1 + G ( i ) C d − 1 G ( i ) T ] ( 1 &lt; i &lt; K ) C s − 1 + G ( K ) C d − 1 G ( K ) T       ( i = K )</p><p>B = − C s − 1 D (20)</p><p>The vector, a , on the right-hand side of the Gram equation, is:</p><p>a i = { [ − D T C s − 1 s ( 1 ) + C A 1 − 1 m A ( 1 ) ]       ( i = 1 ) [ − D T C s − 1 s ( i ) + C s − 1 s ( i − 1 ) + G ( i ) T C d − 1 d ( i ) ] ( 1 &lt; i &lt; K ) C s − 1 s ( i − 1 ) + G ( i ) T C d − 1 d ( i )       ( i = K ) (21)</p><p>Here, we define s ( K ) to be zero.</p></sec><sec id="s3"><title>3. Recursive Solution Using the Thomas Method</title><p>Insight into the behavior of the GLS solution can be gained by solving the Gram equation iteratively [<xref ref-type="bibr" rid="scirp.118128-ref10">10</xref>]. We the Thomas method [<xref ref-type="bibr" rid="scirp.118128-ref11">11</xref>] [<xref ref-type="bibr" rid="scirp.118128-ref12">12</xref>] (see Appendix A.3), though other methods are viable alternatives. It consists of a forward-in-time pass through the system that recursively calculates two quantities, A ^ i and a ^ i :</p><p>A ^ i − 1 ≡ { A 1 − 1 ( i = 1 ) [ A i − B A ^ i − 1 − 1 B T ] − 1 ( i &gt; 1 )</p><p>a ^ i ≡ { a 1 ( i = 1 ) [ a i − B A ^ i − 1 − 1 a ^ i − 1 ] ( i &gt; 1 ) (22)</p><p>After the forward recursion, the system is block-upper-bidiagonal with row i having elements A ^ i and B T (except for the last row, which lacks the B T ) and the modified right-hand side is a ^ i . The solution, m G ( i ) , is achieved through a backward recursion:</p><p>m G ( i ) = { A ^ K − 1 a ^ K ( i = K ) A ^ i − 1 [ a ^ i − B T m ( i + 1 ) ] ( i &lt; K ) (23)</p><p>It is evident that information is propagated both forward and backward in time during the solution process. Furthermore, computation time grows no faster than the number of steps, K, in the recursion.</p><p>The Thomas method has a disadvantage in the common case where the covariances matrices, C A 1 , C s and C d are diagonal and when D and G ( i ) are sparse, because although F is then also sparse, the matrices, A ^ i − 1 , are in general not sparse, so the effort needed to compute them scales with M 3 . Other direct methods share this limitation, too. Consequently, the overall calculation scales with K M 3 . An iterative method [<xref ref-type="bibr" rid="scirp.118128-ref10">10</xref>], such conjugate gradient method, applied to the Gram equation, F T F m = F T h , is usually a better choice. This method requires that the quantity, u = ( F T F ) v , be calculated for an arbitrary vector, v , and this quantity can be very efficiently calculated as u = F T ( F v ) [<xref ref-type="bibr" rid="scirp.118128-ref13">13</xref>]. In cases in which the dynamical equation approximates a partial differential equation, the number of non-zero elements in the matrix, F , scale with K M . The conjugate gradient algorithm requires no more than K M iterations (and often much fewer), each requiring K M multiplications. Thus, the overall solution time scales with K 2 M 2 . Consequently, the conjugate gradient method has a speed advantage when M &gt; K .</p></sec><sec id="s4"><title>4. Present-Time Solution</title><p>Suppose that the analysis focuses on the “present-time”, i, in the sense that only the solution, m P ( i ) , determined using data up to and including time, i, is of interest. One can assemble a sequence of present-time solutions during the forward recursion, using the fact that the ith solution can always be considered to be the final one. No backwards recursion is needed to compute the solution for the final time. However, the forms of the “final” A ^ i and a ^ i differ from that of the previous A s in the recursion, so a separate computation is needed:</p><p>m P ( i ) = A ^ ′ i − 1 a ^ ′ i</p><p>A ^ ′ i = A ^ i − D T C s − 1 D = C s i − 1 + G ( i ) C d − 1 G ( i ) T     and     a ^ ′ i = a ^ i + D T C s − 1 s ( i ) (24)</p><p>Consequently, in order to create a sequence of present-time solution, the two linear systems must be solved at each step in the forward recursion. The present-time solution is the same as the reference solution, m D ( i ) , defined in the previous section.</p><p>Kalman Filtering</p><p>Kalman Filtering (KF) is a solution method with an algorithm that, like the Thomas present-time solution, is forward-in-time, only [<xref ref-type="bibr" rid="scirp.118128-ref2">2</xref>] [<xref ref-type="bibr" rid="scirp.118128-ref3">3</xref>]. It consists of four steps, the final three which are iterated.</p><p>Step 1 assigns the i = 1 solution, m K ( 1 ) , and its covariance, C m ( 1 ) :</p><p>m K ( 1 ) = m A ( 1 )     and     C m ( 1 ) = C A 1 (25)</p><p>Step 2 propagates the solution and its covariance forward in time using the dynamical equation, and considers it to be prior information.</p><p>m A ( i ) = ( D m K ( i − 1 ) + s ( i − 1 ) )     and     C A ( i ) = D C m ( i − 1 ) D T + C s (26)</p><p>Step 3 uses GLS to combine the prior information, m A ( i ) , with covariance C A ( i ) , and data, d ( i ) , with covariance C d , into a solution, m K ( i ) , with covariance, C m . Any of the equivalent forms of the GLS solutions described in Section 2 can be used in this step.</p><p>Step 4, increments i and returns to Step 2, creating a recursion.</p><p>Most often, the GLS solution and its variance are written as:</p><p>m K ( i ) = m A ( i ) + G i − g Δ d ( i )     with       Δ d ( i ) = d ( i ) − G ( i ) m A ( i )</p><p>with     G i − g = C A ( i ) G ( i ) T [ C d ( i ) + G ( i ) C A ( i ) G ( i ) T ] − 1       and       C A ( i ) = D C m ( i − 1 ) D T + C s</p><p>and     C m ( i ) = [ I − G i − g G ( i ) ] C A ( i ) = P G ( i ) C A ( i )     and     P G ( i ) ≡ [ I − G i − g G ( i ) ] (27)</p><p>However, any of the equivalent forms described above can substitute, such as:</p><p>m K ( i ) = A ˜ i − 1 a ˜ i</p><p>C A ( i ) = D C m ( i − 1 ) D T + C s     and     A ˜ i − 1 = [ G ( i ) T C d − 1 G ( i ) + [ C A ( i ) ] − 1 ] − 1 = C m ( i )</p><p>a ˜ i = G ( i ) T C d − 1 d + [ C A ( i ) ] − 1 D m K ( i − 1 ) + [ C A ( i ) ] − 1 s ( i − 1 ) (28)</p></sec><sec id="s5"><title>5. Kalman Filtering Is Not “Filtering” in the Strict Sense</title><p>A standard Infinite Impulse Response (IIR) filter has the form v ∗ m = u ∗ z , where z is the “input” timeseries, m is the “output” timeseries, u and v are filters (with v 1 = 1 ) and ∗ signifies convolution [<xref ref-type="bibr" rid="scirp.118128-ref14">14</xref>]. Key to this formulation is that the filter coefficients are constants; that is, they are not a function of time.</p><p>If the generalized inverse in KF was time-independent, so that G i − g = G − g , then KF could be put into form of an IIR filter:</p><p>[ [ I ] [ G − g G D − I ] ] ∗ [ ⋮ m ( i − 1 ) m ( i ) ⋮ ] = [ G − g − G − g G ] ∗ [ ⋮ [ d ( i − 1 ) s ( i − 2 ) ] [ d ( i ) s ( i − 1 ) ] ⋮ ] (29)</p><p>because the convolution reproduces the KF solution:</p><p>m ( i ) = m ( i − 1 ) + G − g ( d ( i ) − G ( D m ( i − 1 ) + s ( i − 1 ) ) ) (30)</p><p>So, from this point of view, the KF has a v of length 2 (each element of which is a matrix), and a u of length 1 (each element of which is a row vector of two matrices). However, this formulation does not really correspond to a standard IIR filter, because the filter coefficients, which depend upon the generalized inverse, G i − g , depend upon time, i. Hence, the word “filter”, though generally indicative of the KF process, oversimplifies the actual operation being performed. KF is not filtering in the strict sense.</p></sec><sec id="s6"><title>6. The Present-Time Thomas and Kalman Filtering Solutions Are Equal</title><p>We will now demonstrate that the present-time Thomas solution, m P ( i ) , and the Kalman filtering solution, m K ( i ) , are equal. We will make use of an identity, abbreviated TV82-B, that is due to [<xref ref-type="bibr" rid="scirp.118128-ref6">6</xref>], which shows that for invertible symmetric matrices C 1 and C 2 and arbitrary matrix, M :</p><p>C 2 − C 2 M T [ M C 2 M T + C 1 ] − 1 M C 2 = [ M T C 1 − 1 M + C 2 − 1 ] − 1 (31)</p><p>Thus, for instance, when C 1 − 1 = C A ( i ) , C 2 − 1 = C s and M = D T :</p><p>[ C A ( i ) ] − 1 = [ D C m ( i − 1 ) D T + C s ] − 1 = C s − 1 − C s − 1 D [ D T C s − 1 D + [ C m ( i − 1 ) ] − 1 ] − 1 D T C s − 1 (32)</p><p>The PT and KF recursions both start with m K ( 1 ) = m P T ( 1 ) = m A and C K A ( 1 ) = C P T A ( 1 ) = C A . The ( i = 2 ) case in irregular, and must be examine separately. The KF solution is:</p><p>m K ( 2 ) = A ˜ 2 − 1 a ˜ 2</p><p>with     A ˜ 2 = G ( 2 ) T C d − 1 G ( 2 ) + [ C A ( 2 ) ] − 1     with     C A ( 2 ) = D C A 1 D T + C s</p><p>and     a ˜ 2 = G ( 2 ) T C d − 1 d ( 2 ) + [ C A ( 2 ) ] − 1 D m K ( 1 ) + [ C A ( 2 ) ] − 1 s ( 1 ) (33)</p><p>This can be compared with the present-time Thomas solution:</p><p>m P ( 2 ) = A ^ ′ 2 − 1 a ^ ′ 2</p><p>with     A ^ ′ 2 = G ( 2 ) C d − 1 G ( 2 ) T + C s − 1 − C s − 1 D [ D T C s − 1 D + C A 1 − 1 ] − 1 D T C s − 1                       = G ( 2 ) C d − 1 G ( 2 ) T + [ D C A 1 D T + C s ] − 1</p><p>a ^ ′ 2 = G ( 2 ) T C d − 1 d ( 2 ) + C s − 1 s ( 1 ) − C s − 1 D A ^ 1 − 1 D T C s − 1 s ( 1 ) + C s − 1 D A ^ 1 − 1 C s − 1 D C A 1 − 1 m A (34)</p><p>Note that we used TV82-B to simplify the expression for A ^ 2 . By inspection, A ^ ′ 2 = A ˜ 2 . Thus, the two solutions are equal if a ^ ′ 2 = a ˜ 2 . The terms involving d ( 2 ) match. The terms involving s ( 1 ) would match if it could be shown that:</p><p>C s − 1 − C s − 1 D [ D T C s − 1 D + C A 1 − 1 ] − 1 D T C s − 1 = ? [ D C A 1 D T + C s ] − 1 (35)</p><p>But this equation is true by TV82-B. The terms m A also match, because of the matrix identity</p><p>[ D C A 1 D T + C s ] − 1 D = C s − 1 D [ D T C s − 1 D + C A 1 − 1 ] − 1 D C A 1 − 1 (36)</p><p>derived in Appendix A.4. Consequently, the solutions, m K ( 2 ) = m P ( 2 ) , and their posterior covariances, C K m ( 2 ) = A ^ ′ 2 − 1 = C P m ( 2 ) = A ˜ 2 − 1 are equal. Applying C K m ( i ) = A ˜ i − 1 to the KF recursion, and TV82-B and A ^ i = A ^ ′ i + D T C s − 1 D T to present-time Thomas recursion, leads to:</p><p>A ˜ i + 1 = G ( i + 1 ) T C d − 1 G ( i + 1 ) + [ D A ˜ i − 1 D T + C s ] − 1</p><p>A ^ ′ i + 1 = G ( i + 1 ) C d − 1 G ( i + 1 ) T + C s − 1 − C s − 1 D [ D T C s − 1 D + A ^ ′ i ] − 1 D T C s − 1 = G ( i + 1 ) C d − 1 G ( i + 1 ) T + [ D A ^ ′ i − 1 D T + C s ] − 1 (37)</p><p>Thus, A ˜ i + 1 = A ^ ′ i + 1 as long as A ˜ i − 1 = A ^ ′ i − 1 . Because the latter is true for i = 2 , so the formula can be successively applied to show A ˜ i + 1 = A ^ ′ i + 1 for all i &gt; 2 . Similarly, the procedure that demonstrated the equality of a ^ 2 and a ˜ 2 can be extended to</p><p>a ˜ i + 1 = G ( i + 1 ) T C d − 1 d ( i + 1 ) + [ D A ˜ i − 1 D T + C s ] − 1 s ( i ) + [ D A ˜ i − 1 D T + C s ] − 1 D m K ( i )</p><p>and</p><p>a ^ i + 1 = C s − 1 s ( i ) + G ( i + 1 ) T C d − 1 d ( i + 1 ) + C s − 1 D A ^ i − 1 [ a ^ ′ i − D T C s − 1 s ( i ) ] = G ( i + 1 ) T C d − 1 d ( i + 1 ) + [ C s − 1 − C s − 1 D A ^ i − 1 D T C s − 1 ] s ( i ) + C s − 1 D A ^ i − 1 a ^ ′ i = G ( i + 1 ) T C d − 1 d ( i + 1 ) + [ C s + D A ^ ′ i − 1 D T ] − 1 s ( i ) + C s − 1 D A ^ i − 1 A ^ ′ i m P ( i ) (38)</p><p>Here we have used TV82-B and m P ( i ) = A ^ i − 1 a ^ ′ i . The terms ending in d ( i + 1 ) match. The terms ending in s ( i ) also match, since it has been established previously that A ˜ i − 1 = A ^ ′ i − 1 . In order for a ˜ i + 1 to equal a ^ i + 1 , we must have m P ( i ) = m K ( i ) and:</p><p>[ D A ˜ i − 1 D T + C s ] − 1 D = ? C s − 1 D [ A ^ ′ i − 1 + D T C s − 1 D ] A ^ ′ i (39)</p><p>This equation has the same form as identity (36), where the equality has been demonstrated. Starting with i = 2 , we have m P ( 2 ) = m K ( 2 ) and A ^ ′ 2 = A ˜ 2 , which implies A ^ ′ 3 = A ˜ 3 and a ^ 3 = a ˜ 3 , which implies m P ( 3 ) = m K ( 3 ) . This process can be iterated indefinitely, establishing that the present-time Thomas and Kalman solutions, and their posterior variance, are equal.</p></sec><sec id="s7"><title>7. Comparison between the Present-Time Solution and GLS</title><p>The present-time solutions at time, j, depends on information available for times, ( i ≤ j ) but not upon information that subsequently becomes available (that is, for times ( i &gt; j ) . This limitation is necessary in a real-time scenario. However, the lack of future data leads to a solution that is poorer estimate of the true solution, than a GLS solution in which the state vectors at all time are globally adjusted to best-fit all the prior information and data.</p><p>An outlier that occurs at, or immediately before, the present moment can cause large error in the present-time solution. The full GLS solution is less affected because measurements in the near future may compensate (<xref ref-type="fig" rid="fig1">Figure 1</xref>).</p><p>Having established links between KF and GLS, we are now able to apply several useful inverse theory concepts and especially resolution [<xref ref-type="bibr" rid="scirp.118128-ref4">4</xref>] [<xref ref-type="bibr" rid="scirp.118128-ref5">5</xref>] [<xref ref-type="bibr" rid="scirp.118128-ref7">7</xref>]. In a GLS problem, model resolution refers to the ability of the data assimilation process to reconstruct deviations of the true model from the one predicted by the prior information, alone [<xref ref-type="bibr" rid="scirp.118128-ref1">1</xref>]. Data resolution refers to its ability to reconstruct deviations of the data from the one predicted by the prior information [<xref ref-type="bibr" rid="scirp.118128-ref9">9</xref>]. Model resolution is quantified by a ( K M ) &#215; ( K M ) matrix, R = G − g G and data resolution by a ( K N ) &#215; ( K N ) matrix, N = G G − g , that satisfy.</p><p>( m − m A ) = R ( m t r u e − m A )     and     ( d − G m A ) = N ( m t r u e − G m A ) (40)</p><p>Resolution is perfect when R = N = I . When R ≠ I , an element of the reconstructed state vector is a weighted average of all elements of the true state vector. When N ≠ I , an element of the reconstructed data vector is a weighted average of all elements of the true data vector. The resolution matrices quantify resolution in both time and space. As we will show in the example, below, the model resolution (or data resolution) can be temporally poor, even when it is spatially good.</p><p>Another important quantity is the full ( K M ) &#215; ( K M ) posterior covariance matrix, C m = [ F T F ] − 1 . In addition to correlations between elements of the state vector at a given time, it contains the correlations between elements of the state vectors at different times. These coefficients are needed for computing confidence intervals of quantities that depend of the state vectors at two or more times. Although R , N and C m are large matrices, methods are available for efficiently computing selected elements of them using the conjugate gradient method [<xref ref-type="bibr" rid="scirp.118128-ref1">1</xref>].</p><p>Finally, there may be instances in which the covariance, C s ( p ) , may be known only up to some scalar parameter, p ∈ ℝ . For example, the choice [ C s ( p ) ] i j = γ 2 exp { − 1 2 p − 2 ( x i − x j ) 2 } expresses decorrelation over a scale length, p. An initially poor guess of the parameter, p, can be “tuned” using partial derivatives based on the Bayesian principle [<xref ref-type="bibr" rid="scirp.118128-ref15">15</xref>] and the gradient descent method [<xref ref-type="bibr" rid="scirp.118128-ref16">16</xref>].</p></sec><sec id="s8"><title>8. Example</title><p>We consider a data assimilation problem based on the heat diffusion Equation (3), with Δ t = Δ x = 1 and c 0 = 0.4 . The state vector, m ( i ) , represents temperature at position, x , and is of length M = 31 . The source is Gaussian function in space and impulsive function in time, plus additive noise:</p><p>s j ( i ) = exp [ − 1 2 s x − 2 ( x j − x &#175; ) 2 ] δ i 1 + n s ( j ) (41)</p><p>with scale length, s x = 5 , and peak position, x &#175; = 1 2 M Δ x . Here, n s ( j ) is a Normally-distributed random variable with zero mean and variance, σ s 2 = 0.05 . The initial condition is:</p><p>[ m A ( 0 ) ] i = m 0 + n A ( j ) (42)</p><p>where m 0 = 0.1 and n A ( j ) is a Normally-distributed random variable with zero mean and variance, σ d 2 = 0.07 . The dynamical Equation (2) is iterated for K = 61 time steps, to provide the “true” state vector. As expected, the solution has a Gaussian shape with a width that increases, and an amplitude that decreased, with time (<xref ref-type="fig" rid="fig2">Figure 2</xref>(A)). The data are a total of N = 10 temperature measurement at each time, i ≥ 2 , made at randomly-selected positions (without duplications) and perturbed with Normally-distributed random noise with zero mean and variance, σ d 2 = 0.10 .</p><p>GLS solutions were computed by both the full Thomas algorithm (<xref ref-type="fig" rid="fig2">Figure 2</xref>(B)) and by solving the [ F T F ] m = F T f system by the conjugate gradient method (not</p><p>shown). They were found to be identical to machine precision. Present-time solutions were computed for both the present-time Thomas (<xref ref-type="fig" rid="fig2">Figure 2</xref>(C)) and KF algorithms (<xref ref-type="fig" rid="fig2">Figure 2</xref>(D)). They were also found to be identical to machine precision.</p><p>In general, both the GLS and present-time solutions fit the data well. However, the present-time solution matches the true model more poorly than does the GLS solution (<xref ref-type="fig" rid="fig3">Figure 3</xref>). In this numerical experiment, the present-time solution is about 10% poorer than the GLS solution, quantified with the root mean squared deviation from the true solution. However, the percentage, while always positive, varies considerably when the underlying parameters are changed.</p><p>Both the Thomas and Kalman versions of the present-time algorithm are well suited for providing ongoing diagnostic information, such as posterior covariance and root mean squared data fitting error (<xref ref-type="fig" rid="fig4">Figure 4</xref>), which can provide quality control in real time applications.</p><p>The model resolution matrix, R , (<xref ref-type="fig" rid="fig5">Figure 5</xref>(A)) for this exemplary problem has a poorly-populated central diagonal, meaning that some elements of the state vector, m j ( i ) , are well-resolved from their spatial neighbors, m j &#177; 1 ( i ) while others are very poorly resolved. The matrix large elements along other diagonals, corresponding to, offset from the main diagonal by M rows, indicating that some elements are not well resolved from their temporal neighbors, m j ( i &#177; 1 ) . The temporal width of the resolving kernel is about &#177;4, and the shape is very irregular, indicating very uneven averaging is taking place (<xref ref-type="fig" rid="fig6">Figure 6</xref>). The data resolution matrix, N , (<xref ref-type="fig" rid="fig5">Figure 5</xref>(A)) has a well-populated central diagonal, meaning that mosts elements of the predicted data vector, d j p r e ( i ) , are well-resolved from their spatial neighbors, d j &#177; 1 p r e ( i ) . Like R , it also has elements along other diagonals, corresponding to, offset from the main diagonal by M rows, indicating that some elements are not well resolved from their temporal neighbors, d j p r e ( i &#177; 1 ) .</p></sec><sec id="s9"><title>9. Conclusion</title><p>In this paper, we examine a data processing scenario in which real-time data assimilation is performed using Kalman Filtering, and then reanalysis is performed using generalized least squares (GLS). In this problem, spatial characteristics of the system are described by a state vector (mode parameter vector), and its temporal characteristics by the evolution of the state-vector with time. We explore the relationship between the real-time Kalman Filter estimate and the GLS reanalysis estimate of the state vector. We show that the KF solution at a given time is equal to the GLS solution that one would obtain if it excluded data for future times. Furthermore, we show that the recursive procedure in KF is exactly equivalent to the solution of the GLS problem via Thomas’ algorithm for solving the block-tridiagonal matrix that arises in the reanalysis problem. This connection indicates that GLS reanalysis is better considered the final step of a single process, rather than a “different method” arbitrarily being applied, post factor. Now that this connection between KF and GLS has seen identified, the familiar GLS concepts of model and data resolution can be applied to KF. We provide an exemplary problem, based on thermal diffusion. In addition to showcasing our result, the example demonstrates that the state vector and vector of predicted data can be poorly-resolved in time, even when they are well resolved in space.</p></sec><sec id="s10"><title>Acknowledgements</title><p>The author thanks the students in his 2022 Geophysical Inverse Theory course at Columbia University (New York, USA), and especially George Lu, for helpful discussion.</p></sec><sec id="s11"><title>Conflicts of Interest</title><p>The author declares no conflicts of interest regarding the publication of this paper.</p></sec><sec id="s12"><title>Cite this paper</title><p>Menke, W. (2022) Links between Kalman Filtering and Data Assimilation with Generalized Least Squares. Applied Mathematics, 13, 566-584. https://doi.org/10.4236/am.2022.136036</p></sec><sec id="s13"><title>Appendix</title><p>A.1) Proof that C m = [ I − G − g G ] C A ≡ P G C A . This derivation is well known [<xref ref-type="bibr" rid="scirp.118128-ref6">6</xref>], but is presented here for completeness and to point out potential pitfalls should it be applied incorrectly. Because d and m A are independent of one another, m G = G − g ( d − G m A ) + m A = G − g d + ( I − G − g G ) m A , the normal rules of error propagation apply. The posterior covariance, C m , is:</p><p>C m = G − g C d G − g T + ( I − G − g G ) C A ( I − G T G − g T ) = C A + G − g C d G − g T + G − g G C A G T G − g T − C A G T G − g T − G − g G C A = C A + G − g [ G C A G T + C d ] G − g T − C A G T G − g T − G − g G C A = C A + C A G T A − 1 A A − 1 G C A − C A G T A − 1 G C A − C A G T A − 1 G C A = C A + C A G T A − 1 G C A − C A G T A − 1 G C A − C A G T A − 1 G C A = C A − C A G T A − 1 G C A = C A − G − g G C A = [ I − G − g G ] C A ≡ P G C A (A.1)</p><p>Here, we have used the fact that G − g = C A G T A − 1 , with A = [ G C A G T + C d ] . Although P G , has the form of a projection operator, it is a function of C A , and has deceptive properties. Consider the case in which C A = ε S , where S is an invertible symmetric matrix. In the limit of the parameter, ε , becoming indefinitely large, C m does not also become indefinitely large, but rather tends to [ G T C d − 1 G ] − 1 , which is the posterior variance that arises from the data, only. The zero limit tends to zero; that is, when m A is known very accurately, so is m G .</p><p>A.2) Derivation of the TV82-A and TV82-B identities, following [<xref ref-type="bibr" rid="scirp.118128-ref6">6</xref>]. Consider invertible symmetric matrices, C 1 and C 2 , and arbitrary matrix, M . The expression M T + M T C 1 − 1 M C 2 M T can alternately be factored:</p><p>M T C 1 − 1 [ C 1 + M C 2 M T ] = [ C 2 − 1 + M T C 1 − 1 M ] C 2 M T (A.2)</p><p>Multiplying by the inverses yields identity TV82-A:</p><p>C 2 M T [ C 1 + M C 2 M T ] − 1 = [ C 2 − 1 + M T C 1 − 1 M ] − 1 M T C 1 − 1 (A.3)</p><p>Now consider the expression C 2 − C 2 M T [ C 1 + M C 2 M T ] − 1 M C 2 , which by the above identity equals C 2 − [ C 2 + M T C 1 − 1 M ] − 1 M T C 1 − 1 M C 2 . Factoring out the term in brackets</p><p>[ C 2 − 1 + M T C 1 − 1 M ] − 1 [ [ C 2 − 1 + M T C 1 − 1 M ] C 2 − M T C 1 − 1 M C 2 ] (A.4)</p><p>cancelling terms yields identity TV82-B:</p><p>C 2 − C 2 M T [ C 1 + M C 2 M T ] − 1 M C 2 = [ C 2 − 1 + M T C 1 − 1 M ] − 1 (A.5)</p><p>A.3) The Thomas algorithm [<xref ref-type="bibr" rid="scirp.118128-ref10">10</xref>] for a symmetric block-diagonal matrix is well-known; we reproduce it here for completeness. The th row of the matrix has elements B , A i , B T and the right-hand size is a i . Consider the step in the upper-triangularization process when rows ( i − 1 ) and above have been triangularized, but rows ( i − 1 ) and below have not:</p><p>A ^ i − 1 m ( i − 1 ) + B T m ( i ) = a ^ i − 1</p><p>B m ( i − 1 ) + A i m ( i ) + B T m ( i + 1 ) = a i (A.6)</p><p>The second row is modified by multiply the top row by − B A ^ i − 1 − 1 and adding the result to the second, which eliminates the first term, yielding:</p><p>[ A i − B A ^ i − 1 − 1 B T ] m ( i ) + B T m ( i + 1 ) = a i − B A ^ i − 1 − 1 a ^ i − 1 (A.7)</p><p>Note that the new row has two terms, and that the coefficient of the second is always B T , which is the same pattern as the first row. Thus, the bottom row becomes a new top row, and the recursion is</p><p>A ^ 1 − 1 = A 1 − 1     followed   by   A ^ i − 1 = [ A i − B A ^ i − 1 − 1 B T ] − 1</p><p>a ^ 1 = a 1     followed   by     a ^ i = a i − B A ^ i − 1 − 1 a ^ i − 1 (A.8)</p><p>After the recursion, the matrix upper-bidiagonal with diagonals, A ^ i and B T and the right-hand side is a ^ i . It is back-solved as:</p><p>m ( K ) = A ^ K − 1 a ^ K     followed   by     m ( i ) = A ^ i − 1 [ a ^ i − B T m ( i + 1 ) ] (A.9)</p><p>A.4) Verification of the identity in (36):</p><p>[ D C A 1 D T + C s ] − 1 D = ? C s − 1 D [ D T C s − 1 D + C A 1 − 1 ] − 1 D C A 1 − 1</p><p>[ D C A 1 D T + C s ] − 1 D C A 1 = ? C s − 1 D [ D T C s − 1 D + C A 1 − 1 ] − 1</p><p>D C A 1 [ D T C s − 1 D + C A 1 − 1 ] = ? [ D C A 1 D T + C s ] C s − 1 D</p><p>[ D C A 1 D T C s − 1 D + D ] = [ D C A 1 D T C s − 1 D + D ] (A.10)</p></sec></body><back><ref-list><title>References</title><ref id="scirp.118128-ref1"><label>1</label><mixed-citation publication-type="other" xlink:type="simple">Menke, W. (2014) Review of the Generalized Least Squares Method. Surveys in Geophysics, 36, 1-25. https://doi.org/10.1007/s10712-014-9303-1</mixed-citation></ref><ref id="scirp.118128-ref2"><label>2</label><mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Stratonovich</surname><given-names> R.L. </given-names></name>,<etal>et al</etal>. (<year>1959</year>)<article-title>Optimum Nonlinear Systems Which Bring about a Separation of a Signal with Constant Parameters from Noise</article-title><source> Radiofizika</source><volume> 2</volume>,<fpage> 892</fpage>-<lpage>901</lpage>.<pub-id pub-id-type="doi"></pub-id></mixed-citation></ref><ref id="scirp.118128-ref3"><label>3</label><mixed-citation publication-type="other" xlink:type="simple">Zarchan, P. and Musoff, H. (2000) Fundamentals of Kalman Filtering: A Practical Approach. American Institute of Aeronautics and Astronautics, Reston.</mixed-citation></ref><ref id="scirp.118128-ref4"><label>4</label><mixed-citation publication-type="other" xlink:type="simple">Backus, G. and Gilbert, F. (1968) The Resolving Power of Gross Earth Data. Geophysical Journal International, 16, 169-205. https://doi.org/10.1111/j.1365-246X.1968.tb00216.x</mixed-citation></ref><ref id="scirp.118128-ref5"><label>5</label><mixed-citation publication-type="other" xlink:type="simple">Wiggins, R.A. (1972) The General Linear Inverse Problem: Implications of Surface Waves and free Oscillations for Earth Structure. Reviews of Geophysics, 10, 251-285. https://doi.org/10.1029/RG010i001p00251</mixed-citation></ref><ref id="scirp.118128-ref6"><label>6</label><mixed-citation publication-type="other" xlink:type="simple">Tarantola, A. and Valette, B. (1982) Inverse Problems = Quest for Information. Journal of Geophysics, 50, 159-170.</mixed-citation></ref><ref id="scirp.118128-ref7"><label>7</label><mixed-citation publication-type="other" xlink:type="simple">Menke, W. (2018) Geophysical Data Analysis: Discrete Inverse Theory. 4th Edition (Textbook), Elsevier, Amsterdam, 350.</mixed-citation></ref><ref id="scirp.118128-ref8"><label>8</label><mixed-citation publication-type="other" xlink:type="simple">Menke, W. and Menke, J. (2016) Environmental Data Analysis with MATLAB. 2nd Edition, Academic Press (Elsevier), Cambridge, MA, 342 p. https://doi.org/10.1016/C2015-0-01993-1</mixed-citation></ref><ref id="scirp.118128-ref9"><label>9</label><mixed-citation publication-type="other" xlink:type="simple">Menke, W. and Creel, R. (2021) Gaussian Process Regression Reviewed in the Context of Inverse Theory. Surveys in Geophysics, 42, 473-503. https://doi.org/10.1007/s10712-021-09640-w</mixed-citation></ref><ref id="scirp.118128-ref10"><label>10</label><mixed-citation publication-type="other" xlink:type="simple">Dax, A. (2020) The Equivalence between Orthogonal Iterations and Alternating Least Squares. Advances in Linear Algebra &amp; Matrix Theory, 10, 7-21. https://doi.org/10.4236/alamt.2020.102002</mixed-citation></ref><ref id="scirp.118128-ref11"><label>11</label><mixed-citation publication-type="other" xlink:type="simple">Thomas, L.H. (1949) Elliptic Problems in Linear Difference Equations over a Network. Watson Scientific Computing Laboratory Report, Columbia University, New York, 71.</mixed-citation></ref><ref id="scirp.118128-ref12"><label>12</label><mixed-citation publication-type="other" xlink:type="simple">Press, W.H., Teukolsky, S.A., Vetterling, W.T. and Flannery, B.P. (2007) Numerical Recipes: The Art of Scientific Computing. 3rd Edition, Cambridge University Press, New York.</mixed-citation></ref><ref id="scirp.118128-ref13"><label>13</label><mixed-citation publication-type="book" xlink:type="simple">Menke, W. (2005) Case Studies of Seismic Tomography and Earthquake Location in a Regional Context. In: Levander, A. and Nolet, G., Eds., Seismic Earth: Array Analysis of Broadband Seismograms, Vol. 157, American Geophysical Union, Washington DC, 7-36. https://doi.org/10.1029/157GM02</mixed-citation></ref><ref id="scirp.118128-ref14"><label>14</label><mixed-citation publication-type="other" xlink:type="simple">Schlichth&amp;#228;rle, D. (2000) Digital Filters: Basics and Design. Springer, Berlin, 361 p. https://doi.org/10.1007/978-3-662-04170-3</mixed-citation></ref><ref id="scirp.118128-ref15"><label>15</label><mixed-citation publication-type="other" xlink:type="simple">Menke, W. (2021) Tuning of Prior Covariance in Generalized Least Squares. Applied Mathematics, 12, 157-170. https://doi.org/10.4236/am.2021.123011</mixed-citation></ref><ref id="scirp.118128-ref16"><label>16</label><mixed-citation publication-type="other" xlink:type="simple">Snyman, J.A. and Wilke, D.N. (2018) Practical Mathematical Optimization: Basic Optimization Theory and Gradient-Based Algorithms. Springer Optimization and Its Applications, Vol. 133, 2nd Edition, Springer, New York, 340 p. https://doi.org/10.1007/978-3-319-77586-9_9</mixed-citation></ref></ref-list></back></article>