<?xml version="1.0" encoding="UTF-8"?><!DOCTYPE article  PUBLIC "-//NLM//DTD Journal Publishing DTD v3.0 20080202//EN" "http://dtd.nlm.nih.gov/publishing/3.0/journalpublishing3.dtd"><article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" dtd-version="3.0" xml:lang="en" article-type="research article"><front><journal-meta><journal-id journal-id-type="publisher-id">AM</journal-id><journal-title-group><journal-title>Applied Mathematics</journal-title></journal-title-group><issn pub-type="epub">2152-7385</issn><publisher><publisher-name>Scientific Research Publishing</publisher-name></publisher></journal-meta><article-meta><article-id pub-id-type="doi">10.4236/am.2017.811117</article-id><article-id pub-id-type="publisher-id">AM-80553</article-id><article-categories><subj-group subj-group-type="heading"><subject>Articles</subject></subj-group><subj-group subj-group-type="Discipline-v2"><subject>Physics&amp;Mathematics</subject></subj-group></article-categories><title-group><article-title>
 
 
  Stability Analysis of a Neurocontroller with Kalman Estimator for the Inverted Pendulum Case
 
</article-title></title-group><contrib-group><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Julián</surname><given-names>Antonio Pucheta</given-names></name><xref ref-type="aff" rid="aff1"><sup>1</sup></xref><xref ref-type="corresp" rid="cor1"><sup>*</sup></xref></contrib><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Cristian</surname><given-names>Rodríguez Rivero</given-names></name><xref ref-type="aff" rid="aff1"><sup>1</sup></xref><xref ref-type="corresp" rid="cor1"><sup>*</sup></xref></contrib><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Carlos</surname><given-names>Alberto Salas</given-names></name><xref ref-type="aff" rid="aff2"><sup>2</sup></xref><xref ref-type="corresp" rid="cor1"><sup>*</sup></xref></contrib><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Martín</surname><given-names>Herrera</given-names></name><xref ref-type="aff" rid="aff2"><sup>2</sup></xref><xref ref-type="corresp" rid="cor1"><sup>*</sup></xref></contrib><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Sergio</surname><given-names>Oscar Laboret</given-names></name><xref ref-type="aff" rid="aff1"><sup>1</sup></xref><xref ref-type="corresp" rid="cor1"><sup>*</sup></xref></contrib></contrib-group><aff id="aff2"><addr-line>Departamento de Ingeniería Electrónica, Facultad de Tecnología y Ciencias Aplicadas, Universidad Nacional de Catamarca,
Catamarca, Argentina</addr-line></aff><aff id="aff1"><addr-line>Departamentos de Electrotecnia y de Electrónica, Laboratorio de Investigación Matemática Aplicada a Control (LIMAC),
Facultad de Ciencias Exactas, Físicas y Naturales-Universidad Nacional de Córdoba, Córdoba, Argentina</addr-line></aff><author-notes><corresp id="cor1">* E-mail:<email>jpucheta@unc.edu.ar(JAP)</email>;<email>cristian.rodriguezrivero@gmail.com(CRR)</email>;<email>calberto.salas@gmail.com(CAS)</email>;<email>ing_martin_herrera@yahoo.com.ar(MH)</email>;<email>slaboret@yahoo.com.ar(SOL)</email>;</corresp></author-notes><pub-date pub-type="epub"><day>01</day><month>11</month><year>2017</year></pub-date><volume>08</volume><issue>11</issue><fpage>1602</fpage><lpage>1618</lpage><history><date date-type="received"><day>28,</day>	<month>September</month>	<year>2017</year></date><date date-type="rev-recd"><day>21,</day>	<month>November</month>	<year>2017</year>	</date><date date-type="accepted"><day>24,</day>	<month>November</month>	<year>2017</year></date></history><permissions><copyright-statement>&#169; Copyright  2014 by authors and Scientific Research Publishing Inc. </copyright-statement><copyright-year>2014</copyright-year><license><license-p>This work is licensed under the Creative Commons Attribution International License (CC BY). http://creativecommons.org/licenses/by/4.0/</license-p></license></permissions><abstract><p>
 
 
  
    In this paper, a practical analysis of stability by simulation for the effect of incorporating a Kalman estimator in the control loop of the inverted pendulum with a neurocontroller is presented. The neurocontroller is calculated by approximate optimal control, without considering the Kalman estimator in the loop following the Theorem of the separation. The results are compared with a time-varying linear controller, which in noiseless conditions in the state or in the measurement has an acceptable performance, but when it is under noise conditions its operation closes into a state space range more limited than the one proposed here. 
  
 
</p></abstract><kwd-group><kwd>Optimal Control</kwd><kwd> Neurocontroller</kwd><kwd> Inverted Pendulum</kwd><kwd> State Estimation</kwd><kwd> Stability</kwd></kwd-group></article-meta></front><body><sec id="s1"><title>1. Introduction</title><p>The motivation for this case study known as inverted pendulum control arises from the need to obtain robust controller systems to implement in situations where it is desired to maintain equilibrium of an unstable system. A direct related situation is the attitude control of a booster rocket at takeoff for sending a payload to space. It is a well-known problem in the control theory literature [<xref ref-type="bibr" rid="scirp.80553-ref1">1</xref>] [<xref ref-type="bibr" rid="scirp.80553-ref2">2</xref>] and machine learning [<xref ref-type="bibr" rid="scirp.80553-ref3">3</xref>] [<xref ref-type="bibr" rid="scirp.80553-ref4">4</xref>] [<xref ref-type="bibr" rid="scirp.80553-ref5">5</xref>] [<xref ref-type="bibr" rid="scirp.80553-ref6">6</xref>] . However, such problems are challenging since real systems are difficult to control and this is to some extent due to the fact that redundant feedback systems must be considered by the controller, as an effect similar to that of incorporating an estimator into the variables controller status. This fact causes instability in the closed loop, which must be foreseen and analyzed. The analysis can be done through simulations with the estimator-controller system, in order to establish some stability domain. In this work we opt for the control based on optimization [<xref ref-type="bibr" rid="scirp.80553-ref7">7</xref>] [<xref ref-type="bibr" rid="scirp.80553-ref8">8</xref>] [<xref ref-type="bibr" rid="scirp.80553-ref9">9</xref>] , where the optimal control problem is formulated. To solve the problem of optimal control, a very powerful tool is the Dynamic Programming technique [<xref ref-type="bibr" rid="scirp.80553-ref10">10</xref>] [<xref ref-type="bibr" rid="scirp.80553-ref11">11</xref>] implemented with approximations [<xref ref-type="bibr" rid="scirp.80553-ref3">3</xref>] [<xref ref-type="bibr" rid="scirp.80553-ref4">4</xref>] [<xref ref-type="bibr" rid="scirp.80553-ref5">5</xref>] in a machine learning scheme [<xref ref-type="bibr" rid="scirp.80553-ref12">12</xref>] , since it allows dealing with constrained, nonlinear processes and non-quadratic performance indexes. However, it is often difficult to achieve a methodology for the implementation of controllers based on machine learning, since they require heuristic and a good knowledge of the involved adaptation mechanisms [<xref ref-type="bibr" rid="scirp.80553-ref3">3</xref>] . In this work, a methodology to determine the conditions that achieve good results to implement in simulation is shown. A controller consisting of a compact function called neurocontroller is achieved.</p><p>In this paper, cases with and without estimator of a model that represents an inverted pendulum are studied. When a time varying linear quadratic regulator (TVLQR) with direct state measurement is used, good performance can be achieved. However, it can be improved with a neurocontroller. The obtained performances by using linear and neurocontroller are shown in <xref ref-type="fig" rid="fig1">Figure 1</xref> and <xref ref-type="fig" rid="fig2">Figure 2</xref>, respectively. Note that the cumulative cost of the linear controller</p><p>(3760.7) is 28% higher than that of the Neurocontroller (2691.3). However, when the controller is used in more realistic situations using a state estimator and considering noisy conditions in the measurements, the performance of the linear controller deteriorates more than the performance of the neurocontroller even until fails to stabilize the system for the same initial conditions. In this paper, an analysis of the system performance deterioration is shown when it requires a state estimator.</p><p>This paper is organized as follows. After this Introduction, the problem is detailed and expressed as mathematical equation in Section 2. In Section 3 is detailed the proposed solution. In Section 4 the implementation of the obtained solution and another one with classical methods for comparison purposes is developed. The obtained results are discussed in Section 5, with its pros and cons.</p></sec><sec id="s2"><title>2. Problem Formulation</title><p>The dynamic programming approach assumes that the process evolution can be split in stages [<xref ref-type="bibr" rid="scirp.80553-ref10">10</xref>] , so take the version of dynamic systems in discrete time [<xref ref-type="bibr" rid="scirp.80553-ref9">9</xref>] is straightforward. The problem formulation puts in formal terms the optimal control elements. These elements are the cost function to minimize, the control law and the dynamic system model with its constraints. If the system model cannot be or is not feasible to express it in closed analytical form through a differential equation, it is useful to generate a black box model [<xref ref-type="bibr" rid="scirp.80553-ref13">13</xref>] .</p><p>This section introduces the nomenclature used along the article. The symbols are listed and explained.</p><p>i. k ∈ ℕ describes the discrete time variable.</p><p>ii. x(k), with x ( k ) ∈ ℜ n is the time dependent n-dimensional state vector whose components has the system’s state variables. These variables describe the process dynamics over discrete time.</p><p>iii. f is a nonlinear continuous function, f : ℜ n &#215; ℜ m → ℜ n that describes the relation between the state vector for two time instants.</p><p>iv. u ( k ) ∈ ℜ m is the system input or manipulated vector.</p><p>v. I is a convex function, I : ℜ n &#215; ℜ m &#215; ℕ → ℜ + called the performance index designed by the control engineer.</p><p>vi. J is a convex function, J : ℜ n &#215; ℜ m &#215; ℕ → ℜ + named the cost function designed by the control engineer.</p><p>vii. X is a bounded and closed set, X ⊂ ℜ n .</p><p>viii. U is a bounded and closed set, U ⊂ ℜ m .</p><p>ix. m is the control law or decision policy μ : ℜ n → ℜ m , this function maps each state vector value with a control action u.</p><p>x. r and v are real value arrays of parameters r , v ∈ ℜ n + h + 2 , where h ∈ ℕ is determined by the control engineer.</p><p>xi. J ˜ is the approximation of the cost function J, and its domains includes the parameter vector r, J ˜ : ℜ n &#215; ℜ m &#215; ℜ n + h + 2 → ℜ + .</p><p>xii. μ ˜ is a function whose behavior approximates the function m, includes the parameter vector v, μ ˜ : ℜ n &#215; ℜ n + h + 2 → ℜ m .</p><p>xiii. J * ( x ( k ) ) is the minimum cost to go from the state x at time k up to the terminal state at time N.</p><p>xiv. u o ( k ) ∈ ℜ m is the optimal control action at time k.</p><p>xv. ξ 1 , ξ 2 , ⋯ , ξ h are real scalar values.</p><p>xvi. S ˜ data set of samples from the process under study.</p><p>xvii. C<sub>(</sub><sub>i</sub><sub>) </sub>is the cost associate to a control law evolving from the state i up to the terminal process state.</p><p>xviii. J ( i ) μ is the value of the cost to go function obtained after use the control law m starting at state i up to terminal state at time N.</p><p>xix. &#209; gradient operator.</p><p>xx. Q ( i , u ) , real valued function associated at state i and action u.</p><p>xxi. η n is a function that varies with iteration number n, bounded between 0 and 1.</p><p>xxii. Q ˜ ( i , u ) is the approximate version of the factor Q ( i , u ) .</p><p>xxiii. γ n is the discount factor, variable with iteration n and bounded between 0 and 1.</p><p>xxiv. μ &#175; control action expressed as look up table from every state.</p><p>xxv. J ˜ μ ( ⋅ , r ) approximate cost to go function associated with the control law m.</p><p>xxvi. A ∈ ℜ n &#215; n is the state matrix for the linear dynamic model.</p><p>xxvii. B ∈ ℜ n &#215; 1 is the input matrix of the linear dynamic model.</p><p>xxviii. F ∈ ℜ n &#215; n is the additive noise model at the state variables.</p><p>xxix. v ( k ) ∈ ℜ n is the random sequence with Gaussian distribution in each variable, zero mean and unit variance.</p><p>xxx. y ( k ) ∈ ℜ is the linear model output.</p><p>xxxi. C ∈ ℜ 1 &#215; n is the linear model output matrix.</p><p>xxxii. G ∈ ℜ is the noise model at the measured variable.</p><p>xxxiii. w ( k ) ∈ ℜ is a white noise sequence with zero mean and unit variance.</p><p>xxxiv. δ ∈ ℜ is longitudinal displacement of the cart.</p><p>xxxv. δ ˙ ∈ ℜ is longitudinal velocity of the cart.</p><p>xxxvi. ϕ ∈ ℜ is the angle of the inverted pendulum bar.</p><p>xxxvii. ϕ ˙ ∈ ℜ is the angular velocity of the inverted pendulum bar.</p><p>xxxviii. M<sub>P</sub> is the cart concentrated mass, whose value here is 0.5 Kgr.</p><p>xxxix. m<sub>P</sub> is the bar concentrated mass, valued here is 0.1 Kgr.</p><p>xl. F<sub>P</sub> is the displacement friction constant assigned 0.1 N∙m<sup>−1</sup>∙s.</p><p>xli. l<sub>P</sub> is the size of the pendulum bar, 0.6 m.</p><p>xlii. g<sub>P</sub> is the standard acceleration due to gravity, 9.81 m∙s<sup>−2</sup>.</p><p>xliii. Q ∈ ℜ 4 &#215; 4 is the weighing matrix for the state vector from k = 0 to k = N − 1 with N the terminal state time.</p><p>xliv. S ∈ ℜ 4 &#215; 4 is the weighing matrix for the state vector at the terminal state time N.</p><p>xlv. R ∈ ℜ is the weighing matrix for the control action variable.</p><p>Thus, to formulate the optimal control problem the expressions of the process model in discrete time, the restrictions in the variables and the cost function to be minimized are presented. Next, the problem of minimizing the separable cost function is considered by</p><p>x ( k + 1 ) = f ( x ( k ) , u ( k ) , k ) ,   k = 0 , 1 , ⋯ , N − 1 (1)</p><p>where x(0) has a fixed value and the constraints must be satisfied together with the system equation,</p><p>J ( x ( k ) , u ( k ) ) = ∑ k = 0 N I ( x ( k ) , u ( k ) ) (2)</p><p>where the constraints on state and manipulated variables are</p><p>x ∈ X ⊂ ℜ n , u ∈ U ⊂ ℜ m . (3)</p><p>The function I(&#215;) is defined by the control engineer which must be convex but not necessarily quadratic, and f(&#215;) is the nonlinear relationship between instants k and k + 1 of the state and manipulated variables. Moreover, they are bounded and continuous functions of their arguments, and both x and u belong to closed and bounded subsets of &#194;<sup>n</sup> and &#194;<sup>m</sup>, respectively. Then, the Weierstrass theorem asserts that there exists a minimization policy also called control law. Therefore, it is desired to find a correspondence relation</p><p>μ ( x ( k ) ) : ℜ n → ℜ m (4)</p><p>that makes evolve the processes modelled by (2) from any initial condition to the final terminal state x(N) satisfying constraints (3), and minimizing the cost function (1). The implementation is shown in <xref ref-type="fig" rid="fig3">Figure 3</xref>, where the flow of information between the controller and the closed-loop system is stated. Note that the behavior of the closed loop system is done by designing the performance index which is added at each stage in the cost function (1).</p></sec><sec id="s3"><title>3. Proposed Solution</title><p>In order to solve the formulated problem, the proposed solution is by using dynamic programming and then approximations are introduced through functions</p><p>μ ˜ ( x ( k ) , v ) : ℜ n → ℜ m , (5)</p><p>J ˜ ( x , u , r ) : ℜ n &#215; m &#215; N → ℜ + (6)</p><p>where the parameter vectors v and r must be determined.</p><p>The procedure to solve the optimal control problem for both continuous and discrete time dynamic systems is well known [<xref ref-type="bibr" rid="scirp.80553-ref7">7</xref>] [<xref ref-type="bibr" rid="scirp.80553-ref8">8</xref>] [<xref ref-type="bibr" rid="scirp.80553-ref14">14</xref>] , and consists of analytically minimize the proposed cost function (1) and from this minimization achieve an expression for function μ. When the system is linear and the cost function is quadratic, the optimal control problem has unique solution through the Riccati Equation. However, when the system is nonlinear the solution of the Hamilton-Jacobi-Bellman equation [<xref ref-type="bibr" rid="scirp.80553-ref14">14</xref>] must be found, whose solution is restricted to a certain class of nonlinear systems. Here, an optimization principle to solve the same control problem that allows to use any cost function and respecting the constraints in the state variables and in the control variables in a natural way is used.</p><p>The principle of optimality [<xref ref-type="bibr" rid="scirp.80553-ref10">10</xref>] allows solving an optimization problem in</p><p>which a dynamic process evolves over time through stages. Applying the principle of optimality in (1), we obtain</p><p>J * ( x ( k ) ) = min u ( k ) J ( x ( k ) , u ( k ) )                           = min u ( k ) { I ( x ( k ) , u ( k ) ) + J * ( f ( x ( k ) , u ( k ) ) ) } , (7)</p><p>called the Bellman’s Equation. Therefore, the optimal control action u<sup>o</sup> will be</p><p>u o ( k ) = arg min u ( k ) { I ( x ( k ) , u ( k ) ) + J * ( f ( x ( k ) , u ( k ) ) ) } , (8)</p><p>which is the optimal policy of decisions or optimal control law. Note that J<sup>*</sup> does not depend explicitly on u(k), as shows Equation (8).</p><p>To obtain the control law or the decision policy, there exists numerical methods, [<xref ref-type="bibr" rid="scirp.80553-ref3">3</xref>] [<xref ref-type="bibr" rid="scirp.80553-ref4">4</xref>] [<xref ref-type="bibr" rid="scirp.80553-ref10">10</xref>] [<xref ref-type="bibr" rid="scirp.80553-ref11">11</xref>] and approximations [<xref ref-type="bibr" rid="scirp.80553-ref3">3</xref>] [<xref ref-type="bibr" rid="scirp.80553-ref5">5</xref>] [<xref ref-type="bibr" rid="scirp.80553-ref12">12</xref>] which are detailed below. Now, an approximation function for values of Equation (1) in a compact domain is introduced. Thus, a compact representation of the cost associated with each state of the process is obtained.</p><p>The approximation function incorporates a set of vectors of parameters r, which is defined as a partitioned vector whose structure defines the function structure,</p><p>r = { r 1 1 , r 1 2 , ⋯ , r 1 h , r 2 } (9)</p><p>where each vector r<sub>1</sub> has the same dimension, which is the number of inputs of the function plus one to consider a static scalar unit parameter. So, h intermediate scalar values ξ are computed as the scalar product between the input vector x and the corresponding parameters as</p><p>ξ 1 = [ x T 1 ] ⋅ r 1 1 , (10)</p><p>ξ 2 = [ x T 1 ] ⋅ r 1 2 , (11)</p><p>ξ h = [ x T 1 ] ⋅ r 1 h , (12)</p><p>every single value are processed through the hyperbolic tangent function, avoiding large numbers by</p><p>f ( ξ ) = exp ( ξ ) − exp ( − ξ ) exp ( ξ ) + exp ( − ξ ) = 1 − 2 1 + exp ( 2 ξ ) (13)</p><p>where the right side has only one exp(&#215;) computation for improving calculation time. So, with these h values together with the polarization 1 the inner product is implemented with the rest of the r parameter vector which is r<sub>2</sub>, and must be consistent in its dimension to be able to perform the product</p><p>μ ˜ ( x ( k ) , r ) = [ f ( ξ 1 ) f ( ξ 2 ) ⋯ f ( ξ h ) 1 ] T ⋅ r 2 . (14)</p><p>This approximation function has the parameter h as a tuning parameter for the dimension of vector r, in terms of the structure of the approximation function.</p><p>Finding a suitable value for vector r, one have the approximated value of the minimum cost that is incurred to reach the terminal state from the current state x(k), and with the model of the system can be found the control policy using the argument u(k) that minimizes</p><p>J * ( x ( k ) ) = min u k { I ( x ( k ) , u ( k ) ) + J ˜ * ( f ( x ( k ) , u ( k ) ) , r ) } . (15)</p><p>For finding r, the search process that finds the policy function is divided into two tasks, as shown in <xref ref-type="fig" rid="fig4">Figure 4</xref>. One of them, is the evaluation of a defined control policy or control law, from which the costs for all stages of the process are calculated. The other one, is the policy improvement procedure. Both tasks are done approximately with respect to the original system, because an approximation function is used that tune its behavior.</p><p>A set of representative data S ˜ in the state space in a domain is available and for each state i ∈ S ˜ the cost values C(i) are calculated. To this end, an initial control law or control policy is proposed, and the system (1) is evolved from the given state i to the terminal stage, evaluating the performance index by expression (2) of the cost function J ( i ) μ . This procedure is performed for every state i ∈ S ˜ . Then the approximated cost function is tuned by minimizing in r</p><p>min r ∑ i ∈ S ˜ ( J ˜ ( i , r ) − C ( i ) ) 2 (16)</p><p>an approximation function for the cost associated to the evaluated policy is obtained. The parameters vector r is obtained by minimizing expression (16). The incremental gradient iteration is</p><p>r : = r + η n ∇ J ˜ ( i , r ) ( C ( i ) − J ˜ ( i , r ) ) , ∀ i ∈ S ˜ (17)</p><p>where η fulfills the conditions</p><p>∑ n = 0 ∞ η n ( i , u ) = ∞ , ∑ n = 0 ∞ η n 2 ( i , u ) &lt; ∞ , ∀ i , u ∈ U ( i ) , (18)</p><p>such that the algorithm converges, where n refers to the tuning iteration n. For computing C(i) the performance evaluation is implemented through the system model (1) and the cost function (2).</p><p>Then the costs associated with each state-action pair are computed, by using</p><p>the auxiliary cost function Q(i, u), which in its approximate version is</p><p>Q ˜ ( i , u ) = I ( i , u ) + γ n J ˜ ( j , r ) (19)</p><p>where g<sub>n</sub> is a discount factor that can vary from iteration to iteration up to reach unity. Then, the improved policy is obtained by the table</p><p>μ &#175; ( i ) = arg min u ∈ U ( i ) ( I ( i , u ) + γ n J ˜ ( j , r ) ) , ∀ i ∈ S ˜ . (20)</p><p>Once available J ˜ μ ( ⋅ , r ) , it can be obtained μ &#175; ( ⋅ ) from Equation (20). Then the costs associated with each state, symbolized by C(i), are evaluated by Equation (17), and the r parameters are tuned by obtaining a new version of the approximation function J ˜ μ ( i , r ) , ∀ i ∈ S ˜ . Then, the policy improvement task is carried out, in which a new tabulated control policy μ &#175; ( ⋅ ) expressed as (20) is obtained. After that, the calculation of the costs for each state i starts, and in each iteration the function g<sub>n</sub> is updated.</p><p>Simultaneously with the described tasks, an approximation for the improved control law μ &#175; ( ⋅ ) is introduced, by a function with parameters v as shown <xref ref-type="fig" rid="fig5">Figure 5</xref> following the same structure as that described by Equations (9)-(14).</p><p>Thus, since the function μ ( ⋅ ) is the analytical solution of the optimal control problem, it is intended to obtain an approximation μ ˜ ( ⋅ , v ) of the function μ ( ⋅ ) -which is expressed as table, where v is the parameter vector.</p><p>To find the approximation function μ ˜ ( ⋅ , v ) , using the data of the improved policy μ &#175; ( ⋅ ) defined in (20), it is proposed to minimize the expression</p><p>min v ∑ i ∈ S ( μ ˜ ( i , v ) − μ &#175; ( i ) ) 2 (21)</p><p>within the set S ˜ , where the control law is represented by μ ˜ ( ⋅ , v ) with the tuning parameters vector v. A solution for Equation (21) is obtained by the incremental gradient method [<xref ref-type="bibr" rid="scirp.80553-ref13">13</xref>] , which is expressed as iterations on n by</p><p>v : = v + η n ∇ μ ˜ ( i , v ) ( μ &#175; ( i ) − μ ˜ ( i , v ) ) , ∀ i ∈ S ˜ (22)</p><p>where η<sub>n</sub> fulfills the conditions (18). A summary of the algorithm is detailed in <xref ref-type="table" rid="table1">Table 1</xref>.</p><p>Note that two approximation problems are solved at the same time, since given μ &#175; ( ⋅ ) it is evaluated to find J ˜ μ ( ⋅ , r ) of J μ ( ⋅ ) defined by C(i) with i ∈ S ˜ .</p><p>Then, given J ˜ μ ( ⋅ , r ) , the improved policy μ &#175; ( i ) is computed for i ∈ S ˜ and then find the new policy μ ˜ ( ⋅ , v ) . Once the function μ ˜ ( ⋅ , v ) is available, the control actions are obtained as shown in <xref ref-type="fig" rid="fig5">Figure 5</xref>. The control scheme is shown in <xref ref-type="fig" rid="fig6">Figure 6</xref>.</p><p>The algorithm to solve the optimal control problem for nonlinear processes with non-quadratic cost function and constraints was detailed. Given the employment of approximations, the topic of approximation function in dynamic systems [<xref ref-type="bibr" rid="scirp.80553-ref13">13</xref>] must be well mastered to obtain suitable result in the closed loop system.</p><p>As general suggestions, it must be mentioned that as in many nonlinear system, the algorithm is strongly dependent on the initial conditions. Thus, its dependence lies on the initial policy and on the states used to compute μ ˜ ( ⋅ , v ) , represented by set S ˜ .</p><p>The parameter tune speed with respect of the iterations, is fixed by the function γ, and the method is sensitive to this parameter. Usually one can make the first attempts setting γ = 1 constant, with few iterations, and then begin to modify it to converge to 1 with the iterations, always verifying that the performance of the controller improves at the long term. The adjustment parameters amount in each approximation function depends on the data complexity, which generally are conditioned by implementing some normalization or feature extraction techniques.</p></sec><sec id="s4"><title>4. Control of the Inverted Pendulum</title><p>The inverted pendulum can be represented as shown <xref ref-type="fig" rid="fig7">Figure 7</xref>. For this cart-bar</p><p>system, a controller will be designed using the algorithm of <xref ref-type="table" rid="table1">Table 1</xref>. Knowing that the equations that describe the angle and the linear displacement dynamics are,</p><p>{ ( M P + m P ) δ &#168; + m P l P ϕ &#168; cos φ − m P l P ϕ ˙ 2 sin ϕ + F P δ ˙ = u l P ϕ &#168; − g P sin ϕ + δ &#168; cos ϕ = 0 (23)</p><p>whereas the controller is designed the system trajectories are generated by simulation for initial angle f of 0.2 radians. It is considered that the force u must fulfill with the constraint − 30 ≤ u ≤ 30 .</p><p>The proposed cost function is composed by</p><p>J ( x , u ) = ∑ k = 0 N x k T [ 5 0 0 0 0 0 0 0 0 0 50 0 0 0 0 0 ] x k + 0.001 ⋅ θ u (24)</p><p>where q<sub>u</sub> is defined to constrain the values of u<sub>k</sub> by</p><p>θ u = { | u k − 30 | ,if u k &gt; 30 , | − 30 − u k | ,if u k &lt; − 30. (25)</p><p>The continuous time model is discretized at a rate of 0.1 Section.</p><p>In order to retrieve the system state vector, a Kalman estimator is used, where the discrete-time linearized version estimate of (23) is given by</p><p>x k + 1 = A x k + B u k + F v k (26)</p><p>y k = C x k + G w k (27)</p><p>where x T = [ δ , δ ˙ , ϕ , ϕ ˙ ] , u<sub>k</sub> is the operating force as shows <xref ref-type="fig" rid="fig7">Figure 7</xref>, and the matrices A ∈ ℜ 4 &#215; 4 and B ∈ ℜ 4 &#215; 1 are obtained by discretizing the continuous time linear version of (23) with a sampling time of 0.1 Section For the case of the pendulum v ∈ ℜ 4 , w ∈ ℜ 1 , assuming that Gaussian sequences have zero mean and unit variance. The matrices F and G are defined as</p><p>F = [ σ 11 0 0 0 0 σ 22 0 0 0 0 σ 33 0 0 0 0 σ 44 ] (28)</p><p>with σ i i = 1 &#215; 10 − 3 for the state vector x, and</p><p>G = [ ς ] (29)</p><p>with ς = 1 &#215; 10 − 3 for the measure y(k). Furthermore, C = [1 0 0 0].</p><p>To find the x(k) estimate, it used a priori estimate of the observed states by means of</p><p>x ^ ( k + 1 ) − = A x ^ ( k ) + B u ( k ) (30)</p><p>and these states x ^ are obtained from measurements of the system output</p><p>x ^ ( k ) = x ^ ( k ) − + K O ( y ( k ) − C x ^ ( k ) − ) (31)</p><p>where K<sub>O</sub> is the Kalman gain [<xref ref-type="bibr" rid="scirp.80553-ref15">15</xref>] . This gain is calculated by using a Gaussian noise model in the state x(k) and in the measurement y(k) given by (26) and (27).</p><p>The tune of the <xref ref-type="table" rid="table1">Table 1</xref> algorithm was done to achieve the control objective that is to bring the bar to the vertical position, starting from positions smaller than 1 radian. <xref ref-type="fig" rid="fig1">Figure 1</xref>0 shows the evolution of the algorithm parameters which are the cost to go function from the initial condition [0 0 0.2 0]<sup>T</sup> to the final time, set in 10 sec. Note that the behavior is not stable in first half of the performed iterations, but then stabilizes and the control objective is achieved.</p><p>The set S ˜ was of 3000 samples in the &#194;<sup>4</sup> space of the state variables, with the range shown in <xref ref-type="fig" rid="fig1">Figure 1</xref>1. The control law is</p><p>u ( k ) = μ ˜ ( x ^ ( k ) , v ) (32)</p><p>where x ^ ( k ) is obtained from Equation (31), and v contains the parameters corresponding to a set as (9), where 7 hidden nodes were used, which gives 7 vectors r 1 i ∈ ℜ 5 , and vector r 2 ∈ ℜ 8 , implementing the Equation (14). For the approximation of the function J(⋅) defined in Equation (24) it is used the same structure of the approximation function as for the control law. The parameters tuning was performed by the Levenberg Marquardt algorithm [<xref ref-type="bibr" rid="scirp.80553-ref5">5</xref>] .</p><p>In order to perform the comparison of the neurocontroller performance, a time variant-discrete linear quadratic regulator controller with the classical LQR theory in discrete time [<xref ref-type="bibr" rid="scirp.80553-ref16">16</xref>] (TVDLQR) was implemented. Here, the design matrices were</p><p>Q = [ 10 0 0 0 0 10 3 0 0 0 0 10 0 0 0 0 10 3 ] , S = [ 10 0 0 0 0 10 0 0 0 0 10 0 0 0 0 10 2 ] and   R = 1 . (33)</p><p><xref ref-type="fig" rid="fig8">Figure 8</xref> and <xref ref-type="fig" rid="fig1">Figure 1</xref>2 show both performance results of the neurocontroller and the TVDLQR, with the same Kalman estimator. Under no-noise conditions, the performance of the system in both cases are quite similar, as shown in <xref ref-type="fig" rid="fig1">Figure 1</xref> and <xref ref-type="fig" rid="fig2">Figure 2</xref>. Nevertheless, under noise conditions the TVDLQR does not achieve the same performance as that of the neurocontroller since the last allows to increase the range of initial conditions from 0.19 to 0.47 rad as seen in <xref ref-type="fig" rid="fig1">Figure 1</xref>3 and <xref ref-type="fig" rid="fig1">Figure 1</xref>2. Furthermore, estimated variables used by the LQR controller are shown in <xref ref-type="fig" rid="fig9">Figure 9</xref>.</p><p>Since the control objective of the system with estimator is that the pendulum does not fall, a qualitative analysis of the performance of the TVLQR and NC controllers can be inferred. As can be seen in the examples shown, in <xref ref-type="fig" rid="fig8">Figure 8</xref> can be seen that the linear controller meets the control objective for initial conditions of 0.19 rad or less. In contrast, for the case of the NC in <xref ref-type="fig" rid="fig1">Figure 1</xref>3 and <xref ref-type="fig" rid="fig1">Figure 1</xref>2 shows that with initial conditions of 0.19 radians the linear controller does not meet the control objective, but the NC does. In <xref ref-type="fig" rid="fig1">Figure 1</xref>0 can be see that the tuning parameter’s procedure is erratic and difficult to adjust since the value of the costs associated with the control policies are not necessarily monotonically decreasing. This means that is hard to tune the algorithm of <xref ref-type="table" rid="table1">Table 1</xref>, which must be tuned by trial-test and error for each particular system. Also, in <xref ref-type="fig" rid="fig1">Figure 1</xref>1 the range of the system samples used in the calculation of the NC to</p><p>implement the algorithm of <xref ref-type="table" rid="table1">Table 1</xref> is detailed. Note that the system evolution can be out of range and the policy function must be able to give a response that stabilizes the system. This was achieved because of the suitable relation between the data set S ˜ and the problem complexity. Therefore, the results are encouraging since in the examples the control objective is met, which is to obtain zero error in the output with respect to the desired value that is the origin. In a comparison of the performances from both control systems can be seen in <xref ref-type="table" rid="table2">Table 2</xref>. It is important to state that the linear controller is unable to keep the pendulum vertical when the angle initial condition exceeds 0.2 rad whereas the NC achieves it even up to 0.47 rad.</p><table-wrap id="table1" ><label><xref ref-type="table" rid="table2">Table 2</xref></label><caption><title> Comparative figures of performance results obtained by both controller estimator systems. The Monte Carlo simulation was realized with 150 trays, along 1000 time steps of 0.1sec. <xref ref-type="fig" rid="fig8">Figure 8</xref> and <xref ref-type="fig" rid="fig1">Figure 1</xref>2 show the temporal evolution of these indices</title></caption><table><tbody><thead><tr><th align="center" valign="middle" >Item</th><th align="center" valign="middle" >Do</th></tr></thead><tr><td align="center" valign="middle" >1</td><td align="center" valign="middle" >Initialize: g<sub>n</sub>, Iterations, h, parameters r and v.</td></tr><tr><td align="center" valign="middle" >2</td><td align="center" valign="middle" >Evaluate the initial policy μ ˜ ( ⋅ , v ) through (16).</td></tr><tr><td align="center" valign="middle" >3</td><td align="center" valign="middle" >Update the parameters r via (17)</td></tr><tr><td align="center" valign="middle" >4</td><td align="center" valign="middle" >Compute functions Q(&#215;) by (19)</td></tr><tr><td align="center" valign="middle" >5</td><td align="center" valign="middle" >Update policy μ &#175; ( ⋅ ) using (20)</td></tr><tr><td align="center" valign="middle" >6</td><td align="center" valign="middle" >Update the parameters v of μ ˜ ( ⋅ , v ) by (22)</td></tr><tr><td align="center" valign="middle" >7</td><td align="center" valign="middle" >Update function g<sub>n</sub>.</td></tr><tr><td align="center" valign="middle" >8</td><td align="center" valign="middle" >Go to 2 and repeat items 2, 3, 4, 5, 6 and 7 until the complete the Iterations</td></tr></tbody></table></table-wrap><p>The temporal evolution of these indices is shown in <xref ref-type="fig" rid="fig8">Figure 8</xref> and <xref ref-type="fig" rid="fig1">Figure 1</xref>2, where the mean value and the 66% quota from the 150 trajectories of the Monte Carlo simulation are highlighted. In <xref ref-type="table" rid="table2">Table 2</xref> are resumed the performance achieved by each controller. Note that even with higher cost to go figures, the NC gives robust behavior with regards to the initial conditions and noise.</p></sec><sec id="s5"><title>5. Conclusions</title><p>In this paper, a stability analysis of a neurocontroller with Kalman estimator for the inverted pendulum case was presented. The NC performance was compared against the TVLQR controller with the same Kalman estimator.</p><p>The obtained results can be stated that the use of approximate optimal control implemented through <xref ref-type="table" rid="table1">Table 1</xref>, is a very powerful and promising tool to deal with controllers for nonlinear processes with Kalman estimator.</p><p>The initial angle range was possible to extend from 0.19 rad for the case of linear controller with estimator, to 0.47 rad in the present scheme simulating a Monte Carlo with 150 trajectories.</p><p>It is important to highlight that the technique requires a good mastery for functions approximation for dynamic processes, and simulation of natural process through numerical methods.</p></sec><sec id="s6"><title>Cite this paper</title><p>Pucheta, J., Rodr&#237;guez Rivero, C., Salas, C., Herrera, M. and Laboret, S. (2017) Stability Analysis of a Neurocontroller with Kalman Estimator for the Inverted Pendulum Case. Applied Mathematics, 8, 1602-1618. https://doi.org/10.4236/am.2017.811117</p></sec></body><back><ref-list><title>References</title><ref id="scirp.80553-ref1"><label>1</label><mixed-citation publication-type="other" xlink:type="simple">Sontag, E. (1998) Mathematical Control Theory, Deterministic Finite Dimensional Systems. Springer, USA, 104-115.</mixed-citation></ref><ref id="scirp.80553-ref2"><label>2</label><mixed-citation publication-type="other" xlink:type="simple">Almobaied, M., Eksin, I. and Guzelkaya, M. (2016) Design of LQR Controller with Big Bang-Big Crunch Optimization Algorithm Based on Time Domain Criteria. Control and Automation (Med) 2016, 24th Mediterranean Conference On, Athens, Greece, June 21-24 2016, 1192-1197.</mixed-citation></ref><ref id="scirp.80553-ref3"><label>3</label><mixed-citation publication-type="other" xlink:type="simple">Zhang, H.G., Liu, D.R., Luo, Y.H. and Wang, D. (2013) Adaptive Dynamic Programming for Control. Algorithms and Stability. Springer-Verlag, London.  
&lt;br /&gt;https://doi.org/10.1007/978-1-4471-4757-2</mixed-citation></ref><ref id="scirp.80553-ref4"><label>4</label><mixed-citation publication-type="other" xlink:type="simple">Liu, D.R., Wei, Q.L., Wang, D., Yang, X. and Li, H.L. (2017) Adaptive Dynamic Programming with Applications in Optimal Control. Springer International Publishing AG, London. &lt;br /&gt;https://doi.org/10.1007/978-3-319-50815-3</mixed-citation></ref><ref id="scirp.80553-ref5"><label>5</label><mixed-citation publication-type="other" xlink:type="simple">Bertsekas, D. and Tsitsiklis, J. (1996) Neuro-Dynamic Programming. Athena Scientific. The MIT Press, Cambridge, MA.</mixed-citation></ref><ref id="scirp.80553-ref6"><label>6</label><mixed-citation publication-type="other" xlink:type="simple">Sutton, R.S. and Barto, A.G. (1998) Reinforcement Learning: An Introduction. In: Adaptive Computation and Machine Learning. The MIT Press, Cambridge, MA.  
&lt;br /&gt;https://doi.org/10.1109/TNN.1998.712192</mixed-citation></ref><ref id="scirp.80553-ref7"><label>7</label><mixed-citation publication-type="other" xlink:type="simple">Anderson, B. and Moore, J. (1971) Linear Optimal Control. Prentice-Hall International Inc., London. https://doi.org/10.1115/1.3426525</mixed-citation></ref><ref id="scirp.80553-ref8"><label>8</label><mixed-citation publication-type="other" xlink:type="simple">Ogata, K. (1997) Modern Control Engineering. Fifth Edition, Prentice Hall Inc., Englewood Cliffs, NJ, ISBN 0-13-615673-8.</mixed-citation></ref><ref id="scirp.80553-ref9"><label>9</label><mixed-citation publication-type="book" xlink:type="simple">Ogata, K. (1995) Discrete-Time Control Systems. 2nd Ed., Prentice Hall, Englewood Cliffs, NJ.</mixed-citation></ref><ref id="scirp.80553-ref10"><label>10</label><mixed-citation publication-type="other" xlink:type="simple">Bellman, R. and Dreyfus, S. (1962) Applied Dynamic Programming. Princeton University Press, Princeton, NJ. https://doi.org/10.1515/9781400874651</mixed-citation></ref><ref id="scirp.80553-ref11"><label>11</label><mixed-citation publication-type="other" xlink:type="simple">Luus, R. (2000) Iterative Dynamic Programming. CRC Press, Inc., Boca Raton, FL, USA. https://doi.org/10.1201/9781420036022</mixed-citation></ref><ref id="scirp.80553-ref12"><label>12</label><mixed-citation publication-type="other" xlink:type="simple">Bertsekas, D. and Tsitsiklis, J. (2006) Slides MIT.  
&lt;br /&gt;http://web.mit.edu/dimitrib/www/NDP_Review.pdf</mixed-citation></ref><ref id="scirp.80553-ref13"><label>13</label><mixed-citation publication-type="other" xlink:type="simple">Kirk, D.E. (2004) Optimal Control Theory: An Introduction. Dover Publications, USA.</mixed-citation></ref><ref id="scirp.80553-ref14"><label>14</label><mixed-citation publication-type="other" xlink:type="simple">Ljung, L. (1998) System Identification: Theory for the User. Prentice Hall PTR, USA. https://doi.org/10.1007/978-1-4612-1768-8_11</mixed-citation></ref><ref id="scirp.80553-ref15"><label>15</label><mixed-citation publication-type="other" xlink:type="simple">Naidu, D.S. (2003) Optimal Control Systems. Electrical Engineering Handbook, CRC Press, Boca Raton, Florida, 275.</mixed-citation></ref><ref id="scirp.80553-ref16"><label>16</label><mixed-citation publication-type="other" xlink:type="simple">Kalman, R.E. (1960) A New Approach to Linear Filtering and Prediction Problems. Journal of Basic Engineering, 82, 35-45. https://doi.org/10.1115/1.3662552</mixed-citation></ref></ref-list></back></article>