<?xml version="1.0" encoding="UTF-8"?><!DOCTYPE article  PUBLIC "-//NLM//DTD Journal Publishing DTD v3.0 20080202//EN" "http://dtd.nlm.nih.gov/publishing/3.0/journalpublishing3.dtd"><article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" dtd-version="3.0" xml:lang="en" article-type="research article"><front><journal-meta><journal-id journal-id-type="publisher-id">JMF</journal-id><journal-title-group><journal-title>Journal of Mathematical Finance</journal-title></journal-title-group><issn pub-type="epub">2162-2434</issn><publisher><publisher-name>Scientific Research Publishing</publisher-name></publisher></journal-meta><article-meta><article-id pub-id-type="doi">10.4236/jmf.2019.93029</article-id><article-id pub-id-type="publisher-id">JMF-94637</article-id><article-categories><subj-group subj-group-type="heading"><subject>Articles</subject></subj-group><subj-group subj-group-type="Discipline-v2"><subject>Business&amp;Economics</subject><subject> Physics&amp;Mathematics</subject></subj-group></article-categories><title-group><article-title>
 
 
  Derivatives Pricing via Machine Learning
 
</article-title></title-group><contrib-group><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Tingting</surname><given-names>Ye</given-names></name><xref ref-type="aff" rid="aff1"><sup>1</sup></xref></contrib><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Liangliang</surname><given-names>Zhang</given-names></name><xref ref-type="aff" rid="aff2"><sup>2</sup></xref></contrib></contrib-group><aff id="aff2"><addr-line>Bank of America, Charlotte, NC, USA</addr-line></aff><aff id="aff1"><addr-line>Department of Accounting, Questrom School of Business, Boston University, Boston, MA, USA</addr-line></aff><pub-date pub-type="epub"><day>19</day><month>06</month><year>2019</year></pub-date><volume>09</volume><issue>03</issue><fpage>561</fpage><lpage>589</lpage><history><date date-type="received"><day>22,</day>	<month>May</month>	<year>2019</year></date><date date-type="rev-recd"><day>24,</day>	<month>August</month>	<year>2019</year>	</date><date date-type="accepted"><day>27,</day>	<month>August</month>	<year>2019</year></date></history><permissions><copyright-statement>&#169; Copyright  2014 by authors and Scientific Research Publishing Inc. </copyright-statement><copyright-year>2014</copyright-year><license><license-p>This work is licensed under the Creative Commons Attribution International License (CC BY). http://creativecommons.org/licenses/by/4.0/</license-p></license></permissions><abstract><p>
 
 
  In this paper, we combine the theory of stochastic process and techniques of machine learning with the regression analysis, first proposed by [1] to solve for American option prices, and apply the new methodologies on financial derivatives pricing. Rigorous convergence proofs are provided for some of the methods we propose. Numerical examples show good applicability of the algorithms. More applications in finance are discussed in the Appendices.
 
</p></abstract><kwd-group><kwd>Machine Learning</kwd><kwd> Regression Analysis</kwd><kwd> Jump-Diffusion</kwd><kwd> Derivatives Pricing</kwd><kwd> Hilbert Space</kwd><kwd> Orthogonal Projection</kwd></kwd-group></article-meta></front><body><sec id="s1"><title>1. Introduction</title><p>Theoretical and empirical finance research involves the evaluation of conditional expectations, which, in a continuous time jump-diffusion setting, can be related to second order partial integral differential equations of parabolic type (PIDEs) by the Feynman-Kac theorem, and other types of equations such as backward stochastic differential equations with jumps (BSDEJs) or quasi-linear PIDEs in more complicated settings. In theoretical continuous-time finance, many problems, such as asset pricing with market frictions, dynamic hedging or dynamic portfolio-consumption choice problems, can be related to Hamilton-Jacobi-Bellman (HJB) equations via dynamic programming techniques. The HJB equations, from another perspective, are equivalent to BSDEs derived from a probabilistic approach. The nonlinear BSDEs, studied in [<xref ref-type="bibr" rid="scirp.94637-ref2">2</xref>], can be decomposed into a sequence of linear equations, which can be solved by taking conditional expectations, via Picard iteration. For empirical studies, the focus of the literature has been the evaluation of the cross sectional conditional risk-adjusted expected returns and the explanation of them using factors. See [<xref ref-type="bibr" rid="scirp.94637-ref3">3</xref>] [<xref ref-type="bibr" rid="scirp.94637-ref4">4</xref>] and [<xref ref-type="bibr" rid="scirp.94637-ref5">5</xref>] as good illustrations. It is easily seen that, regardless of the fact whether the underlying models are continuous-time or discrete-time, evaluating conditional expectations is inevitable in finance literature. Moreover, in order to perform XVA computations for the measurement of counterparty credit risk, we need to evaluate the conditional expectations, i.e., the derivative prices, on a future simulation grid, as outlined in [<xref ref-type="bibr" rid="scirp.94637-ref6">6</xref>]. These facts call for efficient methods to compute the quantities aforementioned.</p><p>In this paper, we extend the basis function expansion approach proposed in [<xref ref-type="bibr" rid="scirp.94637-ref1">1</xref>] with machine learning techniques. Specifically, we propose new efficient methods to evaluate conditional expectations, regardless of the dynamics of the underlying stochastic process, as long as they can be simulated. Rigorous convergence proofs are given using Hilbert space theory. The methodologies can be applied to time zero pricing as well as pricing on a future simulation grid, with the advantage of ANN approximation most prominent in high dimensional problems. In the sequel, we show applications of our methodologies on the pricing of European derivatives and extension to contracts with optimal stopping feature is straightforward through either [<xref ref-type="bibr" rid="scirp.94637-ref1">1</xref>] approach or reflected-BSDEs.</p><p>Compared to the literature on traditional stochastic analysis, our methodologies are able to handle large data sets and high-dimensional problems, therefore suffering much less from the curse of dimensionality due to the nature of ANN methods. Moreover, our methodologies are very efficient when evaluating solutions of BSDEJs and PIDEs on a future simulation grid, where none of the traditional methodologies applies. With respect to recent machine learning literature on numerical solutions to BSDEs and PDEs, our methodologies enjoy the theoretical advantage of being able to handle equations with jump-diffusion and convergence results are provided. When applied to the solutions of BSDEJs and PIDEs, our methodologies require much less number of parameters, as compared to the current machine learning based methods to be mentioned below. At any step in the solution process, only one ANN is needed and we do not require nested optimization. In terms of application, not all the prices of OTC derivatives can be easily translated into BSDEJs and PIDEs, for example, a range accrual with both American and barrier (knock-out, for example) feature. However, our methodologies are naturally suitable in those situations. To conclude, our methods enjoy many theoretical and empirical advantages, which makes them attractive and novel.</p><p>There has been a huge literature on applications of machine learning techniques to financial research. Classical applications focus on the prediction of market variables such as equity indexes or FX rates and the detection of market anomalies, for example, [<xref ref-type="bibr" rid="scirp.94637-ref7">7</xref>] and [<xref ref-type="bibr" rid="scirp.94637-ref8">8</xref>]. Option pricing via a brute-force curving fitting by ANNs dates back to [<xref ref-type="bibr" rid="scirp.94637-ref9">9</xref>]. More applications of machine learning in finance, especially option pricing prediction, are surveyed in [<xref ref-type="bibr" rid="scirp.94637-ref10">10</xref>]. See references therein. Pricing of American options in high dimensions can be found in [<xref ref-type="bibr" rid="scirp.94637-ref11">11</xref>], which is closest to our method 1. However, there are several improvements of our methods compared to this reference. First of all, we enable deep neural network (DNN) approximation and show convergence. Second, we can incorporate constraints in DNN approximation estimation and prove the mathematical validity of this approach. Third, we propose two more efficient methods to complement the first method of ours. Our treatment of constraints in the estimation of DNNs extends the work of [<xref ref-type="bibr" rid="scirp.94637-ref12">12</xref>] in that we can deal with a larger class of constraints by specifying a general Hilbert subspace as the constrained set. Risk measure computation using machine learning can be found in [<xref ref-type="bibr" rid="scirp.94637-ref13">13</xref>]. Applications of machine learning function approximation on financial econometrics can be found in [<xref ref-type="bibr" rid="scirp.94637-ref14">14</xref>], [<xref ref-type="bibr" rid="scirp.94637-ref15">15</xref>], [<xref ref-type="bibr" rid="scirp.94637-ref16">16</xref>] and [<xref ref-type="bibr" rid="scirp.94637-ref17">17</xref>]. Recent applications include empirical and theoretical asset pricing, reinforcement learning and Q-learning in solving dynamic programming problems such as optimal investment-consumption choice, option pricing and optimal trading strategies construction, e.g., [<xref ref-type="bibr" rid="scirp.94637-ref18">18</xref>], [<xref ref-type="bibr" rid="scirp.94637-ref19">19</xref>], [<xref ref-type="bibr" rid="scirp.94637-ref20">20</xref>], [<xref ref-type="bibr" rid="scirp.94637-ref21">21</xref>], [<xref ref-type="bibr" rid="scirp.94637-ref22">22</xref>], [<xref ref-type="bibr" rid="scirp.94637-ref23">23</xref>], [<xref ref-type="bibr" rid="scirp.94637-ref24">24</xref>], [<xref ref-type="bibr" rid="scirp.94637-ref25">25</xref>], [<xref ref-type="bibr" rid="scirp.94637-ref26">26</xref>], [<xref ref-type="bibr" rid="scirp.94637-ref27">27</xref>], [<xref ref-type="bibr" rid="scirp.94637-ref28">28</xref>] and references therein. Numerical methods to solve PDEs and BSDEs or the related inverse problems can be found in [<xref ref-type="bibr" rid="scirp.94637-ref29">29</xref>], [<xref ref-type="bibr" rid="scirp.94637-ref30">30</xref>], [<xref ref-type="bibr" rid="scirp.94637-ref31">31</xref>], [<xref ref-type="bibr" rid="scirp.94637-ref32">32</xref>], [<xref ref-type="bibr" rid="scirp.94637-ref33">33</xref>], [<xref ref-type="bibr" rid="scirp.94637-ref34">34</xref>], [<xref ref-type="bibr" rid="scirp.94637-ref35">35</xref>], [<xref ref-type="bibr" rid="scirp.94637-ref36">36</xref>], [<xref ref-type="bibr" rid="scirp.94637-ref37">37</xref>], [<xref ref-type="bibr" rid="scirp.94637-ref38">38</xref>] and [<xref ref-type="bibr" rid="scirp.94637-ref39">39</xref>]. Machine learning based methods enjoy the advantage of being fast, able to handle large data sets and high dimensional problems.</p><p>Our methodologies are combinations of traditional statistical learning theory and stochastic analysis with advanced machine learning techniques, introducing powerful function approximation method via the universal approximation theorem and artificial neural networks (ANNs), while preserving the regression-type analysis documented in [<xref ref-type="bibr" rid="scirp.94637-ref1">1</xref>]. The methods are very easy to use, effective, accurate as illustrated by numerical experiments and time efficient. They are different from the convergent expansion method, e.g., [<xref ref-type="bibr" rid="scirp.94637-ref40">40</xref>], simulation methods such as [<xref ref-type="bibr" rid="scirp.94637-ref41">41</xref>], [<xref ref-type="bibr" rid="scirp.94637-ref42">42</xref>], [<xref ref-type="bibr" rid="scirp.94637-ref43">43</xref>] and [<xref ref-type="bibr" rid="scirp.94637-ref44">44</xref>] or the asymptotic expansion method proposed by [<xref ref-type="bibr" rid="scirp.94637-ref45">45</xref>], [<xref ref-type="bibr" rid="scirp.94637-ref46">46</xref>], [<xref ref-type="bibr" rid="scirp.94637-ref47">47</xref>], [<xref ref-type="bibr" rid="scirp.94637-ref48">48</xref>] [<xref ref-type="bibr" rid="scirp.94637-ref49">49</xref>] [<xref ref-type="bibr" rid="scirp.94637-ref50">50</xref>] [<xref ref-type="bibr" rid="scirp.94637-ref51">51</xref>] [<xref ref-type="bibr" rid="scirp.94637-ref52">52</xref>], in that we no longer resort to polynomial basis function expansion or small-diffusion type analysis. Our methods are also different from the pure machine learning based ones documented in [<xref ref-type="bibr" rid="scirp.94637-ref29">29</xref>], [<xref ref-type="bibr" rid="scirp.94637-ref30">30</xref>], [<xref ref-type="bibr" rid="scirp.94637-ref31">31</xref>], [<xref ref-type="bibr" rid="scirp.94637-ref32">32</xref>], [<xref ref-type="bibr" rid="scirp.94637-ref33">33</xref>], [<xref ref-type="bibr" rid="scirp.94637-ref34">34</xref>] and [<xref ref-type="bibr" rid="scirp.94637-ref35">35</xref>], in that we utilize the lead-lag regression formula to evaluate the conditional expectations, preserving the time dependent structure and our methods are able to handle jump-diffusion processes easily.</p><p>The organization of this paper is as follows. Section 2 documents the methodologies. Section 3 illustrates the usefulness of our methods by considering European and American derivatives pricing. Section 4 considers numerical experiments and Section 5 concludes. An outline of the proofs and other applications can be found in the appendices.</p></sec><sec id="s2"><title>2. The Methodology</title><p>Mathematical Setup</p><p>We use a Markov process modeled by a jump-diffusion as illustration. Suppose that we have a stochastic differential equation with jumps</p><p>  d X t = μ ( t , X t ) d t + σ ( t , X t ) d W t + ∫ E     γ ( t , X t , e ) N ˜ ( d t , d e ) ,       X 0 = x 0 (1)</p><p>where X ∈ ℝ r , W ∈ ℝ d is a standard d-dimensional Brownian motion and N ˜ is a q-dimensional compensated Poisson random measure, with the compensator ν   ( d t , d e ) : = ν ( d e ) d t . Information filtration F t = F t W , N is generated by ( W , N ) . We hope to evaluate the conditional expectation E t [ ψ ( X T ) ] for any 0 &lt; t &lt; T , e.g., see [<xref ref-type="bibr" rid="scirp.94637-ref53">53</xref>]. Assumptions on ψ and X are stated below.</p><p>Assumption 1 (On Growth Condition of ψ). ψ has polynomial growth in its argument x, i.e., there exists a positive integer P, independent of x, such that for all | x | &gt; 1 , we have, for constant C independent of x</p><p>| ψ ( x ) | ≤ C | x | P . (2)</p><p>The following assumption is w.r.t. X.</p><p>Assumption 2 (On X). There exists a unique strong solution to Equation (1) and X has finite polynomial moments of all orders.</p><p>The General Approximation Theory</p><p>First, we need the following assumptions, definitions and results. Please note that, some of the spaces we introduce are actually conditional ones. The discussions of conditional Hilbert spaces can be found in [<xref ref-type="bibr" rid="scirp.94637-ref54">54</xref>], e.g., L 2 ( F t ) is a conditional Hilbert space for all t ∈ [ 0, T ] .</p><p>Definition 3 (Projection Operator). For Hilbert spaces X and H , where H ⊂ X . Define PROJ H x as the projection of x ∈ X onto H .</p><p>Definition 4 (Orthogonal Space). For Hilbert spaces X and H , where H ⊂ X . Define ORTH H X as the orthogonal space of H in X .</p><p>Definition 5 (Spanning the Hilbert Space). Assume that E = { e j } j ∈ Λ is a set of elements in Hilbert space X and Λ is an index set. Define H E as the intersection of all Hilbert subspaces of X containing E .</p><p>Assumption 6 (On Joint Continuity). X and H are two Hilbert spaces and H ⊂ X . Moreover, { H n } n = 1 ∞ is a sequence of Hilbert sub-spaces of H satisfying H n ⊂ H n + 1 for any n ≥ 1 and ∪ n = 1 ∞ H n &#175; = H . We have lim n → ∞ ‖ h − PROJ H n h n ‖ H = 0 for any h ∈ H and lim n → ∞ h n = h .</p><p>The next two theorems are well-known in the literature.</p><p>Theorem 7 (Hilbert Projection Theorem). Let H ⊂ X be two Hilbert spaces and let x ∈ X . Then, PROJ H x exists and is unique. Moreover, it is characterized uniquely by x − PROJ H x ∈ ORTH H X .</p><p>Theorem 8 (Repeated Projection Theorem). Let G ⊂ H ⊂ X be three Hilbert spaces. Then, for any x ∈ X , PROJ G x = PROJ G ( PROJ H x ) .</p><p>Remark 9 The conditions of Theorems 7 and 8 on G and H can be relaxed to convexity and completeness instead of Hilbert sub-spaces.</p><p>Finally, we have the result below.</p><p>Theorem 10. Suppose X is a Hilbert space, { H n } n = 1 ∞ and H are Hilbert subspaces of X satisfying H n ⊂ H n + 1 and ∪ n = 1 ∞ H n &#175; = H ⊂ X . x ∈ X , define h n = PROJ H n x and h = PROJ H x . Then we have lim n → ∞ h n = h w.r.t. the norm topology in X , if Assumption 6 is satisfied.</p><p>Sometimes we need to add constraints on the calibrated ANN, e.g., the shape constraints. The following assumption and theorem deal with this situation.</p><p>Assumption 11 (On Constrained Sub-space). Suppose that Ψ ⊂ X such that { Ψ ∩ H n } n = 1 ∞ is a sequence of non-empty convex and complete subspaces of X satisfying Assumption 6, where X and { H n } n = 1 ∞ are described.</p><p>The following theorem handles the constrained approximation and its convergence.</p><p>Theorem 12 (On Constrained Approximation). Under Assumptions 6 and 11, for x ∈ X , if h = PROJ H x ∈ Ψ , then, we have lim n → ∞ PROJ Ψ ∩ H n x = h .</p><p>Remark 13 (On ψ). In Theorem 12, the set Ψ represents prior knowledge on constraints that h satisfies. It can be represented by a set of non-linear inequalities or equalities on functionals of h. Common constraints for option pricing include non-negativity constraint and the positiveness constraint on the second order derivatives. The verification of { Ψ ∩ H n } n = 1 ∞ satisfying Assumption 6 should be based on a case-by-case manner.</p><p>To proceed further, we need the following assumptions.</p><p>Assumption 14 (On Some Spaces). { H t J } J = 1 ∞ is an increasing sequence of Hilbert sub-spaces of L 2 ( F t ) , H t J ⊂ H t J + 1 , ∪ J = 1 ∞ H t J &#175; = H t ⊂ L 2 ( F t ) . Moreover, { E t [ ξ T ] | ξ T ∈ L 2 ( F T ) , E t [ ξ T ] ∈ L 2 ( F t ) } &#175; ⊂ H t ⊂ L 2 ( F t ) ⊂ L 2 ( F T ) = X T .</p><p>Assumption 15 (On Structure of H t J ). { e t j } j ∈ Λ is a set of elements of L 2 ( F t ) , such that H t J = H { e t j } j ∈ Λ J , where Λ J ⊂ Λ J + 1 ⊂ Λ for any J ≥ 1 and ∪ J = 1 ∞ Λ J = Λ , satisfies Assumption 14<sup>1</sup>.</p><p>Then, we have the following results.</p><p>Lemma 1. For any adapted stochastic process ξ such that ξ T ∈ L 2 ( F T ) , if E t [ ξ T ] ∈ L 2 ( F t ) , we have</p><p>E t [ ξ T ] = arg min η t ∈ L 2 ( F t ) E [ ( ξ T − η t ) 2 ] . (3)</p><p>The following proposition is a natural extension of Lemma 1.</p><p>Proposition 16. For any measurable function ψ and stochastic process X such that ψ ( X T ) ∈ L 2 ( F T ) and E t [ ψ ( X T ) ] ∈ L 2 ( F t ) , we have</p><p>E t [ ψ ( X T ) ] = arg min ξ t ∈ L 2 ( F t ) E [ ( ψ ( X T ) − ξ t ) 2 ] . (4)</p><p>Here ξ t ∈ F t and the above minimization problem has a unique solution. In particular, if X is a Markov process, then ξ t = ϕ ( t , X t ) , i.e., ξ t is a function of time t and X t .</p><p>We then have the following theorem.</p><p>Theorem 17. Under Assumptions 1, 2, 6, 14 and 15, for any adapted stochastic process ξ such that ξ T ∈ L 2 ( F T ) and E t [ ξ T ] ∈ L 2 ( F t ) , we have</p><p>lim J → ∞ arg min η t ∈ H t J E [ ( ξ T − η t ) 2 ] = L 2 ( F t ) E t [ ξ T ] . (5)</p><p>Further, for any measurable function ψ and stochastic process X such that ψ ( X T ) ∈ L 2 ( F T ) and E t [ ψ ( X T ) ] ∈ L 2 ( F t ) , we have the following equality</p><p>lim J → ∞ arg min ξ t ∈ H t J E [ ( ψ ( X T ) − ξ t ) 2 ] = L 2 ( F t ) E t [ ψ ( X T ) ] . (6)</p><p>If X is Markov, then we have ξ t = ϕ ( t , X t ) , i.e., ξ t is a function of time t and X t .</p><p>The following theorem justifies the Monte Carlo approximation of expectation in the above optimization problems.</p><p>Theorem 18 (On Sequential Convergence). Under Assumptions 1, 2, 6, 14 and 15, suppose that | Λ J | = m J &lt; ∞ for all J ≥ 1 , { X T i } i = 1 M and { e t j , i } j , i = 1 , 1 m J , M are M i.i.d. copies of X T and { e t j } j = 1 m J . Then we have</p><p>lim J → ∞ lim M → ∞ arg min ξ t m ∈ H t J 1 M ∑ m = 1 M ( ψ ( X T m ) − ξ t m ) 2 = ℙ E t [ ψ ( X T ) ] . (7)</p><p>The following results justify the universal approximation and ANN approximation approaches proposed in this paper.</p><p>Proposition 19 (On Universal Approximation Theory). Let σ denote the function in the universal approximation theorem mentioned in [<xref ref-type="bibr" rid="scirp.94637-ref55">55</xref>], [<xref ref-type="bibr" rid="scirp.94637-ref56">56</xref>] and [<xref ref-type="bibr" rid="scirp.94637-ref57">57</xref>]. Define { e t j } j = 1 m n : = { σ ( α j + β j X t ) } j = 1 m n , where X satisfies Equation (1) and Assumption 2, α j and β j have at most n significant digits in total, where n ∈ ℕ , i.e., n belongs to the set of natural numbers, j runs from 1 to m n and</p><p>m n is the number of all related { e t j } , i.e., m n = | { σ ( α + β X t ) | α   and   β haveatmost   n   totalsignificantdigits } | . Then, { H { e t j } j = 1 m n } n ∈ ℕ satisfies Assumptions 6, 14 and 15. Therefore, Theorems 17 and 18 apply.</p><p>Proposition 20 (On Deep Neural Network Approximation). For the DNN defined in ( [<xref ref-type="bibr" rid="scirp.94637-ref58">58</xref>], Definition 1.1], observe that W l ( x ) = α l + β l x . Define</p><p>e t j : = W L , j ∘ ρ ∘ W L − 1, j ∘ ρ ∘ ⋯ ∘ W 1, j ∘ ρ ( X t ) (8)</p><p>where W l , j ( x ) = α l , j + β l , j x satisfies that l = 1 , 2 , ⋯ , L , ( α l , j , β l , j ) have at most n total significant digits and n ∈ ℕ . Then, { H { 1, e t j } j = 1 m n } n ∈ ℕ , where 1 means</p><p>function f ( x ) ≡ 1 for all x, satisfies Assumptions 6, 14 and 15. Therefore, Theorems 17 and 18 apply after a localization argument on ψ and X on a compact sub-domain in ℝ r .</p><p>Remark 21 (On DNN). Please note that, in Proposition 20, we do not intend to prove the convergence when the number of layers goes to infinity. Instead, we show convergence when the number of connections goes to infinity, which can be achieved via enlarging the number of neurons in each layer with the total number of layers remaining fixed.</p><p>Remark 22 (On Euler Time Discretization). [<xref ref-type="bibr" rid="scirp.94637-ref59">59</xref>] proposes an exact simulation method for multi-dimensional stochastic differential equations. The discussion of discretization error, of the regression approach proposed in this paper, with Euler method is not hard if ψ satisfies Assumption 1, in which case the dominated convergence theorem and L 2 convergence of Euler method can be applied to show the convergence.</p><p>The proofs of the above results can be found in Appendix A. In what follows, we will propose three methods to compute, approximately, the function ϕ in Proposition 16.</p><p>Method 1</p><p>In general, ϕ , defined in Proposition 16 and Theorem 17, can not be found in closed-form. A natural thought would be to resort to function expansion representations, i.e., to find the solution to the following problem</p><p>E t [ ψ ( X T ) ] = arg ​ min { a j , θ j } j = 0 ∞ ∈ A E [ ( ψ ( X T ) − ∑ j = 0 ∞     a j e j ( t , X t | θ j ) ) 2 ] (9)</p><p>where A is an appropriate space for coefficients { a j , θ j } j = 0 ∞ and { e j ( θ j ) } j = 0 ∞ is a set of functions, with Span ( { e j ( θ j ) } j = 0 ∞ ) <sup>2</sup> dense in an appropriate function space Φ <sup>3</sup>. To further proceed, we seek a truncation of the function representation formula as follows</p><p>E t [ ψ ( X T ) ] ≅ arg min { a j , θ j } j = 0 J ∈ A J E [ ( ψ ( X T ) − ∑ j = 0 J     a j e j ( t , X t | θ j ) ) 2 ] (10)</p><p>for J sufficiently large, where A J is a compact set in the Euclidean space where { a j , θ j } j = 0 J take values. The last step would be to use Monte Carlo simulation to approximate the unconditional expectation appearing in Equations (9) and (10). Therefore turning the conditional expectation computation problem, into a least-square function regression problem, similar to [<xref ref-type="bibr" rid="scirp.94637-ref1">1</xref>]. An obvious choice of { e j ( θ j ) } j = 0 ∞ is polynomial basis, for example, the set of Fourier-Hermite basis functions. For expansion using Fourier-Hermite basis functions in high dimensions, see [<xref ref-type="bibr" rid="scirp.94637-ref60">60</xref>].</p><p>In fact, Artificial Neural Networks (ANNs) prove to be an efficient and convergent function approximation tool that we can utilize in the above expressions. Write</p><p>E t [ ψ ( X T ) ] ≅ arg min { a j , θ j } j = 0 J ∈ A J E [ ( ψ ( X T ) − ANN J ( { a j , θ j } j = 0 J | t , X t ) ) 2 ] (11)</p><p>where ANN J denotes an ANN with parameters { a j , θ j } j = 0 J .</p><p>Note that, via proper time discretization and fixed point iteration, solving a BSDE with jumps can be decomposed into a series of evaluations of conditional expectations. The machine learning based method outlined above can be applied there. We will write down the algorithm to solve a general Coupled Forward-Backward Stochastic Differential Equation with Jumps (CFBSDEJs) in the appendix. Extensions to other types of BSDEJs are possible.</p><p>Here we assume that X is a Markov process. To handle path dependency or non-Markov processes, we can apply the backward induction method outlined in [<xref ref-type="bibr" rid="scirp.94637-ref1">1</xref>]. With the machine learning approach, it is easy to see that this method enables us to get the values of conditional expectations on a future simulation grid.</p><p>Method 2</p><p>Another method to utilize the idea of [<xref ref-type="bibr" rid="scirp.94637-ref1">1</xref>] is inspired by the boosting random tree method (BRT), see, [<xref ref-type="bibr" rid="scirp.94637-ref61">61</xref>], for example. Partition the domain space ℝ r = ∪ k = 1 K   U t k <sup>4</sup>, where { U t k } k = 1 K is a set of disjoint sets in ℝ r and consider</p><p>E t [ ψ ( X T ) ] = arg min ϕ ∈ Φ E [ ( ψ ( X T ) − ϕ ( t , X t ) ) 2 ] ≅ arg min ∑ k = 1 K ϕ k ( t , x ) 1 x ∈ U t k ∈ Φ E [ ( ψ ( X T ) − ϕ k ( t , X t ) ) 2 1 X t ∈ U t k ] . (12)</p><p>The choice of { U t k } k = 1 K is important and we can use the machine learning classification techniques (or any classification rule), such as kmeans function in R programming language, in Monte Carlo simulation and related computations. Denote d U = sup x , y ∈ U | x − y | . It is possible to show that as long as lim K → ∞ max 1 ≤ k ≤ K d U t k = 0 , we only need finite number of functions, for example, { e j ( θ j ) } j = 0 J , to approximate each { ϕ k } k = 1 K and obtain convergence. In practice, although the domain of X t is ℝ r , it might be centered at a small subspace ℂ t , therefore facilitating the partition process. Note also that this method might require us to mollify the function ψ , if it is not smooth. We adopt finite order Taylor expansion as the function expansion representation approach. The following theorems provide convergence analysis for this method.</p><p>Theorem 23. For an appropriate function space Φ , we have</p><p>E t [ ψ ( X T ) ] = arg min ϕ ∈ Φ E [ ( ψ ( X T ) − ϕ ( t , X t ) ) 2 ] = arg min ϕ ∈ Φ E [ ( ψ ( X T ) − ϕ ( t , X t ) ) 2 ∑ k = 1 K     1 X t ∈ U t k ] = arg min ϕ ∈ Φ E [ ( ψ ( X T ) ∑ k = 1 K 1 X t ∈ U t k − ∑ k = 1 K     ϕ ( t , X t ) 1 X t ∈ U t k ) 2 ] = arg min ∑ k = 1 K ϕ k ( t , x ) 1 x ∈ U t k ∈ Φ E [ ( ψ ( X T ) − ϕ k ( t , X t ) ) 2 1 X t ∈ U t k ] . (13)</p><p>Theorem 24. Let H t J be as described previously and H t = { ϕ ( t , X t ) | ϕ ∈ Φ } . Then, we have</p><p>lim max 1 ≤ k ≤ K d U t k → 0 ‖ ∑ k = 1 K     ϕ ^ k ( t , X t ) 1 X t ∈ U t k − ϕ ( t , X t ) ‖ L 2 ( F t ) = 0 (14)</p><p>with J large enough, fixed, finite and ϕ ^ k is an approximation to ϕ k , which satisfies</p><p>E [ ( ϕ k ( t , X t ) − ϕ ^ k ( t , X t ) ) 2 1 X t ∈ U t k ] ≤ ϵ K (15)</p><p>for any k = 1 , 2 , ⋯ , K , K ∈ ℕ , lim K → ∞ K ϵ K = 0 and ϵ K is independent of k when K is sufficiently large.</p><p>Method 3</p><p>Next, we propose an algorithm combining the ANN and universal approximation theorem (UAT). Suppose that L 2 ( F t ) is the space where we are performing the approximation. Also assume that F t W , N = F t X , i.e., the information filtration is equivalently generated by X. Define an ANN with connection N by ANN ( x , N , θ j , j ) , where x is the state variables that the ANN depends on, θ j is the vector of parameters and j is its label. We define the following nested regression approximation</p><p>ψ ( X T ) = ANN ( X t , N , θ 1 ,1 ) + ϵ t , T 1 (16)</p><p>ϵ t , T 1 = ANN ( X t , N , θ 2 ,2 ) + ϵ t , T 2 (17)</p><p>ϵ t , T 2 = ANN ( X t , N , θ 3 ,3 ) + ϵ t , T 3 (18)</p><p>⋯ = ⋯ (19)</p><p>ϵ t , T J = ANN ( X t , N , θ J + 1 , J + 1 ) + ϵ t , T J + 1 (20)</p><p>⋯ = ⋯ (21)</p><p>where { ∑ j = 1 J + 1     ANN ( X t , N , θ j , j ) } J = 0 ∞ is the approximate sequence of E t [ ψ ( X T ) ] .</p><p>In this paper, we will test and compare the performance of all of the proposed methods. A general discussion and rigorous proofs can be found in Appendix A<sup>5</sup>.</p></sec><sec id="s3"><title>3. Applications in Derivatives Pricing</title><sec id="s3_1"><title>3.1. European Option Pricing</title><p>Suppose that the payoff of a European claim can be written as, similar to [<xref ref-type="bibr" rid="scirp.94637-ref62">62</xref>] and [<xref ref-type="bibr" rid="scirp.94637-ref63">63</xref>], ( f , ψ ) , where f t is a stream of cash flows materialized at each time instance t and ψ T is a one-time terminal payoff at time T. Therefore, under no-arbitrage condition, the price of this European payoff can be written as, under risk neutral measure</p><p>V t e : = E t [ ∫ t T D t , u f u d u + D t , T ψ T ] (22)</p><p>where D t , u : = e − ∫ t u   r v d v is the stochastic discount factor. If we assume a Markov structure f t = f ( t , X t ) and ψ T = ψ ( X T ) , then V t e : = v e ( t , X t ) , i.e., V t e is a function of time t and state vector X t . This problem is a canonical application of the evaluation of conditional expectations and we can apply the methodologies outlined in Section 2 to solve it. European claims with barrier features can be incorporated and priced in a similar way. For example, the price of a knock-in European claim can be written as</p><p>V t e : = E t [ ∫ τ T D t , u f u d u + D τ , T ψ T ] (23)</p><p>where τ = inf v ∈ [ t , T ] { X v ∈ T | X t ∉ T } , where T ⊂ ℝ r . In our setting, the dynamics of X can be arbitrary, possibly stochastic differential equations with jumps, Markov chains, or even non-Markov processes. Previously, Monte Carlo based method for option pricing can be found in [<xref ref-type="bibr" rid="scirp.94637-ref64">64</xref>] and [<xref ref-type="bibr" rid="scirp.94637-ref65">65</xref>], among others.</p></sec><sec id="s3_2"><title>3.2. American Option Pricing</title><p>Still use ( f , ψ ) to denote the payoff structure of an American claim, whose price can be obtained via formula</p><p>V t a : = sup τ ∈ S [ t , T ] E t [ ∫ t τ D t , u f u d u + D t , τ ψ τ ] . (24)</p><p>Here S [ t , T ] is the space of all the stopping times in [ t , T ] . We refer the interested readers to [<xref ref-type="bibr" rid="scirp.94637-ref62">62</xref>] and [<xref ref-type="bibr" rid="scirp.94637-ref66">66</xref>] for general derivation and explanation of Equation (24). It is also possible to derive the general BSDE that an American claim price satisfies, for example [<xref ref-type="bibr" rid="scirp.94637-ref67">67</xref>]. Moreover, in [<xref ref-type="bibr" rid="scirp.94637-ref27">27</xref>] and [<xref ref-type="bibr" rid="scirp.94637-ref1">1</xref>], the authors utilize a backward induction approach to solve optimal stopping problems. The idea can be carried out using the methodologies documented in Section 2. American claims with barrier features can be incorporated and priced in a similar way. It is also known that American option prices can be related to reflected BSDEs (RBSDEs), a rigorous discussion of existence and uniqueness of such equations can be found in [<xref ref-type="bibr" rid="scirp.94637-ref68">68</xref>] and references therein.</p></sec></sec><sec id="s4"><title>4. Numerical Experiments</title><sec id="s4_1"><title>4.1. European Option Pricing</title><p>In this section, we consider a Heston model</p><p>d S t S t = r d t + ν t d W t ,       S 0 = s 0 (25)</p><p>d ν t = κ ( θ − ν t ) d t + σ ν t ( ρ d W t + 1 − ρ 2 d B t ) ,       ν 0 = v 0 (26)</p><p>where ( W , B ) is a two dimensional standard Brownian motion. The parameter values are chosen as r = 0.05 , κ = 1.00 , θ = 0.04 , σ = 0.10 , ρ = − 0.50 , s 0 = 1.00 , K = 1.00 and v 0 = 0.04 . Time to maturity is set to be T = 0.50 ,</p><p>with time discretization step h = 0.01 and N = T h = 50 . The number of</p><p>simulation paths is M = 10000 . We price a plain vanilla European call option ( S T − K ) + as an illustration. The QQ-plots are displayed in Figures 1-10. The first three correspond to a recursive evaluation, i.e., regressing the values at t + 1 on state variables at time t. The rest of the plots correspond to direct regression, i.e., regressing the discounted payoffs at time T on state variables at time t. Figures 10-12 are for the prices of a digital call option under Black-Scholes setting and Figures 13-15 are QQ-plots for Delta values. <xref ref-type="fig" rid="fig16">Figure 16</xref> and <xref ref-type="fig" rid="fig17">Figure 17</xref> show the QQ-plots for method 3 under Heston model with 3 nested ANN approximations of size 4 and one ANN approximation of size 12 using R routine nnet. The absolute RMSE for the former is 0.1938% and latter 0.2581%, with the running time 10.36 seconds compared to 52.31 for ANN approximation with size 12.</p></sec><sec id="s4_2"><title>4.2. American Option Pricing</title><p>Here we refer the readers to [<xref ref-type="bibr" rid="scirp.94637-ref67">67</xref>] for the BSDE satisfied by a plain vanilla American option. For r = 0.03 , d = 0.07 , σ = 0.20 , T = 3.00 , N = 150 , S 0 = 100 and K = 100 , the benchmark American option price at t 0 = 0 is 9.0660 and the relative difference of our Monte-Carlo price is 0.27%. The running time is less than 30 seconds.</p></sec></sec><sec id="s5"><title>5. Conclusion and Future Research</title><p>In this paper, we show how machine learning techniques, specifically, ANN function approximation methods, can be applied to derivatives pricing. We relate pricing problems to the evaluation of conditional expectations via BSDEJs and PIDEs. Future research topics can, potentially, be the development of reinforcement learning methodologies to solve dynamic programming problems and apply them in the context of empirical asset pricing literature. Moreover, the evaluation of energy derivatives calls for SDEJs defined in a Hilbert space. The same theoretical constructions can also be found in the evaluation of fixed income derivatives, such as the random field models proposed and studied in [<xref ref-type="bibr" rid="scirp.94637-ref69">69</xref>]. One can, of course, apply Karhunen-Lo&#233;ve expansion for a dimension reduction to reduce the problem to the evaluation of conditional expectations of regular SDEJs. However, the development of machine learning based methods to solve directly the conditional expectations on the stochastic processes defined in a Hilbert space is important. In addition, stochastic differential games, that arise in the context of American game options, equity swaps, and the related Mckean-Vlasov type FBSDEJs (mean-field FBSDEJ, see [<xref ref-type="bibr" rid="scirp.94637-ref70">70</xref>] ) are important topics in mathematical finance. They are also related to the theoretical analysis of high-frequency trading. Finding machine-learning based numerical methods to solve these equations is of great interest to us. Last, but not least, machine learning methods in asset pricing and portfolio optimization, which can be found in [<xref ref-type="bibr" rid="scirp.94637-ref71">71</xref>], [<xref ref-type="bibr" rid="scirp.94637-ref72">72</xref>], [<xref ref-type="bibr" rid="scirp.94637-ref73">73</xref>], [<xref ref-type="bibr" rid="scirp.94637-ref28">28</xref>], [<xref ref-type="bibr" rid="scirp.94637-ref74">74</xref>] and [<xref ref-type="bibr" rid="scirp.94637-ref75">75</xref>], admit an elegant way to price financial derivatives under ℙ -measure. For example, we can use the method in [<xref ref-type="bibr" rid="scirp.94637-ref72">72</xref>] to calibrate the SDF process and use [<xref ref-type="bibr" rid="scirp.94637-ref75">75</xref>] to generate market scenarios. These methodologies, combined with the methods documented in this paper and [<xref ref-type="bibr" rid="scirp.94637-ref1">1</xref>], have the potential to solve for any derivative price. We leave all the development to future research.</p></sec><sec id="s6"><title>Acknowledgements</title><p>We thank the Editor and the referee for their comments. Moreover, we are grateful to Professor J&#233;r&#244;me Detemple, Professor Marcel Rindisbacher and Professor Weidong Tian for their useful suggestions.</p></sec><sec id="s7"><title>Conflicts of Interest</title><p>The authors declare no conflicts of interest regarding the publication of this paper.</p></sec><sec id="s8"><title>Cite this paper</title><p>Ye, T.T. and Zhang, L.L. (2019) Derivatives Pricing via Machine Learning. Journal of Mathematical Finance, 9, 561-589. https://doi.org/10.4236/jmf.2019.93029</p></sec><sec id="s9"><title>Appendix</title>A. Convergence of the Proposed Methodologies<p>Proof of Theorem 10. It is known from the projection theorem of Hilbert space that { h n } n = 1 ∞ and h actually exist and are unique. Moreover, PROJ H n h = h n as indicated by the repeated projection theorem. It is also known that h − h n ∈ ORTH H n H . As we ask that Assumption 6 hold, we know that ‖ h − h n ‖ X → 0 + as n → ∞ .</p><p>Proof of Theorem 12. The proof follows from Assumption 6 and Theorem 8. We have</p><p>lim n → ∞ PROJ Ψ ∩ H n x (27)</p><p>= lim n → ∞ PROJ Ψ ∩ H n PROJ H n x (28)</p><p>= lim n → ∞ PROJ Ψ ∩ H n h n (29)</p><p>= PROJ Ψ ∩ H h (30)</p><p>= h . (31)</p><p>This concludes the proof.</p><p>Proof of Lemma 1. For any λ t ∈ L 2 ( F t ) , we have</p><p>E [ ( ξ T − λ t ) 2 ] (32)</p><p>= E [ ( ξ T − E t [ ξ T ] ) 2 ] + E [ ( λ t − E t [ ξ T ] ) 2 ] (33)</p><p>        + 2 E [ ( λ t − E t [ ξ T ] ) ( ξ T − E t [ ξ T ] ) ] ︸ = 0 (34)</p><p>= E [ ( ξ T − E t [ ξ T ] ) 2 ] + E [ ( λ t − E t [ ξ T ] ) 2 ] (35)</p><p>≥ E [ ( ξ T − E t [ ξ T ] ) 2 ] . (36)</p><p>Therefore we have the claim announced.</p><p>Proof of Theorem 17. The proof of this theorem follows from Assumptions 1, 2, 6, 14, 15 and Theorem 10, by choosing { E t [ ξ T ] | ξ T ∈ L 2 ( F T ) , E t [ ξ T ] ∈ L 2 ( F t ) } &#175; ⊂ H t ⊂ L 2 ( F t ) ⊂ L 2 ( F T ) = X T .</p><p>Proof of Theorem 18. Essentially, Equation (7) is the result of Gauss-Markov Theorem and the consistency property of OLS estimator.</p><p>Proof of Proposition 19. This is a direct consequence of the discussion in ( [<xref ref-type="bibr" rid="scirp.94637-ref57">57</xref>], Section 3) (see Equation (5)) and Theorem 10. To elaborate, consider X T = L 2 ( F T ) , x = ψ ( X T ) , its projections h and h n on H t = ∪ n = 1 ∞ H t n &#175; ⊂ L 2 ( F t ) and H t n defined in this proposition. Suppose that</p><p>h = ∑ j = 1 ∞     λ j e t j and h n = ∑ j = 1 m n     μ j n e t j ,</p><p>where m n &lt; m n + 1 and { e t j } j = 1 ∞ is a set of orthonormal basis in H t . From the repeated projection theorem, we know that μ j n + 1 = μ j n = λ j for any 1 ≤ j ≤ m n <sup>6</sup> and n ∈ ℕ . From the L 2 property of h, we know that ∑ j = 1 ∞     λ j 2 &lt; ∞ . Therefore, ‖ h − h n ‖ L 2 ( F T ) = ∑ j = n + 1 ∞     λ j 2 → 0 as n → ∞ .</p><p>Proof of Proposition 20. This is a direct consequence of the discussion in ( [<xref ref-type="bibr" rid="scirp.94637-ref58">58</xref>], Theorem 2.2), localization arguments, Theorem 10 and the proof of Proposition 19.</p><p>Proof of Theorem 23. The first, second and third equality are obvious given an appropriate choice of Φ depending on the Markov property of X and its moment conditions in Assumption 2. Actually, because of the existence and uniqueness of ϕ ∈ Φ such that the RHS of the first equality achieves minimum, we know that</p><p>min ϕ ∈ Φ E [ ( ψ ( X T ) ∑ k = 1 K 1 X t ∈ U t k − ∑ k = 1 K     ϕ ( t , X t ) 1 X t ∈ U t k ) 2 ] (37)</p><p>≤ min ∑ k = 1 K ϕ k ( t , x ) 1 x ∈ U t k ∈ Φ E [ ( ψ ( X T ) − ϕ k ( t , X t ) ) 2 1 X t ∈ U t k ] . (38)</p><p>From another perspective, we know that min ∑ k = 1 K     ϕ k ( t , x ) 1 x ∈ U t k ∈ Φ E [ ( ψ ( X T ) − ϕ k ( t , X t ) ) 2 1 X t ∈ U t k ] is a piecewise minimization. Therefore</p><p>min ϕ ∈ Φ E [ ( ψ ( X T ) ∑ k = 1 K 1 X t ∈ U t k − ∑ k = 1 K     ϕ ( t , X t ) 1 X t ∈ U t k ) 2 ] (39)</p><p>≥ min ∑ k = 1 K ϕ k ( t , x ) 1 x ∈ U t k ∈ Φ E [ ( ψ ( X T ) − ϕ k ( t , X t ) ) 2 1 X t ∈ U t k ] . (40)</p><p>The last equality in Equation (13) holds.</p><p>Proof of Theorem 24. The proof of this theorem is a direct consequence of Equations (13), (15) and triangle inequality.</p>B. Other Applications<p>In this section, we document other applications of our methodologies in finance.</p><p>B.1. Joint Valuation and Calibration</p><p>Suppose that there are N derivatives contracts whose prices at time t 0 can be expressed as { V t 0 n } n = 1 N . Their payoffs are { φ n ( X ⋅ ) } n = 1 N , where X is an</p><p>r-dimensional vector of state variables. Sometimes we write X θ to explicitly state dependence of X on its vector of parameters θ . Here suppose X θ satisfies a system of stochastic differential equations with jumps</p><p>d X t θ = μ ( t , X t θ | θ ) d t + σ ( t , X t θ | θ ) d W t + ∫ E     γ ( t , X t θ , e | θ ) N ˜ ( d t , d e ) . (41)</p><p>The main idea is that { V t 0 n } n = 1 N might contain derivatives contracts from different asset classes or hybrid ones. Therefore, we need to model X as a joint high dimensional cross-asset system. One potential problem is that θ is in general a high-dimensional vector, which will be hard to estimate using usual optimization routines in R or MATLAB software system. However, we can apply ADAM method, studied in [<xref ref-type="bibr" rid="scirp.94637-ref76">76</xref>] for the parameter estimation. It is based on a stochastic iteration method via the gradient of the MSE function. The key to evaluate the gradient of the MSE function is to evaluate the dynamics of ∂ θ X t θ . It satisfies the following system of SDEJ</p><p>d ∂ θ X t θ = ∂ θ μ ( t , X t θ | θ ) d t + ∂ x μ ( t , X t θ | θ ) ∂ θ X t θ d t     + ∂ θ σ ( t , X t θ | θ ) d W t + ∂ x σ ( t , X t θ | θ ) ∂ θ X t θ d W t     + ∫ E     ∂ θ γ ( t , X t θ , e | θ ) N ˜ ( d t , d e )     + ∫ E     ∂ x γ ( t , X t θ , e | θ ) ∂ θ X t θ N ˜ ( d t , d e ) . (42)</p><p>The existence and uniqueness of the solution to the SDEJ system (42) can be obtained with necessary regularity conditions on the coefficients.</p><p>B.2. Option Surface Fitting</p><p>There is a strand of literature that strives to fit option panels using different dynamics for the underlying assets, for example, [<xref ref-type="bibr" rid="scirp.94637-ref77">77</xref>] on stochastic volatility models, [<xref ref-type="bibr" rid="scirp.94637-ref78">78</xref>] on local volatility models and [<xref ref-type="bibr" rid="scirp.94637-ref79">79</xref>] on local-stochastic volatility models. Models that incorporate jumps can be found in [<xref ref-type="bibr" rid="scirp.94637-ref80">80</xref>], [<xref ref-type="bibr" rid="scirp.94637-ref81">81</xref>] and references therein.</p><p>Consider the following stochastic differential equation</p><p>d S t S t = r ( t , X t ) d t + σ ( t , S t , X t ) d W t ,       S 0 = s 0 d X t = α ( t , X t ) d t + β ( t , X t ) d W t ,       X 0 = x 0 . (43)</p><p>Here we model σ by a DNN. The advantage of doing so is that it might fully capture the market volatility surface meantime ensuring a good dynamic fit, while still preserving the existence and uniqueness result for the related stochastic differential equation system (43).</p><p>B.3. Credit Risk Management: Evaluation on a Future Simulation Grid</p><p>We refer the problem definition to [<xref ref-type="bibr" rid="scirp.94637-ref6">6</xref>]. It is easy to illustrate that the problem is equivalent to the evaluation of conditional expectations on a future simulation grid and our methods are suitable for this type of problems. Note that, some XVA quantities, such as KVA, require the evaluation of CVA on a future simulation grid. Our methodologies, such as the ones proposed in Sections 2 and B.7, can be applied on the evaluation of KVA, once we obtain future present values of financial claims.</p><p>B.4. Dynamic Hedging</p><p>There are references that utilize machine learning (mainly Reinforcement Learning, or RL) to solve dynamic hedging problems, e.g., [<xref ref-type="bibr" rid="scirp.94637-ref82">82</xref>], [<xref ref-type="bibr" rid="scirp.94637-ref83">83</xref>] and [<xref ref-type="bibr" rid="scirp.94637-ref84">84</xref>]. However, here in this paper we will not follow this route. Instead, we use the BSDE formulation of the problem in [<xref ref-type="bibr" rid="scirp.94637-ref2">2</xref>] and try to solve the BSDE that characterizes the hedging problem. The methodology is outlined in Appendix B.11.</p><p>B.5. Dynamic Portfolio-Consumption Choice</p><p>We use [<xref ref-type="bibr" rid="scirp.94637-ref85">85</xref>] as an example and try to solve the related coupled FBSDE with jumps. The methodology is outlined in Appendix B.11. Other examples of dynamic portfolio optimization can be found in [<xref ref-type="bibr" rid="scirp.94637-ref53">53</xref>], [<xref ref-type="bibr" rid="scirp.94637-ref86">86</xref>], [<xref ref-type="bibr" rid="scirp.94637-ref87">87</xref>], [<xref ref-type="bibr" rid="scirp.94637-ref88">88</xref>], [<xref ref-type="bibr" rid="scirp.94637-ref89">89</xref>], [<xref ref-type="bibr" rid="scirp.94637-ref90">90</xref>], [<xref ref-type="bibr" rid="scirp.94637-ref91">91</xref>], [<xref ref-type="bibr" rid="scirp.94637-ref92">92</xref>], [<xref ref-type="bibr" rid="scirp.94637-ref93">93</xref>] and [<xref ref-type="bibr" rid="scirp.94637-ref94">94</xref>]. Essentially, dynamic portfolio-consumption choice problems are stochastic programming in nature and can be related to HJB equations or BSDEs. An example of using HJB representation of the problem can be found in [<xref ref-type="bibr" rid="scirp.94637-ref95">95</xref>]. The equations can be solved using the methodologies outlined in Section 0 and Appendix B.11.</p><p>B.6. Transition Density Approximation</p><p>We can generalize the theory in [<xref ref-type="bibr" rid="scirp.94637-ref96">96</xref>] and [<xref ref-type="bibr" rid="scirp.94637-ref97">97</xref>] to approximate the transition density of a multivariate time-inhomogeneous stochastic differential equation with jumps. According to [<xref ref-type="bibr" rid="scirp.94637-ref96">96</xref>] and [<xref ref-type="bibr" rid="scirp.94637-ref97">97</xref>], the transition density of a multivariate time-inhomogeneous stochastic differential equation with or without jumps can be approximated by polynomials in a weighted-Hilbert space. See ( [<xref ref-type="bibr" rid="scirp.94637-ref97">97</xref>], Equation (2.1)), for example. The key is to evaluate the coefficients { c α } α , which is, again, the evaluation of conditional expectations. The resulted transition density can be used in option pricing, MLE estimation for MSDEJs and prediction, filtering and smoothing problems for hidden Markov models, see [<xref ref-type="bibr" rid="scirp.94637-ref98">98</xref>].</p><p>B.7. Evaluating Conditional Expectations via a Measure Change</p><p>Consider the following equation</p><p>E t [ ψ ( X τ ) ] = ∫ ℝ r     Γ ( t , x ; τ , y ) ψ ( y ) d y (44)</p><p>= ∫ ℝ r     Γ 0 ( t 0 , x ; τ , y ) Γ ( t , x ; τ , y ) Γ 0 ( t 0 , x ; τ , y ) ψ ( y ) d y (45)</p><p>where Γ 0 is the transition density of a stochastic differential equation with jumps, which can be simulated for arbitrary ( t , τ ) without using time discretization<sup>7</sup> and Γ is the transition density function of X. Γ can be approximated by the method outlined in Appendix B.6. It is immediately obvious that we can generate random numbers from Γ 0 and reuse them for the evaluation of the conditional expectation on the left hand side of Equation (44) for different ( t , τ ) .</p><p>B.8. Empirical Asset Pricing with Factor Models: Evaluating Expected Returns</p><p>In this section, we propose to use machine learning, mainly, ANN techniques, to construct factor models and evaluate the conditional expected asset returns and risk-premium cross-sectionally. Related references are [<xref ref-type="bibr" rid="scirp.94637-ref28">28</xref>] and [<xref ref-type="bibr" rid="scirp.94637-ref74">74</xref>], among others. [<xref ref-type="bibr" rid="scirp.94637-ref3">3</xref>] provide a good example with basis function expansion to capture the non-linearity in asset returns. Specifically, consider the following lead-lag regression</p><p>R t + 1 = f ( t , X t ) + ε t , t + 1 . (46)</p><p>Here E t [ ε t , t + 1 ] = 0 and X is a set of risk factors. Then, E t [ R t + 1 ] = f ( t , X t ) . Linear factor models assume that f ( t , x ) = a t + b t x . f can also be approximated by basis function expansion, using universal approximation theorem, or via ANNs. The fitted conditional expected asset returns can be fed into the mean-variance optimizer, i.e., [<xref ref-type="bibr" rid="scirp.94637-ref99">99</xref>] and construct long-short portfolios or other trading strategies.</p><p>B.9. Recovery and Representation Theorem</p><p>In [<xref ref-type="bibr" rid="scirp.94637-ref100">100</xref>], the authors propose a model-free recovery theorem, based on a series expansion of higher order conditional moments of asset returns. Their work inspires us to exploit the ANN-factor models to represent the higher order conditional moments of the asset returns and therefore validating the recovery theorem proposed there-in. Moreover, similar to [<xref ref-type="bibr" rid="scirp.94637-ref57">57</xref>], our machine learning approximation to the conditional expectations of financial payoffs amounts to a compound option representation of arbitrary L 2 -claims in the financial economic system. Also, the second numerical method means that any financial claim, can be locally approximated by a linear combination of power derivatives, following the same idea.</p><p>B.10. Theoretical Asset Pricing via Dynamic Stochastic General Equilibrium</p><p>Note that, the equation systems proposed in [<xref ref-type="bibr" rid="scirp.94637-ref101">101</xref>], [<xref ref-type="bibr" rid="scirp.94637-ref102">102</xref>] and [<xref ref-type="bibr" rid="scirp.94637-ref103">103</xref>] can be transformed into BSDEs and we can use time discretization and apply the techniques proposed in Section 2 and Appendix B.11 to solve them. In this paper, however, we will not test our methods on this strand of literature.</p><p>B.11. Solving High-Dimensional CFBSDEJs</p><p>A coupled forward-backward stochastic differential equation with jumps (CFBSDEJ) can be written as</p><p>d X t = μ ( t , X t , Y t , Z t , V t ) d t + σ ( t , X t , Y t , Z t , V t ) d W t                   + ∫ E     γ ( t , X t , Y t , Z t , V t , e ) N ˜ ( d t , d e ) X 0 = x 0 d Y t = f ( t , X t , Y t , Z t , V t ) d t + Z t d W t + ∫ E     U t ( e ) N ˜ ( d t , d e ) V t = ∫ E     U t ( e ) ν ( d e ) Y T = ϕ ( X T ) (47)</p><p>where N ˜ ( d t , d e ) = N ( d t , d e ) − ν ( d e ) d t is a compensated Poisson random measure. We take the following steps to solve Equation (47) numerically.</p><p>Time Discretization</p><p>Discretize time interval [   t , T ] into n-equal distance sub-intervals π = { [ t i , t i + 1 ) } i = 0 n − 1 with h = t i + 1 − t i n , t 0 = t and t n = T . Consider the following Euler discretized equation.</p><p>d X t i = μ ( t i , X t i , Y t i , Z t i , V t i ) h + σ ( t i , X t i , Y t i , Z t i , V t i ) d W t i               + ∫ E     γ ( t i , X t i , Y t i , Z t i , V t i , e ) N ˜ ( d t i , d e ) X 0 = x 0 d Y t i = f ( t i , X t i , Y t i , Z t i , V t i ) h + Z t i d W t i + ∫ E     U t i ( e ) N ˜ ( d t i , d e ) V t i = ∫ E     U t i ( e ) ν ( d e ) Y T = ϕ ( X T ) (48)</p><p>where d X t i : = X t i + 1 − X t i and d Y t i : = Y t i + 1 − Y t i . Denote the solution to the time-discretized CFBSDEJ as ( X π , Y π , Z π , U π ) . We need the following assumption.</p><p>Assumption 25. Under the norm ‖   ⋅   ‖ K [ t , T ] 2 introduced in [<xref ref-type="bibr" rid="scirp.94637-ref104">104</xref>], we have</p><p>‖ ( X , Y , Z , U ) − ( X π , Y π , Z π , U π ) ‖ K [ t , T ] 2 → 0 (49)</p><p>as n → ∞ .</p><p>Mollification</p><p>Define a sequence of functions ( μ m , σ m , γ m , f m , ϕ m ) , which are bounded and have bounded derivatives of all orders and</p><p>lim m → ∞ ( μ m , σ m , γ m , f m , ϕ m ) = ( μ , σ , γ , f , ϕ ) (50)</p><p>in a point-wise sense. Also denote the solution to the CFBSDEJ with coefficients ( μ m , σ m , γ m , f m , ϕ m ) as ( X m , Y m , Z m , U m ) . Then, we have the following theorem.</p><p>Theorem 26. Under Assumption 25</p><p>E t [ g ( X u π , m , Y u π , m , Z u π , m , V u π , m ) ] → E t [ g ( X u , Y u , Z u , V u ) ] (51)</p><p>as n , m → ∞ for arbitrary T &gt; u &gt; t &gt; 0 . g is a function with at most polynomial growth in its arguments.</p><p>Picard Iteration</p><p>After the time discretization and mollification are done, we will resort to Picard fixed point iteration technique to decompose the solution ( X π , m , Y π , m , Z π , m , U π , m ) to a sequence of uncoupled FBSDEJs whose solutions are denoted by ( X π , m , k , Y π , m , k , Z π , m , k , U π , m , k ) , where k denotes the index of Picard iteration. For zeroth order, consider</p><p>d X t i π , m ,1 = μ m ( t i , X t i π , m ,1 ,0,0,0 ) h + σ m ( t i , X t i π , m ,1 ,0,0,0 ) d W t i                       + ∫ E     γ m ( t i , X t i π , m ,1 ,0,0,0, e ) N ˜ ( d t i , d e ) X 0 π , m ,1 = x 0</p><p>d Y t i π , m ,1 = f m ( t i , X t i π , m ,1 , Y t i π , m ,1 , Z t i π , m ,1 , V t i π , m ,1 ) h + Z t i π , m ,1 d W t i                           + ∫ E     U t i π , m ,1 ( e ) N ˜ ( d t i , d e ) V t i π , m ,1 = ∫ E     U t i π , m ,1 ( e ) ν ( d e ) Y T π , m ,1 = ϕ ( X T π , m ,1 ) (52)</p><p>For k ≥ 2 , define</p><p>d X t i π , m , k = μ m ( t i , X t i π , m , k , Y t i π , m , k − 1 , Z t i π , m , k − 1 , V t i π , m , k − 1 ) h                       + σ m ( t i , X t i π , m , k , Y t i π , m , k − 1 , Z t i π , m , k − 1 , V t i π , m , k − 1 ) d W t i                       + ∫ E     γ m ( t i , X t i π , m , k , Y t i π , m , k − 1 , Z t i π , m , k − 1 , V t i π , m , k − 1 , e ) N ˜ ( d t i , d e ) X 0 π , m , k = x 0</p><p>d Y t i π , m , k = f m ( t i , X t i π , m , k , Y t i π , m , k , Z t i π , m , k , V t i π , m , k ) h                       + Z t i π , m , k d W t i + ∫ E     U t i π , m , k ( e ) N ˜ ( d t i , d e ) V t i π , m , k = ∫ E     U t i π , m , k ( e ) ν ( d e ) Y T π , m , k = ϕ ( X T π , m , k ) (53)</p><p>Evaluation of Conditional Expectations</p><p>For Equation system (53), we can start from the last time interval and work backwards. The problem is transformed into the evaluation of E t i [ u ( t i + 1 , X t i + 1 π , m , k ) ] , where u is the intermediate solution and satisfies u ( T , ⋅ ) = ϕ ( ⋅ ) .</p><p>B.12. Pricing Kernel Approximation</p><p>A pricing kernel η t is an L 2 ( F t ) stochastic process, adapted to the information filtration { F t } 0 ≤ t ≤ T , such that</p><p>V t = E t [ D t , T η t , T V T ] (54)</p><p>where V T is an F T payoff, D t , T = D T D t = e − ∫ t T r v d v and η t , T = η T η t . It is obvious that η t = E t [ η T ] , i.e., η is a ℙ -martingale. Represent</p><p>D T η T = ∑ j = 0 ∞     a j e T j (θj)</p><p>where { e T j } j = 0 ∞ is a set of orthonormal basis in L 2 ( F T ) space and θ j is the vector of coefficients of e j . Suppose that we have K derivative contracts, denoted by { V T k } k = 1 K , with basis representation V T k = ∑ j = 0 ∞     b k j e T j ( θ j ) . Therefore</p><p>V t 0 k = E t 0 [ ∑ j = 0 ∞     a j e T j ( θ j ) ∑ j = 0 ∞     b k j e T j ( θ j ) ] = ∑ j = 0 ∞     a j b k j . (55)</p><p>Equation (55), if truncated after J terms, formulates a linear equation system and the unknowns { a j } j = 0 J and { θ j } j = 0 J can be recovered from ordinary least square optimization. After we obtain η T , η t can be recovered by η t = E t [ η T ] , via the methodology outlined in Section 2.</p><p>Remark 27. If { e t j ( θ j ) } j = 0 ∞ is not orthonormal, Equation (55) becomes nonlinear in { θ j } j = 0 J . The evaluations remain the same, with only more complicated numerical computations. The basis can also be represented by ANNs.</p><p>Remark 28. For a specific representation via universal approximation theorem, see [<xref ref-type="bibr" rid="scirp.94637-ref55">55</xref>].</p><p>Remark 29. It is possible to allow shape constraints in the estimation (55) and formulate a constrained optimization problem, see [<xref ref-type="bibr" rid="scirp.94637-ref105">105</xref>], for example.</p><p>We can also directly utilize the method proposed in Section 2, when used with time discretization and Monte Carlo simulation. Denote M as the number of</p><p>sample paths and { V T m , k } m = 1 , k = 1 M , K as M simulated final payoffs for each of the K derivatives. Define { a m } m = 1 M as M real numbers. Let { V 0 k } k = 1 K be K derivative prices at time t 0 = 0 . Find the solution to the following optimization problem</p><p>{ a m } m = 1 M = arg min { ϕ m } m = 1 M ​ [ ∑ k = 1 K ( V 0 k − 1 M ∑ m = 1 M     ϕ m V T m , k ) 2 ] . (56)</p><p>After obtaining { a m } m = 1 M , we try to find function relation g such that</p><p>a m = g ( T , X T m ) = D 0 , T m η T m</p><p>where { X T m } m = 1 M is a set of simulated state variables at time T. When fitting g, we can add some shape or no-arbitrage constraints, or other regularization conditions, to the optimization problem and formulate a constrained ANN</p><p>(ACNN). We always assume that the matrix t ( { V T m , k } m = 1 , k = 1 M , K ) { V T m , k } m = 1 , k = 1 M , K is a K &#215; K invertible matrix, where t ( ⋅ ) is the matrix transpose operator.</p>C. Intuition of Convergence Proof for Appendix B.11<p>In Appendix B.11, we propose a method to solve numerically a CFBSDEJ. As long as the time discretization step is convergent, we can argue that the methodology converges, in some sense, to the true one, as outlined above in Appendix B.11. Potentially, we need an a priori estimate formula, similar to the one in [<xref ref-type="bibr" rid="scirp.94637-ref2">2</xref>], for coupled BSDEs, to justify Picard iteration at every time discretization step.</p></sec><sec id="s10"><title>NOTES</title></sec></body><back><ref-list><title>References</title><ref id="scirp.94637-ref1"><label>1</label><mixed-citation publication-type="other" xlink:type="simple">Longstaff, F. and Schwartz, E. (2001) Valuing American Options by Simulation: A Simple Least—Square Approach. The Review of Financial Studies, 14, 113-147.https://doi.org/10.1093/rfs/14.1.113</mixed-citation></ref><ref id="scirp.94637-ref2"><label>2</label><mixed-citation publication-type="other" xlink:type="simple">El Karoui, N., Peng, S. and Quenez, M.C. (1997) Backward Stochastic Differential Equations in Finance. Mathematical Finance, 7, 1-71.https://doi.org/10.1111/1467-9965.00022</mixed-citation></ref><ref id="scirp.94637-ref3"><label>3</label><mixed-citation publication-type="other" xlink:type="simple">Adrian, T., Crump, R. and Vogt, E. (2018) Nonlinearity and Flight-to-Safety in the Risk-Return Trade-Off for Stocks and Bonds. Forthcoming in Journal of Finance, 74, 1931-1973.</mixed-citation></ref><ref id="scirp.94637-ref4"><label>4</label><mixed-citation publication-type="other" xlink:type="simple">Fama, E. and French, K. (1993) Common Risk Factors in the Returns on Stocks and Bonds. Journal of Financial Economics, 33, 3-56.https://doi.org/10.1016/0304-405X(93)90023-5</mixed-citation></ref><ref id="scirp.94637-ref5"><label>5</label><mixed-citation publication-type="other" xlink:type="simple">Fama, E. and French, K. (2015) A Five-Factor Asset Pricing Model. Journal of Financial Economics, 116, 1-22.</mixed-citation></ref><ref id="scirp.94637-ref6"><label>6</label><mixed-citation publication-type="other" xlink:type="simple">Zhu, S. and Pykhtin, M. (2008) A Guide to Modeling Counterparty Credit Risk. Working Paper. https://papers.ssrn.com/sol3/papers.cfm?abstract_id=1032522</mixed-citation></ref><ref id="scirp.94637-ref7"><label>7</label><mixed-citation publication-type="other" xlink:type="simple">Aydogdu, M. (2018) Predicting Stock Returns Using Neural Networks. Working Paper. https://papers.ssrn.com/sol3/papers.cfm?abstract_id=3141492 https://doi.org/10.2139/ssrn.3141492</mixed-citation></ref><ref id="scirp.94637-ref8"><label>8</label><mixed-citation publication-type="other" xlink:type="simple">Voshgha, H. (2008) Early Detection of Defaulting Firms: Artificial Neural Network Application; Australian Context. Working Paper. https://papers.ssrn.com/sol3/papers.cfm?abstract_id=2130505</mixed-citation></ref><ref id="scirp.94637-ref9"><label>9</label><mixed-citation publication-type="other" xlink:type="simple">Hutchinson, J., Lo, A. and Poggio, T. (1994) A Nonparametric Approach to Pricing and Hedging Derivative Securities via Learning Networks. Journal of Finance, 49, 851-889. https://doi.org/10.1111/j.1540-6261.1994.tb00081.x</mixed-citation></ref><ref id="scirp.94637-ref10"><label>10</label><mixed-citation publication-type="other" xlink:type="simple">Hahn, J.T. (2013) Option Pricing Using Artificial Neural Networks: The Australian Perspective. Ph.D. Thesis, Bond University, Queensland.</mixed-citation></ref><ref id="scirp.94637-ref11"><label>11</label><mixed-citation publication-type="other" xlink:type="simple">Kohler, M., Krzyzak, M. and Todorovic, N. (2010) Pricing of High-Dimensional American Options by Neural Networks. Mathematical Finance, 20, 383-410.https://doi.org/10.1111/j.1467-9965.2010.00404.x</mixed-citation></ref><ref id="scirp.94637-ref12"><label>12</label><mixed-citation publication-type="other" xlink:type="simple">Dugas, C., Bengio, Y., Bélisle, F., Nadeau, C. and Garcia, R. (2009) Incorporating Functional Knowledge in Neural Networks. Journal of Machine Learning Research, 10, 1239-1262.</mixed-citation></ref><ref id="scirp.94637-ref13"><label>13</label><mixed-citation publication-type="other" xlink:type="simple">Eckstein, S., Kupper, M. and Pohl, M. (2018) Robust Risk Aggregation with Neural Networks. Quantitative Finance, 1-40. https://arxiv.org/abs/1811.00304</mixed-citation></ref><ref id="scirp.94637-ref14"><label>14</label><mixed-citation publication-type="other" xlink:type="simple">Giovanis, E. (2010) Applications of Neural Network Radial Basis Function in Economics and Financial time Series. SSRN Electronic Journal. https://papers.ssrn.com/sol3/papers.cfm?abstract_id=1667442 https://doi.org/10.2139/ssrn.1667442</mixed-citation></ref><ref id="scirp.94637-ref15"><label>15</label><mixed-citation publication-type="other" xlink:type="simple">Kopitkov, D. and Indelman, V. (2018) Deep PDF: Probabilistic Surface Optimization and Density Estimation. Computer Science, 1-18. https://arxiv.org/abs/1807.10728</mixed-citation></ref><ref id="scirp.94637-ref16"><label>16</label><mixed-citation publication-type="other" xlink:type="simple">Luo, R., Zhang, W., Xu, X. and Wang, J. (2017) A Neural Stochastic Volatility Model. Computer Science, 1-11. https://arxiv.org/pdf/1712.00504.pdf</mixed-citation></ref><ref id="scirp.94637-ref17"><label>17</label><mixed-citation publication-type="other" xlink:type="simple">Sasaki, H. and Hyvarinen, A. (2018) Neural-Kernelized Conditional Density Estimation. Statistics, 1-12. https://arxiv.org/abs/1806.01754</mixed-citation></ref><ref id="scirp.94637-ref18"><label>18</label><mixed-citation publication-type="other" xlink:type="simple">Weissensteiner, A. (2009) AQ-Learning Approach to Derive Optimal Consumption and Investment Strategies. IEEE Transactions on Neural Networks, 20, 1234-1243.https://doi.org/10.1109/TNN.2009.2020850</mixed-citation></ref><ref id="scirp.94637-ref19"><label>19</label><mixed-citation publication-type="other" xlink:type="simple">Casgrain, P. and Jaimungal, S. (2016) Trading Algorithms with Learning in Latent Alpha Models. SSRN Electronic Journal. https://doi.org/10.2139/ssrn.2871403https://papers.ssrn.com/sol3/papers.cfm?abstract_id=2871403</mixed-citation></ref><ref id="scirp.94637-ref20"><label>20</label><mixed-citation publication-type="other" xlink:type="simple">Heaton, J., Polson, N. and Witte, J. (2016) Deep Learning for Finance: Deep Portfolios. Applied Stochastic Models in Business and Industry, 33, 3-12. https://papers.ssrn.com/sol3/papers.cfm?abstract_id=2838013https://doi.org/10.2139/ssrn.2838013</mixed-citation></ref><ref id="scirp.94637-ref21"><label>21</label><mixed-citation publication-type="other" xlink:type="simple">Samo, Y. and Vernuurt, A. (2016) Stochastic Portfolio Theory: A Machine Learning Perspective. Quantitative Finance, 1-9. https://arxiv.org/pdf/1605.02654.pdf</mixed-citation></ref><ref id="scirp.94637-ref22"><label>22</label><mixed-citation publication-type="other" xlink:type="simple">Jiang, Z., Xu, D. and Liang, J. (2017) A Deep Reinforcement Learning Framework for the Financial Portfolio Management Problem. Computational Finance, 1-31. https://arxiv.org/pdf/1706.10059.pdf</mixed-citation></ref><ref id="scirp.94637-ref23"><label>23</label><mixed-citation publication-type="other" xlink:type="simple">Deng, Y., Bao, F., Kong, Y., Ren, Z. and Dai, Q. (2017) Deep Direct Reinforcement Learning for Financial Signal Representation and Trading. IEEE Transactions on Neural Networks and Learning Systems, 28, 653-664.https://doi.org/10.1109/TNNLS.2016.2522401</mixed-citation></ref><ref id="scirp.94637-ref24"><label>24</label><mixed-citation publication-type="other" xlink:type="simple">Halperin, I. (2017) QLBS: Q-Learner in the Black-Scholes(-Merton) Worlds. Quantitative Finance, 1-34. https://arxiv.org/abs/1712.04609v2 https://doi.org/10.2139/ssrn.3087076</mixed-citation></ref><ref id="scirp.94637-ref25"><label>25</label><mixed-citation publication-type="other" xlink:type="simple">Ritter, G. (2017) Machine Learning for Trading. Working Paper. https://papers.ssrn.com/sol3/papers.cfm?abstract_id=3015609 https://doi.org/10.2139/ssrn.3015609</mixed-citation></ref><ref id="scirp.94637-ref26"><label>26</label><mixed-citation publication-type="other" xlink:type="simple">Xing, F., Cambrida, E., Malandri, L. and Vercellis, C. (2018) Discovering Bayesian Market Views for Intelligent Asset Allocatio. https://arxiv.org/pdf/1802.09911.pdf</mixed-citation></ref><ref id="scirp.94637-ref27"><label>27</label><mixed-citation publication-type="other" xlink:type="simple">Becker, S., Cheridito, P. and Jentzen, A. (2018) Deep Optimal Stopping. Mathematics, arXiv: 1804. 05394. https://arxiv.org/abs/1804.05394</mixed-citation></ref><ref id="scirp.94637-ref28"><label>28</label><mixed-citation publication-type="other" xlink:type="simple">Gu, S., Kelly, B. and Xiu, D. (2018) Empirical Asset Pricing via Machine Learning. 31st Australasian Finance and Banking Conference 2018, Sydney, 13-15 December 2018. https://doi.org/10.3386/w25398https://papers.ssrn.com/sol3/papers.cfm?abstract_id=3159577</mixed-citation></ref><ref id="scirp.94637-ref29"><label>29</label><mixed-citation publication-type="other" xlink:type="simple">Weinan, E., Han, J. and Jentzen, A. (2017) Deep Learning-Based Numerical Methods for High-Dimensional Parabolic Partial Differential Equations and Backward Stochastic Differential Equations. Mathematics, 1-39. https://arxiv.org/pdf/1706.04702.pdf</mixed-citation></ref><ref id="scirp.94637-ref30"><label>30</label><mixed-citation publication-type="other" xlink:type="simple">Weinan, E., Hutzenthaler, M., Jentzen, A. and Kruse, T. (2017) On Multilevel Picard Numerical Approximations for High-Dimensional Nonlinear Parabolic Partial Differential Equations and High-Dimensional Nonlinear Backward Stochastic Differential Equations. Mathematics, 1-25. https://arxiv.org/pdf/1708.03223.pdf</mixed-citation></ref><ref id="scirp.94637-ref31"><label>31</label><mixed-citation publication-type="other" xlink:type="simple">Han, J., Jentzen, A. and Weinan, E. (2017) Overcoming the Curse of Dimensionality: Solving High-Dimensional Partial Differential Equations Using Deep Learning. Mathematics, 1-14. https://arxiv.org/pdf/1707.02568.pdf</mixed-citation></ref><ref id="scirp.94637-ref32"><label>32</label><mixed-citation publication-type="other" xlink:type="simple">Khoo, Y., Lu, J. and Ying, L. (2017) Solving Parametric PDE Problems with Artificial Neural Networks. Mathematics, 1-17. https://arxiv.org/pdf/1707.03351.pdf</mixed-citation></ref><ref id="scirp.94637-ref33"><label>33</label><mixed-citation publication-type="other" xlink:type="simple">Beck, C., Weinan, E. and Jentzen, A. (2017) Machine Learning Approximation Algorithms for High-Dimensional Fully Nonlinear Partial Differential Equations and Second-Order Backward Stochastic Differential Equations. Mathematics, 1-56. https://arxiv.org/pdf/1709.05963.pdf</mixed-citation></ref><ref id="scirp.94637-ref34"><label>34</label><mixed-citation publication-type="other" xlink:type="simple">Sirignano, J. and Spiliopoulos, K. (2017) DGM: A Deep Learning Algorithm for Solving Partial Differential Equations. Mathematics, 1-31.https://arxiv.org/pdf/1708.07469.pdf</mixed-citation></ref><ref id="scirp.94637-ref35"><label>35</label><mixed-citation publication-type="other" xlink:type="simple">Long, Z., Lu, Y. and Ma, X. (2018) PDE-Net: Learning PDEs from Data. Mathematics, 1-17. https://arxiv.org/pdf/1710.09668.pdf</mixed-citation></ref><ref id="scirp.94637-ref36"><label>36</label><mixed-citation publication-type="other" xlink:type="simple">Long, Z. and Lu, Y. (2018) PDE-Net 2.0: Learning PDEs from Data with a Numeric Symbolic Hybrid Deep Network. Computer Science, 1-16. https://arxiv.org/pdf/1812.04426.pdf</mixed-citation></ref><ref id="scirp.94637-ref37"><label>37</label><mixed-citation publication-type="other" xlink:type="simple">Haehnel, P., Marecek, J. and Monteil, J. (2018) Scaling up Deep Learning for PDE-Based Models. Computer Science, 1-39. https://arxiv.org/pdf/1810.09425.pdf</mixed-citation></ref><ref id="scirp.94637-ref38"><label>38</label><mixed-citation publication-type="other" xlink:type="simple">Berg, J. and Nystrom, K. (2018) Data-Driven Discovery of PDEs in Complex Datasets. Statistics, 1-22. https://arxiv.org/pdf/1808.10788.pdf</mixed-citation></ref><ref id="scirp.94637-ref39"><label>39</label><mixed-citation publication-type="other" xlink:type="simple">Rudy, S., Alla, A., Brunton, S. and Nathan Kutz, J. (2018) Data-Driven Identification of Parametric Partial Differential Equations. Mathematics, 1-17. https://arxiv.org/pdf/1806.00732.pdf</mixed-citation></ref><ref id="scirp.94637-ref40"><label>40</label><mixed-citation publication-type="other" xlink:type="simple">Detemple, J., Lorig, M., Rindisbacher, M. and Zhang, L. (2018) An Analytical Expansion Method for Forward Backwards to Chastic Differential Equations with Jumps.</mixed-citation></ref><ref id="scirp.94637-ref41"><label>41</label><mixed-citation publication-type="other" xlink:type="simple">Briand, P. and Labart, C. (2012) Simulation of BSDEs by Wiener Chaos Expansion. The Annals of Applied Probability, 24, 1129-1171. https://doi.org/10.1214/13-AAP943</mixed-citation></ref><ref id="scirp.94637-ref42"><label>42</label><mixed-citation publication-type="other" xlink:type="simple">Geiss, C. and Labart, C. (2015) Simulation of BSDEs with Jumps by Wiener Chaos Expansion. Mathematics, arXiv: 1502.05649. http://arxiv.org/abs/1502.05649</mixed-citation></ref><ref id="scirp.94637-ref43"><label>43</label><mixed-citation publication-type="other" xlink:type="simple">Gnameho, K., Stadje, M. and Pelsser, A. (2017) A Regression-Later Algorithm for Backward Stochastic Differential Equations. Mathematics, 1-33. https://arxiv.org/pdf/1706.07986</mixed-citation></ref><ref id="scirp.94637-ref44"><label>44</label><mixed-citation publication-type="other" xlink:type="simple">Gobet, E. and Labart, C. (2007) Error Expansion for the Discretization of Backward Stochastic Differential Equations. Stochastic Processes and Their Applications, 117, 803-829. https://doi.org/10.1016/j.spa.2006.10.007</mixed-citation></ref><ref id="scirp.94637-ref45"><label>45</label><mixed-citation publication-type="other" xlink:type="simple">Takahashi, A. and Yamada, T. (2016) An Asymptotic Expansion for Forward-Backward SDEs: A Malliavin Calculus Approach. Asia-Pacific Financial Markets, 23, 337-373.</mixed-citation></ref><ref id="scirp.94637-ref46"><label>46</label><mixed-citation publication-type="other" xlink:type="simple">Takahashi, A. and Yamada, T. (2015) On the Expansion to Quadratic FBSDEs.</mixed-citation></ref><ref id="scirp.94637-ref47"><label>47</label><mixed-citation publication-type="other" xlink:type="simple">Gobet, E. and Pagliarani, S. (2014) Analytical Approximations of BSDEs with Non-Smooth Driver. SIAM Journal on Financial Mathematics, 6, 919-958. https://doi.org/10.2139/ssrn.2448691</mixed-citation></ref><ref id="scirp.94637-ref48"><label>48</label><mixed-citation publication-type="other" xlink:type="simple">Fujii, M. and Takahashi, A. (2012) Analytical Approximation for Non-Linear FBSDEs with Perturbation Scheme. International Journal of Theoretical and Applied Finance, 15, Article ID: 1250034. https://doi.org/10.1142/S0219024912500343</mixed-citation></ref><ref id="scirp.94637-ref49"><label>49</label><mixed-citation publication-type="other" xlink:type="simple">Fujii, M. and Takahashi, A. (2012) Perturbative Expansion of FBSDE in an Incomplete Market with Stochastic Volatility. The Quarterly Journal of Finance, 2, 1-22.https://doi.org/10.2139/ssrn.1999137</mixed-citation></ref><ref id="scirp.94637-ref50"><label>50</label><mixed-citation publication-type="other" xlink:type="simple">Fujii, M. and Takahashi, A. (2015) Asymptotic Expansion for Forward-Backward SDEs with Jumps. Quantitative Finance, 1-39. https://doi.org/10.2139/ssrn.2672890https://papers.ssrn.com/sol3/papers.cfm?abstract_id=2672890</mixed-citation></ref><ref id="scirp.94637-ref51"><label>51</label><mixed-citation publication-type="other" xlink:type="simple">Fujii, M. and Takahashi, A. (2016) Quadratic-Exponential Growth BSDEs with Jumps and Their Malliavin’s Differentiability. Working Paper.https://doi.org/10.2139/ssrn.2705670http://papers.ssrn.com/sol3/papers.cfm?abstract_id=2705670</mixed-citation></ref><ref id="scirp.94637-ref52"><label>52</label><mixed-citation publication-type="other" xlink:type="simple">Fujii, M. and Takahashi, A. (2016) Solving Backward Stochastic Differential Equations by Connecting the Short-Term Expansions. Quantitative Finance, 1-41. https://papers.ssrn.com/sol3/papers.cfm?abstract_id=2795490</mixed-citation></ref><ref id="scirp.94637-ref53"><label>53</label><mixed-citation publication-type="other" xlink:type="simple">Detemple, J. and Rindisbacher, M. (2005) Closed-Form Solutions for Optimal Portfolio Selection with Stochastic Interest Rate and Investment Constraints. Mathematical Finance, 15, 539-568. https://doi.org/10.1111/j.1467-9965.2005.00250.x</mixed-citation></ref><ref id="scirp.94637-ref54"><label>54</label><mixed-citation publication-type="other" xlink:type="simple">Hansen, L. and Richard, S. (1987) The Role of Conditioning Information in Deducing Testable Restrictions Implied by Dynamic Asset Pricing Models. Econometrica, 55, 587-613. https://doi.org/10.2307/1913601</mixed-citation></ref><ref id="scirp.94637-ref55"><label>55</label><mixed-citation publication-type="other" xlink:type="simple">Jiang, J. and Tian, W. (2018) Semi-Nonparametric Approximation and Index Options. Annals of Finance, 1-38. https://doi.org/10.1007/s10436-018-0341-4</mixed-citation></ref><ref id="scirp.94637-ref56"><label>56</label><mixed-citation publication-type="other" xlink:type="simple">Tian, W. (2014) Spanning with Indexes. Journal of Mathematical Economics, 53, 111-118. https://doi.org/10.1016/j.jmateco.2014.06.007</mixed-citation></ref><ref id="scirp.94637-ref57"><label>57</label><mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Tian</surname><given-names> W. </given-names></name>,<etal>et al</etal>. (<year>2018</year>)<article-title>The Financial Market: Not as Big as You Think</article-title><source> Mathematics and Financial Economics</source><volume> 51</volume>,<fpage> 1</fpage>-<lpage>19</lpage>.<pub-id pub-id-type="doi"></pub-id></mixed-citation></ref><ref id="scirp.94637-ref58"><label>58</label><mixed-citation publication-type="other" xlink:type="simple">Bolcskei, H., Grohs, P., Kutyniok, G. and Petersen, P. (2018) Optimal Approximation with Sparsely Connected Deep Neural Networks. Computer Science, 1-36. https://arxiv.org/abs/1705.01714</mixed-citation></ref><ref id="scirp.94637-ref59"><label>59</label><mixed-citation publication-type="other" xlink:type="simple">Henry-Labordere, P. (2015) Exact Simulation of Multi-Dimensional Stochastic Differential Equations. Working Paper, 1-28. https://doi.org/10.2139/ssrn.2598505https://papers.ssrn.com/sol3/papers.cfm?abstract_id=2598505</mixed-citation></ref><ref id="scirp.94637-ref60"><label>60</label><mixed-citation publication-type="other" xlink:type="simple">Prater, A. (2012) Discrete Sparse Fourier Hermite Approximations in High Dimensions. Doctoral Thesis, Syracuse University, New York.</mixed-citation></ref><ref id="scirp.94637-ref61"><label>61</label><mixed-citation publication-type="other" xlink:type="simple">Fonseca, Y., Medeiros, M., Vasconcelos, G. and Veiga, A. (2018) Boost: Boosting Smooth Trees for Partial Effect Estimation in Nonlinear Regressions. Statistics, 1-30. https://arxiv.org/pdf/1808.03698.pdf</mixed-citation></ref><ref id="scirp.94637-ref62"><label>62</label><mixed-citation publication-type="other" xlink:type="simple">Detemple, J. (2006) American-Style Derivatives: Valuation and Computation. Chapman and Hall/CRC, New York. https://doi.org/10.1201/9781420034868</mixed-citation></ref><ref id="scirp.94637-ref63"><label>63</label><mixed-citation publication-type="other" xlink:type="simple">Guyon, J. and Henry-Labordere, P. (2014) Nonliner Option Pricing. Chapman and Hall, New York. https://doi.org/10.1201/b16332</mixed-citation></ref><ref id="scirp.94637-ref64"><label>64</label><mixed-citation publication-type="other" xlink:type="simple">Detemple, J., Garcia, R. and Rindisbacher, M. (2005) Representation Formulas for Malliavin Derivatives of Diffusion Processes. Finance and Stochastics, 9, 349-367.https://doi.org/10.1007/s00780-004-0151-6</mixed-citation></ref><ref id="scirp.94637-ref65"><label>65</label><mixed-citation publication-type="other" xlink:type="simple">Detemple, J. and Rindisbacher, M. (2005) Asymptotic Properties of Monte Carlo Estimators of Derivatives. Management Science, 51, 1657-1675.https://doi.org/10.1287/mnsc.1050.0398</mixed-citation></ref><ref id="scirp.94637-ref66"><label>66</label><mixed-citation publication-type="other" xlink:type="simple">Detemple, J. (2014) Optimal Exercise for Derivative Securities. Annual Review of Financial Economics, 6, 459-487. https://doi.org/10.1146/annurev-financial-110613-034241</mixed-citation></ref><ref id="scirp.94637-ref67"><label>67</label><mixed-citation publication-type="other" xlink:type="simple">Fujii, M., Sato, S. and Takahashi, A. (2012) An FBSDE Approach to American Option Pricing with an Interacting Particle Method. Quantitative Finance, 1-18. https://arxiv.org/abs/1211.5867 https://doi.org/10.2139/ssrn.2180696</mixed-citation></ref><ref id="scirp.94637-ref68"><label>68</label><mixed-citation publication-type="other" xlink:type="simple">Chassagneux, J., Elie, R. and Kharroubi, I. (2010) A Note on Existence and Uniqueness for Solutions of Multidimensional Reflected BSDEs. Electronic Communications in Probability, 16, 120-128. https://doi.org/10.1214/ECP.v16-1614</mixed-citation></ref><ref id="scirp.94637-ref69"><label>69</label><mixed-citation publication-type="other" xlink:type="simple">Collin-Dufresne, P. and Goldstein, R. (2003) Generalizing the Affine Framework to HJM and Random Field Models. SSRN Electronic Journal. https://doi.org/10.2139/ssrn.410421https://papers.ssrn.com/sol3/papers.cfm?abstract_id=410421</mixed-citation></ref><ref id="scirp.94637-ref70"><label>70</label><mixed-citation publication-type="other" xlink:type="simple">Carmona, R. and Delarue, F. (2015) Forward-Backward Stochastic Differential Equations and Controlled McKean-Vlasov Dynamics. Annals of Probability, 43, 2647-2700. https://doi.org/10.1214/14-AOP946</mixed-citation></ref><ref id="scirp.94637-ref71"><label>71</label><mixed-citation publication-type="other" xlink:type="simple">Bianchi, D., Büchner, M. and Tamoni, A. (2019) Bond Risk Premia with Machine Learning. USC-INET Research Paper No. 19-11. https://doi.org/10.2139/ssrn.3400941https://papers.ssrn.com/sol3/papers.cfm?abstract_id=3232721</mixed-citation></ref><ref id="scirp.94637-ref72"><label>72</label><mixed-citation publication-type="other" xlink:type="simple">Chen, L., Pelger, M. and Zhu, J. (2019) Deep Learning in Asset Pricing. Quantitative Finance, 1-89. https://arxiv.org/abs/1904.00745 https://doi.org/10.2139/ssrn.3350138</mixed-citation></ref><ref id="scirp.94637-ref73"><label>73</label><mixed-citation publication-type="other" xlink:type="simple">Feng, G., Polson, N. and Xu, J. (2019) Deep Learning in Asset Pricing. Statistics, 1-33. https://papers.ssrn.com/sol3/papers.cfm?abstract_id=3350138</mixed-citation></ref><ref id="scirp.94637-ref74"><label>74</label><mixed-citation publication-type="other" xlink:type="simple">Yang, Q., Ye, T. and Zhang, L. (2018) A General Framework of Optimal Investment. Working Paper. https://papers.ssrn.com/sol3/papers.cfm?abstract_id=3136708</mixed-citation></ref><ref id="scirp.94637-ref75"><label>75</label><mixed-citation publication-type="other" xlink:type="simple">Yu, P., Lee, J., Kulyatin, I., Shi, Z. and Dasgupta, S. (2019) Model-Based Deep Reinforcement Learning for Dynamic Portfolio Optimization. Computer Science, 1-21. https://arxiv.org/abs/1901.08740</mixed-citation></ref><ref id="scirp.94637-ref76"><label>76</label><mixed-citation publication-type="other" xlink:type="simple">Kingma, D. and Ba, J.L. (2014) Adam: A Method for Stochastic Optimization. Computer Science, 1-15. https://arxiv.org/abs/1412.6980</mixed-citation></ref><ref id="scirp.94637-ref77"><label>77</label><mixed-citation publication-type="other" xlink:type="simple">Heston, S. (1993) A Closed-Form Solution for Options with Stochastic Volatility with Applications to Bond and Currency Options. The Review of Financial Studies, 6, 327-343. https://doi.org/10.1093/rfs/6.2.327</mixed-citation></ref><ref id="scirp.94637-ref78"><label>78</label><mixed-citation publication-type="other" xlink:type="simple">Dupire, B. (1994) Pricing with a Smile. Risk. http://www.risk.net/data/risk/pdf/technical/2007/risk20_0707_technical_volatility.pdf</mixed-citation></ref><ref id="scirp.94637-ref79"><label>79</label><mixed-citation publication-type="other" xlink:type="simple">Homescu, C. (2014) Local Stochastic Volatility Models: Calibration and Pricing. Working Paper. https://papers.ssrn.com/sol3/papers.cfm?abstract_id=2448098https://doi.org/10.2139/ssrn.2448098</mixed-citation></ref><ref id="scirp.94637-ref80"><label>80</label><mixed-citation publication-type="other" xlink:type="simple">Broadie, M., Chernov, M. and Johannes, M. (2007) Model Specification and Risk Premia: Evidence from futures Options. Journal of Finance, 62, 1453-1490.https://doi.org/10.1111/j.1540-6261.2007.01241.x</mixed-citation></ref><ref id="scirp.94637-ref81"><label>81</label><mixed-citation publication-type="other" xlink:type="simple">Guennon, H. (2016) Local Volatility Models Enhanced with Jumps. Working Paper, 1-11. https://papers.ssrn.com/abstract=2781102https://doi.org/10.2139/ssrn.2781102</mixed-citation></ref><ref id="scirp.94637-ref82"><label>82</label><mixed-citation publication-type="other" xlink:type="simple">Buehler, H., Gonon, L., Teichmann, J. and Wood, B. (2018) Deep Hedging. Working Paper. https://doi.org/10.2139/ssrn.3120710  https://arxiv.org/abs/1802.03042</mixed-citation></ref><ref id="scirp.94637-ref83"><label>83</label><mixed-citation publication-type="other" xlink:type="simple">Halperin, I. (2018) The QLBS Q-Learner Goes NuQLear: Fitted Q Iteration, Inverse RL, and Option Portfolios. Quantitative Finance, 1-18. https://arxiv.org/abs/1801.06077https://doi.org/10.2139/ssrn.3102707</mixed-citation></ref><ref id="scirp.94637-ref84"><label>84</label><mixed-citation publication-type="other" xlink:type="simple">Halperin, I. (2018) QLBS: Q-Learner in the Black-Scholes(-Merton) Worlds. Quantitative Finance, 1-34. https://arxiv.org/abs/1712.04609https://doi.org/10.2139/ssrn.3087076</mixed-citation></ref><ref id="scirp.94637-ref85"><label>85</label><mixed-citation publication-type="other" xlink:type="simple">Schroder, M. and Skiadas, C. (2008) Optimality and State Pricing in Constrained Financial Markets with Recursive Utility under Continuous and Discontinuous Information. Mathematical Finance, 18, 199-238. https://doi.org/10.1111/j.1467-9965.2007.00330.x</mixed-citation></ref><ref id="scirp.94637-ref86"><label>86</label><mixed-citation publication-type="other" xlink:type="simple">Detemple, J. and Zapatero, F. (1991) Asset Prices in an Exchange Economy with Habit Formation. Econometrica, 59, 1633-1657. https://doi.org/10.2307/2938283</mixed-citation></ref><ref id="scirp.94637-ref87"><label>87</label><mixed-citation publication-type="other" xlink:type="simple">Karatzas, I., Lehoczky, J., Shreve, S. and Xu, G. (1991) Martingale and Duality Methods for Utility Maximization in a Incomplete Market. SIAM Journal on Control and Optimization, 29, 702-730. https://doi.org/10.1137/0329039</mixed-citation></ref><ref id="scirp.94637-ref88"><label>88</label><mixed-citation publication-type="other" xlink:type="simple">He, H. and Pearson, N. (1991) Consumption and Portfolio Policies with Incomplete Markets and Short-Sale Constraints: The Infinite Dimensional Case. Journal of Economic Theory, 54, 259-304. https://doi.org/10.1016/0022-0531(91)90123-L</mixed-citation></ref><ref id="scirp.94637-ref89"><label>89</label><mixed-citation publication-type="other" xlink:type="simple">Karatzas, I. and Cvitanic, J. (1992) Convex Duality in Constrained Portfolio Optimization. Annals of Applied Probability, 2, 767-818.https://doi.org/10.1214/aoap/1177005576</mixed-citation></ref><ref id="scirp.94637-ref90"><label>90</label><mixed-citation publication-type="other" xlink:type="simple">Detemple, J., Garcia, R. and Rindisbacher, M. (2003) A Monte Carlo Method for Optimal Portfolios. Journal of Finance, 58, 401-446.https://doi.org/10.1111/1540-6261.00529</mixed-citation></ref><ref id="scirp.94637-ref91"><label>91</label><mixed-citation publication-type="other" xlink:type="simple">Detemple, J., Garcia, R. and Rindisbacher, M. (2005) Intertemporal Asset Allocation: A Comparison of Methods. Journal of Banking and Finance, 29, 2821-2848.https://doi.org/10.1016/j.jbankfin.2005.02.004</mixed-citation></ref><ref id="scirp.94637-ref92"><label>92</label><mixed-citation publication-type="other" xlink:type="simple">Detemple, J. and Rindisbacher, M. (2010) Dynamic Asset Allocation: Portfolio Decomposition Formula and Applications. The Review of Financial Studies, 23, 25-100. https://doi.org/10.1093/rfs/hhp040</mixed-citation></ref><ref id="scirp.94637-ref93"><label>93</label><mixed-citation publication-type="other" xlink:type="simple">Detemple, J. (2012) Portfolio Selection: A Review. Journal of Optimization Theory and Applications, 161, 1-21. https://doi.org/10.1007/s10957-012-0208-1</mixed-citation></ref><ref id="scirp.94637-ref94"><label>94</label><mixed-citation publication-type="other" xlink:type="simple">Matoussi, A. and Xing, H. (2016) Convex Duality for Stochastic Differential Utility. Quantitative Finance, 1-22. http://arxiv.org/pdf/1601.03562.pdfhttps://doi.org/10.2139/ssrn.2715425</mixed-citation></ref><ref id="scirp.94637-ref95"><label>95</label><mixed-citation publication-type="other" xlink:type="simple">Kraft, H., Seiferling, T. and Seifried, F. (2015) Optimal Consumption and Investment with Epstein-Z in Recursive Utility. Working Paper. https://doi.org/10.2139/ssrn.2444747http://papers.ssrn.com/sol3/papers.cfm?abstract_id=2424706</mixed-citation></ref><ref id="scirp.94637-ref96"><label>96</label><mixed-citation publication-type="other" xlink:type="simple">Ait-Sahalia, Y. (2008) Closed-Form Likelihood Expansions for Multivariate Diffusions. Annals of Statistics, 36, 906-937. https://doi.org/10.1214/009053607000000622</mixed-citation></ref><ref id="scirp.94637-ref97"><label>97</label><mixed-citation publication-type="other" xlink:type="simple">Filipovic, D., Mayerhofer, E. and Schneider, P. (2013) Density Approximations for Multivariate Affine Jump Diffusion Processes. Journal of Econometrics, 176, 93-111. https://doi.org/10.1016/j.jeconom.2012.12.003</mixed-citation></ref><ref id="scirp.94637-ref98"><label>98</label><mixed-citation publication-type="other" xlink:type="simple">Van Handel, R. (2008) Hidden Markov Models. Princeton Lecture Notes.</mixed-citation></ref><ref id="scirp.94637-ref99"><label>99</label><mixed-citation publication-type="other" xlink:type="simple">Markowitz, H. (1952) Portfolio Selection. Journal of Finance, 7, 77-91.https://doi.org/10.1111/j.1540-6261.1952.tb01525.x</mixed-citation></ref><ref id="scirp.94637-ref100"><label>100</label><mixed-citation publication-type="other" xlink:type="simple">Schneider, P. and Trojani, F. (2018) (Almost) Model Free Recovery. Forthcoming in Journal of Finance, 74, 323-370. https://doi.org/10.1111/jofi.12737</mixed-citation></ref><ref id="scirp.94637-ref101"><label>101</label><mixed-citation publication-type="other" xlink:type="simple">Chabakauri, G. (2013) Dynamic Equilibrium with Two Stocks, Heterogeneous Investors, and Portfolio Constraints. The Review of Financial Studies, 26, 3104-3141. https://doi.org/10.2139/ssrn.2221073http://papers.ssrn.com/sol3/papers.cfm?abstract_id=2221073</mixed-citation></ref><ref id="scirp.94637-ref102"><label>102</label><mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Chabakauri</surname><given-names> G. </given-names></name>,<etal>et al</etal>. (<year>2015</year>)<article-title>Asset Pricing with Heterogeneous Preferences, Beliefs, and Portfolio Constraints</article-title><source> Journal of Monetary Economics</source><volume> 75</volume>,<fpage> 21</fpage>-<lpage>34</lpage>.<pub-id pub-id-type="doi"></pub-id></mixed-citation></ref><ref id="scirp.94637-ref103"><label>103</label><mixed-citation publication-type="other" xlink:type="simple">Kardaras, C., Xing, H. and Zitkovic, G. (2015) Incomplete Stochastic Equilibria for Dynamic Monetary Utility. Mathematics, 1-33. https://arxiv.org/abs/1505.07224</mixed-citation></ref><ref id="scirp.94637-ref104"><label>104</label><mixed-citation publication-type="other" xlink:type="simple">Halle, J.O. (2010) Backward Stochastic Differential Equations with Jumps. Master Thesis, University of Oslo, Oslo, Norway.</mixed-citation></ref><ref id="scirp.94637-ref105"><label>105</label><mixed-citation publication-type="other" xlink:type="simple">Dalderop, J. (2016) Nonparametric State-Price Density Estimation Using High Frequency Data. Working Paper. https://doi.org/10.2139/ssrn.2718938</mixed-citation></ref></ref-list></back></article>