<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.4 20241031//EN" "JATS-journalpublishing1-4.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article" dtd-version="1.4" xml:lang="en">
  <front>
    <journal-meta>
      <journal-id journal-id-type="publisher-id">jamp</journal-id>
      <journal-title-group>
        <journal-title>Journal of Applied Mathematics and Physics</journal-title>
      </journal-title-group>
      <issn pub-type="epub">2327-4379</issn>
      <issn pub-type="ppub">2327-4352</issn>
      <publisher>
        <publisher-name>Scientific Research Publishing</publisher-name>
      </publisher>
    </journal-meta>
    <article-meta>
      <article-id pub-id-type="doi">10.4236/jamp.2026.147138</article-id>
      <article-id pub-id-type="publisher-id">jamp-152984</article-id>
      <article-categories>
        <subj-group>
          <subject>Article</subject>
        </subj-group>
        <subj-group>
          <subject>Physics</subject>
          <subject>Mathematics</subject>
        </subj-group>
      </article-categories>
      <title-group>
        <article-title>Applications of Large Deviation Theory to Massive Data Models: A Numerical Approach for Measuring the Rarity of Extreme Events</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <contrib-id contrib-id-type="orcid">0009-0009-2741-8953</contrib-id>
          <name name-style="western">
            <surname>Diop</surname>
            <given-names>Bou</given-names>
          </name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
      </contrib-group>
      <aff id="aff1"><label>1</label> Department of Mathematics and Computer Science, Faculty of Technology and Computer Science, University Iba Der Thiam of Thiès, Thiès, Senegal </aff>
      <author-notes>
        <fn fn-type="conflict" id="fn-conflict">
          <p>The author declares no conflicts of interest regarding the publication of this paper.</p>
        </fn>
      </author-notes>
      <pub-date pub-type="epub">
        <day>14</day>
        <month>07</month>
        <year>2026</year>
      </pub-date>
      <pub-date pub-type="collection">
        <month>07</month>
        <year>2026</year>
      </pub-date>
      <volume>14</volume>
      <issue>07</issue>
      <fpage>2773</fpage>
      <lpage>2782</lpage>
      <history>
        <date date-type="received">
          <day>20</day>
          <month>11</month>
          <year>2025</year>
        </date>
        <date date-type="accepted">
          <day>28</day>
          <month>07</month>
          <year>2026</year>
        </date>
        <date date-type="published">
          <day>31</day>
          <month>07</month>
          <year>2026</year>
        </date>
      </history>
      <permissions>
        <copyright-statement>© 2026 by the authors and Scientific Research Publishing Inc.</copyright-statement>
        <copyright-year>2026</copyright-year>
        <license license-type="open-access">
          <license-p> This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license ( <ext-link ext-link-type="uri" xlink:href="https://creativecommons.org/licenses/by/4.0/">https://creativecommons.org/licenses/by/4.0/</ext-link> ). </license-p>
        </license>
      </permissions>
      <self-uri content-type="doi" xlink:href="https://doi.org/10.4236/jamp.2026.147138">https://doi.org/10.4236/jamp.2026.147138</self-uri>
      <abstract>
        <p>The exponential growth of data in modern systems increases the need for robust statistical methods capable of describing rare but critical events. This paper proposes a computational methodology based on large deviation theory to efficiently estimate rare-event probabilities in large-scale systems. We introduce a coherent approach that integrates a solid theoretical framework with advanced numerical algorithms for estimating rate functions. Unlike standard Monte Carlo methods, our approach leverages convex optimization and adaptive importance sampling, which significantly reduces variance and accelerates estimator convergence, as demonstrated by our numerical results. These results attest to the numerical robustness and practical utility of our method in diverse fields such as quantitative finance, telecommunications networks, and machine learning.</p>
      </abstract>
      <kwd-group kwd-group-type="author-generated" xml:lang="en">
        <kwd>Large Deviation Theory</kwd>
        <kwd>Rare-Event Probability Estimation</kwd>
        <kwd>Adaptive Importance Sampling</kwd>
        <kwd>Convex Optimization</kwd>
        <kwd>Rate Function Estimation</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec1">
      <title>1. Introduction</title>
      <p>The surge in data volume within contemporary systems has heightened the demand for resilient statistical methodologies adept at delineating infrequent yet pivotal occurrences. Initiated by Cramér [<xref ref-type="bibr" rid="B1">1</xref>] and developed by Donsker-Varadhan [<xref ref-type="bibr" rid="B2">2</xref>], Varadhan [<xref ref-type="bibr" rid="B3">3</xref>], among others, large deviation theory provides a powerful mathematical framework for assessing the probability of extreme events. It seeks to calculate the likelihood of highly infrequent occurrences by characterizing their exponential decay in relation to system size. In contrast to the foundational works of Dembo-Zeitouni [<xref ref-type="bibr" rid="B4">4</xref>] and recent advances by Touchette [<xref ref-type="bibr" rid="B5">5</xref>][<xref ref-type="bibr" rid="B6">6</xref>], our contribution introduces an innovative computational methodology for efficiently estimating rate functions in massive data contexts. The novelty lies in the systematic integration of convex optimization algorithms, ensuring enhanced numerical stability and faster convergence compared to classical approaches.</p>
      <p>As illustrated in <xref ref-type="fig" rid="fig1">Figure 1</xref>, our framework establishes a comprehensive approach for estimating rate functions from empirical data. The numerical results presented in <bold>Table 1</bold> demonstrate the effectiveness of our methodology compared to traditional approaches.</p>
      <p>The following sections detail the theoretical framework, the associated numerical methods, as well as concrete applications.</p>
      <fig id="fig1">
        <label>Figure 1</label>
        <graphic xlink:href="https://html.scirp.org/file/1724455-rId15.jpeg?20260731025242" />
      </fig>
      <p><bold>Figure 1.</bold>Comparison between estimated rate function (blue solid line) and theoretical Cramér rate function (red dashed line) for standard normal distribution. The close agreement validates our numerical estimation methodology.</p>
      <p><bold>Table 1</bold><bold>.</bold>Comparison of estimation methods for N = 1000. Tests were done on an Intel i7 processor with 16 GB of RAM.</p>
      <table-wrap id="tbl1">
        <label>Table 1</label>
        <table>
          <tbody>
            <tr>
              <td>
                <bold>Method</bold>
              </td>
              <td>
                <bold>Estimation</bold>
              </td>
              <td>
                <bold>Variance</bold>
              </td>
              <td>
                <bold>Computation Time (s)</bold>
              </td>
            </tr>
            <tr>
              <td>Naive Monte-Carlo</td>
              <td>
                2.1 × 10
                <sup>−</sup>
                <sup>5</sup>
              </td>
              <td>
                1.8 × 10
                <sup>−</sup>
                <sup>4</sup>
              </td>
              <td>12.3</td>
            </tr>
            <tr>
              <td>Importance sampling</td>
              <td>
                1.9 × 10
                <sup>−</sup>
                <sup>5</sup>
              </td>
              <td>
                2.3 × 10
                <sup>−</sup>
                <sup>7</sup>
              </td>
              <td>8.7</td>
            </tr>
            <tr>
              <td>Saddle-point</td>
              <td>
                2.0 × 10
                <sup>−</sup>
                <sup>5</sup>
              </td>
              <td>-</td>
              <td>0.1</td>
            </tr>
          </tbody>
        </table>
      </table-wrap>
    </sec>
    <sec id="sec2">
      <title>2. Related Work and State of the Art</title>
      <p>Large deviation theory has undergone considerable evolution since the pioneering work of Cramér. The classical texts of Dembo-Zeitouni [<xref ref-type="bibr" rid="B4">4</xref>] and Dupuis-Ellis [<xref ref-type="bibr" rid="B7">7</xref>] lay the foundations of contemporary theory. Touchette [<xref ref-type="bibr" rid="B5">5</xref>][<xref ref-type="bibr" rid="B6">6</xref>] has recently broadened applications to statistical physics and complex systems.</p>
      <p>The work of Bottou and Bousquet [<xref ref-type="bibr" rid="B8">8</xref>] in the field of massive data opened the way for the use of large deviations in machine learning. Recent advances, such as those of Glasserman and He [<xref ref-type="bibr" rid="B9">9</xref>] on large deviations in machine learning, as well as the work of Roberts and Rosenthal [<xref ref-type="bibr" rid="B10">10</xref>] on rare-event simulation in high dimensions, have enriched the methodological framework. Recent research by Chatterjee [<xref ref-type="bibr" rid="B11">11</xref>] on neural networks and that of Chen and Gao [<xref ref-type="bibr" rid="B12">12</xref>] on generative models reinforce these connections.</p>
      <p>These studies provide the theoretical foundation for our contribution, which aims to amalgamate numerical optimization and simulation techniques for calculating rate functions in extensive data systems. Despite these advances, few studies propose a complete integration of large deviation theory and large-scale numerical estimation. Our method stands out by explicitly combining analytical rigor with advanced computational techniques, enabling reliable and efficient estimation of rate functions. This synthesis paves the way for a unified methodology that merges theoretical estimation with extensive numerical techniques.</p>
    </sec>
    <sec id="sec3">
      <title>3. Mathematical Framework</title>
      <sec id="sec3dot1">
        <title>3.1. Notations and Definitions</title>
        <p>Let <inline-formula><mml:math><mml:mrow><mml:msub><mml:mrow><mml:mrow><mml:mo> ( </mml:mo><mml:mrow><mml:msub><mml:mi> X </mml:mi><mml:mi> i </mml:mi></mml:msub></mml:mrow><mml:mo> ) </mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi> i </mml:mi><mml:mo> ≥ </mml:mo><mml:mn> 1 </mml:mn></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula> be a sequence of independent and identically distributed (i.i.d.) random variables and let <inline-formula><mml:math><mml:mrow><mml:msub><mml:mi> S </mml:mi><mml:mi> N </mml:mi></mml:msub><mml:mo> = </mml:mo><mml:mfrac><mml:mn> 1 </mml:mn><mml:mi> N </mml:mi></mml:mfrac><mml:mstyle displaystyle="true"><mml:msubsup><mml:mo> ∑ </mml:mo><mml:mrow><mml:mi> i </mml:mi><mml:mo> = </mml:mo><mml:mn> 1 </mml:mn></mml:mrow><mml:mi> N </mml:mi></mml:msubsup><mml:mrow><mml:msub><mml:mi> X </mml:mi><mml:mi> i </mml:mi></mml:msub></mml:mrow></mml:mstyle></mml:mrow></mml:math></inline-formula> be their empirical mean.</p>
        <p><bold>Definition 3.1 (Principle of Large Deviations)</bold><bold>.</bold><italic>Let</italic><inline-formula><mml:math><mml:mrow><mml:msub><mml:mrow><mml:mrow><mml:mo> { </mml:mo><mml:mrow><mml:msub><mml:mi> ℙ </mml:mi><mml:mi> N </mml:mi></mml:msub></mml:mrow><mml:mo> } </mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi> N </mml:mi><mml:mo> ≥ </mml:mo><mml:mn> 1 </mml:mn></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula><italic>be a sequence of probability measures on a Polish space</italic><inline-formula><mml:math><mml:mi mathvariant="script"> X </mml:mi></mml:math></inline-formula><italic>. We say that</italic><inline-formula><mml:math><mml:mrow><mml:mrow><mml:mo> { </mml:mo><mml:mrow><mml:msub><mml:mi> ℙ </mml:mi><mml:mi> N </mml:mi></mml:msub></mml:mrow><mml:mo> } </mml:mo></mml:mrow></mml:mrow></mml:math></inline-formula><italic>satisfies the large deviation principle with rate function</italic><inline-formula><mml:math><mml:mrow><mml:mi> I </mml:mi><mml:mo> : </mml:mo><mml:mi mathvariant="script"> X </mml:mi><mml:mo> → </mml:mo><mml:mrow><mml:mo> [ </mml:mo><mml:mrow><mml:mn> 0 </mml:mn><mml:mo> , </mml:mo><mml:mo> + </mml:mo><mml:mi> ∞ </mml:mi></mml:mrow><mml:mo> ] </mml:mo></mml:mrow></mml:mrow></mml:math></inline-formula><italic>and speed</italic><inline-formula><mml:math><mml:mrow><mml:msub><mml:mi> a </mml:mi><mml:mi> N </mml:mi></mml:msub><mml:mo> → </mml:mo><mml:mo> + </mml:mo><mml:mi> ∞ </mml:mi></mml:mrow></mml:math></inline-formula><italic>if for every Borel set</italic><inline-formula><mml:math><mml:mrow><mml:mi> E </mml:mi><mml:mo> ⊂ </mml:mo><mml:mi mathvariant="script"> X </mml:mi></mml:mrow></mml:math></inline-formula></p>
        <disp-formula id="FD1">
          <mml:math display="inline">
            <mml:mrow>
              <mml:mo>−</mml:mo>
              <mml:msub>
                <mml:mrow>
                  <mml:mi>inf</mml:mi>
                </mml:mrow>
                <mml:mrow>
                  <mml:mi>x</mml:mi>
                  <mml:mo>∈</mml:mo>
                  <mml:msup>
                    <mml:mi>E</mml:mi>
                    <mml:mo>∘</mml:mo>
                  </mml:msup>
                </mml:mrow>
              </mml:msub>
              <mml:mi>I</mml:mi>
              <mml:mrow>
                <mml:mo>(</mml:mo>
                <mml:mi>x</mml:mi>
                <mml:mo>)</mml:mo>
              </mml:mrow>
              <mml:mo>≤</mml:mo>
              <mml:mi>lim</mml:mi>
              <mml:msub>
                <mml:mrow>
                  <mml:mi>inf</mml:mi>
                </mml:mrow>
                <mml:mrow>
                  <mml:mi>N</mml:mi>
                  <mml:mo>→</mml:mo>
                  <mml:mi>∞</mml:mi>
                </mml:mrow>
              </mml:msub>
              <mml:mfrac>
                <mml:mn>1</mml:mn>
                <mml:mrow>
                  <mml:msub>
                    <mml:mi>a</mml:mi>
                    <mml:mi>N</mml:mi>
                  </mml:msub>
                </mml:mrow>
              </mml:mfrac>
              <mml:mi>ln</mml:mi>
              <mml:msub>
                <mml:mi>ℙ</mml:mi>
                <mml:mi>N</mml:mi>
              </mml:msub>
              <mml:mrow>
                <mml:mo>(</mml:mo>
                <mml:mi>E</mml:mi>
                <mml:mo>)</mml:mo>
              </mml:mrow>
              <mml:mo>≤</mml:mo>
              <mml:mi>lim</mml:mi>
              <mml:msub>
                <mml:mrow>
                  <mml:mi>sup</mml:mi>
                </mml:mrow>
                <mml:mrow>
                  <mml:mi>N</mml:mi>
                  <mml:mo>→</mml:mo>
                  <mml:mi>∞</mml:mi>
                </mml:mrow>
              </mml:msub>
              <mml:mfrac>
                <mml:mn>1</mml:mn>
                <mml:mrow>
                  <mml:msub>
                    <mml:mi>a</mml:mi>
                    <mml:mi>N</mml:mi>
                  </mml:msub>
                </mml:mrow>
              </mml:mfrac>
              <mml:mi>ln</mml:mi>
              <mml:msub>
                <mml:mi>ℙ</mml:mi>
                <mml:mi>n</mml:mi>
              </mml:msub>
              <mml:mrow>
                <mml:mo>(</mml:mo>
                <mml:mi>E</mml:mi>
                <mml:mo>)</mml:mo>
              </mml:mrow>
              <mml:mo>≤</mml:mo>
              <mml:mo>−</mml:mo>
              <mml:msub>
                <mml:mrow>
                  <mml:mi>inf</mml:mi>
                </mml:mrow>
                <mml:mrow>
                  <mml:mi>x</mml:mi>
                  <mml:mo>∈</mml:mo>
                  <mml:mover accent="true">
                    <mml:mi>E</mml:mi>
                    <mml:mo>¯</mml:mo>
                  </mml:mover>
                </mml:mrow>
              </mml:msub>
              <mml:mi>I</mml:mi>
              <mml:mrow>
                <mml:mo>(</mml:mo>
                <mml:mi>x</mml:mi>
                <mml:mo>)</mml:mo>
              </mml:mrow>
              <mml:mo>,</mml:mo>
            </mml:mrow>
          </mml:math>
        </disp-formula>
        <p>where <inline-formula><mml:math><mml:mrow><mml:msup><mml:mi> E </mml:mi><mml:mo> ∘ </mml:mo></mml:msup></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math><mml:mover accent="true"><mml:mi> E </mml:mi><mml:mo> ¯ </mml:mo></mml:mover></mml:math></inline-formula> denote the interior and closure of <inline-formula><mml:math><mml:mi> E </mml:mi></mml:math></inline-formula> , respectively.</p>
        <p><bold>Remark 3.2</bold><bold>.</bold><italic>The rate function</italic><inline-formula><mml:math><mml:mrow><mml:mi> I </mml:mi><mml:mrow><mml:mo> ( </mml:mo><mml:mi> x </mml:mi><mml:mo> ) </mml:mo></mml:mrow></mml:mrow></mml:math></inline-formula><italic>quantifies the exponential cost of shifting the empirical mean towards the value</italic><inline-formula><mml:math><mml:mi> x </mml:mi></mml:math></inline-formula><italic>. A high value of</italic><inline-formula><mml:math><mml:mrow><mml:mi> I </mml:mi><mml:mrow><mml:mo> ( </mml:mo><mml:mi> x </mml:mi><mml:mo> ) </mml:mo></mml:mrow></mml:mrow></mml:math></inline-formula><italic>indicates a lower probability of the corresponding event.</italic></p>
        <p><bold>Remark 3.3</bold><bold>.</bold><italic>The function</italic><inline-formula><mml:math><mml:mrow><mml:mi> I </mml:mi><mml:mrow><mml:mo> ( </mml:mo><mml:mi> x </mml:mi><mml:mo> ) </mml:mo></mml:mrow></mml:mrow></mml:math></inline-formula><italic>can be interpreted as a convex distance between the true distribution and the empirical distribution. This information-theoretic interpretation</italic>,<italic>consistent with Sanov’s theorem</italic>,<italic>makes the rate function equivalent to the</italic><italic>Kullback-Leibler</italic><italic>divergence.</italic></p>
      </sec>
      <sec id="sec3dot2">
        <title>3.2. Cramér Transform and Rate Function</title>
        <p><bold>Theorem 3.4 (Cramér’s Theorem)</bold><italic>Let</italic><inline-formula><mml:math><mml:mrow><mml:msub><mml:mi> X </mml:mi><mml:mn> 1 </mml:mn></mml:msub><mml:mo> , </mml:mo><mml:msub><mml:mi> X </mml:mi><mml:mn> 2 </mml:mn></mml:msub><mml:mo> , </mml:mo><mml:mo> ⋯ </mml:mo></mml:mrow></mml:math></inline-formula><italic>be a sequence of</italic><italic>i.i.d.</italic><italic>random variables in</italic><inline-formula><mml:math><mml:mrow><mml:msup><mml:mi> ℝ </mml:mi><mml:mi> d </mml:mi></mml:msup></mml:mrow></mml:math></inline-formula><italic>. The Cramér transform is defined by:</italic></p>
        <disp-formula id="FD2">
          <mml:math>
            <mml:mrow>
              <mml:mi>λ</mml:mi>
              <mml:mrow>
                <mml:mo>(</mml:mo>
                <mml:mi>θ</mml:mi>
                <mml:mo>)</mml:mo>
              </mml:mrow>
              <mml:mo>=</mml:mo>
              <mml:mi>ln</mml:mi>
              <mml:mi mathvariant="double-struck">E</mml:mi>
              <mml:mrow>
                <mml:mo>[</mml:mo>
                <mml:mrow>
                  <mml:msup>
                    <mml:mtext>e</mml:mtext>
                    <mml:mrow>
                      <mml:mrow>
                        <mml:mo>〈</mml:mo>
                        <mml:mrow>
                          <mml:mi>θ</mml:mi>
                          <mml:mo>,</mml:mo>
                          <mml:msub>
                            <mml:mi>X</mml:mi>
                            <mml:mn>1</mml:mn>
                          </mml:msub>
                        </mml:mrow>
                        <mml:mo>〉</mml:mo>
                      </mml:mrow>
                    </mml:mrow>
                  </mml:msup>
                </mml:mrow>
                <mml:mo>]</mml:mo>
              </mml:mrow>
              <mml:mo>,</mml:mo>
            </mml:mrow>
          </mml:math>
        </disp-formula>
        <p>where <inline-formula><mml:math><mml:mrow><mml:mrow><mml:mo> 〈 </mml:mo><mml:mrow><mml:mo> ⋅ </mml:mo><mml:mo> , </mml:mo><mml:mo> ⋅ </mml:mo></mml:mrow><mml:mo> 〉 </mml:mo></mml:mrow></mml:mrow></mml:math></inline-formula> denotes the inner product. The Legendre-Fenchel transform gives the rate function for the empirical mean <inline-formula><mml:math><mml:mrow><mml:msub><mml:mi> S </mml:mi><mml:mi> N </mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> :</p>
        <disp-formula id="FD3">
          <mml:math>
            <mml:mrow>
              <mml:mi>I</mml:mi>
              <mml:mrow>
                <mml:mo>(</mml:mo>
                <mml:mi>x</mml:mi>
                <mml:mo>)</mml:mo>
              </mml:mrow>
              <mml:mo>=</mml:mo>
              <mml:munder>
                <mml:mrow>
                  <mml:mtext>sup</mml:mtext>
                </mml:mrow>
                <mml:mrow>
                  <mml:mi>θ</mml:mi>
                  <mml:mo>∈</mml:mo>
                  <mml:msup>
                    <mml:mi>ℝ</mml:mi>
                    <mml:mi>d</mml:mi>
                  </mml:msup>
                </mml:mrow>
              </mml:munder>
              <mml:mrow>
                <mml:mo>[</mml:mo>
                <mml:mrow>
                  <mml:mrow>
                    <mml:mo>〈</mml:mo>
                    <mml:mrow>
                      <mml:mi>θ</mml:mi>
                      <mml:mo>,</mml:mo>
                      <mml:mi>x</mml:mi>
                    </mml:mrow>
                    <mml:mo>〉</mml:mo>
                  </mml:mrow>
                  <mml:mo>−</mml:mo>
                  <mml:mi>λ</mml:mi>
                  <mml:mrow>
                    <mml:mo>(</mml:mo>
                    <mml:mi>θ</mml:mi>
                    <mml:mo>)</mml:mo>
                  </mml:mrow>
                </mml:mrow>
                <mml:mo>]</mml:mo>
              </mml:mrow>
              <mml:mo>.</mml:mo>
            </mml:mrow>
          </mml:math>
        </disp-formula>
        <p><bold>Example 3.5 (Gaussian case)</bold><bold>.</bold><italic>If</italic><inline-formula><mml:math><mml:mrow><mml:msub><mml:mi> X </mml:mi><mml:mi> i </mml:mi></mml:msub><mml:mo> ~ </mml:mo><mml:mi mathvariant="script"> N </mml:mi><mml:mrow><mml:mo> ( </mml:mo><mml:mrow><mml:mn> 0 </mml:mn><mml:mo> , </mml:mo><mml:mn> 1 </mml:mn></mml:mrow><mml:mo> ) </mml:mo></mml:mrow></mml:mrow></mml:math></inline-formula> ,<italic>then</italic><inline-formula><mml:math><mml:mrow><mml:mi> λ </mml:mi><mml:mrow><mml:mo> ( </mml:mo><mml:mi> θ </mml:mi><mml:mo> ) </mml:mo></mml:mrow><mml:mo> = </mml:mo><mml:mrow><mml:mrow><mml:msup><mml:mi> θ </mml:mi><mml:mn> 2 </mml:mn></mml:msup></mml:mrow><mml:mo> / </mml:mo><mml:mn> 2 </mml:mn></mml:mrow></mml:mrow></mml:math></inline-formula><italic>and the rate function is</italic><inline-formula><mml:math><mml:mrow><mml:mi> I </mml:mi><mml:mrow><mml:mo> ( </mml:mo><mml:mi> x </mml:mi><mml:mo> ) </mml:mo></mml:mrow><mml:mo> = </mml:mo><mml:mrow><mml:mrow><mml:msup><mml:mi> x </mml:mi><mml:mn> 2 </mml:mn></mml:msup></mml:mrow><mml:mo> / </mml:mo><mml:mn> 2 </mml:mn></mml:mrow></mml:mrow></mml:math></inline-formula><italic>.</italic></p>
      </sec>
      <sec id="sec3dot3">
        <title>3.3. Convergence of the Numerical Estimator</title>
        <p><bold>Lemma 3.6</bold><bold>.</bold><italic>Let</italic><inline-formula><mml:math><mml:mrow><mml:msub><mml:mover accent="true"><mml:mi> I </mml:mi><mml:mo> ^ </mml:mo></mml:mover><mml:mi> N </mml:mi></mml:msub><mml:mrow><mml:mo> ( </mml:mo><mml:mi> x </mml:mi><mml:mo> ) </mml:mo></mml:mrow></mml:mrow></mml:math></inline-formula><italic>be the numerical estimate of the rate function obtained by convex optimization. Under standard regularity conditions</italic>,<italic>we obtain almost sure convergence:</italic></p>
        <disp-formula id="FD4">
          <mml:math>
            <mml:mrow>
              <mml:munder>
                <mml:mrow>
                  <mml:mtext>lim</mml:mtext>
                </mml:mrow>
                <mml:mrow>
                  <mml:mi>N</mml:mi>
                  <mml:mo>→</mml:mo>
                  <mml:mi>∞</mml:mi>
                </mml:mrow>
              </mml:munder>
              <mml:msub>
                <mml:mover accent="true">
                  <mml:mi>I</mml:mi>
                  <mml:mo>^</mml:mo>
                </mml:mover>
                <mml:mi>N</mml:mi>
              </mml:msub>
              <mml:mrow>
                <mml:mo>(</mml:mo>
                <mml:mi>x</mml:mi>
                <mml:mo>)</mml:mo>
              </mml:mrow>
              <mml:mo>=</mml:mo>
              <mml:mi>I</mml:mi>
              <mml:mrow>
                <mml:mo>(</mml:mo>
                <mml:mi>x</mml:mi>
                <mml:mo>)</mml:mo>
              </mml:mrow>
              <mml:mo>,</mml:mo>
            </mml:mrow>
          </mml:math>
        </disp-formula>
        <p>with convergence rate <inline-formula><mml:math><mml:mrow><mml:msub><mml:mi> O </mml:mi><mml:mi> p </mml:mi></mml:msub><mml:mrow><mml:mo> ( </mml:mo><mml:mrow><mml:mrow><mml:mn> 1 </mml:mn><mml:mo> / </mml:mo><mml:mrow><mml:msqrt><mml:mi> N </mml:mi></mml:msqrt></mml:mrow></mml:mrow></mml:mrow><mml:mo> ) </mml:mo></mml:mrow></mml:mrow></mml:math></inline-formula> .</p>
        <p><italic>Proof sketch.</italic> This convergence result is derived from a concentration argument using Bernstein inequalities, combined with the convexity of <inline-formula><mml:math><mml:mrow><mml:mi> λ </mml:mi><mml:mrow><mml:mo> ( </mml:mo><mml:mi> θ </mml:mi><mml:mo> ) </mml:mo></mml:mrow></mml:mrow></mml:math></inline-formula> . This result can be viewed as an application of the delta theorem within the framework of Cramér transform convexity. The key steps are:</p>
        <p>1) Estimation of the Cramér transform: <inline-formula><mml:math><mml:mrow><mml:msub><mml:mi> λ </mml:mi><mml:mi> N </mml:mi></mml:msub><mml:mrow><mml:mo> ( </mml:mo><mml:mi> θ </mml:mi><mml:mo> ) </mml:mo></mml:mrow><mml:mo> = </mml:mo><mml:mfrac><mml:mn> 1 </mml:mn><mml:mi> N </mml:mi></mml:mfrac><mml:mi> ln </mml:mi><mml:mi mathvariant="double-struck"> E </mml:mi><mml:mrow><mml:mo> [ </mml:mo><mml:mrow><mml:msup><mml:mtext> e </mml:mtext><mml:mrow><mml:mrow><mml:mo> 〈 </mml:mo><mml:mrow><mml:mi> θ </mml:mi><mml:mo> , </mml:mo><mml:msub><mml:mi> S </mml:mi><mml:mi> N </mml:mi></mml:msub></mml:mrow><mml:mo> 〉 </mml:mo></mml:mrow></mml:mrow></mml:msup></mml:mrow><mml:mo> ] </mml:mo></mml:mrow></mml:mrow></mml:math></inline-formula> .</p>
        <p>2) Verification of convergence: <inline-formula><mml:math><mml:mrow><mml:mi> λ </mml:mi><mml:mrow><mml:mo> ( </mml:mo><mml:mi> θ </mml:mi><mml:mo> ) </mml:mo></mml:mrow><mml:mo> = </mml:mo><mml:msub><mml:mrow><mml:mi> lim </mml:mi></mml:mrow><mml:mrow><mml:mi> N </mml:mi><mml:mo> → </mml:mo><mml:mi> ∞ </mml:mi></mml:mrow></mml:msub><mml:msub><mml:mi> λ </mml:mi><mml:mi> N </mml:mi></mml:msub><mml:mrow><mml:mo> ( </mml:mo><mml:mi> θ </mml:mi><mml:mo> ) </mml:mo></mml:mrow></mml:mrow></mml:math></inline-formula> .</p>
        <p>3) Application of the Legendre-Fenchel transform.</p>
        <p>4) Verification of convexity and differentiability conditions.</p>
        <p>For technical details, see [<xref ref-type="bibr" rid="B4">4</xref>]. □</p>
      </sec>
    </sec>
    <sec id="sec4">
      <title>4. Applications for Large Data</title>
      <p><bold>Remark 4.1</bold><bold>.</bold><italic>While the law of large numbers describes the typical behavior of systems in the presence of large amounts of data</italic>,<italic>large deviation theory quantifies deviations from this behavior.</italic><italic>It thus</italic><italic>allows describing the probability of unusual configurations in high-dimensional systems</italic>,<italic>making it a valuable tool for anomaly detection or extreme risk management. These approximations are particularly useful in contexts where accurate measurement of extreme observations is required.</italic></p>
      <sec id="sec4dot1">
        <title>4.1. Estimating Rare Probabilities</title>
        <p>For large data sets, the probability of rare events can be approximated by:</p>
        <disp-formula id="FD5">
          <mml:math>
            <mml:mrow>
              <mml:mi>ℙ</mml:mi>
              <mml:mrow>
                <mml:mo>(</mml:mo>
                <mml:mrow>
                  <mml:msub>
                    <mml:mi>S</mml:mi>
                    <mml:mi>N</mml:mi>
                  </mml:msub>
                  <mml:mo>∈</mml:mo>
                  <mml:mi>A</mml:mi>
                </mml:mrow>
                <mml:mo>)</mml:mo>
              </mml:mrow>
              <mml:mo>≈</mml:mo>
              <mml:mtext>exp</mml:mtext>
              <mml:mrow>
                <mml:mo>(</mml:mo>
                <mml:mrow>
                  <mml:mo>−</mml:mo>
                  <mml:mi>N</mml:mi>
                  <mml:munder>
                    <mml:mrow>
                      <mml:mtext>inf</mml:mtext>
                    </mml:mrow>
                    <mml:mrow>
                      <mml:mi>x</mml:mi>
                      <mml:mo>∈</mml:mo>
                      <mml:mi>A</mml:mi>
                    </mml:mrow>
                  </mml:munder>
                  <mml:mi>I</mml:mi>
                  <mml:mrow>
                    <mml:mo>(</mml:mo>
                    <mml:mi>x</mml:mi>
                    <mml:mo>)</mml:mo>
                  </mml:mrow>
                </mml:mrow>
                <mml:mo>)</mml:mo>
              </mml:mrow>
              <mml:mo>,</mml:mo>
            </mml:mrow>
          </mml:math>
        </disp-formula>
        <p>for sets <inline-formula><mml:math><mml:mi> A </mml:mi></mml:math></inline-formula> where <inline-formula><mml:math display="inline"><mml:mrow><mml:msub><mml:mrow><mml:mi> inf </mml:mi></mml:mrow><mml:mrow><mml:mi> x </mml:mi><mml:mo> ∈ </mml:mo><mml:msup><mml:mi> A </mml:mi><mml:mo> ∘ </mml:mo></mml:msup></mml:mrow></mml:msub><mml:mi> I </mml:mi><mml:mrow><mml:mo> ( </mml:mo><mml:mi> x </mml:mi><mml:mo> ) </mml:mo></mml:mrow><mml:mo> = </mml:mo><mml:msub><mml:mrow><mml:mi> inf </mml:mi></mml:mrow><mml:mrow><mml:mi> x </mml:mi><mml:mo> ∈ </mml:mo><mml:mover accent="true"><mml:mi> A </mml:mi><mml:mo> ¯ </mml:mo></mml:mover></mml:mrow></mml:msub><mml:mi> I </mml:mi><mml:mrow><mml:mo> ( </mml:mo><mml:mi> x </mml:mi><mml:mo> ) </mml:mo></mml:mrow></mml:mrow></mml:math></inline-formula> .</p>
      </sec>
      <sec id="sec4dot2">
        <title>4.2. Stochastic Optimization</title>
        <p>For optimization problems with rarity constraints:</p>
        <disp-formula id="FD6">
          <mml:math>
            <mml:mrow>
              <mml:munder>
                <mml:mrow>
                  <mml:mtext>min</mml:mtext>
                </mml:mrow>
                <mml:mrow>
                  <mml:mi>x</mml:mi>
                  <mml:mo>∈</mml:mo>
                  <mml:mi mathvariant="script">X</mml:mi>
                </mml:mrow>
              </mml:munder>
              <mml:mi>f</mml:mi>
              <mml:mrow>
                <mml:mo>(</mml:mo>
                <mml:mi>x</mml:mi>
                <mml:mo>)</mml:mo>
              </mml:mrow>
            </mml:mrow>
          </mml:math>
        </disp-formula>
        <disp-formula id="FD7">
          <mml:math>
            <mml:mrow>
              <mml:mtext>s</mml:mtext>
              <mml:mo>.</mml:mo>
              <mml:mtext>t</mml:mtext>
              <mml:mo>.</mml:mo>
              <mml:mtext>
                 
              </mml:mtext>
              <mml:mtext>
                 
              </mml:mtext>
              <mml:mi>ℙ</mml:mi>
              <mml:mrow>
                <mml:mo>(</mml:mo>
                <mml:mrow>
                  <mml:mi>g</mml:mi>
                  <mml:mrow>
                    <mml:mo>(</mml:mo>
                    <mml:mrow>
                      <mml:mi>x</mml:mi>
                      <mml:mo>,</mml:mo>
                      <mml:mi>ξ</mml:mi>
                    </mml:mrow>
                    <mml:mo>)</mml:mo>
                  </mml:mrow>
                  <mml:mo>≤</mml:mo>
                  <mml:mn>0</mml:mn>
                </mml:mrow>
                <mml:mo>)</mml:mo>
              </mml:mrow>
              <mml:mo>≥</mml:mo>
              <mml:mn>1</mml:mn>
              <mml:mo>−</mml:mo>
              <mml:mi>α</mml:mi>
              <mml:mo>,</mml:mo>
            </mml:mrow>
          </mml:math>
        </disp-formula>
        <p>large deviation theory allows approximating the constraint as follows [<xref ref-type="bibr" rid="B13">13</xref>]:</p>
        <disp-formula id="FD8">
          <mml:math>
            <mml:mrow>
              <mml:mi>ℙ</mml:mi>
              <mml:mrow>
                <mml:mo>(</mml:mo>
                <mml:mrow>
                  <mml:mi>g</mml:mi>
                  <mml:mrow>
                    <mml:mo>(</mml:mo>
                    <mml:mrow>
                      <mml:mi>x</mml:mi>
                      <mml:mo>,</mml:mo>
                      <mml:mi>ξ</mml:mi>
                    </mml:mrow>
                    <mml:mo>)</mml:mo>
                  </mml:mrow>
                  <mml:mo>≤</mml:mo>
                  <mml:mn>0</mml:mn>
                </mml:mrow>
                <mml:mo>)</mml:mo>
              </mml:mrow>
              <mml:mo>≈</mml:mo>
              <mml:mn>1</mml:mn>
              <mml:mo>−</mml:mo>
              <mml:mtext>exp</mml:mtext>
              <mml:mrow>
                <mml:mo>(</mml:mo>
                <mml:mrow>
                  <mml:mo>−</mml:mo>
                  <mml:mi>N</mml:mi>
                  <mml:mi>I</mml:mi>
                  <mml:mrow>
                    <mml:mo>(</mml:mo>
                    <mml:mi>x</mml:mi>
                    <mml:mo>)</mml:mo>
                  </mml:mrow>
                </mml:mrow>
                <mml:mo>)</mml:mo>
              </mml:mrow>
              <mml:mo>.</mml:mo>
            </mml:mrow>
          </mml:math>
        </disp-formula>
        <p>This approximation is accurate in the asymptotic limit <inline-formula><mml:math><mml:mrow><mml:mi> N </mml:mi><mml:mo> → </mml:mo><mml:mi> ∞ </mml:mi></mml:mrow></mml:math></inline-formula> , offering a practical constraint for reliability limits.</p>
        <p><bold>Remark 4.2</bold><bold>.</bold><italic>This approximation is valid for large</italic><inline-formula><mml:math><mml:mi> N </mml:mi></mml:math></inline-formula><italic>and regular functions</italic><inline-formula><mml:math><mml:mi> g </mml:mi></mml:math></inline-formula><italic>. For smaller sample sizes</italic>, <italic>finite-sample</italic><italic>corrections may be necessary. For example</italic>,<italic>approaches based on Edgeworth expansions or bootstrap methods can be employed to refine the estimate. The reader is referred to</italic>[<xref ref-type="bibr" rid="B14">14</xref>]<italic>for a discussion of finite-sample corrections in rare-event simulation. This method is particularly well-suited for machine learning problems where</italic><inline-formula><mml:math><mml:mi> N </mml:mi></mml:math></inline-formula><italic>is generally large.</italic></p>
      </sec>
    </sec>
    <sec id="sec5">
      <title>5. Numerical Algorithms</title>
      <p>We present two classical numerical methods tailored for rare-event quantification: saddle-point approximation and importance sampling simulation. Both methods have been implemented in Python; the following study demonstrates their numerical performance.</p>
      <sec id="sec5dot1">
        <title>5.1. Saddle-Point Method</title>
        <p>The saddle-point approximation, derived from Laplace’s theorem, facilitates the estimation of rare probabilities with significant asymptotic precision for large <inline-formula><mml:math><mml:mi> N </mml:mi></mml:math></inline-formula> . Under regularity conditions (convex function, valid Laplace asymptotic), we obtain:</p>
        <disp-formula id="FD9">
          <mml:math>
            <mml:mrow>
              <mml:mi>ℙ</mml:mi>
              <mml:mrow>
                <mml:mo>(</mml:mo>
                <mml:mrow>
                  <mml:msub>
                    <mml:mi>S</mml:mi>
                    <mml:mi>N</mml:mi>
                  </mml:msub>
                  <mml:mo>≥</mml:mo>
                  <mml:mi>a</mml:mi>
                </mml:mrow>
                <mml:mo>)</mml:mo>
              </mml:mrow>
              <mml:mo>~</mml:mo>
              <mml:mfrac>
                <mml:mn>1</mml:mn>
                <mml:mrow>
                  <mml:msup>
                    <mml:mi>θ</mml:mi>
                    <mml:mo>′</mml:mo>
                  </mml:msup>
                  <mml:msqrt>
                    <mml:mrow>
                      <mml:mn>2</mml:mn>
                      <mml:mi>π</mml:mi>
                      <mml:mi>N</mml:mi>
                      <mml:msup>
                        <mml:mi>λ</mml:mi>
                        <mml:mo>″</mml:mo>
                      </mml:msup>
                      <mml:mrow>
                        <mml:mo>(</mml:mo>
                        <mml:msup>
                          <mml:mi>θ</mml:mi>
                          <mml:mo>′</mml:mo>
                        </mml:msup>
                        <mml:mo>)</mml:mo>
                      </mml:mrow>
                    </mml:mrow>
                  </mml:msqrt>
                </mml:mrow>
              </mml:mfrac>
              <mml:mtext>exp</mml:mtext>
              <mml:mrow>
                <mml:mo>(</mml:mo>
                <mml:mrow>
                  <mml:mo>−</mml:mo>
                  <mml:mi>N</mml:mi>
                  <mml:mi>I</mml:mi>
                  <mml:mrow>
                    <mml:mo>(</mml:mo>
                    <mml:mi>a</mml:mi>
                    <mml:mo>)</mml:mo>
                  </mml:mrow>
                </mml:mrow>
                <mml:mo>)</mml:mo>
              </mml:mrow>
              <mml:mo>,</mml:mo>
            </mml:mrow>
          </mml:math>
        </disp-formula>
        <p>where <inline-formula><mml:math><mml:mrow><mml:msup><mml:mi> θ </mml:mi><mml:mtext> * </mml:mtext></mml:msup></mml:mrow></mml:math></inline-formula> is the solution of <inline-formula><mml:math><mml:mrow><mml:msup><mml:mi> λ </mml:mi><mml:mo> ′ </mml:mo></mml:msup><mml:mrow><mml:mo> ( </mml:mo><mml:mrow><mml:msup><mml:mi> θ </mml:mi><mml:mtext> * </mml:mtext></mml:msup></mml:mrow><mml:mo> ) </mml:mo></mml:mrow><mml:mo> = </mml:mo><mml:mi> a </mml:mi></mml:mrow></mml:math></inline-formula> . This approximation comes from applying the Laplace asymptotic expansion to the Cramér transform.</p>
        <p><bold>Example 5.1 (Application to the Gaussian case)</bold><italic>For</italic><inline-formula><mml:math><mml:mrow><mml:msub><mml:mi> X </mml:mi><mml:mi> i </mml:mi></mml:msub><mml:mo> ~ </mml:mo><mml:mi mathvariant="script"> N </mml:mi><mml:mrow><mml:mo> ( </mml:mo><mml:mrow><mml:mn> 0 </mml:mn><mml:mo> , </mml:mo><mml:mn> 1 </mml:mn></mml:mrow><mml:mo> ) </mml:mo></mml:mrow></mml:mrow></mml:math></inline-formula><italic>and</italic><inline-formula><mml:math><mml:mrow><mml:mi> a </mml:mi><mml:mo> = </mml:mo><mml:mn> 2 </mml:mn></mml:mrow></mml:math></inline-formula> ,<italic>we f</italic><italic>ind</italic><inline-formula><mml:math><mml:mrow><mml:msup><mml:mi> θ </mml:mi><mml:mo> * </mml:mo></mml:msup><mml:mo> = </mml:mo><mml:mn> 2 </mml:mn></mml:mrow></mml:math></inline-formula><italic>and</italic><inline-formula><mml:math><mml:mrow><mml:msup><mml:mi> λ </mml:mi><mml:mo> ″ </mml:mo></mml:msup><mml:mrow><mml:mo> ( </mml:mo><mml:msup><mml:mi> θ </mml:mi><mml:mo> ′ </mml:mo></mml:msup><mml:mo> ) </mml:mo></mml:mrow><mml:mo> = </mml:mo><mml:mn> 1 </mml:mn></mml:mrow></mml:math></inline-formula> ,<italic>hence:</italic></p>
        <disp-formula id="FD10">
          <mml:math>
            <mml:mrow>
              <mml:mi>ℙ</mml:mi>
              <mml:mrow>
                <mml:mo>(</mml:mo>
                <mml:mrow>
                  <mml:msub>
                    <mml:mi>S</mml:mi>
                    <mml:mi>N</mml:mi>
                  </mml:msub>
                  <mml:mo>≥</mml:mo>
                  <mml:mn>2</mml:mn>
                </mml:mrow>
                <mml:mo>)</mml:mo>
              </mml:mrow>
              <mml:mo>~</mml:mo>
              <mml:mfrac>
                <mml:mn>1</mml:mn>
                <mml:mrow>
                  <mml:mn>2</mml:mn>
                  <mml:msqrt>
                    <mml:mrow>
                      <mml:mn>2</mml:mn>
                      <mml:mi>π</mml:mi>
                      <mml:mi>N</mml:mi>
                    </mml:mrow>
                  </mml:msqrt>
                </mml:mrow>
              </mml:mfrac>
              <mml:mtext>exp</mml:mtext>
              <mml:mrow>
                <mml:mo>(</mml:mo>
                <mml:mrow>
                  <mml:mo>−</mml:mo>
                  <mml:mn>2</mml:mn>
                  <mml:mi>N</mml:mi>
                </mml:mrow>
                <mml:mo>)</mml:mo>
              </mml:mrow>
              <mml:mo>.</mml:mo>
            </mml:mrow>
          </mml:math>
        </disp-formula>
      </sec>
      <sec id="sec5dot2">
        <title>5.2. Importance Sampling (Monte Carlo)</title>
        <fig id="fig2">
          <label>Figure 2</label>
          <graphic xlink:href="https://html.scirp.org/file/1724455-rId116.jpeg?20260731025244" />
        </fig>
        <p><bold>Remark 5.2</bold><bold>.</bold><italic>The aim of importance sampling is to focus simulations on regions of the state space where rare events are most likely to occur</italic>,<italic>thereby reducing the variance of the estimator. The effectiveness</italic><italic>of importance</italic><italic>sampling critically depends on the choice of the optimal parameter</italic><inline-formula><mml:math><mml:mrow><mml:msup><mml:mi> θ </mml:mi><mml:mo> * </mml:mo></mml:msup></mml:mrow></mml:math></inline-formula> ,<italic>the solution to the dual problem</italic><inline-formula><mml:math><mml:mrow><mml:msup><mml:mi> λ </mml:mi><mml:mo> ′ </mml:mo></mml:msup><mml:mrow><mml:mo> ( </mml:mo><mml:mrow><mml:msup><mml:mi> θ </mml:mi><mml:mo> * </mml:mo></mml:msup></mml:mrow><mml:mo> ) </mml:mo></mml:mrow><mml:mo> = </mml:mo><mml:mi> a </mml:mi></mml:mrow></mml:math></inline-formula><italic>. For a comprehensive implementation of optimized importance sampling</italic>,<italic>the reader is referred to</italic>[<xref ref-type="bibr" rid="B14">14</xref>]<italic>. The algorithmic complexity grows linearly with</italic><inline-formula><mml:math><mml:mi> N </mml:mi></mml:math></inline-formula> ,<italic>the</italic><italic>most costly</italic><italic>steps being sample generation and weight computation.</italic></p>
      </sec>
    </sec>
    <sec id="sec6">
      <title>6. Numerical Results</title>
      <sec id="sec6dot1">
        <title>Case Study: Estimating the Rate Function</title>
        <p>We provide a numerical analysis for rate function estimation. Experiments were performed on a sample size of 10<sup>4</sup>, utilizing <inline-formula><mml:math><mml:mrow><mml:mi> N </mml:mi><mml:mo> = </mml:mo><mml:mn> 1000 </mml:mn></mml:mrow></mml:math></inline-formula> for probability estimations. <bold>App</bold><bold>endix A</bold> shows the full implementation.</p>
      </sec>
    </sec>
    <sec id="sec7">
      <title>7. Practical Applications</title>
      <p>The following applications illustrate the interdisciplinary breadth of large deviation theory.</p>
      <sec id="sec7dot1">
        <title>7.1. Quantitative Finance</title>
        <p><bold>Example 7.1 (High Credit Risk)</bold><italic>The probability of significant losses in a credit portfolio can be estimated by:</italic></p>
        <disp-formula id="FD11">
          <mml:math>
            <mml:mrow>
              <mml:mi>ℙ</mml:mi>
              <mml:mrow>
                <mml:mo>(</mml:mo>
                <mml:mrow>
                  <mml:mtext>Losses</mml:mtext>
                  <mml:mo>≥</mml:mo>
                  <mml:mi>u</mml:mi>
                </mml:mrow>
                <mml:mo>)</mml:mo>
              </mml:mrow>
              <mml:mo>≈</mml:mo>
              <mml:mtext>exp</mml:mtext>
              <mml:mrow>
                <mml:mo>(</mml:mo>
                <mml:mrow>
                  <mml:mo>−</mml:mo>
                  <mml:mi>N</mml:mi>
                  <mml:mi>I</mml:mi>
                  <mml:mrow>
                    <mml:mo>(</mml:mo>
                    <mml:mrow>
                      <mml:mfrac>
                        <mml:mi>u</mml:mi>
                        <mml:mi>N</mml:mi>
                      </mml:mfrac>
                    </mml:mrow>
                    <mml:mo>)</mml:mo>
                  </mml:mrow>
                </mml:mrow>
                <mml:mo>)</mml:mo>
              </mml:mrow>
              <mml:mo>.</mml:mo>
            </mml:mrow>
          </mml:math>
        </disp-formula>
        <p>This estimation is useful for calculating Value at Risk (VaR) and Expected Shortfall. For recent applications of large deviation theory in finance, see [<xref ref-type="bibr" rid="B9">9</xref>] and [<xref ref-type="bibr" rid="B10">10</xref>].</p>
      </sec>
      <sec id="sec7dot2">
        <title>7.2. Telecommunications Networks</title>
        <p><bold>Example 7.2 (Network Congestion)</bold><bold>.</bold><italic>Assuming data flows follow Poisson processes</italic>,<italic>the probability of exceeding a network’s capacity is given by:</italic></p>
        <disp-formula id="FD12">
          <mml:math>
            <mml:mrow>
              <mml:mi>ℙ</mml:mi>
              <mml:mrow>
                <mml:mo>(</mml:mo>
                <mml:mrow>
                  <mml:mi>Q</mml:mi>
                  <mml:mo>≥</mml:mo>
                  <mml:mi>B</mml:mi>
                </mml:mrow>
                <mml:mo>)</mml:mo>
              </mml:mrow>
              <mml:mo>≈</mml:mo>
              <mml:mtext>exp</mml:mtext>
              <mml:mrow>
                <mml:mo>(</mml:mo>
                <mml:mrow>
                  <mml:mo>−</mml:mo>
                  <mml:mi>δ</mml:mi>
                  <mml:mi>B</mml:mi>
                </mml:mrow>
                <mml:mo>)</mml:mo>
              </mml:mrow>
              <mml:mo>,</mml:mo>
            </mml:mrow>
          </mml:math>
        </disp-formula>
        <p>where <inline-formula><mml:math><mml:mi> δ </mml:mi></mml:math></inline-formula> is the deviation parameter from large deviation theory, indicating the exponential decay rate. For case studies in this area, the reader may consult [<xref ref-type="bibr" rid="B13">13</xref>] and [<xref ref-type="bibr" rid="B14">14</xref>].</p>
      </sec>
      <sec id="sec7dot3">
        <title>7.3. Machine Learning</title>
        <p><bold>Example 7.3 (High Generalization Error)</bold><bold>.</bold><italic>According to</italic>[<xref ref-type="bibr" rid="B8">8</xref>],<italic>the probability of an unusually high generalization error in a classification model can be assessed by:</italic></p>
        <disp-formula id="FD13">
          <mml:math>
            <mml:mrow>
              <mml:mi>ℙ</mml:mi>
              <mml:mrow>
                <mml:mo>(</mml:mo>
                <mml:mrow>
                  <mml:mtext>Error</mml:mtext>
                  <mml:mo>≥</mml:mo>
                  <mml:mi>ϵ</mml:mi>
                  <mml:mo>+</mml:mo>
                  <mml:mi>δ</mml:mi>
                </mml:mrow>
                <mml:mo>)</mml:mo>
              </mml:mrow>
              <mml:mo>≈</mml:mo>
              <mml:mtext>exp</mml:mtext>
              <mml:mrow>
                <mml:mo>(</mml:mo>
                <mml:mrow>
                  <mml:mo>−</mml:mo>
                  <mml:mi>n</mml:mi>
                  <mml:mi>I</mml:mi>
                  <mml:mrow>
                    <mml:mo>(</mml:mo>
                    <mml:mi>δ</mml:mi>
                    <mml:mo>)</mml:mo>
                  </mml:mrow>
                </mml:mrow>
                <mml:mo>)</mml:mo>
              </mml:mrow>
              <mml:mo>,</mml:mo>
            </mml:mrow>
          </mml:math>
        </disp-formula>
        <p>where <inline-formula><mml:math><mml:mi> ϵ </mml:mi></mml:math></inline-formula> is the typical error and <inline-formula><mml:math><mml:mrow><mml:mi> I </mml:mi><mml:mrow><mml:mo> ( </mml:mo><mml:mi> δ </mml:mi><mml:mo> ) </mml:mo></mml:mrow></mml:mrow></mml:math></inline-formula> quantifies the rarity of deviations. This method can be connected to Vapnik-Chervonenkis theory for examining PAC generalization limits. Recent works, such as those in [<xref ref-type="bibr" rid="B11">11</xref>] and [<xref ref-type="bibr" rid="B12">12</xref>], explore these connections.</p>
      </sec>
      <sec id="sec7dot4">
        <title>7.4. Other Potential Areas</title>
        <p>The presented methodology naturally extends to other domains:</p>
        <p><bold>Cybersecurity:</bold> Detection of rare but critical intrusions.<bold>Biostatistics:</bold> Occurrence of rare genetic mutations.<bold>Neural Networks:</bold> Analysis of very high activation states to study the stability and robustness of deep architectures subjected to extreme perturbations [<xref ref-type="bibr" rid="B11">11</xref>].<bold>Energy Systems:</bold> Management of very high demand peaks.</p>
        <p>These results demonstrate the adaptability of large deviation theory as a unified framework for risk analysis across various scientific and technological domains.</p>
      </sec>
    </sec>
    <sec id="sec8">
      <title>8. Conclusions and Perspectives</title>
      <sec id="sec8dot1">
        <title>8.1. Summary of Contributions</title>
        <p>The results obtained show the potential of large deviation theory to serve as a bridge between theoretical probability and computational data science. This work constitutes a step towards the complete integration of large deviation theory into massive data analysis methods. The main contributions are:</p>
        <p>The development of improved numerical algorithms for estimating the rate function, integrating convex optimization techniques for better stability.Application to various practical domains, with empirical validation.The combination of theoretical guarantees with efficient implementations.Demonstration of superior performance compared to traditional methods, particularly in terms of variance reduction and convergence speed.</p>
      </sec>
      <sec id="sec8dot2">
        <title>8.2. Research Perspectives</title>
        <p>This work opens several promising research avenues:</p>
        <p><bold>Extension to correlated data:</bold> Adapting techniques to systems with spatial or temporal dependencies.<bold>Machine learning:</bold> Using the framework to analyze the performance of deep learning models on new data.<bold>High-dimensional data:</bold> Creating efficient techniques for high-dimensional problems.<bold>Robust optimization:</bold> Combining with optimization under uncertainty methods.</p>
        <p>A significant research direction will involve applying these tools to real datasets to assess the numerical robustness and scalability of the proposed methodologies. These findings provide a solid basis for the advancement of hybrid probabilistic methodologies, integrating theoretical analysis with computational intelligence.</p>
      </sec>
    </sec>
    <sec id="sec9">
      <title>Availability of Data and Materials</title>
      <p>The Python code implementation used to generate the numerical results in this study is available in the <bold>Appendix</bold>. Synthetic datasets were generated for demonstration purposes using the numpy.random module as detailed in the implementation section. The code can be reproduced to verify all presented results.</p>
    </sec>
    <sec id="sec10">
      <title>Funding</title>
      <p>This research received no specific grant from any funding agency in the public, commercial, or not-for-profit sectors.</p>
    </sec>
    <sec id="sec11">
      <title>Author Contributions</title>
      <p>BOU DIOP is the sole contributor to this manuscript, responsible for the conceptualization, methodology, formal analysis, investigation, writing, implementation, and validation of all presented results.</p>
    </sec>
    <sec id="sec12">
      <title>Acknowledgements</title>
      <p>The author would like to thank the anonymous reviewers for their valuable comments and suggestions that helped improve this manuscript. Special thanks to the scientific computing community for maintaining the open-source tools (Python, NumPy, SciPy) that made this research possible.</p>
    </sec>
    <sec id="sec13">
      <title>Appendix</title>
      <p><bold>Code Implementation</bold></p>
      <fig id="fig3">
        <label>Figure 3</label>
        <graphic xlink:href="https://html.scirp.org/file/1724455-rId142.jpeg?20260731025246" />
      </fig>
    </sec>
  </body>
  <back>
    <ref-list>
      <title>References</title>
      <ref id="B1">
        <label>1.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Cramér, H. (1938) Sur un nouveau théorème-limite de la théorie des probabilités. <italic>Actualités Scientifiques et Industrielles</italic>, 736, 2-23.</mixed-citation>
          <element-citation publication-type="other">
            <year>1938</year>
            <article-title>Sur un nouveau théorème-limite de la théorie des probabilités</article-title>
            <source>Actualités Scientifiques et Industrielles</source>
            <volume>736</volume>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B2">
        <label>2.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Donsker, M.D. and Varadhan, S.R.S. (1975) Asymptotic Evaluation of Certain Markov Process Expectations for Large Time, I. <italic>Communications</italic><italic>on</italic><italic>Pure</italic><italic>and</italic><italic>Applied</italic><italic>Mathematics</italic>, 28, 1-47. https://doi.org/10.1002/cpa.3160280102 <pub-id pub-id-type="doi">10.1002/cpa.3160280102</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1002/cpa.3160280102">https://doi.org/10.1002/cpa.3160280102</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Donsker, M.D.</string-name>
              <string-name>Varadhan, S.R.S.</string-name>
              <string-name>Time, I.</string-name>
            </person-group>
            <year>1975</year>
            <article-title>Asymptotic Evaluation of Certain Markov Process Expectations for Large Time, I</article-title>
            <source>Communications on Pure and Applied Mathematics</source>
            <volume>28</volume>
            <pub-id pub-id-type="doi">10.1002/cpa.3160280102</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B3">
        <label>3.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Varadhan, S.R.S. (1984) Large Deviations and Applications. Society for Industrial and Applied Mathematics. https://doi.org/10.1137/1.9781611970241 <pub-id pub-id-type="doi">10.1137/1.9781611970241</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1137/1.9781611970241">https://doi.org/10.1137/1.9781611970241</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Varadhan, S.R.S.</string-name>
            </person-group>
            <year>1984</year>
            <article-title>Large Deviations and Applications</article-title>
            <pub-id pub-id-type="doi">10.1137/1.9781611970241</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B4">
        <label>4.</label>
        <citation-alternatives>
          <mixed-citation publication-type="book">Dembo, A. and Zeitouni, O. (1998) Large Deviations Techniques and Applications. 2nd Edition, Springer.</mixed-citation>
          <element-citation publication-type="book">
            <person-group person-group-type="author">
              <string-name>Dembo, A.</string-name>
              <string-name>Zeitouni, O.</string-name>
              <string-name>Edition, S</string-name>
            </person-group>
            <year>1998</year>
            <article-title>Large Deviations Techniques and Applications</article-title>
            <source>2nd Edition</source>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B5">
        <label>5.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Touchette, H. (2009) The Large Deviation Approach to Statistical Mechanics. <italic>Physics</italic><italic>Reports</italic>, 478, 1-69. https://doi.org/10.1016/j.physrep.2009.05.002 <pub-id pub-id-type="doi">10.1016/j.physrep.2009.05.002</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1016/j.physrep.2009.05.002">https://doi.org/10.1016/j.physrep.2009.05.002</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Touchette, H.</string-name>
            </person-group>
            <year>2009</year>
            <article-title>The Large Deviation Approach to Statistical Mechanics</article-title>
            <source>Physics Reports</source>
            <volume>478</volume>
            <pub-id pub-id-type="doi">10.1016/j.physrep.2009.05.002</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B6">
        <label>6.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Touchette, H. (2018) Introduction to Dynamical Large Deviations of Markov Processes. <italic>Physica A</italic>: <italic>Statistical Mechanics and Its Applications</italic>, 504, 5-19. https://doi.org/10.1016/j.physa.2017.10.046 <pub-id pub-id-type="doi">10.1016/j.physa.2017.10.046</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1016/j.physa.2017.10.046">https://doi.org/10.1016/j.physa.2017.10.046</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Touchette, H.</string-name>
            </person-group>
            <year>2018</year>
            <article-title>Introduction to Dynamical Large Deviations of Markov Processes</article-title>
            <source>Physica A: Statistical Mechanics and Its Applications</source>
            <volume>504</volume>
            <pub-id pub-id-type="doi">10.1016/j.physa.2017.10.046</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B7">
        <label>7.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Dupuis, P. and Ellis, R.S. (2011) A Weak Convergence Approach to the Theory of Large Deviations. Wiley.</mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Dupuis, P.</string-name>
              <string-name>Ellis, R.S.</string-name>
            </person-group>
            <year>2011</year>
            <article-title>A Weak Convergence Approach to the Theory of Large Deviations</article-title>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B8">
        <label>8.</label>
        <citation-alternatives>
          <mixed-citation publication-type="confproc">Bottou, L. and Bousquet, O. (2012) The Tradeoffs of Large-Scale Learning. <italic>Proceedings of the</italic>25 <italic>th International Conference on Neural Information Processing Systems</italic>, Siem Reap, 13-16 December 2018, 1617-1624.</mixed-citation>
          <element-citation publication-type="confproc">
            <person-group person-group-type="author">
              <string-name>Bottou, L.</string-name>
              <string-name>Bousquet, O.</string-name>
              <string-name>Systems, S</string-name>
            </person-group>
            <year>2012</year>
            <article-title>The Tradeoffs of Large-Scale Learning</article-title>
            <source>Proceedings of the 25th International Conference on Neural Information Processing Systems</source>
            <volume>13</volume>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B9">
        <label>9.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Glasserman, P. and He, P. (2018) Large Deviations in Machine Learning. <italic>Journal of Applied Probability</italic>, 55, 1099-1116.</mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Glasserman, P.</string-name>
              <string-name>He, P.</string-name>
            </person-group>
            <year>2018</year>
            <article-title>Large Deviations in Machine Learning</article-title>
            <source>Journal of Applied Probability</source>
            <volume>55</volume>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B10">
        <label>10.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Roberts, G.O. and Rosenthal, J.S. (2020) Rare Event Simulation in High-Dimensional Systems. <italic>Statistical Science</italic>, 35, 539-560.</mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Roberts, G.O.</string-name>
              <string-name>Rosenthal, J.S.</string-name>
            </person-group>
            <year>2020</year>
            <article-title>Rare Event Simulation in High-Dimensional Systems</article-title>
            <source>Statistical Science</source>
            <volume>35</volume>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B11">
        <label>11.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Chatterjee, S. (2023) Large Deviation Principles for Random Neural Networks. Journal <italic>of Theoretical Probability</italic>, 36, 789-823.</mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Chatterjee, S.</string-name>
            </person-group>
            <year>2023</year>
            <article-title>Large Deviation Principles for Random Neural Networks</article-title>
            <source>Journal of Theoretical Probability</source>
            <volume>36</volume>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B12">
        <label>12.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Chen, X. and Gao, J. (2024) Rare Event Analysis in Deep Generative Models via Large Deviations. <italic>Neural Computation</italic>, 36, 45-78.</mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Chen, X.</string-name>
              <string-name>Gao, J.</string-name>
            </person-group>
            <year>2024</year>
            <article-title>Rare Event Analysis in Deep Generative Models via Large Deviations</article-title>
            <source>Neural Computation</source>
            <volume>36</volume>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B13">
        <label>13.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Dupuis, P., Leder, K. and Wang, H. (2007) Importance Sampling for Sums of Random Variables with Regularly Varying Tails. <italic>ACM Transactions on Modeling and Computer Simul</italic><italic>ation</italic>, 17, 14-es. https://doi.org/10.1145/1243991.1243995 <pub-id pub-id-type="doi">10.1145/1243991.1243995</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1145/1243991.1243995">https://doi.org/10.1145/1243991.1243995</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Dupuis, P.</string-name>
              <string-name>Leder, K.</string-name>
              <string-name>Wang, H.</string-name>
            </person-group>
            <year>2007</year>
            <article-title>Importance Sampling for Sums of Random Variables with Regularly Varying Tails</article-title>
            <source>ACM Transactions on Modeling and Computer Simulation</source>
            <volume>17</volume>
            <pub-id pub-id-type="doi">10.1145/1243991.1243995</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B14">
        <label>14.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Bucklew, J. (2013) Introduction to Rare Event Simulation. Springer.</mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Bucklew, J.</string-name>
            </person-group>
            <year>2013</year>
            <article-title>Introduction to Rare Event Simulation</article-title>
          </element-citation>
        </citation-alternatives>
      </ref>
    </ref-list>
  </back>
</article>