<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.4 20241031//EN" "JATS-journalpublishing1-4.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article" dtd-version="1.4" xml:lang="en">
  <front>
    <journal-meta>
      <journal-id journal-id-type="publisher-id">jcc</journal-id>
      <journal-title-group>
        <journal-title>Journal of Computer and Communications</journal-title>
      </journal-title-group>
      <issn pub-type="epub">2327-5227</issn>
      <issn pub-type="ppub">2327-5219</issn>
      <publisher>
        <publisher-name>Scientific Research Publishing</publisher-name>
      </publisher>
    </journal-meta>
    <article-meta>
      <article-id pub-id-type="doi">10.4236/jcc.2026.145002</article-id>
      <article-id pub-id-type="publisher-id">jcc-151428</article-id>
      <article-categories>
        <subj-group>
          <subject>Article</subject>
        </subj-group>
        <subj-group>
          <subject>Computer Science</subject>
          <subject>Communications</subject>
        </subj-group>
      </article-categories>
      <title-group>
        <article-title>A Review of Data-Driven Prediction of Undesirable Events in Offshore Oil Wells —Based on the Public 3W Benchmark and Time-Series Deep Learning</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <name name-style="western">
            <surname>Liu</surname>
            <given-names>Yiguan</given-names>
          </name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <name name-style="western">
            <surname>Liu</surname>
            <given-names>Xinru</given-names>
          </name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <name name-style="western">
            <surname>Yao</surname>
            <given-names>Jingru</given-names>
          </name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <name name-style="western">
            <surname>Wang</surname>
            <given-names>Yufan</given-names>
          </name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <name name-style="western">
            <surname>Guo</surname>
            <given-names>Beining</given-names>
          </name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author" corresp="yes">
          <name name-style="western">
            <surname>Wang</surname>
            <given-names>Yicheng</given-names>
          </name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
      </contrib-group>
      <aff id="aff1"><label>1</label> College of Science, North China University of Science and Technology, Tangshan, China </aff>
      <author-notes>
        <fn fn-type="conflict" id="fn-conflict">
          <p>The authors declare no conflicts of interest regarding the publication of this paper.</p>
        </fn>
      </author-notes>
      <pub-date pub-type="epub">
        <day>14</day>
        <month>05</month>
        <year>2026</year>
      </pub-date>
      <pub-date pub-type="collection">
        <month>05</month>
        <year>2026</year>
      </pub-date>
      <volume>14</volume>
      <issue>05</issue>
      <fpage>17</fpage>
      <lpage>31</lpage>
      <history>
        <date date-type="received">
          <day>22</day>
          <month>04</month>
          <year>2026</year>
        </date>
        <date date-type="accepted">
          <day>22</day>
          <month>05</month>
          <year>2026</year>
        </date>
        <date date-type="published">
          <day>25</day>
          <month>05</month>
          <year>2026</year>
        </date>
      </history>
      <permissions>
        <copyright-statement>© 2026 by the authors and Scientific Research Publishing Inc.</copyright-statement>
        <copyright-year>2026</copyright-year>
        <license license-type="open-access">
          <license-p> This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license ( <ext-link ext-link-type="uri" xlink:href="https://creativecommons.org/licenses/by/4.0/">https://creativecommons.org/licenses/by/4.0/</ext-link> ). </license-p>
        </license>
      </permissions>
      <self-uri content-type="doi" xlink:href="https://doi.org/10.4236/jcc.2026.145002">https://doi.org/10.4236/jcc.2026.145002</self-uri>
      <abstract>
        <p>The production and operation of offshore oil wells present typical characteristics of strong coupling, high nonlinearity, obvious time-varying behavior, and high operational risks. The occurrence of adverse events is rarely triggered by the over-limit behavior of a single variable; rather, it usually stems from the gradual deviation of multivariate correlation patterns in downhole systems from normal operating conditions. Traditional anomaly recognition methods based on threshold rules and manual experience are widely used in engineering practice and have practical value. Nevertheless, under complex operating-condition switching, weak precursor anomalies, and cross-well operational discrepancies, these methods are often affected by high false-alarm rates, poor transferability, and delayed responses. With the rapid development of technology, data-driven methods have gradually evolved into a core technical solution for anomaly detection and adverse-event prediction in offshore oil wells. Effective anomaly identification and early prediction based on production monitoring data have become important research directions in intelligent oilfield construction. Targeting research progress in predictive analytics for the oil and gas industry and state-of-the-art multivariate time-series anomaly detection methods, this paper conducts a systematic review of adverse-event prediction in offshore oil wells. It summarizes existing studies based on the public 3W benchmark dataset and analyzes key data challenges, including structural missingness, cross-well heterogeneity, class imbalance, and weak precursor features. Furthermore, this paper presents a review-oriented conceptual design of a missing-aware multi-scale time-frequency graph. Transformer framework tailored to offshore oil well scenarios, which is intended to support anomaly detection, adverse-event recognition, and early prediction in future empirical validation. From a design-rationale perspective, this study compares the proposed framework with conventional methods and existing deep models, and discusses its potential advantages in four aspects: missing-aware input encoding, dual-domain time-frequency representation, physics-prior-initialized learnable graph structure, and multi-task learning.</p>
      </abstract>
      <kwd-group kwd-group-type="author-generated" xml:lang="en">
        <kwd>Offshore Oil Wells</kwd>
        <kwd>Adverse Event Prediction</kwd>
        <kwd>3W Dataset</kwd>
        <kwd>Deep Learning</kwd>
        <kwd>Transformer</kwd>
        <kwd>Data-Driven</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec1">
      <title>1. Introduction</title>
      <p>Offshore oil well production systems operate for long periods under the combined effects of high temperature, high pressure, long-distance transportation, and complex phase behavior. Significant coupling exists among downhole safety valves, choke valves, gas-lift systems, subsea manifolds, and surface processing facilities. When pressure fluctuation, flow instability, choke abnormality, severe blockage, or hydrate formation occurs in a local part of the system, it may not only reduce single-well productivity but also trigger severe production fluctuation, well shut-in, and even safety incidents. Because offshore operation and maintenance costs are extremely high, the production losses and maintenance costs caused by anomalies are often much higher than those of onshore wells. Therefore, the prediction of undesirable events in offshore oil wells has long been an important research topic.</p>
      <p>From the perspective of industrial digitalization, oil and gas production is gradually shifting from an operation and maintenance mode that relies on experience and fixed rules to a proactive management mode based on monitoring data, predictive analytics, and intelligent early warning. Azmi [<xref ref-type="bibr" rid="B1">1</xref>], D’Almeida [<xref ref-type="bibr" rid="B2">2</xref>], and Tariq [<xref ref-type="bibr" rid="B3">3</xref>] present similar observations. With the advancement of digitalization, predictive analytics, condition monitoring, and intelligent warning based on multi-source monitoring data are becoming important directions for the development of oil well adverse-event detection and digital transformation.</p>
      <p>The emergence of public data benchmarks has further promoted the development of data-driven research. The 3W Dataset released by Vargas provides a standardized benchmark for the study of rare undesirable events in offshore oil wells [<xref ref-type="bibr" rid="B4">4</xref>]. Subsequently, Marins [<xref ref-type="bibr" rid="B5">5</xref>], Machado [<xref ref-type="bibr" rid="B6">6</xref>], Fernandes [<xref ref-type="bibr" rid="B7">7</xref>], and Bayazitova [<xref ref-type="bibr" rid="B8">8</xref>] conducted studies on this benchmark from the perspectives of tree models, one-class classifiers, autoencoders, and deep recurrent networks, respectively, accelerating the transition of this field from experience-based diagnosis to data-driven analysis. Meanwhile, review studies in the oil and gas industry have also pointed out that predictive analytics, condition monitoring, and intelligent warning are becoming key supports for digital transformation [<xref ref-type="bibr" rid="B9">9</xref>].</p>
      <p>Based on these observations and the special data structure of the dataset, this paper conducts a technology-oriented review of existing studies and proposes a design hypothesis to be validated, namely SDG-Former. Compared with existing studies, its expected benefits come from missing-aware input, time-frequency dual-domain representation, physics-prior-initialized learnable graph structure, and a multi-task output mechanism [<xref ref-type="bibr" rid="B10">10</xref>].</p>
    </sec>
    <sec id="sec2">
      <title>2. Related Work</title>
      <sec id="sec2dot1">
        <title>2.1. Task Definitions under the 3W Dataset Label Structure</title>
        <p>Research on undesirable events in offshore oil wells is not a single anomaly detection problem, but involves three levels: anomaly detection, event recognition, and early prediction. This paper defines these tasks under the 3W label structure. The 3W data are organized as multivariate time-series instances. Version-1.0.0 of the 3W Dataset contains 1984 CSV instances and 8 process variables, and includes missing, frozen, and outlier characteristics from real industrial data.</p>
        <p>An undesirable event refers to an operating condition that causes an oil well to deviate from normal production and may lead to production loss, increased maintenance cost, or safety risk. In the 3W label system, normal operation is coded as 0, steady-state undesirable events are usually coded as 1 - 8, transition states are coded as 101 - 108, and unknown states are coded as NaN. Among them, 101 - 108 can be mapped to the transient precursors of the corresponding steady-state events 1 - 8 (<bold>Table 1</bold>).</p>
        <p>Anomaly detection corresponds to a binary classification task: given a time window, the model determines whether it deviates from the normal production pattern. A window containing only valid label 0 is labeled normal; a window containing steady-state event labels 1 - 8 or transition labels 101-108 is labeled anomalous. Windows containing a large proportion of NaN values are not used for supervised loss, or they participate only in unsupervised representation learning through masking.</p>
        <p>Event recognition corresponds to a multi-class classification task: under the condition that an anomaly has been detected, the model further determines which type of undesirable event is more likely. Specifically, steady-state labels are used as classes, while transition label 100+e is mapped to the same event e so that the transient segment and the steady-state segment remain semantically consistent.</p>
        <p><bold>Table 1</bold><bold>.</bold> Abnormal state types in the 3W dataset [<xref ref-type="bibr" rid="B4">4</xref>].</p>
        <table-wrap id="tbl1">
          <label>Table 1</label>
          <table>
            <tbody>
              <tr>
                <td>Code</td>
                <td>Event Type</td>
                <td>Number of Instances</td>
              </tr>
              <tr>
                <td>0</td>
                <td>Normal</td>
                <td>597</td>
              </tr>
              <tr>
                <td>1</td>
                <td>Abrupt increase of BSW</td>
                <td>129</td>
              </tr>
              <tr>
                <td>2</td>
                <td>Spurious closure of DHSV</td>
                <td>38</td>
              </tr>
              <tr>
                <td>3</td>
                <td>Severe slugging</td>
                <td>106</td>
              </tr>
              <tr>
                <td>4</td>
                <td>Flow instability</td>
                <td>344</td>
              </tr>
              <tr>
                <td>5</td>
                <td>Rapid productivity loss</td>
                <td>451</td>
              </tr>
              <tr>
                <td>6</td>
                <td>Rapid closure of PCK</td>
                <td>221</td>
              </tr>
              <tr>
                <td>7</td>
                <td>Scaling in PCK</td>
                <td>14</td>
              </tr>
              <tr>
                <td>8</td>
                <td>Hydrate in production line</td>
                <td>84</td>
              </tr>
            </tbody>
          </table>
        </table-wrap>
        <p>Early prediction corresponds to an early-warning task: before the steady-state undesirable event begins, the model predicts event risk and type based on multivariate changes in the late normal stage or in the transition segment. Its core indicators are not ordinary window accuracy, but the lead time of the first valid alarm relative to the onset of the steady-state event, the early-detection rate, and the false-alarm rate.</p>
      </sec>
      <sec id="sec2dot2">
        <title>2.2. Traditional Machine Learning Methods</title>
        <p>From the perspective of traditional machine learning, the survey by Chandola <italic>et al.</italic> [<xref ref-type="bibr" rid="B11">11</xref>] and the review of time-series anomaly detection by Blázquez-García <italic>et al.</italic> [<xref ref-type="bibr" rid="B12">12</xref>] indicate that anomaly detection methods can be divided into four categories: statistical-boundary-based methods, density-estimation-based methods, distance-measurement-based methods, and normal-domain-learning methods. In offshore oil well scenarios, these methods have long occupied an important position. They require relatively low sample sizes and annotation quality, and they are also easier to transform into prediction rules in industrial systems.</p>
        <p>The Local Outlier Factor proposed by Breunig <italic>et al.</italic> [<xref ref-type="bibr" rid="B13">13</xref>] identifies outliers based on local density differences. The Isolation Forest proposed by Liu <italic>et al.</italic> [<xref ref-type="bibr" rid="B14">14</xref>] amplifies the separability of abnormal points through random partitioning, whereas the one-class support vector machine proposed by Schölkopf <italic>et al.</italic> [<xref ref-type="bibr" rid="B15">15</xref>] realizes novelty detection by learning the support domain of normal samples. These models jointly form the baseline system for early 3W anomaly detection and also explain why Fernandes. [<xref ref-type="bibr" rid="B7">7</xref>] found that LOF and OCSVM remain competitive in multiple groups of experiments. However, the review of condition monitoring and prognostics by Jardine <italic>et al.</italic> [<xref ref-type="bibr" rid="B16">16</xref>] has pointed out that reliable diagnosis of complex industrial systems ultimately requires stronger time-series representation and higher-level anomaly detection.</p>
      </sec>
      <sec id="sec2dot3">
        <title>2.3. Deep Learning Methods</title>
        <p>Deep learning has promoted the transformation of industrial time-series anomaly detection from shallow discriminators to end-to-end representation learning. The review of time-series classification by Fawaz <italic>et al.</italic> [<xref ref-type="bibr" rid="B17">17</xref>] indicates that convolutional networks are better at extracting local patterns, recurrent networks are more suitable for modeling sequential dependence, and hybrid networks often show stronger adaptability to complex industrial data. Malhotra. [<xref ref-type="bibr" rid="B18">18</xref>] and Hundman. [<xref ref-type="bibr" rid="B19">19</xref>] demonstrate, respectively, from general industrial sequences and spacecraft telemetry scenarios, that LSTM-based prediction error or reconstruction error can effectively characterize abnormal behavior.</p>
        <p>In unsupervised and semi-supervised directions, OmniAnomaly proposed by Su <italic>et al.</italic> [<xref ref-type="bibr" rid="B20">20</xref>], MAD-GAN proposed by Li <italic>et al.</italic> [<xref ref-type="bibr" rid="B21">21</xref>], DAGMM proposed by Zong <italic>et al.</italic> [<xref ref-type="bibr" rid="B22">22</xref>], and USAD proposed by Audibert <italic>et al.</italic> [<xref ref-type="bibr" rid="B23">23</xref>] have become representative methods for industrial multivariate anomaly detection. These methods improve the modeling capability of complex normal patterns from the perspectives of stochastic recurrent modeling, generative adversarial learning, joint density estimation, and dual-autoencoder adversarial training, respectively. However, they also face problems such as training instability and sensitivity to missing values.</p>
        <p>In recent years, Transformers and graph neural networks have further expanded the modeling boundary of anomaly detection. The Transformer proposed by Vaswani <italic>et al.</italic> [<xref ref-type="bibr" rid="B24">24</xref>] provides a general attention framework for long-dependency modeling. Informer proposed by Zhou <italic>et al.</italic> [<xref ref-type="bibr" rid="B25">25</xref>] reduces the computational overhead of long-sequence processing through sparse attention. Anomaly Transformer proposed by Xu <italic>et al.</italic> [<xref ref-type="bibr" rid="B26">26</xref>] uses association discrepancy to characterize anomalous points. TranAD proposed by Tuli <italic>et al.</italic> [<xref ref-type="bibr" rid="B27">27</xref>] enhances the efficiency and stability of multivariate anomaly diagnosis. DCdetector proposed by Yang <italic>et al.</italic> [<xref ref-type="bibr" rid="B28">28</xref>] strengthens anomaly representation from the perspective of contrastive learning and the modeling of interactions among variables. Compared with models that encode only along the time dimension, such methods are more suitable for industrial scenarios because sensors often correspond to clear physical coupling relationships.</p>
        <p>For offshore oil wells, this means that future high-performance models are unlikely to arise from the extreme deepening of a single paradigm. Instead, they are more likely to come from systematic designs that integrate missing awareness, dual-domain representation, graph-temporal fusion, and multi-task prediction.</p>
      </sec>
    </sec>
    <sec id="sec3">
      <title>3. SDG-Former Conceptual Framework</title>
      <sec id="sec3dot1">
        <title>3.1. Design Objectives and Overall Idea</title>
        <p>This paper proposes SDG-Former (Missing-aware Multi-scale Time-frequency Graph Transformer) as a conceptual framework for predicting undesirable events in offshore oil wells. Its objective is to translate the data challenges of the 3W Dataset into verifiable model design hypotheses: the coexistence of long-term dependence and local abrupt changes, the coexistence of variable coupling and structural missingness, the coexistence of anomaly scarcity and event diversity, and the coexistence of engineering deployment requirements and deep-model complexity.</p>
        <p>SDG-Former is constructed around the 3W Dataset and offshore oil well scenarios with four core modules: first, missing-aware input encoding based on sliding windows; second, dual-domain representation that preserves both time-domain evolution and frequency-domain disturbances; third, a graph-temporal Transformer backbone for learning variable associations and cross-scale dependencies; and fourth, multi-task output heads for anomaly detection, event recognition, and lead-time evaluation.</p>
      </sec>
      <sec id="sec3dot2">
        <title>3.2. Missing-Aware Input</title>
        <p>For the problem of inconsistent sensor availability in offshore oil wells, this paper does not simply impute all missing values as a single value. Instead, it constructs a value matrix X, a missing mask, and a time-interval input Delta X for each time step. The purpose is to enable the model to distinguish between a genuinely low observation and the temporary unavailability of a variable, thereby avoiding the mislearning of missingness patterns as stable discriminative features.</p>
        <p>Considering further that the available sensor sets may differ significantly across wells and files, this paper first generates structural missingness based on the available sensor set, and then uses a lightweight router to determine which parameter subspace the sample should enter. Compared with training a complete model separately for each missingness pattern, this shared-backbone and missingness-adaptive strategy preserves structural differences, avoids the unlimited expansion of the model pool, and is more suitable for the large sample size and significant heterogeneity of the 3W data.</p>
      </sec>
      <sec id="sec3dot3">
        <title>3.3. Time-Frequency Dual-Domain Representation</title>
        <p>If modeling is performed only in the time domain, short-term oscillations, phase shifts, and multi-period disturbances may easily be compressed into superficially smooth trends. If modeling is performed only in the frequency domain, the order and duration structure of anomaly development may be ignored. Therefore, this paper adopts a dual-domain representation strategy. On the time-domain side, the window sequence is divided into several temporal patches to reduce the computational overhead of long-sequence attention. On the frequency-domain side, short-time Fourier transform (STFT) and three-band statistics are applied to each variable window to explicitly retain frequency clues related to typical anomalies such as mechanical oscillation, choke fluctuation, and flow instability.</p>
        <p>Temporal patches emphasize the process of anomaly development, whereas frequency-band statistics emphasize the form of abnormal vibration. These two representations are highly complementary in industrial scenarios. In offshore oil wells, for example, severe slugging and flow instability often show stronger periodicity in the frequency domain, whereas choke restriction and rapid productivity decline rely more heavily on time-domain trends and variable linkage relationships. Therefore, dual-domain representation is closer to the detection logic of practical engineering than single-domain modeling. Nevertheless, dual-domain representation remains a design hypothesis to be validated by experiments, rather than a conclusion empirically verified in this paper.</p>
      </sec>
      <sec id="sec3dot4">
        <title>3.4. Graph-Temporal Transformer</title>
        <p>At the backbone level, this paper adopts a three-stage structure consisting of graph structure learning, temporal self-attention, and cross-domain fusion. A learnable adjacency matrix and a graph initialized by physical priors are used to model the relationships among downhole pressure, upstream and downstream choke pressure, gas-lift flow rate, and related temperature variables. The graph structure is not directly fixed; instead, it is allowed to undergo lightweight adjustment during training so as to adapt to dynamic associations across different wells and operating conditions.</p>
        <p>In the temporal dimension, a hybrid encoder combining sparse self-attention and local convolution is introduced. Sparse self-attention follows the efficient design of Informer [<xref ref-type="bibr" rid="B25">25</xref>] for long-sequence processing, while local convolution is used to preserve short-term abrupt changes and edge information. This avoids the over-smoothing of pure Transformer models in industrial short-window tasks and mitigates the insufficient long-distance dependency capture of pure convolutional models.</p>
        <p>Finally, two gating units, namely sensor attention and frequency-band attention, are designed in the fusion layer. The former dynamically measures which variables are more important for the current judgment, while the latter distinguishes the frequency region in which the current anomaly is mainly expressed. Compared with Anomaly Transformer [<xref ref-type="bibr" rid="B26">26</xref>], which focuses on locating anomalies by association discrepancy, this design embeds interpretability into the model structure, thereby providing variable-level and frequency-band-level auxiliary evidence for subsequent engineering analysis.</p>
      </sec>
      <sec id="sec3dot5">
        <title>3.5. Design Choices and Rationale</title>
        <p>To demonstrate that the core modules of SDG-Former are not simply superimposed, this paper further summarizes the basis for key design choices and the hypotheses to be verified from three perspectives: data characteristics, model structure, and task objectives. Specifically, the window length mainly serves the needs of transition process coverage and real-time early warning; the time-frequency dual-domain representation is used to simultaneously characterize trend changes and oscillation disturbances; the learnable adjacency structure is used to express the potential correlations between multivariable sensors; and the multi-task loss is used to unify the objectives of anomaly detection, event recognition, and early prediction(<bold>Table 2</bold>). The relevant design choices, their basis, and the hypotheses to be verified are shown in <bold>Table 2</bold>.</p>
        <p><bold>Table 2</bold><bold>.</bold> Key SDG-former design choices, rationale, and hypotheses to be validated.</p>
        <table-wrap id="tbl2">
          <label>Table 2</label>
          <table>
            <tbody>
              <tr>
                <td>
                  <bold>Key Model Design</bold>
                </td>
                <td>
                  <bold>Design Rationale</bold>
                </td>
                <td>
                  <bold>Hypothesis to Be Validated</bold>
                </td>
              </tr>
              <tr>
                <td>Window length of 120 - 192 sampling points</td>
                <td>
                  With a recording frequency of 1 Hz, this window corresponds to 2 - 3.2 minutes of operating time. This range balances transition-process coverage and computational cost [
                  <xref ref-type="bibr" rid="B4">4</xref>
                  ].
                </td>
                <td>Compared with overly short windows, it may better capture slow precursors; compared with overly long windows, it can reduce label mixing and real-time latency.</td>
              </tr>
              <tr>
                <td>STFT + three-band statistics</td>
                <td>
                  Severe slugging and flow instability have oscillatory or periodic characteristics, and STFT is suitable for representing local frequency changes in non-stationary windows [
                  <xref ref-type="bibr" rid="B17">17</xref>
                  ].
                </td>
                <td>Time-domain trends and frequency-domain disturbances are complementary and may enhance the expression of weak precursors.</td>
              </tr>
              <tr>
                <td>Learnable adjacency initialized by physical priors</td>
                <td>
                  Studies such as GDN show that learning sensor graph structures helps capture variable correlations and supports anomaly interpretation [
                  <xref ref-type="bibr" rid="B25">25</xref>
                  ].
                </td>
                <td>Physical priors may reduce the risk of arbitrary connections under small samples, while learnable edge weights can adapt to cross-well differences.</td>
              </tr>
            </tbody>
          </table>
        </table-wrap>
        <p><bold>Continued</bold></p>
        <table-wrap id="tbl3">
          <label>Table 3</label>
          <table>
            <tbody>
              <tr>
                <td>Multi-task loss</td>
                <td>
                  The 3W Dataset provides normal, transition, and steady-state event labels, and multi-task learning can use related tasks to share representations [
                  <xref ref-type="bibr" rid="B4">4</xref>
                  ].
                </td>
                <td>Jointly optimizing detection, recognition, and early prediction may improve task consistency, but loss weights must be adjusted to avoid objective conflicts.</td>
              </tr>
            </tbody>
          </table>
        </table-wrap>
      </sec>
    </sec>
    <sec id="sec4">
      <title>4. W Benchmark Validation Protocol and Implementation Recommendations</title>
      <sec id="sec4dot1">
        <title>4.1. Introduction to the 3W Dataset</title>
        <p>This paper uses the 3W Dataset [<xref ref-type="bibr" rid="B4">4</xref>] as the core benchmark. The dataset contains undesirable-event instances from naturally flowing offshore oil wells. It is multivariate, time-series-based, rare-anomaly-oriented, and clearly grounded in engineering scenarios. Therefore, it is suitable for anomaly detection, event classification, and early-warning research (<bold>Table 2</bold>). Existing studies based on the 3W Dataset indicate that its research value lies not only in the inclusion of multiple anomaly classes but also in its data characteristics: coexistence of missing and frozen values, significant class imbalance, and large differences in duration and dynamic characteristics among different instances (<bold>Table 3</bold>).</p>
        <p><bold>Table 3</bold><bold>.</bold> Core features of the 3W dataset and their modeling implications.</p>
        <table-wrap id="tbl4">
          <label>Table 4</label>
          <table>
            <tbody>
              <tr>
                <td>
                  <bold>Dimension</bold>
                </td>
                <td>
                  <bold>Main Content</bold>
                </td>
                <td>
                  <bold>Modeling Implication</bold>
                </td>
              </tr>
              <tr>
                <td>Monitoring object</td>
                <td>
                  Multivariate monitoring sequences from naturally flowing offshore oil wells [
                  <xref ref-type="bibr" rid="B4">4</xref>
                  ]
                </td>
                <td>Suitable for studying anomaly detection and early warning in real production scenarios</td>
              </tr>
              <tr>
                <td>Variable structure</td>
                <td>Eight key variables, including pressure, temperature, and flow rate</td>
                <td>Variable coupling and asynchronous responses need to be explicitly modeled</td>
              </tr>
              <tr>
                <td>Label form</td>
                <td>Normal, transition, and anomaly-stage information</td>
                <td>Supports detection, classification, and lead-time evaluation</td>
              </tr>
              <tr>
                <td>Data quality</td>
                <td>Missing values, frozen values, noise, and source heterogeneity</td>
                <td>Requires models with missingness robustness and generalization capability</td>
              </tr>
              <tr>
                <td>Sample distribution</td>
                <td>Rare anomalies and imbalanced categories</td>
                <td>Requires reweighting and robust evaluation strategies during training</td>
              </tr>
            </tbody>
          </table>
        </table-wrap>
        <p>Data processing is performed by sliding-window segmentation. For missing values, neighboring-time imputation is used when the missing proportion is low; if missingness shows a structural distribution, the missing mask is preferentially retained. The sliding-window length is set to 120 - 192 sampling points according to anomaly duration, and the step size is set to half of the window length in order to balance the number of samples and temporal resolution.</p>
        <p>The main experiment uses instances with all labels 0 - 8. Ablation experiments may separately report results using only real instances and using real, simulated, and hand-drawn instances together (<bold>Table 3</bold>). All sliding windows must inherit their source file ID and event ID. Training, validation, and test splits should use file ID as the minimum grouping unit. If well IDs can be parsed from file names, well-level grouping should be preferred; otherwise, file-level GroupKFold should be adopted. Five-fold GroupKFold is recommended. All overlapping windows generated from the same file must appear in the same fold, and random splitting by window is prohibited.</p>
      </sec>
      <sec id="sec4dot2">
        <title>4.2. Data Processing and Normalization Strategies</title>
        <p>Missing-value processing follows the rule of “short-gap imputation, long-gap masking, and structural-missingness preservation.” When a continuous missing segment lasts no more than 10 seconds, linear interpolation or forward/backward neighboring imputation within the same file is used. When a continuous missing segment lasts more than 10 seconds but no more than 60 seconds, the imputed value is used only as a numerical placeholder, while the missing mask M = 0 and the time interval are retained. When a continuous missing segment lasts more than 60 seconds, or when the missing proportion of a variable in a file exceeds 30%, it is no longer treated as an ordinary gap but as a structurally unavailable variable. In this case, only zero-filled numerical placeholders, the missing mask, and the time interval are used, and the value is not regarded as a real observation.</p>
        <p>Frozen values need to be treated separately because freezing is not ordinary missingness but information distortion caused by a sensor holding a constant value for a period of time. It is recommended to detect consecutive repeated values for each variable: if a variable remains exactly the same for no fewer than 60 consecutive sampling points, or if the absolute value of the first-order difference after normalization remains continuously lower than 10<sup>−4</sup>, the interval is marked as a frozen segment. Frozen-segment values are not directly deleted. Instead, a frozen mask F is retained in the input, and the segment is down-weighted or excluded when calculating statistical features, STFT, and supervised loss. Normalization parameters must be fitted only on the training set within each cross-validation fold to avoid test leakage. For continuous variables such as pressure, temperature, and flow rate, robust standardization based on normal windows in the training set is recommended:</p>
        <disp-formula id="FD1">
          <mml:math display="inline">
            <mml:mrow>
              <mml:msup>
                <mml:mi>x</mml:mi>
                <mml:mo>′</mml:mo>
              </mml:msup>
              <mml:mo>=</mml:mo>
              <mml:mfrac>
                <mml:mrow>
                  <mml:mi>x</mml:mi>
                  <mml:mo>−</mml:mo>
                  <mml:msub>
                    <mml:mrow>
                      <mml:mtext>median</mml:mtext>
                    </mml:mrow>
                    <mml:mrow>
                      <mml:mtext>train</mml:mtext>
                    </mml:mrow>
                  </mml:msub>
                </mml:mrow>
                <mml:mrow>
                  <mml:msub>
                    <mml:mrow>
                      <mml:mtext>IQR</mml:mtext>
                    </mml:mrow>
                    <mml:mrow>
                      <mml:mtext>train</mml:mtext>
                    </mml:mrow>
                  </mml:msub>
                  <mml:mo>+</mml:mo>
                  <mml:mi>δ</mml:mi>
                </mml:mrow>
              </mml:mfrac>
            </mml:mrow>
          </mml:math>
        </disp-formula>
        <p>where <inline-formula><mml:math display="inline"><mml:mrow><mml:msub><mml:mrow><mml:mtext> median </mml:mtext></mml:mrow><mml:mrow><mml:mtext> train </mml:mtext></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math display="inline"><mml:mrow><mml:msub><mml:mrow><mml:mtext> IQR </mml:mtext></mml:mrow><mml:mrow><mml:mtext> train </mml:mtext></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula> are estimated only from normal windows in the training set, and epsilon is used to avoid division by zero. For variables with natural ranges, such as valve opening, min-max scaling based on the training-set range or physical range can be used. For binary or multi-valued valve states, the original discrete meaning is retained and a state mask is added. The validation and test sets only apply the parameters fitted on the training set and do not re-estimate them.</p>
      </sec>
      <sec id="sec4dot3">
        <title>4.3. Baseline Model and Evaluation Metrics Setting Recommendations</title>
        <p>Comparative models cover three categories. The first category consists of traditional machine learning baselines, including LOF [<xref ref-type="bibr" rid="B13">13</xref>], Isolation Forest [<xref ref-type="bibr" rid="B14">14</xref>], OCSVM [<xref ref-type="bibr" rid="B15">15</xref>], and Random Forest [<xref ref-type="bibr" rid="B5">5</xref>]. The second category consists of common deep learning baselines for industrial anomaly detection, including OmniAnomaly [<xref ref-type="bibr" rid="B20">20</xref>], MAD-GAN [<xref ref-type="bibr" rid="B21">21</xref>], DAGMM [<xref ref-type="bibr" rid="B22">22</xref>], USAD [<xref ref-type="bibr" rid="B23">23</xref>], and LSTM autoencoders [<xref ref-type="bibr" rid="B18">18</xref>][<xref ref-type="bibr" rid="B19">19</xref>]. The third category consists of recent graph-temporal and Transformer baselines, including Anomaly Transformer [<xref ref-type="bibr" rid="B26">26</xref>], TranAD [<xref ref-type="bibr" rid="B27">27</xref>] DCdetector [<xref ref-type="bibr" rid="B28">28</xref>], and recent improved graph-temporal models [<xref ref-type="bibr" rid="B29">29</xref>].</p>
        <p>The evaluation metrics include Accuracy, Recall, F1, ROC-AUC, and PR-AUC. For rare-anomaly scenarios, PR-AUC often better reflects the model’s recognition quality for minority-class windows than ROC-AUC. For engineering early-warning tasks, the average early-detection ratio and the first detection time of the anomaly should also be added.</p>
        <p>In terms of experimental settings and evaluation, the training stage adopts the AdamW optimizer, cosine-annealing learning rate, class reweighting, and early stopping. During inference, anomaly scores, variable attention, and frequency-band attention are retained for subsequent manual review and engineering interpretation. For edge-deployment scenarios, lower latency can be obtained by reducing the number of Transformer layers, shortening the patch length, and pruning the frequency-band branch.</p>
      </sec>
    </sec>
    <sec id="sec5">
      <title>5. Comparative Discussion Based on Existing Studies</title>
      <p>Existing studies show that Random Forest methods have verified the feasibility of traditional supervised classification for clear fault patterns in early 3W tasks. Comparisons of one-class classifiers indicate that LOF and OCSVM remain competitive when normal samples are sufficient and anomalous samples are limited. Studies on deep recurrent networks and open-world learning [<xref ref-type="bibr" rid="B10">10</xref>] further show that deep models and open-set strategies can provide additional benefits under more complex anomaly-class conditions. The evaluation should not focus only on small improvements in a single metric; instead, it should emphasize stability under complex combinations, rare classes, and cross-well generalization. The following content compares only the method routes in existing literature with the design hypotheses of SDG-Former.</p>
      <sec id="sec5dot1">
        <title>5.1. Discussion Compared with Traditional Methods</title>
        <p>Traditional methods usually perform well in two types of scenarios: cases where anomaly boundaries are clear and statistical features are obvious, and cases where sample sizes are limited and deep models are prone to overfitting. Random Forest [<xref ref-type="bibr" rid="B5">5</xref>], LOF [<xref ref-type="bibr" rid="B13">13</xref>], and Isolation Forest [<xref ref-type="bibr" rid="B14">14</xref>] have clear advantages in this regard because they have low training costs, are friendly to small samples, and are less sensitive to hyperparameters and optimization procedures.</p>
        <p>However, for adverse-event prediction in offshore oil wells, the limitations of traditional methods are also clear. They have difficulty directly using the continuously evolving temporal structure in transition stages and often rely on window-level summary statistics. Their ability to express variable coupling is limited, and they may simplify multivariate linkage anomalies into several univariate shifts. When structural missingness and cross-well operating-condition changes exist, shallow models are more likely to mistake distribution habits in the training set for generalizable rules.</p>
        <p>In comparison, the SDG-Former proposed in this paper has the following expected design advantages: it explicitly exposes weak precursors from both the time and frequency domains through dual-domain representation, thereby reducing reliance on manually designed statistics; it learns dynamic relationships among variables through graph-temporal fusion rather than treating each sensor as an independent channel; and it incorporates structural missingness into the model design through missingness routing instead of passively masking it during preprocessing.</p>
      </sec>
      <sec id="sec5dot2">
        <title>5.2. Discussion Compared with Deep Learning Methods</title>
        <p>Compared with deep learning models, this paper does not regard the reconstruction of normal patterns as the only objective. OmniAnomaly [<xref ref-type="bibr" rid="B20">20</xref>], MAD-GAN [<xref ref-type="bibr" rid="B21">21</xref>], DAGMM [<xref ref-type="bibr" rid="B22">22</xref>], and USAD [<xref ref-type="bibr" rid="B23">23</xref>] are representative models in industrial multivariate anomaly detection. However, such models usually rely on reconstruction error or energy functions to characterize anomalies. Therefore, when normal operating conditions themselves fluctuate substantially and normal patterns differ significantly across wells, threshold transfer becomes difficult.</p>
        <p>Compared with graph-temporal and Transformer models, Anomaly Transformer [<xref ref-type="bibr" rid="B26">26</xref>], TranAD [<xref ref-type="bibr" rid="B27">27</xref>], DCdetector [<xref ref-type="bibr" rid="B28">28</xref>], and recent improved graph-temporal models have significantly improved multivariate anomaly detection performance. Nevertheless, most of them still assume fixed input dimensions and give insufficient consideration to changes in sensor availability. In addition, many models focus on detecting anomalous points, while unified modeling of event-type recognition and early-warning lead time remains insufficient.</p>
      </sec>
      <sec id="sec5dot3">
        <title>5.3. Discussion Compared with Deep Learning Methods</title>
        <p>For offshore oil well scenarios, the most interpretable result is often not a small improvement in a single AUC value, but whether the model can remain stable in cross-well testing, provide sufficiently early prediction before steady-state anomalies occur, and output key variables that are consistent with engineering knowledge (<bold>Table 4</bold>).</p>
        <p>However, the proposed method also has limitations. Dual-domain representation and graph-temporal fusion increase implementation complexity and impose higher requirements on data preprocessing protocols. If the training data cover too few well types and operating conditions, the model may still be affected by distribution shift. Although multi-task learning improves task completeness, it may also introduce objective conflicts, requiring careful adjustment of loss weights and training strategies.</p>
        <p><bold>Table 4</bold><bold>.</bold>Comparison of existing model routes and SDG-former design hypotheses.</p>
        <table-wrap id="tbl5">
          <label>Table 5</label>
          <table>
            <tbody>
              <tr>
                <td>
                  <bold>Model Route</bold>
                </td>
                <td>
                  <bold>Representative Methods</bold>
                </td>
                <td>
                  <bold>Main Advantages</bold>
                </td>
                <td>
                  <bold>Main Limitations</bold>
                </td>
                <td>
                  <bold>Improvements in This Paper</bold>
                </td>
              </tr>
              <tr>
                <td>Shallow anomaly detection</td>
                <td>
                  LOF [
                  <xref ref-type="bibr" rid="B13">13</xref>
                  ], IF [
                  <xref ref-type="bibr" rid="B14">14</xref>
                  ], OCSVM [
                  <xref ref-type="bibr" rid="B15">15</xref>
                  ]
                </td>
                <td>Fast training and friendly to small samples</td>
                <td>Difficult to model long-term dependence and variable coupling</td>
                <td>Introduces graph-temporal modeling and dual-domain representation</td>
              </tr>
              <tr>
                <td>Supervised classification</td>
                <td>
                  RF, XGBoost, etc [
                  <xref ref-type="bibr" rid="B5">5</xref>
                  ].
                </td>
                <td>Strong capability for anomaly recognition</td>
                <td>Strong label dependence and limited generalization</td>
                <td>Enhances early prediction and robustness</td>
              </tr>
              <tr>
                <td>Reconstruction-based deep models</td>
                <td>
                  OmniAnomaly [
                  <xref ref-type="bibr" rid="B20">20</xref>
                  ], USAD [
                  <xref ref-type="bibr" rid="B22">22</xref>
                  ], DAGMM [
                  <xref ref-type="bibr" rid="B23">23</xref>
                  ]
                </td>
                <td>Can learn complex normal patterns</td>
                <td>Difficult threshold transfer and sensitivity to missingness</td>
                <td>Uses multi-task discrimination instead of a single reconstruction path</td>
              </tr>
              <tr>
                <td>
                  Graph/Transformer models [
                  <xref ref-type="bibr" rid="B30">30</xref>
                  ]
                </td>
                <td>
                  GDN [
                  <xref ref-type="bibr" rid="B27">27</xref>
                  ], TranAD [
                  <xref ref-type="bibr" rid="B28">28</xref>
                  ], DCdetector [
                  <xref ref-type="bibr" rid="B30">30</xref>
                  ]
                </td>
                <td>Can learn long dependencies and variable relationships</td>
                <td>Many assume fixed input dimensions</td>
                <td>Adds missingness routing and missing-aware modeling</td>
              </tr>
              <tr>
                <td>Proposed method</td>
                <td>SDG-Former</td>
                <td>Balances detection, recognition, early warning, and interpretation</td>
                <td>Relatively complex structure and high requirements for implementation protocols</td>
                <td>Controls complexity through modular design</td>
              </tr>
            </tbody>
          </table>
        </table-wrap>
      </sec>
    </sec>
    <sec id="sec6">
      <title>6. Conclusions and Prospects</title>
      <p>This paper focuses on data-driven prediction of undesirable events in offshore oil wells. It reorganizes the research background, public data, related work, algorithm design, experimental settings, and discussion of results, and discusses the potential applicability of SDG-Former from the perspective of a conceptual framework. The difficulty of predicting undesirable events in offshore oil wells lies not only in the scarcity of anomalous samples, but also in complex variable coupling, obvious structural missingness, strong cross-well heterogeneity, and the need for early prediction of undesirable events.</p>
      <p>Based on a systematic analysis of the 3W Dataset and related literature, the SDG-Former proposed in this paper forms an algorithmic design route that is more suitable for offshore oil well scenarios through missing-aware input, time-frequency dual-domain representation, a graph-temporal Transformer backbone, and a multi-task output mechanism. It provides a structured research framework that unifies anomaly detection and undesirable-event prediction.</p>
      <p>Future research can proceed in three directions. First, self-supervised pretraining and foundation-modeling ideas can be introduced to improve transferability in small-sample and multi-well scenarios. Second, physical mechanisms, wellbore-flow constraints, and operating procedures can be incorporated into neural-network training to enhance model interpretability and credibility. Third, missing, frozen, and drifting data should be further studied to avoid the implicit amplification of data-quality problems within the model.</p>
    </sec>
    <sec id="sec7">
      <title>Acknowledgements</title>
      <p>The authors would like to extend their sincere gratitude to North China University of Science and Technology in Tangshan, China, as well as to Petrobras for supplying the data used in this study.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <title>References</title>
      <ref id="B1">
        <label>1.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Azmi, P.A.R., Yusoff, M. and Mohd Sallehud-din, M.T. (2024) A Review of Predictive Analytics Models in the Oil and Gas Industries. <italic>Sensors</italic>, 24, Article 4013. https://doi.org/10.3390/s24124013 <pub-id pub-id-type="doi">10.3390/s24124013</pub-id><pub-id pub-id-type="pmid">38931798</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.3390/s24124013">https://doi.org/10.3390/s24124013</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Azmi, P.A.R.</string-name>
              <string-name>Yusoff, M.</string-name>
              <string-name>Sallehud-din, M.T.</string-name>
            </person-group>
            <year>2024</year>
            <article-title>A Review of Predictive Analytics Models in the Oil and Gas Industries</article-title>
            <source>Sensors</source>
            <volume>24</volume>
            <elocation-id>4013</elocation-id>
            <pub-id pub-id-type="doi">10.3390/s24124013</pub-id>
            <pub-id pub-id-type="pmid">38931798</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B2">
        <label>2.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">D’Almeida, A.L., Bergiante, N.C.R., de Souza Ferreira, G., Leta, F.R., de Campos Lima, C.B. and Lima, G.B.A. (2022) Digital Transformation: A Review on Artificial Intelligence Techniques in Drilling and Production Applications. <italic>The International Journal of Advanced Manufacturing Technology</italic>, 119, 5553-5582. https://doi.org/10.1007/s00170-021-08631-w <pub-id pub-id-type="doi">10.1007/s00170-021-08631-w</pub-id><pub-id pub-id-type="pmid">35095165</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1007/s00170-021-08631-w">https://doi.org/10.1007/s00170-021-08631-w</ext-link></mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Almeida, A.L.</string-name>
              <string-name>Bergiante, N.C.R.</string-name>
              <string-name>Ferreira, G.</string-name>
              <string-name>Leta, F.R.</string-name>
              <string-name>Lima, C.B.</string-name>
              <string-name>Lima, G.B.A.</string-name>
            </person-group>
            <year>2022</year>
            <article-title>Digital Transformation: A Review on Artificial Intelligence Techniques in Drilling and Production Applications</article-title>
            <source>The International Journal of Advanced Manufacturing Technology</source>
            <volume>119</volume>
            <pub-id pub-id-type="doi">10.1007/s00170-021-08631-w</pub-id>
            <pub-id pub-id-type="pmid">35095165</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B3">
        <label>3.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Tariq, Z., Aljawad, M.S., Hasan, A., Murtaza, M., Mohammed, E., El-Husseiny, A., <italic>et al.</italic> (2021) A Systematic Review of Data Science and Machine Learning Applications to the Oil and Gas Industry. <italic>Journal of Petroleum Exploration and Production Technology</italic>, 11, 4339-4374. https://doi.org/10.1007/s13202-021-01302-2 <pub-id pub-id-type="doi">10.1007/s13202-021-01302-2</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1007/s13202-021-01302-2">https://doi.org/10.1007/s13202-021-01302-2</ext-link></mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Tariq, Z.</string-name>
              <string-name>Aljawad, M.S.</string-name>
              <string-name>Hasan, A.</string-name>
              <string-name>Murtaza, M.</string-name>
              <string-name>Mohammed, E.</string-name>
              <string-name>El-Husseiny, A.</string-name>
            </person-group>
            <year>2021</year>
            <article-title>A Systematic Review of Data Science and Machine Learning Applications to the Oil and Gas Industry</article-title>
            <source>Journal of Petroleum Exploration and Production Technology</source>
            <volume>11</volume>
            <pub-id pub-id-type="doi">10.1007/s13202-021-01302-2</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B4">
        <label>4.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Vargas, R.E.V., Munaro, C.J., Ciarelli, P.M., Medeiros, A.G., Amaral, B.G.D., Barrionuevo, D.C., <italic>et al.</italic> (2019) A Realistic and Public Dataset with Rare Undesirable Real Events in Oil Wells. <italic>Journal of Petroleum Science and Engineering</italic>, 181, Article ID: 106223. https://doi.org/10.1016/j.petrol.2019.106223 <pub-id pub-id-type="doi">10.1016/j.petrol.2019.106223</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1016/j.petrol.2019.106223">https://doi.org/10.1016/j.petrol.2019.106223</ext-link></mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Vargas, R.E.V.</string-name>
              <string-name>Munaro, C.J.</string-name>
              <string-name>Ciarelli, P.M.</string-name>
              <string-name>Medeiros, A.G.</string-name>
              <string-name>Amaral, B.G.D.</string-name>
              <string-name>Barrionuevo, D.C.</string-name>
            </person-group>
            <year>2019</year>
            <article-title>A Realistic and Public Dataset with Rare Undesirable Real Events in Oil Wells</article-title>
            <source>Journal of Petroleum Science and Engineering</source>
            <volume>181</volume>
            <fpage>106223</fpage>
            <elocation-id>ID</elocation-id>
            <pub-id pub-id-type="doi">10.1016/j.petrol.2019.106223</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B5">
        <label>5.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Marins, M.A., Barros, B.D., Santos, I.H., Barrionuevo, D.C., Vargas, R.E.V., de M. Prego, T., <italic>et al.</italic> (2021) Fault Detection and Classification in Oil Wells and Production/Service Lines Using Random Forest. <italic>Journal of Petroleum Science and Engineering</italic>, 197, Article ID: 107879. https://doi.org/10.1016/j.petrol.2020.107879 <pub-id pub-id-type="doi">10.1016/j.petrol.2020.107879</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1016/j.petrol.2020.107879">https://doi.org/10.1016/j.petrol.2020.107879</ext-link></mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Marins, M.A.</string-name>
              <string-name>Barros, B.D.</string-name>
              <string-name>Santos, I.H.</string-name>
              <string-name>Barrionuevo, D.C.</string-name>
              <string-name>Vargas, R.E.V.</string-name>
              <string-name>Prego, T.</string-name>
            </person-group>
            <year>2021</year>
            <article-title>Fault Detection and Classification in Oil Wells and Production/Service Lines Using Random Forest</article-title>
            <source>Journal of Petroleum Science and Engineering</source>
            <volume>197</volume>
            <fpage>107879</fpage>
            <elocation-id>ID</elocation-id>
            <pub-id pub-id-type="doi">10.1016/j.petrol.2020.107879</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B6">
        <label>6.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Machado, A.P.F., Vargas, R.E.V., Ciarelli, P.M. and Munaro, C.J. (2022) Improving Performance of One-Class Classifiers Applied to Anomaly Detection in Oil Wells. <italic>Journal of Petroleum Science and Engineering</italic>, 218, Article ID: 110983. https://doi.org/10.1016/j.petrol.2022.110983 <pub-id pub-id-type="doi">10.1016/j.petrol.2022.110983</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1016/j.petrol.2022.110983">https://doi.org/10.1016/j.petrol.2022.110983</ext-link></mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Machado, A.P.F.</string-name>
              <string-name>Vargas, R.E.V.</string-name>
              <string-name>Ciarelli, P.M.</string-name>
              <string-name>Munaro, C.J.</string-name>
            </person-group>
            <year>2022</year>
            <article-title>Improving Performance of One-Class Classifiers Applied to Anomaly Detection in Oil Wells</article-title>
            <source>Journal of Petroleum Science and Engineering</source>
            <volume>218</volume>
            <fpage>110983</fpage>
            <elocation-id>ID</elocation-id>
            <pub-id pub-id-type="doi">10.1016/j.petrol.2022.110983</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B7">
        <label>7.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Fernandes, W., Komati, K.S. and Assis de Souza Gazolli, K. (2023) Anomaly Detection in Oil-Producing Wells: A Comparative Study of One-Class Classifiers in a Multivariate Time Series Dataset. <italic>Journal of Petroleum Exploration and Production Technolog</italic><italic>y</italic>, 14, 343-363. https://doi.org/10.1007/s13202-023-01710-6 <pub-id pub-id-type="doi">10.1007/s13202-023-01710-6</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1007/s13202-023-01710-6">https://doi.org/10.1007/s13202-023-01710-6</ext-link></mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Fernandes, W.</string-name>
              <string-name>Komati, K.S.</string-name>
              <string-name>Gazolli, K.</string-name>
            </person-group>
            <year>2023</year>
            <article-title>Anomaly Detection in Oil-Producing Wells: A Comparative Study of One-Class Classifiers in a Multivariate Time Series Dataset</article-title>
            <source>Journal of Petroleum Exploration and Production Technology</source>
            <volume>14</volume>
            <pub-id pub-id-type="doi">10.1007/s13202-023-01710-6</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B8">
        <label>8.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Bayazitova, G., Anastasiadou, M. and dos Santos, V.D. (2024) Oil and Gas Flow Anomaly Detection on Offshore Naturally Flowing Wells Using Deep Neural Networks. <italic>Ge</italic><italic>oenergy</italic><italic>Science and Engineering</italic>, 242, Article ID: 213240. https://doi.org/10.1016/j.geoen.2024.213240 <pub-id pub-id-type="doi">10.1016/j.geoen.2024.213240</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1016/j.geoen.2024.213240">https://doi.org/10.1016/j.geoen.2024.213240</ext-link></mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Bayazitova, G.</string-name>
              <string-name>Anastasiadou, M.</string-name>
              <string-name>Santos, V.D.</string-name>
            </person-group>
            <year>2024</year>
            <article-title>Oil and Gas Flow Anomaly Detection on Offshore Naturally Flowing Wells Using Deep Neural Networks</article-title>
            <source>Geoenergy Science and Engineering</source>
            <volume>242</volume>
            <fpage>213240</fpage>
            <elocation-id>ID</elocation-id>
            <pub-id pub-id-type="doi">10.1016/j.geoen.2024.213240</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B9">
        <label>9.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Soriano-Vargas, A., Werneck, R., Moura, R., Mendes Júnior, P., Prates, R., Castro, M., <italic>et al.</italic> (2021) A Visual Analytics Approach to Anomaly Detection in Hydrocarbon Reservoir Time Series Data. <italic>Journal of Petroleum Science and Engineering</italic>, 206, Article ID: 108988. https://doi.org/10.1016/j.petrol.2021.108988 <pub-id pub-id-type="doi">10.1016/j.petrol.2021.108988</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1016/j.petrol.2021.108988">https://doi.org/10.1016/j.petrol.2021.108988</ext-link></mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Soriano-Vargas, A.</string-name>
              <string-name>Werneck, R.</string-name>
              <string-name>Moura, R.</string-name>
              <string-name>Prates, R.</string-name>
              <string-name>Castro, M.</string-name>
            </person-group>
            <year>2021</year>
            <article-title>A Visual Analytics Approach to Anomaly Detection in Hydrocarbon Reservoir Time Series Data</article-title>
            <source>Journal of Petroleum Science and Engineering</source>
            <volume>206</volume>
            <fpage>108988</fpage>
            <elocation-id>ID</elocation-id>
            <pub-id pub-id-type="doi">10.1016/j.petrol.2021.108988</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B10">
        <label>10.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Lopes, L.G.O., Vieira, T.M.d.A., Aranha, P.E., Lima Junior, E.T.D. and Lira, W.W.M. (2025) Detection and Classification of Anomalies in Oil Well Production Using Open-World Learning. <italic>Engineering Applications of Artificial Intelligence</italic>, 159, Article ID: 111514. https://doi.org/10.1016/j.engappai.2025.111514 <pub-id pub-id-type="doi">10.1016/j.engappai.2025.111514</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1016/j.engappai.2025.111514">https://doi.org/10.1016/j.engappai.2025.111514</ext-link></mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Lopes, L.G.O.</string-name>
              <string-name>Vieira, T.M.</string-name>
              <string-name>Aranha, P.E.</string-name>
              <string-name>Junior, E.T.D.</string-name>
              <string-name>Lira, W.W.M.</string-name>
            </person-group>
            <year>2025</year>
            <article-title>Detection and Classification of Anomalies in Oil Well Production Using Open-World Learning</article-title>
            <source>Engineering Applications of Artificial Intelligence</source>
            <volume>159</volume>
            <fpage>111514</fpage>
            <elocation-id>ID</elocation-id>
            <pub-id pub-id-type="doi">10.1016/j.engappai.2025.111514</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B11">
        <label>11.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Chandola, V., Banerjee, A. and Kumar, V. (2009) Anomaly Detection. <italic>ACM Computing Surveys</italic>, 41, 1-58. https://doi.org/10.1145/1541880.1541882 <pub-id pub-id-type="doi">10.1145/1541880.1541882</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1145/1541880.1541882">https://doi.org/10.1145/1541880.1541882</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Chandola, V.</string-name>
              <string-name>Banerjee, A.</string-name>
              <string-name>Kumar, V.</string-name>
            </person-group>
            <year>2009</year>
            <article-title>Anomaly Detection</article-title>
            <source>ACM Computing Surveys</source>
            <volume>41</volume>
            <pub-id pub-id-type="doi">10.1145/1541880.1541882</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B12">
        <label>12.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Blázquez-García, A., Conde, A., Mori, U. and Lozano, J.A. (2021) A Review on Outlier/Anomaly Detection in Time Series Data. <italic>ACM Computing Surveys</italic>, 54, 1-33. https://doi.org/10.1145/3444690 <pub-id pub-id-type="doi">10.1145/3444690</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1145/3444690">https://doi.org/10.1145/3444690</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Conde, A.</string-name>
              <string-name>Mori, U.</string-name>
              <string-name>Lozano, J.A.</string-name>
            </person-group>
            <year>2021</year>
            <article-title>A Review on Outlier/Anomaly Detection in Time Series Data</article-title>
            <source>ACM Computing Surveys</source>
            <volume>54</volume>
            <pub-id pub-id-type="doi">10.1145/3444690</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B13">
        <label>13.</label>
        <citation-alternatives>
          <mixed-citation publication-type="confproc">Breunig, M.M., Kriegel, H., Ng, R.T. and Sander, J. (2000) LOF: Identifying Density-based Local Outliers. <italic>Proceedings of the</italic> 2000 <italic>ACM SIGMOD International Conference on Management of Data</italic>, Dallas, 15-18 May 2000, 93-104. https://doi.org/10.1145/342009.335388 <pub-id pub-id-type="doi">10.1145/342009.335388</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1145/342009.335388">https://doi.org/10.1145/342009.335388</ext-link></mixed-citation>
          <element-citation publication-type="confproc">
            <person-group person-group-type="author">
              <string-name>Breunig, M.M.</string-name>
              <string-name>Kriegel, H.</string-name>
              <string-name>Ng, R.T.</string-name>
              <string-name>Sander, J.</string-name>
              <string-name>Data, D</string-name>
            </person-group>
            <year>2000</year>
            <article-title>LOF: Identifying Density-based Local Outliers</article-title>
            <source>Proceedings of the 2000 ACM SIGMOD International Conference on Management of Data</source>
            <volume>15</volume>
            <pub-id pub-id-type="doi">10.1145/342009.335388</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B14">
        <label>14.</label>
        <citation-alternatives>
          <mixed-citation publication-type="confproc">Liu, F.T., Ting, K.M. and Zhou, Z. (2008. Isolation Forest. 2008 <italic>Eighth IEEE International Conference on Data Mining</italic>, Pisa, 15-19 December 2008, 413-422. https://doi.org/10.1109/icdm.2008.17 <pub-id pub-id-type="doi">10.1109/icdm.2008.17</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1109/icdm.2008.17">https://doi.org/10.1109/icdm.2008.17</ext-link></mixed-citation>
          <element-citation publication-type="confproc">
            <person-group person-group-type="author">
              <string-name>Liu, F.T.</string-name>
              <string-name>Ting, K.M.</string-name>
              <string-name>Zhou, Z.</string-name>
              <string-name>Mining, P</string-name>
            </person-group>
            <year>2008</year>
            <pub-id pub-id-type="doi">10.1109/icdm.2008.17</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B15">
        <label>15.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Schölkopf, B., Platt, J.C., Shawe-Taylor, J., Smola, A.J. and Williamson, R.C. (2001) Estimating the Support of a High-Dimensional Distribution. <italic>Neural Computation</italic>, 13, 1443-1471. https://doi.org/10.1162/089976601750264965 <pub-id pub-id-type="doi">10.1162/089976601750264965</pub-id><pub-id pub-id-type="pmid">11440593</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1162/089976601750264965">https://doi.org/10.1162/089976601750264965</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Platt, J.C.</string-name>
              <string-name>Shawe-Taylor, J.</string-name>
              <string-name>Smola, A.J.</string-name>
              <string-name>Williamson, R.C.</string-name>
            </person-group>
            <year>2001</year>
            <article-title>Estimating the Support of a High-Dimensional Distribution</article-title>
            <source>Neural Computation</source>
            <volume>13</volume>
            <pub-id pub-id-type="doi">10.1162/089976601750264965</pub-id>
            <pub-id pub-id-type="pmid">11440593</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B16">
        <label>16.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Jardine, A.K.S., Lin, D. and Banjevic, D. (2006) A Review on Machinery Diagnostics and Prognostics Implementing Condition-Based Maintenance. <italic>Mechanical Systems and Signal Processing</italic>, 20, 1483-1510. https://doi.org/10.1016/j.ymssp.2005.09.012 <pub-id pub-id-type="doi">10.1016/j.ymssp.2005.09.012</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1016/j.ymssp.2005.09.012">https://doi.org/10.1016/j.ymssp.2005.09.012</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Jardine, A.K.S.</string-name>
              <string-name>Lin, D.</string-name>
              <string-name>Banjevic, D.</string-name>
            </person-group>
            <year>2006</year>
            <article-title>A Review on Machinery Diagnostics and Prognostics Implementing Condition-Based Maintenance</article-title>
            <source>Mechanical Systems and Signal Processing</source>
            <volume>20</volume>
            <pub-id pub-id-type="doi">10.1016/j.ymssp.2005.09.012</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B17">
        <label>17.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Ismail Fawaz, H., Forestier, G., Weber, J., Idoumghar, L. and Muller, P. (2019) Deep Learning for Time Series Classification: A Review. <italic>Data Mining and Knowledge Discovery</italic>, 33, 917-963. https://doi.org/10.1007/s10618-019-00619-1 <pub-id pub-id-type="doi">10.1007/s10618-019-00619-1</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1007/s10618-019-00619-1">https://doi.org/10.1007/s10618-019-00619-1</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Fawaz, H.</string-name>
              <string-name>Forestier, G.</string-name>
              <string-name>Weber, J.</string-name>
              <string-name>Idoumghar, L.</string-name>
              <string-name>Muller, P.</string-name>
            </person-group>
            <year>2019</year>
            <article-title>Deep Learning for Time Series Classification: A Review</article-title>
            <source>Data Mining and Knowledge Discovery</source>
            <volume>33</volume>
            <pub-id pub-id-type="doi">10.1007/s10618-019-00619-1</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B18">
        <label>18.</label>
        <citation-alternatives>
          <mixed-citation publication-type="confproc">Malhotra, P,. Vig, L., Shroff, G., <italic>et al.</italic> (2015) Long Short Term Memory Networks for Anomaly Detection in Time Series. <italic>ESANN</italic>2015 <italic>Proceedings</italic>, Bruges, 22-24 April 2015, 89-94.</mixed-citation>
          <element-citation publication-type="confproc">
            <person-group person-group-type="author">
              <string-name>Malhotra, P</string-name>
              <string-name>Vig, L.</string-name>
              <string-name>Shroff, G.</string-name>
              <string-name>Proceedings, B</string-name>
            </person-group>
            <year>2015</year>
            <article-title>Long Short Term Memory Networks for Anomaly Detection in Time Series</article-title>
            <source>ESANN 2015 Proceedings</source>
            <volume>22</volume>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B19">
        <label>19.</label>
        <citation-alternatives>
          <mixed-citation publication-type="confproc">Hundman, K., Constantinou, V., Laporte, C., Colwell, I. and Soderstrom, T. (2018) Detecting Spacecraft Anomalies Using Lstms and Nonparametric Dynamic Thresholding. <italic>Proceedings of the</italic> 24 <italic>th ACM SIGKDD International Conference on Knowledge D</italic><italic>iscovery &amp; Data Mining</italic>, London, 19-23 August 2018, 387-395. https://doi.org/10.1145/3219819.3219845 <pub-id pub-id-type="doi">10.1145/3219819.3219845</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1145/3219819.3219845">https://doi.org/10.1145/3219819.3219845</ext-link></mixed-citation>
          <element-citation publication-type="confproc">
            <person-group person-group-type="author">
              <string-name>Hundman, K.</string-name>
              <string-name>Constantinou, V.</string-name>
              <string-name>Laporte, C.</string-name>
              <string-name>Colwell, I.</string-name>
              <string-name>Soderstrom, T.</string-name>
              <string-name>Mining, L</string-name>
            </person-group>
            <year>2018</year>
            <article-title>Detecting Spacecraft Anomalies Using Lstms and Nonparametric Dynamic Thresholding</article-title>
            <source>Proceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery &amp; Data Mining</source>
            <volume>19</volume>
            <pub-id pub-id-type="doi">10.1145/3219819.3219845</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B20">
        <label>20.</label>
        <citation-alternatives>
          <mixed-citation publication-type="confproc">Su, Y., Zhao, Y., Niu, C., Liu, R., Sun, W. and Pei, D. (2019) Robust Anomaly Detection for Multivariate Time Series through Stochastic Recurrent Neural Network. <italic>Proceedings of the</italic> 25 <italic>th ACM SIGKDD International Conference on Knowledge Discovery &amp; Data Mining</italic>, Anchorage, 4-8 August 2019, 2828-2837. https://doi.org/10.1145/3292500.3330672 <pub-id pub-id-type="doi">10.1145/3292500.3330672</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1145/3292500.3330672">https://doi.org/10.1145/3292500.3330672</ext-link></mixed-citation>
          <element-citation publication-type="confproc">
            <person-group person-group-type="author">
              <string-name>Su, Y.</string-name>
              <string-name>Zhao, Y.</string-name>
              <string-name>Niu, C.</string-name>
              <string-name>Liu, R.</string-name>
              <string-name>Sun, W.</string-name>
              <string-name>Pei, D.</string-name>
              <string-name>Mining, A</string-name>
            </person-group>
            <year>2019</year>
            <article-title>Robust Anomaly Detection for Multivariate Time Series through Stochastic Recurrent Neural Network</article-title>
            <source>Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery &amp; Data Mining</source>
            <volume>4</volume>
            <pub-id pub-id-type="doi">10.1145/3292500.3330672</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B21">
        <label>21.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Li, D., Chen, D., Jin, B., Shi, L., Goh, J. and Ng, S. (2019) MAD-GAN: Multivariate Anomaly Detection for Time Series Data with Generative Adversarial Networks. In: Tetko, I., Kůrková, V., Karpov, P. and Theis, F., Eds., <italic>Artificial Neural Networks and Machine Learning</italic>— <italic>ICANN</italic> 2019: <italic>Text and Time Series</italic>, Springer, 703-716. https://doi.org/10.1007/978-3-030-30490-4_56 <pub-id pub-id-type="doi">10.1007/978-3-030-30490-4_56</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1007/978-3-030-30490-4_56">https://doi.org/10.1007/978-3-030-30490-4_56</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Li, D.</string-name>
              <string-name>Chen, D.</string-name>
              <string-name>Jin, B.</string-name>
              <string-name>Shi, L.</string-name>
              <string-name>Goh, J.</string-name>
              <string-name>Ng, S.</string-name>
              <string-name>Tetko, I.</string-name>
              <string-name>Karpov, P.</string-name>
              <string-name>Theis, F.</string-name>
              <string-name>Series, S</string-name>
            </person-group>
            <year>2019</year>
            <article-title>MAD-GAN: Multivariate Anomaly Detection for Time Series Data with Generative Adversarial Networks</article-title>
            <source>In: Tetko</source>
            <volume>703</volume>
            <pub-id pub-id-type="doi">10.1007/978-3-030-30490-4_56</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B22">
        <label>22.</label>
        <citation-alternatives>
          <mixed-citation publication-type="confproc">Zong, B., Song, Q., Min, M.R., <italic>et al.</italic> (2018) Deep Autoencoding Gaussian Mixture Model for Unsupervised Anomaly Detection. 2018 <italic>International Conference on Learning</italic><italic>Representations</italic>, Vancouver, 30 April-3 May 2018, 1-19.</mixed-citation>
          <element-citation publication-type="confproc">
            <person-group person-group-type="author">
              <string-name>Zong, B.</string-name>
              <string-name>Song, Q.</string-name>
              <string-name>Min, M.R.</string-name>
              <string-name>Representations, V</string-name>
            </person-group>
            <year>2018</year>
            <article-title>Deep Autoencoding Gaussian Mixture Model for Unsupervised Anomaly Detection</article-title>
            <source>2018 International Conference on Learning Representations</source>
            <volume>30</volume>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B23">
        <label>23.</label>
        <citation-alternatives>
          <mixed-citation publication-type="confproc">Audibert, J., Michiardi, P., Guyard, F., Marti, S. and Zuluaga, M.A. (2020) USAD: Unsupervised Anomaly Detection on Multivariate Time Series. <italic>Proceedings of the</italic>26 <italic>th ACM SIGKDD International Conference on Knowledge Discovery &amp; Data Mining</italic>, 6-10 July 2020, 3395-3404. https://doi.org/10.1145/3394486.3403392 <pub-id pub-id-type="doi">10.1145/3394486.3403392</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1145/3394486.3403392">https://doi.org/10.1145/3394486.3403392</ext-link></mixed-citation>
          <element-citation publication-type="confproc">
            <person-group person-group-type="author">
              <string-name>Audibert, J.</string-name>
              <string-name>Michiardi, P.</string-name>
              <string-name>Guyard, F.</string-name>
              <string-name>Marti, S.</string-name>
              <string-name>Zuluaga, M.A.</string-name>
            </person-group>
            <year>2020</year>
            <article-title>USAD: Unsupervised Anomaly Detection on Multivariate Time Series</article-title>
            <source>Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery &amp; Data Mining</source>
            <volume>6</volume>
            <pub-id pub-id-type="doi">10.1145/3394486.3403392</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B24">
        <label>24.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Vaswani, A., Shazeer, N., Parmar, N., <italic>et al.</italic> (2017) Attention Is All You Need. arXiv: 1706.03762</mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Vaswani, A.</string-name>
              <string-name>Shazeer, N.</string-name>
              <string-name>Parmar, N.</string-name>
            </person-group>
            <year>2017</year>
            <article-title>Attention Is All You Need</article-title>
            <fpage>1706</fpage>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B25">
        <label>25.</label>
        <citation-alternatives>
          <mixed-citation publication-type="confproc">Zhou, H., Zhang, S., Peng, J., Zhang, S., Li, J., Xiong, H., <italic>et al.</italic> (2021) Informer: Beyond Efficient Transformer for Long Sequence Time-Series Forecasting. <italic>Proceedings of the AAAI Conference on Artificial Intelligence</italic>, 35, 11106-11115. https://doi.org/10.1609/aaai.v35i12.17325 <pub-id pub-id-type="doi">10.1609/aaai.v35i12.17325</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1609/aaai.v35i12.17325">https://doi.org/10.1609/aaai.v35i12.17325</ext-link></mixed-citation>
          <element-citation publication-type="confproc">
            <person-group person-group-type="author">
              <string-name>Zhou, H.</string-name>
              <string-name>Zhang, S.</string-name>
              <string-name>Peng, J.</string-name>
              <string-name>Zhang, S.</string-name>
              <string-name>Li, J.</string-name>
              <string-name>Xiong, H.</string-name>
            </person-group>
            <year>2021</year>
            <article-title>Informer: Beyond Efficient Transformer for Long Sequence Time-Series Forecasting</article-title>
            <source>Proceedings of the AAAI Conference on Artificial Intelligence</source>
            <volume>35</volume>
            <pub-id pub-id-type="doi">10.1609/aaai.v35i12.17325</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B26">
        <label>26.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Xu, J,. Wu, H., Wang, J., <italic>et al.</italic> (2022) Anomaly Transformer: Time Series Anomaly Detection with Association Discrepancy. arXiv: 2110.02642</mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Xu, J</string-name>
              <string-name>Wu, H.</string-name>
              <string-name>Wang, J.</string-name>
            </person-group>
            <year>2022</year>
            <article-title>Anomaly Transformer: Time Series Anomaly Detection with Association Discrepancy</article-title>
            <fpage>2110</fpage>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B27">
        <label>27.</label>
        <citation-alternatives>
          <mixed-citation publication-type="confproc">Tuli, S., Casale, G. and Jennings, N.R. (2022) TranAD: Deep Transformer Networks for Anomaly Detection in Multivariate Time Series Data. <italic>Proceedings of the VLDB Endowment</italic>, 15, 1201-1214. https://doi.org/10.14778/3514061.3514067 <pub-id pub-id-type="doi">10.14778/3514061.3514067</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.14778/3514061.3514067">https://doi.org/10.14778/3514061.3514067</ext-link></mixed-citation>
          <element-citation publication-type="confproc">
            <person-group person-group-type="author">
              <string-name>Tuli, S.</string-name>
              <string-name>Casale, G.</string-name>
              <string-name>Jennings, N.R.</string-name>
            </person-group>
            <year>2022</year>
            <article-title>TranAD: Deep Transformer Networks for Anomaly Detection in Multivariate Time Series Data</article-title>
            <source>Proceedings of the VLDB Endowment</source>
            <volume>15</volume>
            <pub-id pub-id-type="doi">10.14778/3514061.3514067</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B28">
        <label>28.</label>
        <citation-alternatives>
          <mixed-citation publication-type="confproc">Yang, Y., Zhang, C., Zhou, T., Wen, Q. and Sun, L. (2023) DCdetector: Dual Attention Contrastive Representation Learning for Time Series Anomaly Detection. <italic>Proceedings of the</italic> 29 <italic>th ACM SIGKDD Conference on Knowledge Discovery and Data Mining</italic>, Long Beach, 6-10 August 2023, 3033-3045. https://doi.org/10.1145/3580305.3599295 <pub-id pub-id-type="doi">10.1145/3580305.3599295</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1145/3580305.3599295">https://doi.org/10.1145/3580305.3599295</ext-link></mixed-citation>
          <element-citation publication-type="confproc">
            <person-group person-group-type="author">
              <string-name>Yang, Y.</string-name>
              <string-name>Zhang, C.</string-name>
              <string-name>Zhou, T.</string-name>
              <string-name>Wen, Q.</string-name>
              <string-name>Sun, L.</string-name>
              <string-name>Mining, L</string-name>
            </person-group>
            <year>2023</year>
            <article-title>DCdetector: Dual Attention Contrastive Representation Learning for Time Series Anomaly Detection</article-title>
            <source>Proceedings of the 29th ACM SIGKDD Conference on Knowledge Discovery and Data Mining</source>
            <volume>6</volume>
            <pub-id pub-id-type="doi">10.1145/3580305.3599295</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B29">
        <label>29.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Zhao, M., Peng, H., Li, L. and Ren, Y. (2024) Graph Attention Network and Informer for Multivariate Time Series Anomaly Detection. <italic>Sensors</italic>, 24, Article 1522. https://doi.org/10.3390/s24051522 <pub-id pub-id-type="doi">10.3390/s24051522</pub-id><pub-id pub-id-type="pmid">38475058</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.3390/s24051522">https://doi.org/10.3390/s24051522</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Zhao, M.</string-name>
              <string-name>Peng, H.</string-name>
              <string-name>Li, L.</string-name>
              <string-name>Ren, Y.</string-name>
            </person-group>
            <year>2024</year>
            <article-title>Graph Attention Network and Informer for Multivariate Time Series Anomaly Detection</article-title>
            <source>Sensors</source>
            <volume>24</volume>
            <elocation-id>1522</elocation-id>
            <pub-id pub-id-type="doi">10.3390/s24051522</pub-id>
            <pub-id pub-id-type="pmid">38475058</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B30">
        <label>30.</label>
        <citation-alternatives>
          <mixed-citation publication-type="confproc">Deng, A. and Hooi, B. (2021) Graph Neural Network-Based Anomaly Detection in Multivariate Time Series. <italic>Proceedings of the AAAI Conference on Artificial Intelligence</italic>, 35, 4027-4035. https://doi.org/10.1609/aaai.v35i5.16523 <pub-id pub-id-type="doi">10.1609/aaai.v35i5.16523</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1609/aaai.v35i5.16523">https://doi.org/10.1609/aaai.v35i5.16523</ext-link></mixed-citation>
          <element-citation publication-type="confproc">
            <person-group person-group-type="author">
              <string-name>Deng, A.</string-name>
              <string-name>Hooi, B.</string-name>
            </person-group>
            <year>2021</year>
            <article-title>Graph Neural Network-Based Anomaly Detection in Multivariate Time Series</article-title>
            <source>Proceedings of the AAAI Conference on Artificial Intelligence</source>
            <volume>35</volume>
            <pub-id pub-id-type="doi">10.1609/aaai.v35i5.16523</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
    </ref-list>
  </back>
</article>