<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.4 20241031//EN" "JATS-journalpublishing1-4.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article" dtd-version="1.4" xml:lang="en">
  <front>
    <journal-meta>
      <journal-id journal-id-type="publisher-id">jcc</journal-id>
      <journal-title-group>
        <journal-title>Journal of Computer and Communications</journal-title>
      </journal-title-group>
      <issn pub-type="epub">2327-5227</issn>
      <issn pub-type="ppub">2327-5219</issn>
      <publisher>
        <publisher-name>Scientific Research Publishing</publisher-name>
      </publisher>
    </journal-meta>
    <article-meta>
      <article-id pub-id-type="doi">10.4236/jcc.2025.1312005</article-id>
      <article-id pub-id-type="publisher-id">jcc-148094</article-id>
      <article-categories>
        <subj-group>
          <subject>Article</subject>
        </subj-group>
        <subj-group>
          <subject>Computer Science</subject>
          <subject>Communications</subject>
        </subj-group>
      </article-categories>
      <title-group>
        <article-title>SHAP-Driven Interpretability in Financial Fraud Detection: A Multimodal Data Approach</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <name name-style="western">
            <surname>Hui</surname>
            <given-names>Nie</given-names>
          </name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
      </contrib-group>
      <aff id="aff1"><label>1</label> School of Information Management, Sun Yat-Sen University, Guangzhou, China </aff>
      <author-notes>
        <fn fn-type="conflict" id="fn-conflict">
          <p>The author declares no conflicts of interest regarding the publication of this paper.</p>
        </fn>
      </author-notes>
      <pub-date pub-type="epub">
        <day>04</day>
        <month>12</month>
        <year>2025</year>
      </pub-date>
      <pub-date pub-type="collection">
        <month>12</month>
        <year>2025</year>
      </pub-date>
      <volume>13</volume>
      <issue>12</issue>
      <fpage>80</fpage>
      <lpage>99</lpage>
      <history>
        <date date-type="received">
          <day>09</day>
          <month>11</month>
          <year>2025</year>
        </date>
        <date date-type="accepted">
          <day>16</day>
          <month>12</month>
          <year>2025</year>
        </date>
        <date date-type="published">
          <day>19</day>
          <month>12</month>
          <year>2025</year>
        </date>
      </history>
      <permissions>
        <copyright-statement>© 2025 by the authors and Scientific Research Publishing Inc.</copyright-statement>
        <copyright-year>2025</copyright-year>
        <license license-type="open-access">
          <license-p> This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license ( <ext-link ext-link-type="uri" xlink:href="https://creativecommons.org/licenses/by/4.0/">https://creativecommons.org/licenses/by/4.0/</ext-link> ). </license-p>
        </license>
      </permissions>
      <self-uri content-type="doi" xlink:href="https://doi.org/10.4236/jcc.2025.1312005">https://doi.org/10.4236/jcc.2025.1312005</self-uri>
      <abstract>
        <p>Improved accuracy in predicting corporate financial fraud significantly enhances regulatory efficiency and market stability. However, detecting increasingly sophisticated fraud patterns remains challenging due to data heterogeneity and model opacity. This study proposes an innovative, multimodal framework that integrates corporate financial indicators, organizational structures, and semantic features from annual reports. We employ BERT-based semantic extraction on the Management Discussion &amp; Analysis (MD&amp;A) section of an annual report, reduce dimensionality via PCA, and fuse features with corporate financial/organization structural metrics. Ensemble tree models (CatBoost/XGBoost/LightGBM) are optimized for fraud prediction, while SHAP values quantify the contributions of individual features. Experimental results demonstrate a peak ROC-AUC of 0.859, with key findings revealing that: 1) Significant asset transactions, future outlook, and operational overview are the most predictive MD&amp;A contents; 2) Financial indicators dominate feature importance (50% of top predictors); 3) Annual report similarity and tone serve as critical textual red flags. This framework provides regulators with actionable insights through model interpretability, thereby advancing early-warning systems for financial misconduct.</p>
      </abstract>
      <kwd-group kwd-group-type="author-generated" xml:lang="en">
        <kwd>Financial Fraud Detection</kwd>
        <kwd>Multimodal Data Fusion</kwd>
        <kwd>Explainable AI</kwd>
        <kwd>Textual Semantic Analysis</kwd>
        <kwd>Ensemble Learning</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec1">
      <title>1. Introduction</title>
      <p>Listed companies constitute a critical component of the market economy, with their financial health and operational performance exerting a direct influence on the stability and development of a nation’s economy. In recent years, however, the frequent occurrence of financial fraud incidents among Chinese listed companies has imposed severe negative repercussions on the market. As the market economy grows increasingly complex and corporate fraud becomes more sophisticated and concealed, regulatory agencies face unprecedented challenges in detecting financial fraud. Traditional financial fraud detection (FFD) methods, which rely heavily on limited datasets, are proving inadequate in addressing the vast and heterogeneous data generated by modern business operations [<xref ref-type="bibr" rid="B1">1</xref>][<xref ref-type="bibr" rid="B2">2</xref>]. Consequently, exploring novel methodologies to detect increasingly complex and hidden corporate fraud has become imperative.</p>
      <p>Machine learning algorithms have emerged as a promising solution due to their high predictive accuracy and adaptability to diverse data types, distributions, and volumes. These algorithms are particularly well-suited for large-scale, complex data prediction tasks. However, the black-box nature of machine learning models has hindered their practical application in real-world scenarios. Current research on machine learning-based financial fraud prediction models primarily focuses on enhancing predictive performance, with limited attention paid to the role of predictive factors. As a result, the relationship between these factors and corporate financial fraud risk remains inadequately explored [<xref ref-type="bibr" rid="B3">3</xref>][<xref ref-type="bibr" rid="B4">4</xref>]. In practice, interpretable models are often preferred due to their transparency and credibility, making them more readily accepted by stakeholders [<xref ref-type="bibr" rid="B3">3</xref>].</p>
      <p>The rapid expansion of diverse data sources has pushed financial fraud prediction beyond traditional metrics. Predictive models increasingly incorporate external information—organizational charts, annual reports, financial news, and stock commentaries—to boost accuracy and reveal latent fraud signals, offering a fuller view of corporate operations. However, research has not yet fully realized the synergistic integration of external and financial data or exploited multimodal collective effects; analysis and use of textual information remain fragmented and underdeveloped.</p>
      <p>This study addresses those gaps by applying multimodal data fusion and explainable machine learning to corporate financial fraud prediction. Using NLP, we analyze annual reports and combine report-derived features with financial indicators to predict fraud risk in listed companies. We also use explainable models to identify latent risk factors from publicly disclosed information, maximizing the value of multilayered data and offering new perspectives and technical solutions for multimodal fraud detection.</p>
    </sec>
    <sec id="sec2">
      <title>2. Literature Review</title>
      <sec id="sec2dot1">
        <title>2.1. Indicators for Detecting Corporate Financial Fraud</title>
        <p>Research on corporate financial fraud detection has established a comprehensive framework of indicators. Early studies primarily focused on financial metrics, including profitability, solvency, growth potential, operational efficiency, and cash flow [<xref ref-type="bibr" rid="B4">4</xref>][<xref ref-type="bibr" rid="B5">5</xref>]. However, as market environments have grown more complex and fraudulent practices more sophisticated, reliance on financial metrics alone has proven insufficient for effective risk identification. Consequently, researchers have expanded the scope of fraud detection to include non-financial indicators, like board structure or executive team characteristics [<xref ref-type="bibr" rid="B6">6</xref>][<xref ref-type="bibr" rid="B7">7</xref>]. These indicators offer insights into the quality of corporate governance. For instance, Liu <italic>et</italic><italic>al.</italic> [<xref ref-type="bibr" rid="B8">8</xref>] demonstrated that the shareholding ratio of major shareholders serves as a significant predictor of fraudulent behavior in Chinese firms.</p>
        <p>Beyond the tabular-style metrics, textual corporate annual reports have emerged as a critical resource for fraud detection. Annual reports, as the primary medium for corporate disclosure, are instrumental for investors in assessing financial risks. Studies reveal that firms engaging in financial misconduct often employ a dual-fraud strategy, manipulating both financial data and report narratives [<xref ref-type="bibr" rid="B9">9</xref>]. For example, Xu &amp; Zhang [<xref ref-type="bibr" rid="B10">10</xref>] found that high-risk firms tend to adopt an overly optimistic tone in their reports to obscure irregularities. Similarly, Wang <italic>et</italic><italic>al.</italic> [<xref ref-type="bibr" rid="B11">11</xref>] found a negative correlation between a firm’s financial status and report readability, suggesting that obfuscation is a common tactic used conceal financial distress. Zhang <italic>et</italic><italic>al.</italic> [<xref ref-type="bibr" rid="B12">12</xref>] incorporated readability metrics into fraud prediction models, achieving a 26.33% improvement in F1-score, underscoring the predictive value of textual. </p>
        <p>Furthermore, an increasing number of advancements have deepened the analysis of the report content. Brown <italic>et</italic><italic>al.</italic> [<xref ref-type="bibr" rid="B13">13</xref>] and Craja <italic>et</italic><italic>al.</italic> [<xref ref-type="bibr" rid="B14">14</xref>] extracted direct fraud signals from report paragraphs, while Wu and Du [<xref ref-type="bibr" rid="B1">1</xref>] input text feature vectors to LSTM models to achieve a 94.98% prediction accuracy. Liu <italic>et</italic><italic>al.</italic> [<xref ref-type="bibr" rid="B15">15</xref>] further integrated latent semantic features of reports with accounting metrics, demonstrating the synergistic benefits of multimodal data fusion.</p>
      </sec>
      <sec id="sec2dot2">
        <title>2.2. Models for Detecting Financial Fraud</title>
        <p>Methodologies for fraud prediction can generally be classified into three categories: statistical models, machine learning approaches, and deep learning techniques. Statistical models, such as logistic regression (LR), are valued for their predictive power and strong interpretability; however, they often encounter difficulties when dealing with high-dimensional or nonlinear data relationships [<xref ref-type="bibr" rid="B16">16</xref>]. By contrast, machine learning models are well-suited for capturing complex patterns in data. Notably, ensemble methods—such as random forests (RF), XGBoost, and LightGBM—have delivered outstanding performance in fraud detection tasks. For instance, Ali <italic>et</italic><italic>al.</italic> [<xref ref-type="bibr" rid="B17">17</xref>] optimized an XGBoost model to achieve an accuracy rate of 96.05%, while Zhao &amp; Bai [<xref ref-type="bibr" rid="B18">18</xref>] combined LR with XGBoost to further enhance predictive performance. Additionally, Yadav [<xref ref-type="bibr" rid="B19">19</xref>] demonstrated that deep neural networks (DNNs), when applied to textual data with thorough feature engineering, can achieve accuracy of up to 95%. Even so, the black-box nature of machine learning models limits their adoption in practical applications, as most studies prioritize accuracy over interpretability [<xref ref-type="bibr" rid="B17">17</xref>]. Moreover, ensemble methods predominantly rely on structured data, leaving textual information underexplored [<xref ref-type="bibr" rid="B15">15</xref>].</p>
      </sec>
      <sec id="sec2dot3">
        <title>2.3. Emerging Directions and Research Gaps</title>
        <p>Breakthroughs in NLP have significantly expanded the possibilities for fraud detection research. For example, Wu &amp; Du [<xref ref-type="bibr" rid="B1">1</xref>] utilized word embeddings to vectorize annual reports, providing a robust foundation for subsequent analyses. Craja <italic>et</italic><italic>al.</italic> [<xref ref-type="bibr" rid="B14">14</xref>] applied hierarchical attention networks to extract nuanced semantic features from financial texts. Building on these advances, Bhattacharya &amp; Mickovic [<xref ref-type="bibr" rid="B20">20</xref>] fine-tuned BERT for analyzing 10-K reports, achieving a 15% performance improvement over traditional models. Additionally, Wang <italic>et</italic><italic>al.</italic> [<xref ref-type="bibr" rid="B21">21</xref>] proposed a multimodal model that incorporates attention mechanisms to jointly analyze textual and financial data, thereby offering deeper insights into the factors influencing fraud risk.</p>
        <p>Despite these advancements, three critical gaps persist in the literature. First, regarding multimodal data fusion, although Wang <italic>et</italic><italic>al.</italic> [<xref ref-type="bibr" rid="B21">21</xref>] demonstrated that multimodal models outperform unimodal approaches, there is a paucity of research examining the effects of data clustering within such frameworks. Second, in terms of model interpretability, previous studies have not systematically explored the specific mechanisms through which various predictors impact model outputs [<xref ref-type="bibr" rid="B22">22</xref>]. Third, regarding textual depth, current analyses of annual reports primarily focus on lexical, sentiment, and readability features, with limited attention to more nuanced semantic representations.</p>
        <p>To address these gaps, this study integrates advanced deep learning models, multimodal data fusion, and explainable machine learning methodologies. This approach not only aims to improve the accuracy of fraud risk predictions but also seeks to elucidate the key textual and financial determinants underlying these risks. By enhancing both semantic modeling and interpretability, this research aspires to provide methodological innovations for the development of more effective fraud early-warning systems.</p>
      </sec>
    </sec>
    <sec id="sec3">
      <title>3. Research Design</title>
      <sec id="sec3dot1">
        <title>3.1. Research Framework</title>
        <p>This study aims to develop a multimodal, data-driven model for predicting corporate fraud by integrating predictive factors across three dimensions: financial indicators, organizational structure, and annual reports. While financial indicators and organizational structure are represented as tabular data, annual reports are processed as textual data. The primary objective is to uncover latent fraud-related signals embedded within annual reports and to improve prediction accuracy through the integration of these diverse data sources. To this end, we first perform a comprehensive analysis of annual report content to establish a semantic description framework. Next, we combine features derived from financial indicators and organizational structure to construct the predictive model. Furthermore, explainable machine learning techniques are applied to quantitatively evaluate the contribution of each feature to fraud risk, thereby clarifying the model’s reasoning process and providing actionable guidance for regulatory authorities and stakeholders.</p>
        <p>The research workflow is depicted in <xref ref-type="fig" rid="fig1">Figure 1</xref>. Feature extraction is conducted in the initial stage: financial and organizational structure indicators are obtained from the CSMAR database (www.gtarsc.com). Meanwhile, annual reports are parsed to generate four categories of textual features: tone, readability, content similarity, and semantics. These multimodal features are subsequently integrated into the predictive modeling process. Focusing on ensemble tree models, we utilize three widely adopted algorithms—LightGBM, XGBoost, and CatBoost—alongside LR as a baseline. After model training, hyperparameter optimization, and evaluation, the model with the best performance is selected. Finally, the SHAP (Shapley Additive Explanations) method is employed to interpret the model, assessing feature importance and their relationships with fraud risk, with particular attention given to fraud-related elements identified within annual reports.</p>
        <fig id="fig1">
          <label>Figure 1</label>
          <graphic xlink:href="https://html.scirp.org/file/1733376-rId13.jpeg?20251219032617" />
        </fig>
        <p><bold>Figure 1.</bold> Research workflow for financial fraud detection.</p>
      </sec>
      <sec id="sec3dot2">
        <title>3.2. Annual Report Text Analysis</title>
        <p>3.2.1. Semantic Features of Annual Reports</p>
        <p>Annual reports, required by regulators, reveal a company’s operational status. Because worsening financial conditions often trigger corporate fraud, annual reports may contain signals of fraudulent behavior. To identify these signals systematically, we analyze the content of the report—focusing exclusively on the Management Discussion and Analysis (MD&amp;A) section, which is authored by management and discusses causes, risks, and outlooks. The MD&amp;A’s language reflects management’s mindset and strategy; under performance pressure or fraud incentives, it often grows more complex, ambiguous, or overly positive, indicating possible information manipulation. Therefore, the MD&amp;A is a key textual source for detecting corporate fraud and management intent.</p>
        <p>The MD&amp;A section is lengthy but adheres to standardized disclosure rules, so we can split the section into nine definite thematic modules: 1) overall operational overview, 2) core business analysis, 3) non-core business analysis, 4) asset and liability status, 5) investment activities, 6) significant asset &amp; equity transactions, 7) analysis of major subsidiaries and affiliates, 8) structured entities, and 9) future outlook. Each module is extracted using keyword-based regular expressions and truncated to fit the 512-character input limitation of the BERT model.</p>
        <p>Subsequently, we conduct semantic encoding and dimensionality reduction for each module. For semantic encoding, we utilize the pre-trained BERT-Base-Chinese model, which features a 12-layer Transformer encoder with 12 self-attention heads per layer (<bold>Table 1</bold> presents details on fine-tuning parameters). BERT generates a 768-dimensional semantic vector for each module, which is then reduced to between 3 and 20 dimensions through Principal Component Analysis (PCA) to enable integration with structured data features. The optimal dimensionality is determined by model performance across various settings. This reduction in feature dimensionality not only alleviates issues associated with high-dimensional data but also improves computational efficiency and enhances the model’s interpretability.</p>
        <p><bold>Table 1.</bold> BERT fine-tuning parameters. </p>
        <table-wrap id="tbl1">
          <label>Table 1</label>
          <table>
            <tbody>
              <tr>
                <td>Parameters</td>
                <td>Value</td>
                <td>Interpretation</td>
              </tr>
              <tr>
                <td>num_hidden_layers</td>
                <td>12</td>
                <td>number of hidden layers</td>
              </tr>
              <tr>
                <td>hidden_size</td>
                <td>768</td>
                <td>hidden layer dimension</td>
              </tr>
              <tr>
                <td>max_length</td>
                <td>512</td>
                <td>maximum length of input character sequence</td>
              </tr>
              <tr>
                <td>num_attention_heads</td>
                <td>12</td>
                <td>number of attention heads</td>
              </tr>
              <tr>
                <td>epoch</td>
                <td>10</td>
                <td>training epochs</td>
              </tr>
            </tbody>
          </table>
        </table-wrap>
        <p>3.2.2. Tone and Readability</p>
        <p>The tone of an annual report corresponds to its textual sentiment features. In this study, we consider two indicators: the overall tone of the report and the tone of the MD&amp;A section, both obtained from the CNRDS Annual Report Sentiment Database (www.cnrds.com). Consistent with the method proposed by Zeng <italic>et</italic><italic>al.</italic> [<xref ref-type="bibr" rid="B23">23</xref>], the tone value is calculated using the Loughran-McDonald Financial Sentiment Dictionary. Specifically, as shown in Equation (1), the tone is determined based on the counts of positive words (<italic>POSword</italic>) and negative words (<italic>NEGword</italic>) within the text.</p>
        <disp-formula id="FD1">
          <label>(1)</label>
          <mml:math display="inline">
            <mml:mrow>
              <mml:mi>T</mml:mi>
              <mml:mi>o</mml:mi>
              <mml:mi>n</mml:mi>
              <mml:mi>e</mml:mi>
              <mml:mo>=</mml:mo>
              <mml:mfrac>
                <mml:mrow>
                  <mml:mi>P</mml:mi>
                  <mml:mi>O</mml:mi>
                  <mml:mi>S</mml:mi>
                  <mml:mi>w</mml:mi>
                  <mml:mi>o</mml:mi>
                  <mml:mi>r</mml:mi>
                  <mml:mi>d</mml:mi>
                  <mml:mo>−</mml:mo>
                  <mml:mi>N</mml:mi>
                  <mml:mi>E</mml:mi>
                  <mml:mi>G</mml:mi>
                  <mml:mi>w</mml:mi>
                  <mml:mi>o</mml:mi>
                  <mml:mi>r</mml:mi>
                  <mml:mi>d</mml:mi>
                </mml:mrow>
                <mml:mrow>
                  <mml:mi>P</mml:mi>
                  <mml:mi>O</mml:mi>
                  <mml:mi>S</mml:mi>
                  <mml:mi>w</mml:mi>
                  <mml:mi>o</mml:mi>
                  <mml:mi>r</mml:mi>
                  <mml:mi>d</mml:mi>
                  <mml:mo>+</mml:mo>
                  <mml:mi>N</mml:mi>
                  <mml:mi>E</mml:mi>
                  <mml:mi>G</mml:mi>
                  <mml:mi>w</mml:mi>
                  <mml:mi>o</mml:mi>
                  <mml:mi>r</mml:mi>
                  <mml:mi>d</mml:mi>
                </mml:mrow>
              </mml:mfrac>
            </mml:mrow>
          </mml:math>
        </disp-formula>
        <p>Readability, a linguistic feature that reflects text complexity, is typically inversely related to text length and the frequency of technical terms. While financial research uses various methods to assess readability, we simplify measurement by using the MD&amp;A file size as a proxy, since text length correlates closely with cognitive load—longer texts require greater processing effort and thus indicate higher reading difficulty. In financial reports, which follow standardized formats, excessive length often signals structural complexity or redundancy, implying lower readability. Given this study’s focus on textual content across a large sample of financial documents, file size is employed as a simple and accessible readability proxy.</p>
        <p>3.2.3. Content Similarity</p>
        <p>Content similarity measures the degree of change in annual reports across consecutive fiscal years and is typically calculated using cosine similarity (see Equation (2)). Higher similarity indicates fewer new disclosures [<xref ref-type="bibr" rid="B24">24</xref>], a pattern that previous research has associated with increased fraud risk (e.g., Qian &amp; Zhu [<xref ref-type="bibr" rid="B25">25</xref>], 2020). In this study, cosine similarity is computed using document vectors generated by Doc2Vec, a widely used technique that effectively captures contextual information for representing texts in similarity analyses. In Equation (2), <inline-formula><mml:math display="inline"><mml:mrow><mml:msub><mml:mi> V </mml:mi><mml:mi> t </mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> represents the text vector for the annual report in year <italic>t</italic>. </p>
        <disp-formula id="FD2">
          <label>(2)</label>
          <mml:math display="inline">
            <mml:mrow>
              <mml:mtext>Similarity</mml:mtext>
              <mml:mrow>
                <mml:mo>(</mml:mo>
                <mml:mrow>
                  <mml:msub>
                    <mml:mi>V</mml:mi>
                    <mml:mi>t</mml:mi>
                  </mml:msub>
                  <mml:mo>,</mml:mo>
                  <mml:msub>
                    <mml:mi>V</mml:mi>
                    <mml:mrow>
                      <mml:mi>t</mml:mi>
                      <mml:mo>−</mml:mo>
                      <mml:mn>1</mml:mn>
                    </mml:mrow>
                  </mml:msub>
                </mml:mrow>
                <mml:mo>)</mml:mo>
              </mml:mrow>
              <mml:mo>=</mml:mo>
              <mml:mfrac>
                <mml:mrow>
                  <mml:msub>
                    <mml:mi>V</mml:mi>
                    <mml:mi>t</mml:mi>
                  </mml:msub>
                  <mml:mo>⋅</mml:mo>
                  <mml:msub>
                    <mml:mi>V</mml:mi>
                    <mml:mrow>
                      <mml:mi>t</mml:mi>
                      <mml:mo>−</mml:mo>
                      <mml:mn>1</mml:mn>
                    </mml:mrow>
                  </mml:msub>
                </mml:mrow>
                <mml:mrow>
                  <mml:mo>
                  </mml:mo>
                  <mml:mrow>
                    <mml:mo>‖</mml:mo>
                    <mml:mrow>
                      <mml:msub>
                        <mml:mi>V</mml:mi>
                        <mml:mi>t</mml:mi>
                      </mml:msub>
                    </mml:mrow>
                    <mml:mo>‖</mml:mo>
                  </mml:mrow>
                  <mml:mo>⋅</mml:mo>
                  <mml:mrow>
                    <mml:mo>‖</mml:mo>
                    <mml:mrow>
                      <mml:msub>
                        <mml:mi>V</mml:mi>
                        <mml:mrow>
                          <mml:mi>t</mml:mi>
                          <mml:mo>−</mml:mo>
                          <mml:mn>1</mml:mn>
                        </mml:mrow>
                      </mml:msub>
                    </mml:mrow>
                    <mml:mo>‖</mml:mo>
                  </mml:mrow>
                </mml:mrow>
              </mml:mfrac>
            </mml:mrow>
          </mml:math>
        </disp-formula>
      </sec>
      <sec id="sec3dot3">
        <title>3.3. Research Models</title>
        <p>3.3.1. Ensemble Tree-Based ML Algorithms</p>
        <p>In this study, financial fraud detection models are constructed using logistic regression (LR), XGBoost, LightGBM, and CatBoost. Logistic regression serves as the baseline model due to its simplicity and interpretability; however, it is limited in capturing complex, nonlinear relationships. In addition, the performance of gradient-boosted decision tree (GBDT) algorithms is evaluated. As noted by Chen &amp; Guestrin [<xref ref-type="bibr" rid="B26">26</xref>], GBDT integrates a set of weak learners through iterative optimization to minimize loss functions. Previous studies have demonstrated that tree-based ensemble methods consistently outperform deep neural networks on structured tabular data [<xref ref-type="bibr" rid="B22">22</xref>], primarily due to their ability to model nonlinear relationships, robustness during training, and adaptability in terms of parameter sensitivity, computational efficiency, and handling of missing values.</p>
        <p>The GBDT variants employed in this study each offer distinct advantages: XGBoost enhances generalization and computational speed through regularization and parallelization; LightGBM reduces prediction error using histogram-based algorithms and leaf-wise tree growth; and CatBoost offers robust handling of categorical features while delivering high precision and stability.</p>
        <p>3.3.2. Interpretable ML with SHAP</p>
        <p>For the financial fraud detection model built using ensemble tree algorithms, we use the highly interpretable SHAP method to analyze the importance and mechanisms of various features. Rooted in Shapley value theory from game theory [<xref ref-type="bibr" rid="B27">27</xref>], SHAP decomposes model predictions into weighted contributions from input features, thereby enhancing the interpretability of complex machine learning models. Under the SHAP framework, the prediction output <inline-formula><mml:math display="inline"><mml:mrow><mml:msub><mml:mi> y </mml:mi><mml:mi> i </mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> for a given sample <inline-formula><mml:math display="inline"><mml:mrow><mml:msub><mml:mi> x </mml:mi><mml:mi> i </mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> can be expressed as a linear combination of all feature contributions, as shown in Equation (3): </p>
        <disp-formula id="FD3">
          <label>(3)</label>
          <mml:math display="inline">
            <mml:mrow>
              <mml:msub>
                <mml:mi>y</mml:mi>
                <mml:mi>i</mml:mi>
              </mml:msub>
              <mml:mo>=</mml:mo>
              <mml:msub>
                <mml:mi>y</mml:mi>
                <mml:mrow>
                  <mml:mi>b</mml:mi>
                  <mml:mi>a</mml:mi>
                  <mml:mi>s</mml:mi>
                  <mml:mi>e</mml:mi>
                </mml:mrow>
              </mml:msub>
              <mml:mo>+</mml:mo>
              <mml:mstyle displaystyle="true">
                <mml:msubsup>
                  <mml:mo>∑</mml:mo>
                  <mml:mrow>
                    <mml:mi>j</mml:mi>
                    <mml:mo>=</mml:mo>
                    <mml:mn>1</mml:mn>
                  </mml:mrow>
                  <mml:mi>M</mml:mi>
                </mml:msubsup>
                <mml:mrow>
                  <mml:mi>f</mml:mi>
                  <mml:mrow>
                    <mml:mo>(</mml:mo>
                    <mml:mrow>
                      <mml:msub>
                        <mml:mi>x</mml:mi>
                        <mml:mrow>
                          <mml:mi>i</mml:mi>
                          <mml:mi>j</mml:mi>
                        </mml:mrow>
                      </mml:msub>
                    </mml:mrow>
                    <mml:mo>)</mml:mo>
                  </mml:mrow>
                </mml:mrow>
              </mml:mstyle>
            </mml:mrow>
          </mml:math>
        </disp-formula>
        <p>Here, <inline-formula><mml:math display="inline"><mml:mrow><mml:msub><mml:mi> y </mml:mi><mml:mrow><mml:mi> b </mml:mi><mml:mi> a </mml:mi><mml:mi> s </mml:mi><mml:mi> e </mml:mi></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula> represents the global mean of the target variable (baseline value), <inline-formula><mml:math display="inline"><mml:mrow><mml:msub><mml:mi> x </mml:mi><mml:mrow><mml:mi> i </mml:mi><mml:mi> j </mml:mi></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula> denotes the <italic>j</italic>-th feature of sample <inline-formula><mml:math display="inline"><mml:mrow><mml:msub><mml:mi> x </mml:mi><mml:mi> i </mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> , and <inline-formula><mml:math display="inline"><mml:mrow><mml:mi> f </mml:mi><mml:mrow><mml:mo> ( </mml:mo><mml:mrow><mml:msub><mml:mi> x </mml:mi><mml:mrow><mml:mi> i </mml:mi><mml:mi> j </mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo> ) </mml:mo></mml:mrow></mml:mrow></mml:math></inline-formula> corresponds to the SHAP value of the feature, reflecting its marginal contribution to the prediction. A positive SHAP value indicates a promotive effect on fraud risk, while a negative value suggests an inhibitory effect. By quantifying these directional impacts, SHAP elucidates the model’s decision logic. As illustrated in <xref ref-type="fig" rid="fig2">Figure 2</xref>, the baseline value <inline-formula><mml:math><mml:mrow><mml:msub><mml:mi> y </mml:mi><mml:mrow><mml:mi> b </mml:mi><mml:mi> a </mml:mi><mml:mi> s </mml:mi><mml:mi> e </mml:mi></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula> serves as the initial reference (e.g., average predicted probability). Arrows depict deviations from this baseline, e.g., feature <inline-formula><mml:math display="inline"><mml:mrow><mml:msub><mml:mi> x </mml:mi><mml:mi> i </mml:mi></mml:msub><mml:msub><mml:mrow></mml:mrow><mml:mn> 1 </mml:mn></mml:msub></mml:mrow></mml:math></inline-formula> exerts a positive contribution (<inline-formula><mml:math display="inline"><mml:mrow><mml:mi> f </mml:mi><mml:mrow><mml:mo> ( </mml:mo><mml:mrow><mml:msub><mml:mi> x </mml:mi><mml:mrow><mml:mi> i </mml:mi><mml:mn> 1 </mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mo> ) </mml:mo></mml:mrow><mml:mo> &gt; </mml:mo><mml:mn> 0 </mml:mn></mml:mrow></mml:math></inline-formula> ), elevating the prediction above the baseline, while <inline-formula><mml:math display="inline"><mml:mrow><mml:msub><mml:mi> x </mml:mi><mml:mi> i </mml:mi></mml:msub><mml:msub><mml:mrow></mml:mrow><mml:mn> 3 </mml:mn></mml:msub></mml:mrow></mml:math></inline-formula> exhibits a negative contribution, reducing the prediction. The final prediction results from the algebraic sum of all feature contributions. SHAP provides post-hoc explanations applicable to various machine learning algorithms, particularly excelling in interpreting nonlinear decision processes of tree-based ensemble models.</p>
        <fig id="fig2">
          <label>Figure 2</label>
          <graphic xlink:href="https://html.scirp.org/file/1733376-rId41.jpeg?20251219032619" />
        </fig>
        <p><bold>Figure 2.</bold> SHAP Feature attribution diagram.</p>
        <p>The target variable in this study is corporate fraud risk, while the predictors include financial indicators, organizational structure features, and dimensions derived from annual reports. SHAP analysis is employed to quantify the marginal contributions of each predictor to fraud risk, identify early warning indicators significantly associated with fraudulent activities, and provide deeper insights into the underlying mechanisms of corporate fraud.</p>
      </sec>
    </sec>
    <sec id="sec4">
      <title>4. Experiments and Results</title>
      <sec id="sec4dot1">
        <title>4.1. Dataset Construction</title>
        <p>4.1.1. Data Samples</p>
        <p>This study utilizes violation records from the CSMAR database, focusing on firms listed on the Shanghai and Shenzhen Stock Exchanges between 2015 and 2020. Special attention is given to five prevalent forms of financial fraud: fictitious profits, inflated assets, false entries, major omissions, and disclosure misstatements. To ensure sample consistency and avoid confounding factors, financial firms are excluded due to their unique operational characteristics and higher financial risk. The resulting dataset comprises 1226 firms involved in financial misconduct, accounting for a total of 2652 fraud cases.</p>
        <p>For the control group, non-fraudulent firms are selected from the CNRDS ESG-R database, which has provided ESG (Environmental, Social, and Governance) ratings for all Chinese A-share listed companies since 2007. Employing the ESG ratings as the criterion for non-fraud control firms has clear theoretical and empirical support. The governance dimension comprehensively reflects a company’s standards in risk management, internal control, financial reporting quality, and information transparency. A higher governance score typically indicates a sound governance structure and effective oversight mechanisms, thereby reducing the likelihood of fraud. Cohen <italic>et</italic><italic>al</italic><italic>.</italic> [<xref ref-type="bibr" rid="B28">28</xref>] found that governance quality directly affects the reliability of financial reporting; Dechow <italic>et</italic><italic>al.</italic> [<xref ref-type="bibr" rid="B29">29</xref>] showed that weak governance significantly increases fraud risk; and Tamimi and Sebastianelli [<xref ref-type="bibr" rid="B30">30</xref>] further demonstrated that higher scores under the ESG framework are associated with greater information transparency and lower fraud risk. Governance-related ESG indicators, including risk management capability, financial reporting quality, and disclosure transparency, have been proven to be relevant for fraud detection [<xref ref-type="bibr" rid="B31">31</xref>]. </p>
        <p>In the study, non-fraudulent firms are matched to fraudulent firms in a 1:1 ratio by year and industry. Additional criteria require non-fraudulent samples to have higher governance scores, no violation records from 2015 to 2020, and not be under ST (Special Treatment) status. These systematic matches yield 2652 non-fraudulent firms, thereby providing a robust foundation for comparative analysis.</p>
        <p>4.1.2. Feature Variables</p>
        <p><bold>1)</bold><bold>Financial</bold><bold>Indicators</bold></p>
        <p>Nineteen financial indicators across five dimensions—solvency, profitability, growth potential, operational efficiency, and cash flow capacity—are used (see <bold>Table 2</bold>). Solvency indicators assess default risk, profitability measures earning ability, growth reflects expansion, operational efficiency captures asset utilization, and cash flow evaluates management of liquidity. Unusual variations in these indicators, such as sharp shifts in profitability or cash flow, can signal potential financial fraud.</p>
        <p><bold>Table 2.</bold> Financial indicators (19 Items).</p>
        <table-wrap id="tbl2">
          <label>Table 2</label>
          <table>
            <tbody>
              <tr>
                <td>
                  <bold>Dimensions</bold>
                </td>
                <td>
                  <bold>Indicators</bold>
                </td>
                <td>
                  <bold>Definition</bold>
                  <bold>and</bold>
                  <bold>Operation</bold>
                </td>
              </tr>
              <tr>
                <td rowspan="5">Solvency</td>
                <td>
                  <italic>Aslbrt</italic>
                </td>
                <td>Asset Liability Ratio; Total Liabilities/Total Assets</td>
              </tr>
              <tr>
                <td>
                  <italic>Curtrt</italic>
                </td>
                <td>Flow Rate Ratio; Total Liability Ratio/Total Flow Asset</td>
              </tr>
              <tr>
                <td>
                  <italic>Qikrt</italic>
                </td>
                <td>Quick Ratio; (Current Assets - Inventory)/Current Liabilities</td>
              </tr>
              <tr>
                <td>
                  <italic>Equrt</italic>
                </td>
                <td>Equity Ratio; Total Liabilities/Total Shareholders’ Equity</td>
              </tr>
              <tr>
                <td>
                  <italic>Pmcptdbrt</italic>
                </td>
                <td>Long-term Debt Ratio; Total Non-flowing Debt/(Total Shareholders’ Equity + Total Non-flowing Liabilities)</td>
              </tr>
              <tr>
                <td rowspan="4">Profitability</td>
                <td>
                  <italic>Roe_1</italic>
                </td>
                <td>Return on Equity (ROE); Net Profit/Total Shareholders’ Equity</td>
              </tr>
              <tr>
                <td>
                  <italic>Salnpm</italic>
                </td>
                <td>Net Profit margin; Net Profit/Revenue</td>
              </tr>
              <tr>
                <td>
                  <italic>Salgm</italic>
                </td>
                <td>Gross Profit Margin; Gross Profit/Revenue</td>
              </tr>
              <tr>
                <td>
                  <italic>Salpm</italic>
                </td>
                <td>Profit Margin; Profit/Revenue</td>
              </tr>
              <tr>
                <td rowspan="3">Growth potential</td>
                <td>
                  <italic>Atrt</italic>
                </td>
                <td>Total Assets Growth Rate; (Current Period Adjusted Figure - Prior Year Same Period Adjusted Figure)/Prior Year Same Period Adjusted Figure for ABS</td>
              </tr>
              <tr>
                <td>
                  <italic>Opicrt</italic>
                </td>
                <td>Operating Revenue Growth Rate; (Current Period Adjusted Figure - Prior Year Same Period Adjusted Figure)/Prior Year Same Period Adjusted Figure for ABS</td>
              </tr>
              <tr>
                <td>
                  <italic>Oirt</italic>
                </td>
                <td>Operating Profit Growth Rate, (Current Period Adjusted Figure - Prior Year Same Period Adjusted Figure)/Prior Year Same Period Adjusted Figure for ABS</td>
              </tr>
              <tr>
                <td rowspan="4">Operational Efficiency</td>
                <td>
                  <italic>Actrcbto</italic>
                </td>
                <td>Accounts Receivable Turnover; Revenue/((Beginning Net Accounts Receivable + Ending Net Accounts Receivable)/2)</td>
              </tr>
              <tr>
                <td>
                  <italic>Fxastto</italic>
                </td>
                <td>Fixed Asset Turnover; Total Revenue/((Beginning Fixed Assets + Ending Fixed Assets)/2)</td>
              </tr>
              <tr>
                <td>
                  <italic>Totastto</italic>
                </td>
                <td>Total Asset Turnover; Total Revenue/((Beginning Total Assets + Ending Total Assets)/2)</td>
              </tr>
              <tr>
                <td>
                  <italic>Actpayto</italic>
                </td>
                <td>Accounts Payable Turnover; Cost of Goods Sold (COGS)/((Beginning Accounts Payable + Ending Accounts Payable)/2)</td>
              </tr>
              <tr>
                <td rowspan="3">Cash Flow Capacity</td>
                <td>
                  <italic>Opncf_rev</italic>
                </td>
                <td>Net operating cash flow/Total Revenue</td>
              </tr>
              <tr>
                <td>
                  <italic>Opncfrt</italic>
                </td>
                <td>Percentage of net cash flow from operating activities; Net cash flow from operating activities/(Net cash flow from operating activities + Net cash flow from investing activities + Net cash flow from financing activities)</td>
              </tr>
              <tr>
                <td>
                  <italic>Csopindex</italic>
                </td>
                <td>Cash flow to sales ratio; Net cash flow from operating activities/Cash from operations</td>
              </tr>
            </tbody>
          </table>
        </table-wrap>
        <p><bold>2)</bold><bold>Organization</bold><bold>Structure</bold><bold>Features</bold></p>
        <p>Organizational structure related features are classified into two main categories: ownership concentration and top management characteristics (see <bold>Table 3</bold>). Ownership metrics evaluate the proportion of shares held by major shareholders, with particular emphasis on the concentration among the top five shareholders, a key indicator for fraud detection in China (Qian &amp; Luo [<xref ref-type="bibr" rid="B7">7</xref>]). Additionally, a higher shareholding by the largest shareholder serves as an effective deterrent to fraud, as it strengthens shareholder oversight and control [<xref ref-type="bibr" rid="B32">32</xref>]. Top management characteristics encompass variables such as dual roles, management shareholdings, and other related attributes.</p>
        <p><bold>Table 3</bold><bold>.</bold> Features related to corporate organizational structure (11 Items).</p>
        <table-wrap id="tbl3">
          <label>Table 3</label>
          <table>
            <tbody>
              <tr>
                <td>
                  <bold>Dimension</bold>
                </td>
                <td>
                  <bold>Indicators</bold>
                </td>
                <td>
                  <bold>Definition</bold>
                  <bold>and</bold>
                  <bold>Operation</bold>
                </td>
              </tr>
              <tr>
                <td rowspan="5">Ownership Concentration</td>
                <td>
                  <italic>ShrHolder1</italic>
                </td>
                <td>First Largest Shareholder Ownership Ratio</td>
              </tr>
              <tr>
                <td>
                  <italic>ShrHolder3</italic>
                </td>
                <td>Sum of Ownership Ratios of Top Three Shareholders</td>
              </tr>
              <tr>
                <td>
                  <italic>ShrHolder5</italic>
                </td>
                <td>Sum of Ownership Ratios of Top Five Shareholders</td>
              </tr>
              <tr>
                <td>
                  <italic>ShrHolder10</italic>
                </td>
                <td>Sum of Ownership Ratios of Top Ten Shareholders</td>
              </tr>
              <tr>
                <td>
                  <italic>StOwRt</italic>
                </td>
                <td>State-Owned Share Ratio</td>
              </tr>
              <tr>
                <td rowspan="6">Top Management Characteristics</td>
                <td>
                  <italic>Cmceo_Dum</italic>
                </td>
                <td>Whether serving as both Chairman and CEO (1 = Yes, 0 = No)</td>
              </tr>
              <tr>
                <td>
                  <italic>Cmgm_Dum</italic>
                </td>
                <td>Whether serving as both Chairman and General Manager (1 = Yes, 0 = No)</td>
              </tr>
              <tr>
                <td>
                  <italic>MShrRat</italic>
                </td>
                <td>Management Ownership Ratio: Percentage of company shares held by management</td>
              </tr>
              <tr>
                <td>
                  <italic>BShrRat</italic>
                </td>
                <td>Board of Directors Ownership Ratio: Percentage of company shares held by all board members</td>
              </tr>
              <tr>
                <td>
                  <italic>SShrRat</italic>
                </td>
                <td>Supervisory Board Ownership Ratio: Percentage of company shares held by all supervisory board members</td>
              </tr>
              <tr>
                <td>
                  <italic>InDrcRat</italic>
                </td>
                <td>Proportion of Independent Directors: Ratio of independent directors to the total number of directors</td>
              </tr>
            </tbody>
          </table>
        </table-wrap>
        <p><bold>3)</bold><bold>Annual</bold><bold>Report</bold><bold>Characteristics</bold></p>
        <p>Using the method from Section 3.2, we extract semantic features, tone, readability, and content similarity from the MD&amp;A (see <bold>Table 4</bold>).</p>
        <p><bold>Table 4.</bold> Features related to corporate annual report.</p>
        <table-wrap id="tbl4">
          <label>Table 4</label>
          <table>
            <tbody>
              <tr>
                <td>
                  <bold>Dimension</bold>
                </td>
                <td>
                  <bold>Indicators</bold>
                </td>
                <td>
                  <bold>Definition</bold>
                  <bold>and</bold>
                  <bold>Operation</bold>
                </td>
              </tr>
              <tr>
                <td rowspan="2">Tone</td>
                <td>
                  <italic>LM_tone2</italic>
                </td>
                <td>Annual report tone value</td>
              </tr>
              <tr>
                <td>
                  <italic>LM_tone</italic>
                </td>
                <td>MD&amp;A section tone value</td>
              </tr>
              <tr>
                <td>Readability</td>
                <td>
                  <italic>FileSize</italic>
                </td>
                <td>MD&amp;A text file size</td>
              </tr>
              <tr>
                <td>Similarity</td>
                <td>
                  <italic>Similarity</italic>
                </td>
                <td>Adjacent-year MD&amp;A content similarity</td>
              </tr>
              <tr>
                <td>Semantic Features</td>
                <td>
                  <italic>TextFeatrue</italic>
                  <italic>
                    <sub>1</sub>
                  </italic>
                  <italic>-</italic>
                  <italic>TextFeature</italic>
                  <italic>
                    <sub>n</sub>
                  </italic>
                </td>
                <td>MD&amp;A semantic feature vector (n-dimensional)</td>
              </tr>
            </tbody>
          </table>
        </table-wrap>
      </sec>
      <sec id="sec4dot2">
        <title>4.2. Experimental Design</title>
        <p>Following the overall research framework (see <xref ref-type="fig" rid="fig1">Figure 1</xref>), three main experiments were conducted. First, semantic features were extracted from MD&amp;A texts using BERT-based classification models across nine thematic modules to predict corporate fraud risk. The resulting theme-specific semantic vectors were then subjected to dimensionality reduction, yielding compact representations for the entire MD&amp;A section (denoted as <italic>TextFeature</italic><italic><sub>i</sub></italic>, where <inline-formula><mml:math display="inline"><mml:mrow><mml:mi> i </mml:mi><mml:mo> = </mml:mo><mml:mn> 1 </mml:mn><mml:mo> ⋯ </mml:mo><mml:mi> n </mml:mi></mml:mrow></mml:math></inline-formula> ).</p>
        <p>In the second experiment, feature fusion was performed by combining these semantic vectors with financial indicators, organizational structure metrics, and additional features related to tone, readability, and content similarity from annual reports. An ensemble tree-based model was optimized using grid search to achieve the best predictive performance.</p>
        <p>The third experiment employed the interpretable SHAP method on the best-performing model to clarify the contribution of each feature to fraud risk, thereby identifying key indicators of fraudulent behavior. Model performance was evaluated using accuracy, F1-score, and ROC-AUC.</p>
      </sec>
      <sec id="sec4dot3">
        <title>4.3. Results and Analysis</title>
        <p>4.3.1. Semantic Features of MD&amp;A</p>
        <p><bold>Table 5</bold><bold>.</bold> MD&amp;A content-based corporate fraud detection model performance.</p>
        <table-wrap id="tbl5">
          <label>Table 5</label>
          <table>
            <tbody>
              <tr>
                <td>
                  <bold>MD&amp;A-based</bold>
                  <bold>thematic</bold>
                  <bold>modules</bold>
                </td>
                <td>
                  <bold>Accuracy</bold>
                </td>
                <td>
                  <bold>F1</bold>
                </td>
                <td>
                  <bold>ROC-AUC</bold>
                </td>
              </tr>
              <tr>
                <td>
                  <bold>1</bold>
                  <italic>
                    <bold>Overview</bold>
                  </italic>
                </td>
                <td>
                  <bold>0.654</bold>
                </td>
                <td>
                  <bold>0.654</bold>
                </td>
                <td>
                  <bold>0.695</bold>
                </td>
              </tr>
              <tr>
                <td>
                  2
                  <italic>Core</italic>
                  <italic>business</italic>
                  <italic>analysis</italic>
                </td>
                <td>0.615</td>
                <td>0.611</td>
                <td>0.659</td>
              </tr>
              <tr>
                <td>
                  3
                  <italic>Non-core</italic>
                  <italic>business</italic>
                  <italic>analysis</italic>
                </td>
                <td>0.603</td>
                <td>0.603</td>
                <td>0.639</td>
              </tr>
              <tr>
                <td>
                  4
                  <italic>Asset</italic>
                  <italic>and</italic>
                  <italic>liability</italic>
                  <italic>status</italic>
                </td>
                <td>0.619</td>
                <td>0.616</td>
                <td>0.658</td>
              </tr>
              <tr>
                <td>
                  5
                  <italic>Investment</italic>
                  <italic>activities</italic>
                </td>
                <td>0.589</td>
                <td>0.551</td>
                <td>0.632</td>
              </tr>
              <tr>
                <td>
                  <bold>6</bold>
                  <italic>
                    <bold>Significant</bold>
                  </italic>
                  <italic>
                    <bold>asset</bold>
                  </italic>
                  <italic>
                    <bold>&amp;</bold>
                  </italic>
                  <italic>
                    <bold>equity</bold>
                  </italic>
                  <italic>
                    <bold>transactions</bold>
                  </italic>
                </td>
                <td>
                  <bold>0.634</bold>
                </td>
                <td>
                  <bold>0.631</bold>
                </td>
                <td>
                  <bold>0.697</bold>
                </td>
              </tr>
              <tr>
                <td>
                  7
                  <italic>Analysis</italic>
                  <italic>of</italic>
                  <italic>major</italic>
                  <italic>subsidiaries</italic>
                  <italic>and</italic>
                  <italic>affiliates</italic>
                </td>
                <td>0.619</td>
                <td>0.613</td>
                <td>0.696</td>
              </tr>
              <tr>
                <td>
                  8
                  <italic>Structured</italic>
                  <italic>entities</italic>
                </td>
                <td>0.613</td>
                <td>0.613</td>
                <td>0.637</td>
              </tr>
              <tr>
                <td>
                  <bold>9</bold>
                  <italic>
                    <bold>Future</bold>
                  </italic>
                  <italic>
                    <bold>outlook</bold>
                  </italic>
                </td>
                <td>
                  <bold>0.626</bold>
                </td>
                <td>
                  <bold>0.626</bold>
                </td>
                <td>
                  <bold>0.670</bold>
                </td>
              </tr>
            </tbody>
          </table>
        </table-wrap>
        <p><bold>Table 5</bold> presents the fraud prediction performance of BERT models across nine MD&amp;A thematic modules. The <italic>Significant</italic><italic>asset</italic><italic>&amp;</italic><italic>equity</italic><italic>transactions</italic> module achieves the highest predictive performance (ROC-AUC = 0.697), followed by <italic>Overview</italic> (ROC-AUC = 0.695) and <italic>Future</italic><italic>outlook</italic> (ROC-AUC = 0.670). This performance hierarchy can be explained by the informational value of each section: 1) <italic>Significant</italic><italic>asset</italic><italic>and</italic><italic>equity</italic><italic>transactions</italic> documents substantial transactions with direct financial implications; 2) The <italic>Overview</italic> synthesizes the company’s comprehensive status, effectively encapsulating management’s core messaging; and 3) <italic>Future</italic><italic>Outlook</italic> conveys development expectations and profitability projections, offering critical insights into management intentions and potential fraud motives. Conversely, the <italic>Investment</italic><italic>activities</italic> module demonstrates relatively weaker predictive power (ROC-AUC = 0.632), suggesting its limited relevance to fraud risk assessment. </p>
        <p>Semantic vectors (CLS outputs) from the three top modules—<italic>Overview</italic>, <italic>Significant</italic><italic>Asset</italic><italic>and</italic><italic>Equity</italic><italic>Transactions</italic>, and <italic>Future</italic><italic>Outlook</italic>—were dimensionally reduced to remove redundancy and emphasize key information. These reduced vectors were evaluated with XGBoost, LightGBM, CatBoost, and logistic regression (LR); CatBoost performed best and was selected for the final model. Finally, the reduced semantic vector from the best-performing module (<italic>Significant</italic><italic>Asset</italic><italic>and</italic><italic>Equity</italic><italic>Transactions</italic>) was concatenated with other annual report features and corporate financial and organizational structure indicators to form a multi-source feature vector, which was fed into CatBoost (see <bold>Table 6</bold>).</p>
        <p>4.3.2. Integrated-Data Financial Fraud Detection Model</p>
        <p><bold>Table 6</bold> presents the results of corporate fraud prediction using three ensemble tree-based algorithms and LR, which incorporate features from multiple levels, including financial and organizational structure indicators, semantic features derived from MD&amp;A, as well as tone, similarity, and readability. The ensemble tree algorithms consistently outperform the baseline LR model across all evaluation metrics. CatBoost achieves the highest performance, with ROC-AUC scores of 0.859 for models built on the <italic>Significant</italic><italic>Asset</italic><italic>&amp;</italic><italic>Equity</italic><italic>Transactions</italic> module. The corresponding dimension of semantic features for the module is 7, demonstrating that even low-dimensional semantic representations derived from the MD&amp;A section provide valuable information for detecting financial fraud.</p>
        <p><bold>Table 6.</bold> Performance of integrated-data fraud detection models. </p>
        <table-wrap id="tbl6">
          <label>Table 6</label>
          <table>
            <tbody>
              <tr>
                <td>
                  <bold>Algorithm</bold>
                </td>
                <td>
                  <bold>MD&amp;A-based</bold>
                  <bold>thematic</bold>
                  <bold>modules</bold>
                </td>
                <td>
                  <bold>Reduced</bold>
                  <bold>Semantic</bold>
                  <bold>features</bold>
                  <bold>dimension</bold>
                </td>
                <td>
                  <bold>Accuracy</bold>
                </td>
                <td>
                  <bold>F1</bold>
                </td>
                <td>
                  <bold>ROC-AUC</bold>
                </td>
              </tr>
              <tr>
                <td rowspan="3">CatBoost</td>
                <td>
                  1
                  <italic>Overview</italic>
                </td>
                <td>
                  <bold>3</bold>
                </td>
                <td>
                  <bold>0.786</bold>
                </td>
                <td>
                  <bold>0.780</bold>
                </td>
                <td>
                  <bold>0.858</bold>
                </td>
              </tr>
              <tr>
                <td>
                  6
                  <italic>Significant</italic>
                  <italic>asset</italic>
                  <italic>&amp;</italic>
                  <italic>equity</italic>
                  <italic>transactions</italic>
                </td>
                <td>
                  <bold>7</bold>
                </td>
                <td>
                  <bold>0.798</bold>
                </td>
                <td>
                  <bold>0.793</bold>
                </td>
                <td>
                  <bold>0.859</bold>
                </td>
              </tr>
              <tr>
                <td>
                  9
                  <italic>Future</italic>
                  <italic>outlook</italic>
                </td>
                <td>
                  <bold>6</bold>
                </td>
                <td>
                  <bold>0.784</bold>
                </td>
                <td>
                  <bold>0.778</bold>
                </td>
                <td>
                  <bold>0.857</bold>
                </td>
              </tr>
              <tr>
                <td rowspan="3">XGBoost</td>
                <td>
                  1
                  <italic>Overview</italic>
                </td>
                <td>
                  <bold>4</bold>
                </td>
                <td>
                  <bold>0.781</bold>
                </td>
                <td>
                  <bold>0.779</bold>
                </td>
                <td>
                  <bold>0.856</bold>
                </td>
              </tr>
              <tr>
                <td>
                  6
                  <italic>Significant</italic>
                  <italic>asset</italic>
                  <italic>&amp;</italic>
                  <italic>equity</italic>
                  <italic>transactions</italic>
                </td>
                <td>7</td>
                <td>0.784</td>
                <td>0.779</td>
                <td>0.850</td>
              </tr>
              <tr>
                <td>
                  9
                  <italic>Future</italic>
                  <italic>outlook</italic>
                </td>
                <td>4</td>
                <td>0.780</td>
                <td>0.775</td>
                <td>0.853</td>
              </tr>
              <tr>
                <td rowspan="3">LightGBM</td>
                <td>
                  1
                  <italic>Overview</italic>
                </td>
                <td>
                  <bold>3</bold>
                </td>
                <td>
                  <bold>0.786</bold>
                </td>
                <td>
                  <bold>0.782</bold>
                </td>
                <td>
                  <bold>0.852</bold>
                </td>
              </tr>
              <tr>
                <td>
                  6
                  <italic>Significant</italic>
                  <italic>asset</italic>
                  <italic>&amp;</italic>
                  <italic>equity</italic>
                  <italic>transactions</italic>
                </td>
                <td>5</td>
                <td>0.780</td>
                <td>0.773</td>
                <td>0.847</td>
              </tr>
              <tr>
                <td>
                  9
                  <italic>Future</italic>
                  <italic>outlook</italic>
                </td>
                <td>3</td>
                <td>0.774</td>
                <td>0.768</td>
                <td>0.849</td>
              </tr>
              <tr>
                <td rowspan="3">LR</td>
                <td>
                  1
                  <italic>Overview</italic>
                </td>
                <td>
                  <bold>8</bold>
                </td>
                <td>
                  <bold>0.677</bold>
                </td>
                <td>
                  <bold>0.678</bold>
                </td>
                <td>
                  <bold>0.743</bold>
                </td>
              </tr>
              <tr>
                <td>
                  6
                  <italic>Significant</italic>
                  <italic>asset</italic>
                  <italic>&amp;</italic>
                  <italic>equity</italic>
                  <italic>transactions</italic>
                </td>
                <td>10</td>
                <td>0.666</td>
                <td>0.667</td>
                <td>0.730</td>
              </tr>
              <tr>
                <td>
                  9
                  <italic>Future</italic>
                  <italic>outlook</italic>
                </td>
                <td>15</td>
                <td>0.677</td>
                <td>0.675</td>
                <td>0.734</td>
              </tr>
            </tbody>
          </table>
        </table-wrap>
        <p>4.3.3. Feature Importance Analysis</p>
        <p>We utilized the optimal CatBoost model for SHAP analysis, incorporating semantic features from the <italic>Significant</italic><italic>Asset</italic><italic>&amp;</italic><italic>Equity</italic><italic>Transactions</italic> module alongside 41 additional features. <xref ref-type="fig" rid="fig3">Figure 3</xref> presents the top 20 features ranked by SHAP values, where the X-axis represents feature importance and different colors distinguish between feature categories.</p>
        <fig id="fig3">
          <label>Figure 3</label>
          <graphic xlink:href="https://html.scirp.org/file/1733376-rId44.jpeg?20251219032622" />
        </fig>
        <p><bold>Figure 3.</bold> SHAP-based feature importance analysis (Top 20 features).</p>
        <p>Among the top 20 features, financial indicators constitute 50%, with organization structure and annual report features each accounting for 25%. Financial indicators have a median rank of 7, significantly higher than the other two categories, underscoring their dominant role in fraud detection—particularly metrics related to operational efficiency (e.g., accounts receivable turnover, <italic>Actrcbto</italic>; total asset turnover, <italic>Totastto</italic>) and profitability (e.g., return on equity, <italic>Roe_1</italic>).</p>
        <p>Within the realm of organization structure, the shareholding ratio of the largest shareholder (<italic>ShrHolder1</italic>) and the board’s shareholding ratio (<italic>BShrRat</italic>) are particularly noteworthy. The former reflects the concentration of corporate power, while the latter indicates both the alignment of interests between the board and the company and the board’s capacity for independent oversight and decision-making. These findings highlight the importance of power distribution and interest alignment in predicting corporate fraud risk.</p>
        <p>Among annual report features, similarity (<italic>Similariy</italic>), tone (<italic>LM_tone2</italic>), and two semantic indicators (<italic>TextFeature</italic><italic><sub>6</sub></italic> and <italic>TextFeature</italic><italic><sub>1</sub></italic>) rank among the most important variables. Similarity reveals year-over-year changes in report content, while tone indicates management’s confidence in performance and outlook. Although semantic features are less influential than tone or similarity, they still meaningfully aid in fraud detection, suggesting that annual reports encode information closely tied to corporate fraud.</p>
        <p>4.3.4. Analysis of Impact Factors</p>
        <p>The SHAP beeswarm plot (<xref ref-type="fig" rid="fig4">Figure 4</xref>) visually demonstrates the relationship between the top 20 features and financial fraud risk. The x-axis represents SHAP values, indicating both the magnitude and direction of each feature’s impact. Positive SHAP values suggest an increased fraud risk, while negative values indicate a decreased risk.</p>
        <fig id="fig4">
          <label>Figure 4</label>
          <graphic xlink:href="https://html.scirp.org/file/1733376-rId45.jpeg?20251219032623" />
        </fig>
        <p><bold>Figure 4.</bold> SHAP beeswarm plot of top 20 features.</p>
      </sec>
    </sec>
    <sec id="sec5">
      <title>5. Discussion</title>
      <sec id="sec5dot1">
        <title>5.1. Integrated Modeling and Methodological Advances</title>
        <p>This study demonstrates that integrating semantic extraction, data fusion, and multimodal feature incorporation can substantially improve corporate fraud prediction. The fusion-based model, which combines financial indicators, organizational structure variables, and annual report features, achieves a notable ROC-AUC of 0.859. While financial indicators remain the core predictors, the addition of organizational and textual information significantly enhances predictive accuracy and provides deeper insights into fraud detection.</p>
        <p>Notably, ensemble tree models exhibit a clear advantage over traditional logistic regression, thanks to their ability to capture the complex, nonlinear patterns inherent in real-world fraud data. Consistently attaining ROC-AUC scores above 0.85, these ensemble approaches provide a more robust basis for fraud detection in complex environments.</p>
        <p>The integration between textual and tabular data further strengthens model performance. Leveraging BERT for semantic extraction from MD&amp;A sections, followed by dimensionality reduction via PCA, enhances both computational efficiency and the model’s sensitivity to fraud-related signals. Analysis shows that disclosures in the <italic>Overview</italic>, <italic>Significant</italic><italic>asset</italic><italic>&amp;</italic><italic>equity</italic><italic>transactions</italic>, and <italic>Future</italic><italic>outlook</italic> sections of annual reports are particularly predictive of fraudulent activity, which suggests that companies engaged in fraud often focus manipulation on key narrative and transactional elements within their reporting.</p>
      </sec>
      <sec id="sec5dot2">
        <title>5.2. Feature Interpretation and Forensic Insights</title>
        <p>Beyond overall predictive success, model interpretability—enabled by SHAP analysis—provides important forensic insights into the relationship between individual features and fraud risk. Financial metrics stand out as primary indicators: reduced operating efficiency, profitability, and growth are all closely tied to a heightened risk of fraud. Regarding organizational structure, a higher concentration of shares held by major shareholders is negatively correlated with fraud risk, supporting the idea that concentrated ownership can deter fraudulent actions [<xref ref-type="bibr" rid="B33">33</xref>]. In contrast, board shareholding has a nuanced effect. Moderate ownership aligns directors’ interests with shareholders’, reducing agency problems and fraud risk. But excessive ownership can have the opposite effect: large stakes increase board control, weaken external oversight and the balancing role of independent directors, and raise information asymmetry and opportunities for manipulation. Directors with substantial holdings may conceal poor performance or manipulate financials to protect share prices or personal wealth. Empirical studies support this nonlinearity: Jensen and Meckling [<xref ref-type="bibr" rid="B34">34</xref>] argue managerial ownership reduces agency conflict only up to a point; Larcker <italic>et</italic><italic>al.</italic> [<xref ref-type="bibr" rid="B35">35</xref>] find fraud risk rises when insider ownership exceeds an optimal range; and Yang <italic>et</italic><italic>al.</italic> [<xref ref-type="bibr" rid="B36">36</xref>] report that excessive ownership alignment can introduce decision biases. Overall, board ownership and fraud risk appear to follow an inverted U-shaped relationship: moderate ownership improves governance, while excessive ownership concentrates power and increases fraud risk.</p>
        <p>Linguistic and behavioral signals embedded in annual reports also add another layer of valuable insight. Key features such as year-over-year content similarity, tone, semantics, and readability help surface indications of hidden or manipulative practices. Large changes in report content or a shift away from an optimistic tone often accompany increased fraud risk, perhaps as management attempts to obscure irregularities. Module-level analysis highlights the particular importance of the <italic>Overview</italic>, <italic>Significant</italic><italic>asset</italic><italic>&amp;</italic><italic>equity</italic><italic>transactions</italic> ,and <italic>Future</italic><italic>outlook</italic>, while variations in readability provide additional red flags. Overall, how management adjusts report language and complexity can reflect underlying intentions and psychological states. These patterns in annual report content—whether in narrative style, explanation, or structure—offer meaningful clues, allowing for earlier and more accurate identification of potential corporate fraud. </p>
      </sec>
      <sec id="sec5dot3">
        <title>5.3. Research Contributions and Implications</title>
        <p>This study presents new evidence supporting the effectiveness of machine learning ensemble tree algorithms in identifying corporate financial fraud. By employing an innovative method for semantic feature extraction, joined with feature dimension reduction, we modularized extensive annual report texts. This approach not only facilitated the seamless integration of textual and tabular information but also offered a practical solution for managing complex multimodal financial data.</p>
        <p>Such an approach has meaningful implications for practice: extracting module-level risk signals lets regulators and practitioners efficiently screen many disclosures and target fraud-related content. For example, if the <italic>Future</italic><italic>Outlook</italic> text closely repeats the prior year while financial analysis shows a sharp drop in profitability and cash flow, the system can auto-trigger a risk alert to prompt a focused investigation, enabling earlier detection and more precise allocation of regulatory resources. The study also extends risk identification to linguistic behavior, examining how tone, and readability reflect management’s psychological and behavioral tendencies during report preparation. Quantifying these linguistic features aids investors and regulators in assessing corporate transparency and integrity culture and provides a basis for regulators to refine disclosure-review standards, making oversight more scientific, dynamic, and evidence-based.</p>
      </sec>
    </sec>
    <sec id="sec6">
      <title>6. Limitations and Future Directions</title>
      <p>While this study provides insight into corporate fraud prediction, certain limitations remain that highlight potential areas for further investigation. The current model focuses solely on assessing the likelihood of fraudulent behavior, without distinguishing between specific types of fraud. However, in practice, corporate fraud takes diverse forms and exhibits considerable complexity. Future research could address this by developing automated methods to identify and classify different types of fraud, enabling more precise and actionable predictions.</p>
      <p>Another important aspect concerns the processing of annual report texts. Although the study employs modularization and semantic modeling to preserve critical information, capturing the full depth of semantic meaning remains a challenge. More advanced approaches, such as the use of attention mechanisms, may help pinpoint keywords and phrases that are particularly relevant for fraud detection, thereby refining the model’s sensitivity to subtle textual cues.</p>
      <p>Additionally, the analysis presented here is based on data from a single year, which may not fully capture the dynamic and cumulative nature of financial fraud. Many fraudulent activities unfold over longer periods, making the exploration of multi-year data and dynamic feature extraction an important direction for improving model accuracy and robustness.</p>
      <p>Addressing these limitations will contribute to the practical advancement of corporate fraud detection models. By expanding fraud type classification, enhancing text analysis methods, and incorporating temporal changes, future research can provide stronger tools for identifying and mitigating financial fraud risks.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <title>References</title>
      <ref id="B1">
        <label>1.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Wu, X. and Du, S. (2022) An Analysis on Financial Statement Fraud Detection for Chinese Listed Companies Using Deep Learning. <italic>IEEE</italic><italic>Access</italic>, 10, 22516-22532. https://doi.org/10.1109/access.2022.3153478 <pub-id pub-id-type="doi">10.1109/access.2022.3153478</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1109/access.2022.3153478">https://doi.org/10.1109/access.2022.3153478</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Wu, X.</string-name>
              <string-name>Du, S.</string-name>
            </person-group>
            <year>2022</year>
            <article-title>An Analysis on Financial Statement Fraud Detection for Chinese Listed Companies Using Deep Learning</article-title>
            <source>IEEE Access</source>
            <volume>10</volume>
            <pub-id pub-id-type="doi">10.1109/access.2022.3153478</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B2">
        <label>2.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Sun, J., Fujita, H., Chen, P. and Li, H. (2017) Dynamic Financial Distress Prediction with Concept Drift Based on Time Weighting Combined with Adaboost Support Vector Machine Ensemble. <italic>Knowledge</italic>- <italic>Based</italic><italic>Systems</italic>, 120, 4-14. https://doi.org/10.1016/j.knosys.2016.12.019 <pub-id pub-id-type="doi">10.1016/j.knosys.2016.12.019</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1016/j.knosys.2016.12.019">https://doi.org/10.1016/j.knosys.2016.12.019</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Sun, J.</string-name>
              <string-name>Fujita, H.</string-name>
              <string-name>Chen, P.</string-name>
              <string-name>Li, H.</string-name>
            </person-group>
            <year>2017</year>
            <article-title>Dynamic Financial Distress Prediction with Concept Drift Based on Time Weighting Combined with Adaboost Support Vector Machine Ensemble</article-title>
            <source>Knowledge-Based Systems</source>
            <volume>120</volume>
            <pub-id pub-id-type="doi">10.1016/j.knosys.2016.12.019</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B3">
        <label>3.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Fukas, P., Rebstadt, J., Menzel, L. and Thomas, O. (2022) Towards Explainable Artificial Intelligence in Financial Fraud Detection: Using Shapley Additive Explanations to Explore Feature Importance. In: Franch, X., Poels, G., Gailly, F. and Snoeck, M., Eds., <italic>Advanced Information Systems Engineering</italic>, Springer, 109-126. https://doi.org/10.1007/978-3-031-07472-1_7 <pub-id pub-id-type="doi">10.1007/978-3-031-07472-1_7</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1007/978-3-031-07472-1_7">https://doi.org/10.1007/978-3-031-07472-1_7</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Fukas, P.</string-name>
              <string-name>Rebstadt, J.</string-name>
              <string-name>Menzel, L.</string-name>
              <string-name>Thomas, O.</string-name>
              <string-name>Franch, X.</string-name>
              <string-name>Poels, G.</string-name>
              <string-name>Gailly, F.</string-name>
              <string-name>Snoeck, M.</string-name>
              <string-name>Engineering, S</string-name>
            </person-group>
            <year>2022</year>
            <article-title>Towards Explainable Artificial Intelligence in Financial Fraud Detection: Using Shapley Additive Explanations to Explore Feature Importance</article-title>
            <source>In: Franch</source>
            <volume>109</volume>
            <pub-id pub-id-type="doi">10.1007/978-3-031-07472-1_7</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B4">
        <label>4.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Ashtiani, M.N. and Raahemi, B. (2022) Intelligent Fraud Detection in Financial Statements Using Machine Learning and Data Mining: A Systematic Literature Review. <italic>IEEE</italic><italic>Access</italic>, 10, 72504-72525. https://doi.org/10.1109/access.2021.3096799 <pub-id pub-id-type="doi">10.1109/access.2021.3096799</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1109/access.2021.3096799">https://doi.org/10.1109/access.2021.3096799</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Ashtiani, M.N.</string-name>
              <string-name>Raahemi, B.</string-name>
            </person-group>
            <year>2022</year>
            <article-title>Intelligent Fraud Detection in Financial Statements Using Machine Learning and Data Mining: A Systematic Literature Review</article-title>
            <source>IEEE Access</source>
            <volume>10</volume>
            <pub-id pub-id-type="doi">10.1109/access.2021.3096799</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B5">
        <label>5.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Kirkos, E., Spathis, C. and Manolopoulos, Y. (2007) Data Mining Techniques for the Detection of Fraudulent Financial Statements. <italic>Expert</italic><italic>Systems</italic><italic>with</italic><italic>Applications</italic>, 32, 995-1003. https://doi.org/10.1016/j.eswa.2006.02.016 <pub-id pub-id-type="doi">10.1016/j.eswa.2006.02.016</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1016/j.eswa.2006.02.016">https://doi.org/10.1016/j.eswa.2006.02.016</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Kirkos, E.</string-name>
              <string-name>Spathis, C.</string-name>
              <string-name>Manolopoulos, Y.</string-name>
            </person-group>
            <year>2007</year>
            <article-title>Data Mining Techniques for the Detection of Fraudulent Financial Statements</article-title>
            <source>Expert Systems with Applications</source>
            <volume>32</volume>
            <pub-id pub-id-type="doi">10.1016/j.eswa.2006.02.016</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B6">
        <label>6.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Xiong, F.J. and Zhang, L.P. (2016) Risk Identification and Evidence Collection of Financial Fraud on the Listed Companies. <italic>Research on Economics and Management</italic>, 37, 138-144.</mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Xiong, F.J.</string-name>
              <string-name>Zhang, L.P.</string-name>
            </person-group>
            <year>2016</year>
            <article-title>Risk Identification and Evidence Collection of Financial Fraud on the Listed Companies</article-title>
            <source>Research on Economics and Management</source>
            <volume>37</volume>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B7">
        <label>7.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Qian, P. and Luo, M. (2015) Predicting Accounting Fraud in China. <italic>Accounting Re</italic>- <italic>search</italic>, 7, 18-25, 96.</mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Qian, P.</string-name>
              <string-name>Luo, M.</string-name>
            </person-group>
            <year>2015</year>
            <article-title>Predicting Accounting Fraud in China</article-title>
            <source>Accounting Re-search</source>
            <volume>7</volume>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B8">
        <label>8.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Liu, Y.Q., Wu, B. and Zhang, M. (2022) Financial Fraud Recognition Model and Ap-plication. <italic>Journal of Quantitative &amp; Technological Economics</italic>, 39, 152-175.</mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Liu, Y.Q.</string-name>
              <string-name>Wu, B.</string-name>
              <string-name>Zhang, M.</string-name>
            </person-group>
            <year>2022</year>
            <article-title>Financial Fraud Recognition Model and Ap-plication</article-title>
            <source>Journal of Quantitative &amp; Technological Economics</source>
            <volume>39</volume>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B9">
        <label>9.</label>
        <citation-alternatives>
          <mixed-citation publication-type="report">Tan, J.H. and Wang, X.Y. (2022) Corporate Fraud and Manipulation of Annual Report Text Information. <italic>China Soft Science</italic>, 3, 99-111.</mixed-citation>
          <element-citation publication-type="report">
            <person-group person-group-type="author">
              <string-name>Tan, J.H.</string-name>
              <string-name>Wang, X.Y.</string-name>
            </person-group>
            <year>2022</year>
            <article-title>Corporate Fraud and Manipulation of Annual Report Text Information</article-title>
            <source>China Soft Science</source>
            <volume>3</volume>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B10">
        <label>10.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Xu, C. and Zhang, Y.M. (2021) Can Managers’ Tone Management Predict the Financial Fraud Risk: Based on MD&amp;A Forward-looking Text Information. <italic>Science and Management</italic>, 41, 73-81.</mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Xu, C.</string-name>
              <string-name>Zhang, Y.M.</string-name>
            </person-group>
            <year>2021</year>
            <article-title>Can Managers’ Tone Management Predict the Financial Fraud Risk: Based on MD&amp;A Forward-looking Text Information</article-title>
            <source>Science and Management</source>
            <volume>41</volume>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B11">
        <label>11.</label>
        <citation-alternatives>
          <mixed-citation publication-type="report">Wang, K.M., Wang, H.J., Li, D.D. and Dai, X.Y. (2018) Complexity of Annual Report and Management Self-Interest: Empirical Evidence from Chinese Listed Firms. <italic>Journal of Management World</italic>, 34, 120-132, 194.</mixed-citation>
          <element-citation publication-type="report">
            <person-group person-group-type="author">
              <string-name>Wang, K.M.</string-name>
              <string-name>Wang, H.J.</string-name>
              <string-name>Li, D.D.</string-name>
              <string-name>Dai, X.Y.</string-name>
            </person-group>
            <year>2018</year>
            <article-title>Complexity of Annual Report and Management Self-Interest: Empirical Evidence from Chinese Listed Firms</article-title>
            <source>Journal of Management World</source>
            <volume>34</volume>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B12">
        <label>12.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Zhang, Y., Liu, T. and Li, W. (2024) Corporate Fraud Detection Based on Linguistic Readability Vector: Application to Financial Companies in China. <italic>International</italic><italic>Review</italic><italic>of</italic><italic>Financial</italic><italic>Analysis</italic>, 95, Article ID: 103405. https://doi.org/10.1016/j.irfa.2024.103405 <pub-id pub-id-type="doi">10.1016/j.irfa.2024.103405</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1016/j.irfa.2024.103405">https://doi.org/10.1016/j.irfa.2024.103405</ext-link></mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Zhang, Y.</string-name>
              <string-name>Liu, T.</string-name>
              <string-name>Li, W.</string-name>
            </person-group>
            <year>2024</year>
            <article-title>Corporate Fraud Detection Based on Linguistic Readability Vector: Application to Financial Companies in China</article-title>
            <source>International Review of Financial Analysis</source>
            <volume>95</volume>
            <fpage>103405</fpage>
            <elocation-id>ID</elocation-id>
            <pub-id pub-id-type="doi">10.1016/j.irfa.2024.103405</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B13">
        <label>13.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Brown, N.C., Crowley, R.M. and Elliott, W.B. (2020) What Are You Saying? Using <italic>topic</italic> to Detect Financial Misreporting. <italic>Journal</italic><italic>of</italic><italic>Accounting</italic><italic>Research</italic>, 58, 237-291. https://doi.org/10.1111/1475-679x.12294 <pub-id pub-id-type="doi">10.1111/1475-679x.12294</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1111/1475-679x.12294">https://doi.org/10.1111/1475-679x.12294</ext-link></mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Brown, N.C.</string-name>
              <string-name>Crowley, R.M.</string-name>
              <string-name>Elliott, W.B.</string-name>
            </person-group>
            <year>2020</year>
            <article-title>What Are You Saying? Using topic to Detect Financial Misreporting</article-title>
            <source>Journal of Accounting Research</source>
            <volume>58</volume>
            <pub-id pub-id-type="doi">10.1111/1475-679x.12294</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B14">
        <label>14.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Craja, P., Kim, A. and Lessmann, S. (2020) Deep Learning for Detecting Financial Statement Fraud. <italic>Decision</italic><italic>Support</italic><italic>Systems</italic>, 139, Article ID: 113421. https://doi.org/10.1016/j.dss.2020.113421 <pub-id pub-id-type="doi">10.1016/j.dss.2020.113421</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1016/j.dss.2020.113421">https://doi.org/10.1016/j.dss.2020.113421</ext-link></mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Craja, P.</string-name>
              <string-name>Kim, A.</string-name>
              <string-name>Lessmann, S.</string-name>
            </person-group>
            <year>2020</year>
            <article-title>Deep Learning for Detecting Financial Statement Fraud</article-title>
            <source>Decision Support Systems</source>
            <volume>139</volume>
            <fpage>113421</fpage>
            <elocation-id>ID</elocation-id>
            <pub-id pub-id-type="doi">10.1016/j.dss.2020.113421</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B15">
        <label>15.</label>
        <citation-alternatives>
          <mixed-citation publication-type="report">Liu, W., Wang, Z. and Zhang, X. (2025) Research on Financial Fraud Detection by Integrating Latent Semantic Features of Annual Report Text with Accounting Indicators. <italic>Journal</italic><italic>of</italic><italic>Accounting</italic><italic>&amp;</italic><italic>Organizational</italic><italic>Change</italic>, 21, 841-866. https://doi.org/10.1108/jaoc-06-2024-0199 <pub-id pub-id-type="doi">10.1108/jaoc-06-2024-0199</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1108/jaoc-06-2024-0199">https://doi.org/10.1108/jaoc-06-2024-0199</ext-link></mixed-citation>
          <element-citation publication-type="report">
            <person-group person-group-type="author">
              <string-name>Liu, W.</string-name>
              <string-name>Wang, Z.</string-name>
              <string-name>Zhang, X.</string-name>
            </person-group>
            <year>2025</year>
            <article-title>Research on Financial Fraud Detection by Integrating Latent Semantic Features of Annual Report Text with Accounting Indicators</article-title>
            <source>Journal of Accounting &amp; Organizational Change</source>
            <volume>21</volume>
            <pub-id pub-id-type="doi">10.1108/jaoc-06-2024-0199</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B16">
        <label>16.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Hajek, P. and Henriques, R. (2017) Mining Corporate Annual Reports for Intelligent Detection of Financial Statement Fraud—A Comparative Study of Machine Learning Methods. <italic>Knowledge</italic>- <italic>Based</italic><italic>Systems</italic>, 128, 139-152. https://doi.org/10.1016/j.knosys.2017.05.001 <pub-id pub-id-type="doi">10.1016/j.knosys.2017.05.001</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1016/j.knosys.2017.05.001">https://doi.org/10.1016/j.knosys.2017.05.001</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Hajek, P.</string-name>
              <string-name>Henriques, R.</string-name>
            </person-group>
            <year>2017</year>
            <article-title>Mining Corporate Annual Reports for Intelligent Detection of Financial Statement Fraud—A Comparative Study of Machine Learning Methods</article-title>
            <source>Knowledge-Based Systems</source>
            <volume>128</volume>
            <pub-id pub-id-type="doi">10.1016/j.knosys.2017.05.001</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B17">
        <label>17.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Ali, A.A., Khedr, A.M., El-Bannany, M. and Kanakkayil, S. (2023) A Powerful Predicting Model for Financial Statement Fraud Based on Optimized XGBoost Ensemble Learning Technique. <italic>Applied</italic><italic>Sciences</italic>, 13, Article 2272. https://doi.org/10.3390/app13042272 <pub-id pub-id-type="doi">10.3390/app13042272</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.3390/app13042272">https://doi.org/10.3390/app13042272</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Ali, A.A.</string-name>
              <string-name>Khedr, A.M.</string-name>
              <string-name>El-Bannany, M.</string-name>
              <string-name>Kanakkayil, S.</string-name>
            </person-group>
            <year>2023</year>
            <article-title>A Powerful Predicting Model for Financial Statement Fraud Based on Optimized XGBoost Ensemble Learning Technique</article-title>
            <source>Applied Sciences</source>
            <volume>13</volume>
            <elocation-id>2272</elocation-id>
            <pub-id pub-id-type="doi">10.3390/app13042272</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B18">
        <label>18.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Zhao, Z. and Bai, T. (2022) Financial Fraud Detection and Prediction in Listed Companies Using SMOTE and Machine Learning Algorithms. <italic>Entropy</italic>, 24, Article 1157. https://doi.org/10.3390/e24081157 <pub-id pub-id-type="doi">10.3390/e24081157</pub-id><pub-id pub-id-type="pmid">36010821</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.3390/e24081157">https://doi.org/10.3390/e24081157</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Zhao, Z.</string-name>
              <string-name>Bai, T.</string-name>
            </person-group>
            <year>2022</year>
            <article-title>Financial Fraud Detection and Prediction in Listed Companies Using SMOTE and Machine Learning Algorithms</article-title>
            <source>Entropy</source>
            <volume>24</volume>
            <elocation-id>1157</elocation-id>
            <pub-id pub-id-type="doi">10.3390/e24081157</pub-id>
            <pub-id pub-id-type="pmid">36010821</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B19">
        <label>19.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Yadav, A.K.S. (2024) Financial Statement Fraud Detection Using Optimized Deep Neural Network. In: Asirvatham, D., Gonzalez-Longatt, F.M., Falkowski-Gilski, P. and Kanthavel, R., Eds., <italic>Evolutionary Artificial Intelligence</italic>, Springer, 131-141. https://doi.org/10.1007/978-981-99-8438-1_10 <pub-id pub-id-type="doi">10.1007/978-981-99-8438-1_10</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1007/978-981-99-8438-1_10">https://doi.org/10.1007/978-981-99-8438-1_10</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Yadav, A.K.S.</string-name>
              <string-name>Asirvatham, D.</string-name>
              <string-name>Gonzalez-Longatt, F.M.</string-name>
              <string-name>Falkowski-Gilski, P.</string-name>
              <string-name>Kanthavel, R.</string-name>
              <string-name>Intelligence, S</string-name>
            </person-group>
            <year>2024</year>
            <article-title>Financial Statement Fraud Detection Using Optimized Deep Neural Network</article-title>
            <source>In: Asirvatham</source>
            <volume>131</volume>
            <pub-id pub-id-type="doi">10.1007/978-981-99-8438-1_10</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B20">
        <label>20.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Bhattacharya, I. and Mickovic, A. (2024) Accounting Fraud Detection Using Contextual Language Learning. <italic>International</italic><italic>Journal</italic><italic>of</italic><italic>Accounting</italic><italic>Information</italic><italic>Systems</italic>, 53, Article ID: 100682. https://doi.org/10.1016/j.accinf.2024.100682 <pub-id pub-id-type="doi">10.1016/j.accinf.2024.100682</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1016/j.accinf.2024.100682">https://doi.org/10.1016/j.accinf.2024.100682</ext-link></mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Bhattacharya, I.</string-name>
              <string-name>Mickovic, A.</string-name>
            </person-group>
            <year>2024</year>
            <article-title>Accounting Fraud Detection Using Contextual Language Learning</article-title>
            <source>International Journal of Accounting Information Systems</source>
            <volume>53</volume>
            <fpage>100682</fpage>
            <elocation-id>ID</elocation-id>
            <pub-id pub-id-type="doi">10.1016/j.accinf.2024.100682</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B21">
        <label>21.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Wang, G., Ma, J. and Chen, G. (2023) Attentive Statement Fraud Detection: Distinguishing Multimodal Financial Data with Fine-Grained Attention. <italic>Decision</italic><italic>Support</italic><italic>Systems</italic>, 167, Article ID: 113913. https://doi.org/10.1016/j.dss.2022.113913 <pub-id pub-id-type="doi">10.1016/j.dss.2022.113913</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1016/j.dss.2022.113913">https://doi.org/10.1016/j.dss.2022.113913</ext-link></mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Wang, G.</string-name>
              <string-name>Ma, J.</string-name>
              <string-name>Chen, G.</string-name>
            </person-group>
            <year>2023</year>
            <article-title>Attentive Statement Fraud Detection: Distinguishing Multimodal Financial Data with Fine-Grained Attention</article-title>
            <source>Decision Support Systems</source>
            <volume>167</volume>
            <fpage>113913</fpage>
            <elocation-id>ID</elocation-id>
            <pub-id pub-id-type="doi">10.1016/j.dss.2022.113913</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B22">
        <label>22.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Lundberg, S.M., Erion, G., Chen, H., DeGrave, A., Prutkin, J.M., Nair, B., <italic>et al.</italic> (2020) From Local Explanations to Global Understanding with Explainable AI for Trees. <italic>Nature</italic><italic>Machine</italic><italic>Intelligence</italic>, 2, 56-67. https://doi.org/10.1038/s42256-019-0138-9 <pub-id pub-id-type="doi">10.1038/s42256-019-0138-9</pub-id><pub-id pub-id-type="pmid">32607472</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1038/s42256-019-0138-9">https://doi.org/10.1038/s42256-019-0138-9</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Lundberg, S.M.</string-name>
              <string-name>Erion, G.</string-name>
              <string-name>Chen, H.</string-name>
              <string-name>DeGrave, A.</string-name>
              <string-name>Prutkin, J.M.</string-name>
              <string-name>Nair, B.</string-name>
            </person-group>
            <year>2020</year>
            <article-title>From Local Explanations to Global Understanding with Explainable AI for Trees</article-title>
            <source>Nature Machine Intelligence</source>
            <volume>2</volume>
            <pub-id pub-id-type="doi">10.1038/s42256-019-0138-9</pub-id>
            <pub-id pub-id-type="pmid">32607472</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B23">
        <label>23.</label>
        <citation-alternatives>
          <mixed-citation publication-type="report">Zeng, Q.S., Zhou, B., Zhang, C. and Chen, X.Y. (2018) Annual Report Tone and Insider Trading: “Consistent” or “Deceptive”? <italic>Journal of Management World</italic>, 34, 143-160.</mixed-citation>
          <element-citation publication-type="report">
            <person-group person-group-type="author">
              <string-name>Zeng, Q.S.</string-name>
              <string-name>Zhou, B.</string-name>
              <string-name>Zhang, C.</string-name>
              <string-name>Chen, X.Y.</string-name>
            </person-group>
            <year>2018</year>
            <article-title>Annual Report Tone and Insider Trading: “Consistent” or “Deceptive”? Journal of Management World, 34, 143-160</article-title>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B24">
        <label>24.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Li, J., Li, N., Xia, T. and Guo, J. (2023) Textual Analysis and Detection of Financial Fraud: Evidence from Chinese Manufacturing Firms. <italic>Economic</italic><italic>Modelling</italic>, 126, Article ID: 106428. https://doi.org/10.1016/j.econmod.2023.106428 <pub-id pub-id-type="doi">10.1016/j.econmod.2023.106428</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1016/j.econmod.2023.106428">https://doi.org/10.1016/j.econmod.2023.106428</ext-link></mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Li, J.</string-name>
              <string-name>Li, N.</string-name>
              <string-name>Xia, T.</string-name>
              <string-name>Guo, J.</string-name>
            </person-group>
            <year>2023</year>
            <article-title>Textual Analysis and Detection of Financial Fraud: Evidence from Chinese Manufacturing Firms</article-title>
            <source>Economic Modelling</source>
            <volume>126</volume>
            <fpage>106428</fpage>
            <elocation-id>ID</elocation-id>
            <pub-id pub-id-type="doi">10.1016/j.econmod.2023.106428</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B25">
        <label>25.</label>
        <citation-alternatives>
          <mixed-citation publication-type="report">Qian, A. and Zhu, D. (2020) Financial Report Textual Similarity and Likelihood of Regulatory Penalties: Based on the Empirical Evidence of Textual Analysis. <italic>Accounting Research</italic>, 9, 44-58.</mixed-citation>
          <element-citation publication-type="report">
            <person-group person-group-type="author">
              <string-name>Qian, A.</string-name>
              <string-name>Zhu, D.</string-name>
            </person-group>
            <year>2020</year>
            <article-title>Financial Report Textual Similarity and Likelihood of Regulatory Penalties: Based on the Empirical Evidence of Textual Analysis</article-title>
            <source>Accounting Research</source>
            <volume>9</volume>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B26">
        <label>26.</label>
        <citation-alternatives>
          <mixed-citation publication-type="confproc">Chen, T. and Guestrin, C. (2016) XGBoost: A Scalable Tree Boosting System. <italic>Proceedings of the</italic> 22 <italic>nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining</italic>, San Francisco, 13-17 August 2016, 785-794. https://doi.org/10.1145/2939672.2939785 <pub-id pub-id-type="doi">10.1145/2939672.2939785</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1145/2939672.2939785">https://doi.org/10.1145/2939672.2939785</ext-link></mixed-citation>
          <element-citation publication-type="confproc">
            <person-group person-group-type="author">
              <string-name>Chen, T.</string-name>
              <string-name>Guestrin, C.</string-name>
              <string-name>Mining, S</string-name>
            </person-group>
            <year>2016</year>
            <article-title>XGBoost: A Scalable Tree Boosting System</article-title>
            <source>Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining</source>
            <volume>13</volume>
            <pub-id pub-id-type="doi">10.1145/2939672.2939785</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B27">
        <label>27.</label>
        <citation-alternatives>
          <mixed-citation publication-type="book">Shapley, L.S. (1953) 17. A Value for n-Person Games. In: Kuhn, H.W. and Tucker, A.W., Eds., <italic>Contributions</italic><italic>to</italic><italic>the</italic><italic>Theory</italic><italic>of</italic><italic>Games</italic> ( <italic>AM</italic>-28), <italic>Volume</italic><italic>II</italic>, Princeton University Press, 307-318. https://doi.org/10.1515/9781400881970-018 <pub-id pub-id-type="doi">10.1515/9781400881970-018</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1515/9781400881970-018">https://doi.org/10.1515/9781400881970-018</ext-link></mixed-citation>
          <element-citation publication-type="book">
            <person-group person-group-type="author">
              <string-name>Shapley, L.S.</string-name>
              <string-name>Kuhn, H.W.</string-name>
              <string-name>Tucker, A.W.</string-name>
              <string-name>II, P</string-name>
            </person-group>
            <year>1953</year>
            <article-title>17</article-title>
            <source>A Value for n-Person Games. In: Kuhn</source>
            <volume>307</volume>
            <pub-id pub-id-type="doi">10.1515/9781400881970-018</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B28">
        <label>28.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Cohen, J.R., Krishnamoorthy, G. and Wright, A. (2004) The Corporate Governance Mosaic and Financial Reporting Quality. <italic>Journal of Accounting Literature</italic>, 23, 87-152.</mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Cohen, J.R.</string-name>
              <string-name>Krishnamoorthy, G.</string-name>
              <string-name>Wright, A.</string-name>
            </person-group>
            <year>2004</year>
            <article-title>The Corporate Governance Mosaic and Financial Reporting Quality</article-title>
            <source>Journal of Accounting Literature</source>
            <volume>23</volume>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B29">
        <label>29.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Dechow, P.M., Ge, W., Larson, C.R. and Sloan, R.G. (2011) Predicting Material Accounting Misstatements. <italic>Contemporary</italic><italic>Accounting</italic><italic>Research</italic>, 28, 17-82. https://doi.org/10.1111/j.1911-3846.2010.01041.x <pub-id pub-id-type="doi">10.1111/j.1911-3846.2010.01041.x</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1111/j.1911-3846.2010.01041.x">https://doi.org/10.1111/j.1911-3846.2010.01041.x</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Dechow, P.M.</string-name>
              <string-name>Ge, W.</string-name>
              <string-name>Larson, C.R.</string-name>
              <string-name>Sloan, R.G.</string-name>
            </person-group>
            <year>2011</year>
            <article-title>Predicting Material Accounting Misstatements</article-title>
            <source>Contemporary Accounting Research</source>
            <volume>28</volume>
            <pub-id pub-id-type="doi">10.1111/j.1911-3846.2010.01041.x</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B30">
        <label>30.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Tamimi, N. and Sebastianelli, R. (2017) Transparency among S&amp;P 500 Companies: An Analysis of ESG Disclosure Scores. <italic>Management</italic><italic>Decision</italic>, 55, 1660-1680. https://doi.org/10.1108/md-01-2017-0018 <pub-id pub-id-type="doi">10.1108/md-01-2017-0018</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1108/md-01-2017-0018">https://doi.org/10.1108/md-01-2017-0018</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Tamimi, N.</string-name>
              <string-name>Sebastianelli, R.</string-name>
            </person-group>
            <year>2017</year>
            <article-title>Transparency among S&amp;P 500 Companies: An Analysis of ESG Disclosure Scores</article-title>
            <source>Management Decision</source>
            <volume>55</volume>
            <pub-id pub-id-type="doi">10.1108/md-01-2017-0018</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B31">
        <label>31.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Yu, E.P., Luu, B.V. and Chen, C.H. (2020) Greenwashing in Environmental, Social and Governance Disclosures. <italic>Research</italic><italic>in</italic><italic>International</italic><italic>Business</italic><italic>and</italic><italic>Finance</italic>, 52, Article ID: 101192. https://doi.org/10.1016/j.ribaf.2020.101192 <pub-id pub-id-type="doi">10.1016/j.ribaf.2020.101192</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1016/j.ribaf.2020.101192">https://doi.org/10.1016/j.ribaf.2020.101192</ext-link></mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Yu, E.P.</string-name>
              <string-name>Luu, B.V.</string-name>
              <string-name>Chen, C.H.</string-name>
              <string-name>Environmental, S</string-name>
            </person-group>
            <year>2020</year>
            <article-title>Greenwashing in Environmental, Social and Governance Disclosures</article-title>
            <source>Research in International Business and Finance</source>
            <volume>52</volume>
            <fpage>101192</fpage>
            <elocation-id>ID</elocation-id>
            <pub-id pub-id-type="doi">10.1016/j.ribaf.2020.101192</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B32">
        <label>32.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Chen, G.J., Lin, H. and Wang, L. (2005) Corporate Governance, Reputation Mechanism and the Behavior of Listed Firms in Committing in Fraud. <italic>Nankai Business Review</italic>, 6, 35-40.</mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Chen, G.J.</string-name>
              <string-name>Lin, H.</string-name>
              <string-name>Wang, L.</string-name>
              <string-name>Governance, R</string-name>
            </person-group>
            <year>2005</year>
            <article-title>Corporate Governance, Reputation Mechanism and the Behavior of Listed Firms in Committing in Fraud</article-title>
            <source>Nankai Business Review</source>
            <volume>6</volume>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B33">
        <label>33.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Pan, W.B. and Hong, Y. (2017) Characteristics of Companies with Financial Fraud based on Time-Dependent COX Model. <italic>Journal of University of Science and Technology of China</italic>, 47, 255-261.</mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Pan, W.B.</string-name>
              <string-name>Hong, Y.</string-name>
            </person-group>
            <year>2017</year>
            <article-title>Characteristics of Companies with Financial Fraud based on Time-Dependent COX Model</article-title>
            <source>Journal of University of Science and Technology of China</source>
            <volume>47</volume>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B34">
        <label>34.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Jensen, M.C. and Meckling, W.H. (1976) Theory of the Firm: Managerial Behavior, Agency Costs and Ownership Structure. <italic>Journal</italic><italic>of</italic><italic>Financial</italic><italic>Economics</italic>, 3, 305-360. https://doi.org/10.1016/0304-405x(76)90026-x <pub-id pub-id-type="doi">10.1016/0304-405x(76)90026-x</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1016/0304-405x(76)90026-x">https://doi.org/10.1016/0304-405x(76)90026-x</ext-link></mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Jensen, M.C.</string-name>
              <string-name>Meckling, W.H.</string-name>
              <string-name>Behavior, A</string-name>
            </person-group>
            <year>1976</year>
            <article-title>Theory of the Firm: Managerial Behavior, Agency Costs and Ownership Structure</article-title>
            <source>Journal of Financial Economics</source>
            <volume>3</volume>
            <pub-id pub-id-type="doi">10.1016/0304-405x(76)90026-x</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B35">
        <label>35.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Larcker, D.F., Richardson, S.A. and Tuna, I. (2007) Corporate Governance, Accounting Outcomes, and Organizational Performance. <italic>The</italic><italic>Accounting</italic><italic>Review</italic>, 82, 963-1008. https://doi.org/10.2308/accr.2007.82.4.963 <pub-id pub-id-type="doi">10.2308/accr.2007.82.4.963</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.2308/accr.2007.82.4.963">https://doi.org/10.2308/accr.2007.82.4.963</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Larcker, D.F.</string-name>
              <string-name>Richardson, S.A.</string-name>
              <string-name>Tuna, I.</string-name>
              <string-name>Governance, A</string-name>
            </person-group>
            <year>2007</year>
            <article-title>Corporate Governance, Accounting Outcomes, and Organizational Performance</article-title>
            <source>The Accounting Review</source>
            <volume>82</volume>
            <pub-id pub-id-type="doi">10.2308/accr.2007.82.4.963</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B36">
        <label>36.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Yang, Q., Yu, L. and Chen, N. (2009) Board Characters and Financial Fraud: Empirical Evidence from Chinese Listed Companies. <italic>Accounting Research</italic>, 7, 64-70, 96.</mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Yang, Q.</string-name>
              <string-name>Yu, L.</string-name>
              <string-name>Chen, N.</string-name>
            </person-group>
            <year>2009</year>
            <article-title>Board Characters and Financial Fraud: Empirical Evidence from Chinese Listed Companies</article-title>
            <source>Accounting Research</source>
            <volume>7</volume>
          </element-citation>
        </citation-alternatives>
      </ref>
    </ref-list>
  </back>
</article>